Browse the handbook
Modules

Datasets and provenance

Load Iris and Wine offline with attribution, explicit splits, source hashes, and training only preprocessing.

aner.dataset provides bundled teaching data, its provenance, immutable row views, reproducible splits, and explicit feature scaling. Iris and Wine are available offline through the same API. The first implementation covers finite numeric tables with one categorical target; image and temporal sources need additional readers and types. The separate aner.data module now provides external CSV sources with typed text/numeric/Bool columns, missingness, and budgeted scans; it does not expand this teaching catalogue.

Included datasets

IDRowsNumeric featuresClassesSource license
iris1504 flower measurements in cm3 species, 50 rows eachCC BY 4.0
wine17813 chemical measurements3 cultivar codes, counts 59, 71, 48CC BY 4.0

Source: UCI Iris and UCI Wine. The dataset source assets retain both raw files, original archives, source notes, full license text, and SHA-256 hashes. The selected data files are hash verified, embedded in Aner, and ready to use offline.

Read third party dataset notices and the asset manifest for attribution, exact feature order, class mappings, missing value counts, source dates, and variant details. Run aner --data-notices to print attribution and the full data license from the executable. Dataset licenses are separate from Aner's software license.

Aner preserves original UCI iris.data. UCI documents corrections for records 35 and 38 in another variant; Aner does not silently apply them. The source also has duplicate complete records. Wine's source does not provide feature units or cultivar names: the API returns unknown units and class_1, class_2, class_3. Unknown metadata is not inferred from feature names.

Load and inspect data

aner
import aner.dataset;

fn main() -> Unit {
    let flowers: Dataset = dataset.iris();
    print(dataset.name(flowers));
    print(dataset.rows(flowers));
    print(dataset.feature_name(flowers, 0));
    print(dataset.feature_unit(flowers, 0));
    print(dataset.citation(flowers));
    print(dataset.sha256(flowers));
}

dataset.load("iris") and dataset.iris() return the same dataset. Catalogue discovery uses dataset.count() and zero based dataset.id(index). Run examples/dataset_catalog.aner to inspect both datasets and all feature/class names.

FunctionResult
dataset.load(id: String), dataset.iris(), dataset.wine()Dataset
dataset.count()Int64 catalogue size
dataset.id(index: Int64)String catalogue ID
dataset.name(data), description(data), citation(data), license(data), source(data), version(data), sha256(data)String metadata; each function uses the dataset. prefix
dataset.rows(data), cols(data), classes(data)Int64 dimensions of the current view and its class vocabulary
dataset.feature_name(data, index), feature_unit(data, index), class_name(data, index)String metadata at a zero based index
dataset.features(data)Detached Tensor of shape rows × features, in original units
dataset.targets(data)Detached rows × 1 Tensor of zero based class IDs
dataset.one_hot(data)Detached rows × classes Tensor with exactly one 1 per row
dataset.row_ids(data)Detached rows × 1 Tensor of original zero based source row IDs

Current tensors store Float64, including the exactly represented integer class/row IDs. The dataset's class name mapping defines their meaning. A model must never treat class IDs as continuous measurements or append them to its input features.

Dataset, DatasetSplit, and Standardizer are immutable built in value types. Copies share immutable backing storage; tensor extraction produces detached data. They can be function parameters and return values, with local type inference as usual. Import aner.dataset to use these types. It also exposes Tensor for its data APIs; neutral inspection uses import aner.tensor;, while calls to nn.* still require import aner.nn;. Composite values cannot be printed or compared directly: inspect their fields through the API.

Split before fitting preprocessing

aner
let parts: DatasetSplit = dataset.split(flowers, 0.2, 2026);
let training = dataset.train(parts);
let testing = dataset.test(parts);
let scaler: Standardizer = dataset.fit_standardizer(dataset.features(training));
let x_train = dataset.transform(scaler, dataset.features(training));
let x_test = dataset.transform(scaler, dataset.features(testing));

dataset.split(data, test_fraction: Float64, seed: Int64) stratifies by class. The fraction must be finite and strictly between zero and one. Each class needs at least two rows. For each class of size n, the test count is floor(n * fraction + 0.5), clamped to 1 through n−1. The actual overall test fraction can differ from the requested fraction, especially for small views. Iris at 0.2 yields 120/30; Wine yields 142/36.

SplitMix64, classes in ascending ID order, and rejection sampled Fisher and Yates define the permutation without standard library distribution differences. Returned views keep their selected rows in original source order. Train/test views are disjoint by source row ID and cover the input view. Repeating the seed and source version reproduces membership; dataset.row_ids makes it inspectable. Splitting a view again follows the same rules.

A row split does not deduplicate observations or group related subjects. Iris's repeated records can occur on opposite sides. This teaching example is not a rigorous generalization benchmark. Future clinical datasets require explicit patient/group/time aware splitting and validation; class stratification alone cannot provide that guarantee.

dataset.fit_standardizer(train_features) computes each column's population mean and standard deviation. Constant columns use scale 1. dataset.transform(scaler, features) applies those fixed statistics and never refits. dataset.center(scaler) and dataset.scale(scaler) expose 1 × feature count tensors. Wrong shapes, invalid handles, nonfinite fractions, and unrepresentable finite results produce diagnostics. Preprocessing is detached from automatic differentiation.

Train the Iris classifier

Follow the one time installation guide to put aner on your PATH. Run these commands from the folder containing the binary bundle's examples directory:

sh
aner run examples/dataset_catalog.aner
aner check examples/neural_iris.aner
aner run examples/neural_iris.aner

The same commands work on macOS, Linux, and Windows. Aner interprets the .aner source and performs native CPU Float64 computation; the example does not invoke another ML framework.

The runnable example fixes the split seed at 2026 and initialization seeds at 42 and 43. Its architecture is four standardized inputs → eight tanh hidden units → three linear logits, with 67 trainable scalars. It trains for 2000 full batch epochs using SGD at learning rate 0.1. Here one optimizer step processes the entire training set, so completed steps and completed epochs coincide. No test accuracy is used for updates, early stopping, or hyperparameter selection.

nn.cross_entropy(logits, one_hot_targets) computes the stable mean multiclass loss over rows directly from raw logits. Do not pass softmax probabilities to that function. Targets must be exact one hot constant tensors. nn.softmax(logits) produces row wise probabilities for inspection, nn.argmax(scores) returns class IDs, and nn.accuracy(scores, class_ids) returns a fraction between 0 and 1. The example evaluates the same class IDs through shared metrics.accuracy(nn.argmax(scores), class_ids) from aner.metrics; the older nn.accuracy API remains available. Argmax breaks ties at the first class index. Accuracy uses all rows supplied by the caller and does not create a test split itself.

The program prints source identity, seeds, partition sizes, training settings, training loss, and final train/test accuracy. A single 30 row test result is a teaching sanity check with substantial sampling uncertainty. Bitwise floating point equality across different hardware/compiler/library versions is not promised.

Create a visual training report

sh
aner run examples/neural_iris.aner --debug-viz reports/iris --debug-values

Use a new output directory. Open reports/iris/index.html in a browser or read summary.txt from a terminal. The trace stores 23 training captures: 0, 1, 2, each subsequent hundred, and 2000. It includes only training graph observations; the test set is evaluated after training. In the output layer, the captured values are logits, not probabilities. Cross entropy gradients refer to the full training batch loss.

--debug-values exports bounded raw cells, including standardized features and model derived tensors. Omit it for statistics. Loaded dataset attribution is retained in the report; this records which datasets were loaded, not complete per tensor data lineage. See visual capture limits before expecting every cell of a larger network to appear. VS Code also provides Aner: Debug Iris.

Reuse the data for other algorithms

Neural classification uses features, one_hot, and class IDs for evaluation. CPU KNN and K means are now implemented in aner.ml. KNN Iris fits feature/class ID tensors and uses shared label accuracy and a confusion matrix. K means Iris fits feature tensors without labels and evaluates held out cluster assignments using adjusted Rand index; cluster IDs are not species IDs. Both reuse the neural example's split and training only scaler. See classical ML. SVMs and decision trees remain future algorithms.

External sources and table to model conversion

The Iris/Wine teaching catalogue is built into Aner. Its manifest describes provenance; it is not a third party dataset loader. Use aner.data for your own external CSV files. DataSchema declares types and nullability; DataScan builds a deferred plan; summaries use native batches; DataTable is an explicitly bounded collection. The existing catalogue Dataset and its class metadata remain separate types. CSV files do not automatically acquire target labels, class vocabularies, unit semantics, or stratified splits.

data.to_tensor(table, max_output_bytes) converts a complete numeric table explicitly, enforcing the current Tensor limit and rejecting missing or nonnumeric columns and Int64 values outside ±2^53. Dataset scalers and ML functions may then consume the result. This allocates new numerical storage; the original table can remain alive. The current KNN, K means and neural training routines do not become streaming models through a streaming reader.

For your own data, retain a source identifier, version and content hash, applicable license and notices, feature order and units, missing value policy, target/task description, and label mapping. Keep original source files and record conversions separately. Aner does not automatically create this metadata for external files.

Aner handbook · Guides and API reference