Neural networks
Dense CPU networks, reverse mode differentiation, SGD, and concise train fitted classifier pipelines.
Aner Neural is the built in neural network library, imported as aner.nn. It provides rank two Float64 tensors, native CPU operations, reverse mode automatic differentiation, and stochastic gradient descent. Aner programs describe a model through the concise classifier API or an explicit training loop; Aner's native C++ runtime performs the tensor operations.
The current library supports small dense neural networks. CNNs, RNNs, Transformers, GPU execution, distributed training, streaming training over external datasets, and model checkpoints are future work. The reference kernels establish understandable semantics and correctness before performance work.
Run the first model
Follow the one time installation guide to put aner on your PATH. Open a terminal in the folder containing the binary bundle's examples directory, then check and run the XOR model:
aner check examples/neural_xor.aner
aner run examples/neural_xor.anerUse the same commands on macOS, Linux, or Windows, including VS Code's terminal. No separate neural library installation is needed: aner.nn ships with Aner.
The XOR example trains on four synthetic observations, using two input features, eight hidden units with tanh, and one output with sigmoid. It computes mean squared error, runs 3,000 full batch updates at learning rate 0.5, and prints the initial loss, final loss, and four predictions. The target values are 0, 1, 1, 0, in that order. This demonstrates that gradients and parameter updates compose into a working training loop; it is not a medical data model or a benchmark for large datasets.
A recorded reference run produced these rounded results:
| Result | Value |
|---|---|
| Initial mean squared error | 0.265645 |
| Final mean squared error | 0.000346018 |
Prediction for [0, 0] | 0.011831 |
Prediction for [0, 1] | 0.980513 |
Prediction for [1, 0] | 0.980293 |
Prediction for [1, 1] | 0.021817 |
Your output should show the loss decreasing and predictions approaching the four targets. These reference values are not a guarantee of identical final decimals across platforms.
Concise dense classifier pipelines
The built in classifier API groups layer parameters and the repeated training loop into native values. It uses the same tensor kernels and autograd as the detailed examples.
import aner.dataset
import aner.nn
import aner.ml
let flowers = dataset.iris()
let split = flowers.split(test: 0.2, seed: 2026)
let network = nn.sequential(seed: 42)
.dense(inputs: 4, outputs: 8)
.tanh()
.dense(inputs: 8, outputs: 3)
let pipeline = ml.classifier(network, standardize: true)
let fitted = pipeline.fit(split.train(), epochs: 2000, rate: 0.1)
print(fitted.evaluate(split.test()).summary())Run aner run examples/neural_iris_concise.aner. The Wine example uses 13 → 16 → 3 dimensions, 1,000 epochs, and rate 0.05. The Iris notebook separates data, topology, fitting, and evaluation into persistent cells. The older detailed Iris program retains its explicit forward/backward/update loop and visual capture calls.
| Call | Result | Meaning |
|---|---|---|
nn.sequential(seed: Int64) | NeuralNetwork | Start an empty immutable network configuration with an explicit initialization seed. |
network.dense(inputs: Int64, outputs: Int64) | NeuralNetwork | Return a configuration with a dense stage; adjacent dimensions must agree. |
network.tanh(), network.relu(), network.sigmoid() | NeuralNetwork | Return a configuration with the named activation stage. |
ml.classifier(network, standardize: Bool) | ClassifierPipeline | Configure a multiclass neural classifier with explicit preprocessing policy. |
pipeline.fit(training: Dataset, epochs: Int64, rate: Float64) | FittedClassifier | Train a fresh fitted snapshot with cross entropy and full batch SGD. |
fitted.evaluate(testing: Dataset) | EvaluationReport | Reuse the fitted preprocessing and evaluate the supplied rows without refitting. |
fitted.summary(), report.summary() | String | Return readable configuration, results, and dataset provenance; the evaluation summary includes the fitted summary. |
fitted.initial_loss(), fitted.training_loss(), fitted.training_accuracy() | Float64 | Read detached training metrics without printing the full report. |
report.loss(), report.accuracy(), report.rows() | Float64, Float64, Int64 | Read evaluation metrics and evaluated row count. |
fitted.predict(samples), fitted.probabilities(samples) | Tensor | Apply fitted preprocessing and return class IDs or probabilities for a compatible Dataset. |
Use all three imports above: aner.dataset for teaching data and partitions, aner.nn for topology, and aner.ml for fitted classification. The names are library APIs, not additional language keywords. Instance calls are resolved from the receiver's static type. Named arguments can make positional configuration clearer without changing type checking. A leading . on the next line continues a method chain.
Configurations and fitted snapshots are separate immutable values. dense and activation methods return a new configuration; they do not mutate an existing let binding. Every fit starts fresh from the network configuration, rather than resuming or changing an earlier fitted model. Standardization, when requested, is fitted solely on the supplied training features. Test features reuse exactly those statistics. This protects that preprocessing boundary; it cannot determine whether the caller chose medically independent patients or a valid study partition.
The current trainer uses CPU Float64 tensors, fixed random weight scale 0.5, zero biases, cross entropy over logits, and full batch SGD. Epoch count and finite positive learning rate are explicit; this API does not expose a configurable loss or optimizer yet. Dense stages use consecutive seeds base_seed + dense_index, where the first Dense stage has index zero and activations do not increment the index. Seed overflow is rejected. No parameter tensors are created until fit runs. The architecture must start with a Dense stage and end with Dense raw logits; its input width must match the dataset feature count and its output width must match the dataset class count. Copies of a fitted value share its immutable snapshot efficiently; fitting another model does not mutate that snapshot. The low level Tensor API still has shared gradient metadata, as documented below.
Summaries preserve topology and per layer seeds, training configuration and metrics, CPU/precision information, fitted scaling statistics, source hashes, exact source row IDs, and dataset attribution. Evaluation checks feature names, units, column order, feature count, and class names/order against the fitted schema. Its summary reports overlap with training rows; the API does not assume that every supplied view is held out. Retain the Aner source with the output: split seed and fraction appear in source but are not retained automatically by dataset views. A fixed seed supports repeatability in the same build and environment; it does not guarantee bitwise identity across compilers or future devices. Saved notebook output is not a serialized fitted model or checkpoint.
Training is bounded by 32 stages, 1,000,000 trainable parameters, 100,000 epochs, and 500,000,000 estimated training work units, in addition to the tensor/graph limits below. The work estimate is a conservative guard, not elapsed time, measured FLOPs, or a process memory quota. Raising the interpreter's step limit does not raise these library limits. Library failures use source located R2701 diagnostics. Large data streaming training, live notebook loss charts, checkpointing, GPU execution, and user defined layer types remain future work.
Create a visual training report
aner run examples/neural_iris_concise.aner --debug-viz reports/iris-conciseChoose a new output directory. The CLI records the native training graph after backward and before each selected update: epochs 0, 1, 2, every max(100, ceil(epochs / 60)) epochs, and the final trained state. This keeps one fit within the recorder's 64 snapshot limit; the Iris example records every 100 epochs. Add --debug-values only when you want bounded raw tensor/gradient values in the report. Capture does not change the model's numerical updates. This is an offline HTML/JSON/text report, not a live notebook chart.
Use one fit call per visual report. Each fit's capture steps begin at zero, while a recorder requires strictly increasing steps; multiple fits in the same recorded CLI program currently fail that check. The detailed Iris example remains useful for choosing your own viz.capture points. Dataset attribution travels with recorded training reports.
Source syntax and tensor shapes
import aner.nn;
fn main() -> Unit {
let input: Tensor = [[0.0, 1.0], [1.0, 0.0]];
var weights: Tensor = nn.parameter(nn.random(2, 1, 42, 0.5));
let output = nn.matmul(input, weights);
print(nn.value(output, 0, 0));
}Place import aner.nn once before function declarations or direct statements. A semicolon is optional at a completed line. The built in aner.viz module may also be imported for optional visual captures. Import aliases and user defined modules are not implemented.
Tensor holds a nonempty rectangular array of Float64 values on the CPU. Every tensor has two dimensions: a single number is represented by shape 1 × 1, a row vector by 1 × n, and a column vector by n × 1. Literal rows must have equal lengths and cells may be expressions but must have type Float64; integer values are not implicitly converted. Generic Vector<T> and Matrix<T> are not introduced by this module.
Tensor values are immutable. Copying a tensor binding shares its underlying storage and differentiation graph. Gradients are graph metadata updated by nn.backward, so handles sharing a graph observe the same gradient state. Concurrent differentiation or gradient access on a shared graph is unsupported; this reference backend executes on one CPU thread. There is no writable element indexing. Read a cell with nn.value, extract a scalar with nn.item, or calculate a new tensor with a library operation. Arithmetic operators such as +, *, and @ do not operate on Tensor; use the explicit functions below.
Shapes are checked at runtime. Elementwise operations require equal shapes; matrix multiplication requires matching inner dimensions. There is no general implicit broadcasting. nn.add_bias is the explicit exception for adding a single bias row to every row of an activation tensor.
print accepts scalars, not whole tensors. nn.rows, nn.cols, nn.value, and nn.item provide bounded ways to inspect a result.
Current API
The table uses positional calls with descriptive placeholders. Exact named argument labels are listed in the language reference. In the table, a, b, x, prediction, target, and parameter are Tensor; rows, cols, row, col, and seed are Int64; scale, factor, and learning_rate are Float64.
| Call | Result | Contract |
|---|---|---|
nn.parameter(x) | Tensor | Make an independent trainable leaf from x's values. |
nn.zeros(rows, cols) | Tensor | Non trainable zero tensor with positive dimensions. |
nn.random(rows, cols, seed, scale) | Tensor | Seeded non trainable values in [-scale, scale); finite nonnegative scale. |
nn.add(a, b) | Tensor | Elementwise addition; equal shapes. |
nn.subtract(a, b) | Tensor | Elementwise subtraction; equal shapes. |
nn.multiply(a, b) | Tensor | Elementwise multiplication; equal shapes. |
nn.matmul(a, b) | Tensor | (m × k) × (k × n) produces m × n. |
nn.add_bias(a, b) | Tensor | a is m × n; b must be 1 × n. |
nn.scale(a, factor) | Tensor | Multiply every element by one finite scalar. |
nn.transpose(a) | Tensor | Exchange rows and columns. |
nn.tanh(x) | Tensor | Elementwise hyperbolic tangent. |
nn.sigmoid(x) | Tensor | Elementwise logistic sigmoid. |
nn.relu(x) | Tensor | Elementwise maximum of zero and the value. |
nn.mean(x) | Tensor | Mean of all cells, returned as 1 × 1. |
nn.mse(prediction, target) | Tensor | Mean squared difference over all cells; equal shapes; 1 × 1 result. |
nn.backward(loss) | Unit | Differentiate a trainable 1 × 1 loss; replace gradients in its reachable graph. |
nn.grad(x) | Tensor | Read a valid previously computed gradient as a detached tensor. |
nn.detach(x) | Tensor | Independent value tensor with no differentiation history. |
nn.sgd(parameter, learning_rate) | Tensor | Return a new trainable leaf after one gradient step; finite positive learning rate. |
nn.item(x) | Float64 | Extract the sole value; requires 1 × 1. |
nn.value(x, row, col) | Float64 | Read a cell with zero based, bounds checked indices. |
nn.rows(x) | Int64 | Number of rows. |
nn.cols(x) | Int64 | Number of columns. |
All tensor values, operation results, and gradients must remain finite. Nonfinite input or computed results produce errors instead of silently carrying NaNs or infinities through training. This tensor contract is stricter than Aner's scalar Float64 arithmetic. Shape, bounds, gradient state, and resource limit failures also report errors.
Differentiation and parameter updates
nn.parameter starts a trainable leaf. Operations involving a trainable tensor record the graph needed for a reverse pass. Ordinary literals, zeros, and random are constants until wrapped in parameter. Tensor handles can share graph nodes; when a parameter reaches the loss along multiple paths, its contributions are summed.
nn.backward(loss) requires a differentiable 1 × 1 tensor. Each call clears the previous gradients for nodes reachable from that loss and calculates new ones; it does not accumulate gradients across calls or batches. A failed backward attempt invalidates gradients for its reachable graph. Nodes outside that graph are unaffected. Shared graph state currently supports single thread access; concurrent backward or optimizer operations are not supported. nn.grad(x) requires a valid gradient from a completed reverse pass involving x.
nn.sgd(p, rate) requires a trainable leaf with a valid gradient from a preceding successful backward pass and a finite, positive learning rate. It computes p - rate * gradient and returns a new trainable leaf, disconnected from the old graph. It does not mutate p. Write p = nn.sgd(p, rate) on a var binding, as the XOR example does. All parameters should be updated after computing gradients from the same loss; recompute the forward pass for the next step. Do not retain every iteration's loss or intermediate tensors, since retained handles keep their graphs alive.
nn.detach removes differentiation history. nn.grad, nn.item, and nn.value do not provide higher order differentiation. ReLU uses a zero derivative at zero. Higher order gradients, persistent gradient accumulation, optimizer state such as Adam, and checkpointing are not part of this release.
Limits and reproducibility
The reference implementation is intentionally bounded:
| Limit | Maximum |
|---|---|
| Elements in one tensor | 1,000,000 |
| Multiply add iterations in one matrix multiplication | 50,000,000 |
| Differentiation graph depth | 256 |
| Nodes reachable in one differentiation graph | 4,096 |
| Elements across one reachable differentiation graph | 4,000,000 |
Positive dimensions and checked size arithmetic are required before allocation. These limits constrain individual operations and graphs; they are not a global process memory quota. The interpreter's source size, nesting, call depth, and execution step limits also still apply. Raising --max-steps does not raise tensor or graph limits.
nn.random uses SplitMix64 and maps the upper 53 bits of each generated word into [0, 1), then into [-scale, scale). A signed Int64 seed is mapped modulo 2^64; it determines the initialization, and no hidden global random generator controls the example. A zero scale returns zeros. The CPU kernels use a defined serial traversal, making runs reproducible in the same build and environment. Do not assume bitwise equality across compilers, mathematical libraries, devices, or future parallel backends. Record the source, seeds, training settings, Aner version, and environment when comparing experiments.
Built in modules and current scope
import aner.nn; selects a built in module and exposes the nn namespace. It does not load external packages. Third party native plugins, module aliases, package manifests, dependency resolution, version negotiation, and stable extension ABIs are not supported.
Use the built in sequential/classifier values for common dense classification, or ordinary Aner functions and explicit training loops for lower level experiments. The current backend runs on one CPU thread. Multicore acceleration, GPUs, clusters, and supercomputers are not available yet.
Keep the data schema and partitions, preprocessing, model architecture, seeds, optimization settings, software version, and metrics with your experiment. Model saving and checkpointing are not available yet. The built in teaching catalogue and preprocessing are documented in Aner datasets; aner.data reads external CSV files with explicit table to tensor conversion.
Observe topology and training
The visual training example imports aner.viz and captures the loss graph after backward at selected steps. Its offline report connects the model topology to recorded values, gradients, and loss history. See visual training reports for small neuron views, bounded summaries, and headless execution. Captures add overhead but do not alter numerical tensors, gradients, or parameter updates.
Multiclass classification
The Iris milestone adds nn.softmax(logits: Tensor) -> Tensor, nn.cross_entropy(logits: Tensor, one_hot_targets: Tensor) -> Tensor, nn.argmax(scores: Tensor) -> Tensor, and nn.accuracy(scores: Tensor, labels: Tensor) -> Float64.
Softmax normalizes each row and supports differentiation. Cross entropy accepts raw logits and matching constant, exact one hot targets; its scalar loss is averaged over rows. It combines normalization and logarithms stably rather than taking the log of rounded softmax probabilities. Logit gradients are (probabilities - targets) / rows. Differentiable targets and soft label distributions are rejected in this first version. Nonrepresentable results remain errors.
Argmax returns a detached rows × 1 Float64 tensor of zero based class IDs, choosing the first index on ties. Accuracy accepts one finite integral in range class ID per row and returns the correctly classified fraction; scores may be logits or probabilities. These metrics do not add autograd edges.
Run the Iris classifier, using the licensed dataset module, for a three class example with training only scaling, held out evaluation, and visual captures.
Shared ML ecosystem
Neural architecture, differentiable losses, and optimization remain in aner.nn. Classical fitted models and the new classifier pipeline live in aner.ml; common evaluation lives in aner.metrics. The Iris neural example now uses metrics.accuracy(nn.argmax(scores), labels), the same label contract as KNN. Existing nn.accuracy(scores, labels) is unchanged. metrics.mean_squared_error is a detached evaluation scalar, whereas nn.mse is a differentiable loss tensor.
The neutral aner.tensor module offers creation and inspection of the same native Tensor values. See classical ML APIs for KNN, K means, and shared evaluation metrics.