Browse the handbook
Modules

Classical machine learning

Fit KNN and K means models and evaluate predictions with the shared tensor and metrics modules.

aner.ml provides a CPU K nearest neighbor classifier and a K means clustering model. aner.metrics evaluates ID predictions, numeric predictions, and partitions. aner.tensor provides neutral tensor construction and inspection, so classical ML programs do not need to import the neural API. All use the same rank two Float64 Tensor data as aner.dataset and aner.nn.

These models use native CPU operations through Aner's interpreter. The current implementation does not compile Aner programs or provide GPU execution, general packages, third party model registration, SVM, or a fitted preprocessing pipeline for KNN and K means.

Run the examples

Follow the one time installation guide to put aner on your PATH. From the folder containing the binary bundle's examples directory:

sh
aner check examples/ml_iris.aner
aner run examples/ml_iris.aner
aner check examples/kmeans_iris.aner
aner run examples/kmeans_iris.aner

The same commands work on macOS, Linux, and Windows, including VS Code's terminal. In VS Code, you can also open either file and use the current file check/run tasks. Extension 0.1.16 highlights the modules and fitted types in the active light or dark theme.

Both runnable examples use split seed 2026, 120 fitting rows and 30 held out rows, with training fitted standardization. KNN uses k = 5. K means uses three clusters, initialization seed 42, at most 100 updates, and tolerance 0.000001; it reports a held out adjusted Rand comparison, not classification accuracy.

These examples reuse the licensed offline Iris data. Read dataset provenance and splitting and third party notices. Small historical teaching datasets, including Iris's preserved duplicate observations, do not establish clinical performance or huge dataset scalability.

A classifier from fit to prediction

aner
import aner.dataset;
import aner.tensor;
import aner.ml;
import aner.metrics;

fn main() -> Unit {
    let flowers = dataset.iris();
    let parts = dataset.split(flowers, 0.2, 2026);
    let training = dataset.train(parts);
    let testing = dataset.test(parts);
    let scaler = dataset.fit_standardizer(dataset.features(training));
    let x_train = dataset.transform(scaler, dataset.features(training));
    let x_test = dataset.transform(scaler, dataset.features(testing));
    let y_train = dataset.targets(training);
    let y_test = dataset.targets(testing);

    let model: KNNClassifier = ml.fit_knn(x_train, y_train, 5);
    let predicted = ml.predict(model, x_test);
    print(metrics.accuracy(predicted, y_test));
    let counts = metrics.confusion_matrix(predicted, y_test, dataset.classes(flowers));
    print(tensor.value(counts, 0, 0));
}

Fit preprocessing on training rows and apply those stored statistics to both partitions. The model does not own the standardizer or transform query features automatically. Do not use test results to choose k, scaling, seeds, or other settings without a separate validation design. Distance calculations use the supplied feature representation: changing feature units or scale changes the model.

KNNClassifier and KMeansModel are opaque immutable fitted value types. Import aner.ml for their names. Only fitting functions create them in Aner source; there is no public unfitted constructor. Bindings, function parameters, return types, and local inference work with these values. They cannot be printed or compared directly; use accessors. Copies share immutable fitted state, and fitting retains independent detached copies of the required data. Every tensor returned by this API is detached from automatic differentiation.

ml.predict accepts either fitted model type, selected by static checking. This is a built in overload, not general user defined overloading. Imports precede function definitions. Each namespace requires its own import. aner.ml and aner.metrics expose Tensor for annotations and literals; calling tensor.* still requires import aner.tensor;. Neural calls still require import aner.nn;.

K nearest neighbor classification

Here N is the number of fitted rows, D the feature count, Q query rows, and C distinct fitted class IDs. All dimensions are positive.

FunctionResult and contract
ml.fit_knn(X: Tensor, y: Tensor, k: Int64)KNNClassifier; X has shape N × D, y has shape N × 1, and 1 ≤ k ≤ N
ml.predict(model: KNNClassifier, X: Tensor)Q × 1 predicted class IDs, from Q × D query features
ml.predict_proba(model: KNNClassifier, X: Tensor)Q × C fractions of uniform neighbor votes; each row sums to one up to floating point rounding
ml.classes(model: KNNClassifier)C × 1 fitted class IDs in ascending order; defines the probability column mapping
ml.neighbor_indices(model: KNNClassifier, X: Tensor)Q × k zero based row indices in the matrix supplied to fit_knn, nearest first
ml.neighbor_distances(model: KNNClassifier, X: Tensor)Q × k Euclidean distances in the same order, not squared distances

Class IDs are nonnegative exactly integral Float64 values no larger than 9,007,199,254,740,991. They need not be contiguous. Predicted IDs preserve those labels; probability column index zero is not necessarily class ID zero. Neighbors are found by exhaustive distance comparison. Equal distances are ordered by fitted row index. Each of the k neighbors contributes one vote; tied class votes choose the smallest class ID. There is no distance weighting or approximate neighbor search.

Neighbor row indices refer to the fitted feature matrix, not automatically to original dataset row IDs. For a dataset view, its dataset.row_ids tensor records the corresponding training to source mapping. Querying training rows includes the row itself when its distance makes it a selected neighbor; no leave one out exclusion is implied. Vote fractions are model outputs, not a guarantee of calibrated probabilities.

K means clustering

aner
// Inside main(), with dataset/ml/metrics imports:
let flowers = dataset.iris();
let features = dataset.features(flowers);
let scaler = dataset.fit_standardizer(features);
let x = dataset.transform(scaler, features);
let model: KMeansModel = ml.fit_kmeans(x, 3, 100, 0.0001, 42);
let clusters = ml.labels(model);
print(ml.inertia(model));
print(metrics.adjusted_rand_index(clusters, dataset.targets(flowers)));

This pattern clusters the supplied population and compares its resulting partition with known species afterward. Class labels are not fitting inputs. It is not a held out classification experiment. For prediction on a reserved partition, fit scaling and centers using only the fitting partition, then reuse both unchanged.

FunctionResult and contract
ml.fit_kmeans(X: Tensor, k: Int64, max_iterations: Int64, tolerance: Float64, seed: Int64)KMeansModel; X is N × D, 1 ≤ k ≤ N, 1 ≤ max_iterations ≤ 1000, finite tolerance ≥ 0
ml.predict(model: KMeansModel, X: Tensor)Q × 1 nearest center IDs for Q × D features
ml.centers(model: KMeansModel)k × D returned centers
ml.labels(model: KMeansModel)N × 1 assignments for the fitted rows, recomputed against the returned centers
ml.inertia(model: KMeansModel)Float64 sum of squared distances of fitted rows to their returned assigned centers
ml.iterations(model: KMeansModel)Int64 completed Lloyd updates
ml.converged(model: KMeansModel)Bool indicating that a stopping criterion was met before or on the final allowed update

Initialization chooses distinct fitted row indices with SplitMix64 and an unbiased partial Fisher and Yates shuffle. The seed is mapped modulo 2^64. Distinct row indices may have identical coordinates. This is one seeded run, without K means++ initialization or multiple restarts.

Each Lloyd update assigns rows to their nearest center and replaces occupied centers with assigned row means. Distance ties choose the lowest cluster index; an empty cluster keeps its previous center. The run converges when assignments are unchanged or the largest Euclidean center movement is at most the absolute tolerance, measured in the supplied feature units. Otherwise it stops at the iteration cap with converged == false. Final assignments and inertia are recomputed using the returned centers, including when the cap is reached.

Cluster IDs are arbitrary center indices 0 through k−1, not species names or class predictions. Do not report metrics.accuracy(clusters, class_ids) as K means classification accuracy. Adjusted Rand index compares the two partitions while ignoring label renaming. A requested k does not guarantee k occupied clusters, especially with repeated coordinates. Inertia depends on scale and k and does not by itself select a scientifically suitable model.

Evaluation metrics

All metric inputs must be finite and have positive dimensions. ID metrics require matching N × 1 columns of nonnegative integral Float64 IDs at most 9,007,199,254,740,991.

FunctionReturn value
metrics.accuracy(predicted_ids: Tensor, true_ids: Tensor)Float64 fraction of matching IDs, between 0 and 1
metrics.confusion_matrix(predicted_ids: Tensor, true_ids: Tensor, class_count: Int64)class_count × class_count Tensor of counts, rows = true class, columns = predicted class; class_count is 1 through 1000 and every ID must be below it
metrics.mean_squared_error(prediction: Tensor, truth: Tensor)Float64 mean squared difference over every cell of two matching shapes
metrics.adjusted_rand_index(cluster_ids: Tensor, true_ids: Tensor)Float64 adjusted pair agreement score, independent of numeric ID names; noncontiguous IDs are accepted

Adjusted Rand index is 1 for identical partitions, including a single row and identical trivial partitions; it can be negative. A value near zero corresponds to the adjustment's chance baseline, not a percentage of correct classifications. The implementation uses sparse contingency counts instead of allocating a table indexed by the largest ID. The adjusted Rand reference explains the standard chance adjustment and label permutation invariance; Aner implements its own native calculation.

Mean squared error uses scaled accumulation so a finite mean can be computed even when an individual unnormalized squared difference would overflow; an unrepresentable final mean fails.

metrics.accuracy takes predicted IDs. Existing nn.accuracy takes a score/logit matrix and class IDs, performing argmax internally. For score matrices, explicitly call tensor.argmax(scores) before metrics.accuracy. metrics.mean_squared_error is a detached evaluation scalar; use nn.mse for a differentiable training loss. No metric fits a model or creates a data partition.

Neutral tensor access

FunctionResult
tensor.zeros(rows: Int64, cols: Int64)Detached zero filled Tensor
tensor.random(rows: Int64, cols: Int64, seed: Int64, scale: Float64)Detached seeded Tensor, uniform in [−scale, scale); finite scale ≥ 0
tensor.rows(value: Tensor), tensor.cols(value: Tensor)Int64 dimensions
tensor.value(value: Tensor, row: Int64, col: Int64)Float64 cell at zero based indices
tensor.item(value: Tensor)Float64 value of a 1 × 1 Tensor
tensor.argmax(scores: Tensor)Detached rows × 1 column of the maximum score column index; ties choose the first index

Tensor literals contain rank two rectangular Float64 expressions, such as [[0.0, 1.0], [1.0, 0.0]]. The neutral module does not yet expose a complete linear algebra or transformation API. Existing nn tensor functions remain supported. Shared tensor storage does not make KNN or K means differentiable; fitted state and output tensors do not retain neural computation graphs.

Bounds, errors, and reproducibility

Each input/output Tensor is limited to 1,000,000 cells. KNN query operations allow at most 50,000,000 distance coordinate evaluations, counted as Q × N × D; K means prediction uses Q × k × D. K means fitting checks (max_iterations + 1) × N × k × D against the same 50,000,000 limit before running, even if it might converge early. Probability matrices, neighbor outputs, centers, and confusion matrices must separately fit the tensor cell limit. These formulas bound distance coordinate work; neighbor selection, centroid updates, and movement checks add bounded overhead. K means allows at most 1000 updates. There is no global process memory quota or streaming/incremental fitting yet.

Bad shapes, invalid labels or k, nonfinite inputs, unrepresentable numerical intermediates, and exceeded work limits produce diagnostics. ML, metric, and neutral tensor diagnostics use R2301, R2401, and R2501; Tensor literal construction and other existing neural/storage paths may retain R2001. Imports, fitted model types, and function signatures are checked before execution. Native algorithms also enforce their own work bounds; interpreter step counts are not an elementwise work meter.

Within a pinned software/hardware environment, fixed data order, preprocessing, seed, and settings define a repeatable computation. Seeds alone do not promise identical floating point results across compilers, machines, or future algorithm versions. KNN row order ties and K means initialization depend on the actual fitted row order; keep source version, row membership, scaling, and settings with results. Fitted models do not yet have serialization or resume support.

Current scope

For neural training, aner.nn provides a concise classifier API and explicit forward/backward/update loops. Classical ML uses fitted model values because fitting and later prediction have different inputs and state. aner.viz currently records tensor/NN computation graphs; it does not visualize KNN search or K means updates. No classical ML visualization is inferred from a shared Tensor type.

KNN and K means do not yet provide model serialization, automatic preprocessing pipelines, or third party algorithm registration. Keep fitted preprocessing alongside the model in your program and retain the source, data identity, and settings needed to repeat the fit.

Aner handbook · Guides and API reference