Browse the handbook
Design

Machine learning architecture

A research plan for numerical types, device backends, reproducibility, model families, and packages.

Architecture proposal with current implementation notes, October 10, 2026.

Aner should provide classical and neural machine learning libraries over a native tensor runtime with shared numerical rules, and differentiation where an algorithm needs it. Start with CPU execution, then add Apple GPU and NVIDIA GPU backends behind the same tensor interface. Keep experiment provenance, controlled randomness, data validation, and third party extensibility part of the design from the beginning.

This is the broader architecture proposal. A first subset is now implemented: the scalar interpreter, editor support, built in aner.nn import, rank two Float64 CPU tensors, reverse mode differentiation, SGD for small dense networks, opt in aner.viz training reports, and aner.dataset with offline Iris/Wine data, splits, and preprocessing. The first classical ML subset adds fitted KNN/K means models in aner.ml, evaluation in aner.metrics, and neutral construction/inspection in aner.tensor. The separate aner.data module now adds explicit CSV schemas, lazy scans, bounded batch summaries, and limited table collection. See Aner Neural, classical ML, and external CSV data for their exact current APIs. Float32/Int32, type aliases, function type inference, general modules/packages, checkpointing, and accelerators remain unimplemented. Language behavior is defined by LANGUAGE.md.

Language and library responsibilities

Keep fn and brace delimited blocks. Neural network layers should be library types and functions, not new keywords. Aner owns source semantics, type checking, runtime errors, tensor operation contracts, and compilation. Native numerical libraries supply implementations of operations under those contracts, while Aner’s native runtime manages their execution.

Proposed componentResponsibility
Language and module systemNamespaces, public exports, records/parameter collections, type inference, generic functions, diagnostics, and callable functions where APIs need them
Tensor runtimeTyped buffers, dimensions, checked allocation, ownership, device placement, operation dispatch, differentiation records, and synchronization
aner.tensor and aner.linalgTensor construction, explicit conversions, matrix multiplication, shape operations, reductions, and linear algebra
aner.mlTyped fitted classifier/clusterer values now; future regressors, other model families, public model interfaces and pipelines
aner.nn and aner.optimLayers, parameters, losses, optimizers, and training/evaluation state
aner.dataset and aner.metricsTeaching dataset catalogue, train fitted preprocessing and evaluation now; future group/time partitions
aner.dataExternal CSV schemas, lazy scans, bounded batches, summaries and explicit collection now; future joins, richer types, formats and streaming model input
aner.experimentRun configuration, manifests, random streams, checkpoints, resume, and comparison policies

aner.nn, aner.dataset, aner.data, aner.ml, aner.metrics, aner.tensor, and the optional observation module aner.viz are adopted included modules. Other package names and broader responsibilities remain proposals. Differentiation and SGD remain under nn; the new neutral tensor facade covers only its documented basic operations. First party and third party Aner libraries should use the same public interfaces. A first small network can be composed from ordinary functions and tensors; a large object/class system is not required before that works.

text
Aner source -> shared frontend and type checker
    |-> scalar interpreter --+
    |-> future native code --+-> tensor operations and differentiation
                                  |-> reference CPU kernels
                                  |-> optimized CPU kernels
                                  |-> Metal backend
                                  |-> CUDA backend

Distributed execution coordinates multiple runtime processes:
    dataset shards + gradient communication + checkpoint coordination

Once tensor dispatch exists, the interpreter can call native matrix operations. Compiling all scalar Aner code is therefore not a prerequisite for GPU execution. Later typed intermediate representations can retain tensor operations, enabling fusion and device compilation without interpreting each element. Preserve source locations through both paths.

Current model contracts and the next abstraction

External file ingestion uses DataSchema, DataScan, DataReport, and DataTable from aner.data; it does not extend the catalogue’s Dataset type. DataScan holds a plan, and DataTable holds an explicit bounded collection. The explicit table to Tensor operation is data.to_tensor, with a separate output payload limit and checks for missing values, column types, and integer conversion. Direct model fitting from CSV batches remains future work.

The implemented classical subset uses typed fitted values rather than new syntax for each algorithm. ml.fit_knn(X, y, k) returns KNNClassifier; ml.fit_kmeans(X, k, max_iterations, tolerance, seed) returns KMeansModel. Both accept the same ordinary Tensor features that dataset extraction provides. Their state is immutable and detached from automatic differentiation. Built in ml.predict statically selects the supported fitted type. This does not yet provide general overload resolution, public model traits, records, constructors, or user defined model registration.

Current operation familyContract
KNN classificationFit features and categorical ID targets; predict IDs or uniform neighbor vote fractions with an explicit fitted class mapping
K means clusteringFit features without reference labels; retain centers and assignments; predict arbitrary cluster IDs; compare partitions with adjusted Rand index
Neural learningExplicit Aner functions and forward/backward/update loops over parameters and differentiable Tensor operations
MetricsEvaluate supplied IDs, partitions, or numeric predictions; no hidden fitting or data partitioning
PreprocessingCaller fits a Standardizer on the fitting partition and applies it consistently; fitted classical models do not own or apply this transform

aner.tensor is now the neutral public namespace for basic Tensor construction and inspection. It currently delegates to the existing native Tensor implementation under the neural code; this is an API boundary, not a second numerical backend or a completed storage refactor. New model code should not require neural imports merely to count rows or inspect ordinary tensor cells. Preserve the same finite value, shape, ownership, and dtype contracts when the internal storage boundary is extracted later.

The next model abstraction should separate classifier, regressor, clusterer, and transformer capabilities instead of assuming every model has a neural training loop. A future fitted pipeline should retain feature/schema expectations, train fitted transformations, model settings, class vocabulary where relevant, and source/split provenance. Its prediction path must apply the saved transforms without refitting. That wrapper and its serialization are not implemented by the first two built in fitted types.

Before adding public model interfaces, test the design with one more independently implemented algorithm, explicit failure contracts, repeatability checks, and data only save/load. A public interface should let a third party Aner package add a model without editing the interpreter. Current aner.ml functions are registered built ins and do not meet that acceptance gate. Three part imports such as aner.ml.neighbors, local package manifests, user defined records, and traits remain syntax/design proposals; supported imports are the named two part built in module paths.

The visual debugger still observes tensor computation graphs. Shared Tensor inputs do not make KNN neighbor searches or K means iterations available as neural graph captures. A future general model report needs algorithm specific observations with explicit meanings, such as neighbor membership or center movement, rather than a fabricated neural topology.

Tensor and numeric contract

Begin with contiguous dense CPU tensors with explicit element type and runtime dimensions. Design the descriptor to carry layout and strides, while rejecting unsupported layouts in the first kernels. Keep tensor data in packed native buffers rather than one generic interpreter Value per element.

Prioritize Float32 and Float64 tensor elements; add Int32/Int64 for labels, counts, and indexing use cases. Do not derive allocation or address arithmetic width from the element type: validate shape products, byte counts, strides, and backend dimension limits before allocating or launching a kernel.

Document each operation's accepted shapes/types, output shape/type, errors, empty input behavior, aliasing, and numerical policy. Start with equal shape elementwise operations, matrix multiplication, transpose, explicit reshape, reductions, and explicit bias addition/broadcasting. Reject incompatible shapes instead of relying on accidental broadcasting.

Optional annotations should preserve precise types. Contextual literals such as 1 acquire a type from surrounding constraints before defaulting. Aliases can centralize an intended dtype; they affect code using that alias. Generics can later specialize a reusable routine for different dtypes. These changes require a real type inference design, including recursive functions; they are not syntax only edits.

Separate storage, computation, and accumulation precision. A future low precision training configuration might store some tensors in a smaller format while accumulating in Float32. Specify rounding, conversions, reduction order, fused operations, and loss scaling behavior. Never silently downcast Float64 to satisfy a GPU backend.

Quantization is a distinct, explicit transformation. Initial work should target inference, with calibration data, scale/zero point information where applicable, rounding/clipping policy, and measured accuracy effects recorded. Quantization aware training and custom surrogate gradients are later features. See the ONNX quantization operator contract.

Differentiation and the first network

Implement first order reverse mode automatic differentiation over supported floating point tensor operations. During forward execution, record operation identifiers, dependencies, and the values needed by backward rules. Backward execution applies those rules and accumulates parameter gradients. This differentiates the executed tensor computation; it does not imply that arbitrary IO, integer operations, or branch decisions are differentiable.

Keep gradient rules in the shared operation layer so CPU and GPU implementations follow the same mathematical contract. Initially avoid user visible in place mutation of tensors retained for backward computation; optimizer updates occur after backward and create new parameter values or use controlled runtime updates. Specify graph lifetime explicitly. The first implementation retains graph nodes while tensors reference them so repeated backward is supported; dropping the last handle releases the graph. A later API may offer explicit retention/release choices.

Start with matrix multiplication, addition, multiplication, reductions, and a small activation/loss set. Implement a dense layer and SGD first; add Adam after optimizer state and checkpoint contracts exist. Higher order gradients, convolution, attention, sparse tensors, and giant models can follow demonstrated needs.

Use Float64 CPU finite differences to check gradients of smooth operations on small inputs. Test nonsmooth activation conventions separately. Cover shapes, axes, zero sized inputs where supported, shared parameters, repeated backward policy, and incorrect dtype/device use. Gradient checks are development validation, not the training algorithm.

The first end to end example should train a small network on synthetic data, with fixed initialization, a declared loss target, bounded steps, and a recorded run. A tiny XOR classifier is a useful machinery check, not evidence of real world generalization. Follow it with a synthetic tabular experiment that exercises batching and group aware splits.

Hardware backends

TargetInitial implementation routeConditions
CPU on macOS, Windows, LinuxPortable C++ reference kernels, then optimized BLASKeep the reference implementation and control thread count; avoid nested thread pool oversubscription
Apple Silicon GPUObjective-C++ bridge to Metal and MPS/MPSGraph, with custom Metal kernels as neededBegin with Float32 and an explicit operation capability list
NVIDIA GPUCUDA with cuBLAS/cuDNN and custom kernels as neededDevice specific build/runtime packages; no CUDA assumption in core language semantics
Other GPUsA later backend chosen from real hardware requirementsCPU availability does not establish GPU coverage on an OS
Cluster or supercomputerMultiple Aner processes, data partitioning, collective communication, and a launcherSeparate training strategy and recovery semantics from local device selection

For optimized CPU kernels, evaluate OpenBLAS across platforms and Apple Accelerate on macOS. Select providers through the backend boundary, and record their versions and threading settings. “All cores” should be a configurable execution policy, not an unconditional performance claim.

The proposed Aner metal backend would call Apple's native Metal, MPS, and MPSGraph APIs for GPU execution. Apple demonstrates both training and inference with MPSGraph. Keep Apple graph objects behind the backend interface. Metal has no native double precision type, which makes Float32 support a prerequisite for this target; see the documented MPS precision limitation. The Apple Neural Engine is a separate device and is not part of the initial GPU training promise.

NVIDIA's cuDNN native APIs provide optimized neural network operations. Their supported dtype/layout/algorithm combinations need validation just as CPU and Metal kernels do.

The backend interface needs allocation/release, transfers, operation dispatch, capability queries, and completion/error reporting. CPU can complete synchronously; GPU storage must remain alive until work completes. Even unified memory needs synchronization and ownership rules. Query support for the forward and backward operations before training. An unsupported operation should fail clearly, or use an explicitly enabled fallback that is recorded with its transfers and performance implications.

After the operation set stabilizes, evaluate IREE as an optional compiler/runtime integration for tensor graphs. It exposes native APIs and multiple targets, but adopting it requires lowering and conformance work; it does not automatically implement Aner's semantics. Likewise, ONNX Runtime's C API offers an optional route for imported model inference. Model interchange and external inference are separate from owning Aner's training/autodiff behavior.

Distributed training

A cluster is coordinated processes, not just another value of a GPU device setting. Start with synchronous data parallelism: each worker receives a disjoint batch shard, computes gradients, and combines them before updating identical model replicas. Define whether gradient values are sums or means and account for unequal shard sizes so the effective global batch has the intended weighting.

Keep collective communication behind an interface. CPU/MPI and NVIDIA NCCL are candidate transports. NCCL supplies multi GPU/multi node communication, not Aner's data loader, scheduler, or training policy. A future launcher can integrate with a site's scheduler instead of making Aner implement a job scheduler.

Record world size, rank mapping, sharding, global batch size, reduction policy, and communication library versions. Define worker failure, timeout, cancellation, and restart behavior. Initially require the same distributed configuration for exact checkpoint continuation; changing worker count is a separate numerical/reproducibility mode. Multi node Apple GPU training is not an early support promise.

Reproducibility and scientific safeguards

Use three explicit contracts rather than claiming universal identical results:

ContractRequired evidence
Recorded experimentSource/build/package hashes; runtime, backend and driver versions; hardware; dataset snapshot references; preprocessing and split configuration; RNG algorithm/state; dtype and execution settings
Repeatable executionA tested pinned software/hardware configuration, stable input order and random streams, deterministic operations and reductions; strict execution rejects unsupported operations
Agreement across hardwareDeclared numerical tolerances, gradient/output comparisons, and experiment appropriate statistical acceptance criteria

A seed alone does not establish reproducibility across hardware or releases; execution also depends on operation choices, numerical behavior, and the software and hardware configuration. See the reproducibility constraints. Aner’s first strict reference mode should use a controlled single thread CPU implementation; optimized modes need their own verified contracts. Record relaxed precision and nondeterministic algorithm choices when enabled.

Require explicit RNG streams for initialization, shuffling, dropout, and sampling. Version the RNG algorithm and derive substreams from stable operation/data identities rather than thread scheduling. Pin partitioning and input order for strict replay. A checkpoint must include weights, optimizer/scheduler state, step/epoch, RNG state, data order/cursor, precision state, and distributed configuration. Write checkpoints atomically and validate metadata, tensor lengths, and hashes on load. Use data only serialization; model loading must not implicitly execute serialized code.

Data handling belongs in this design alongside model training. Preserve text identifiers and missing value meaning; validate schemas, shapes, units where declared, and memory budgets. Fit normalizers, imputers, and feature selectors on training partitions, then apply the fitted transformation to validation/test data. Store partition provenance and detect overlap for supported split operations. The importance of separating fitting from test data is documented in preprocessing leakage guidance.

Provide group aware and temporal split operations. When evaluating performance on unseen patients, repeated observations from one patient must stay in the same partition; different study questions can require different split policies. The implemented external CSV reader processes bounded native batches and requires explicit limits for collection. General readers and direct model consumption of batches remain future work. Loss/gradient finite value checks should identify the operation and source location, with an explicit policy for stopping or reporting.

These checks apply to tracked Aner operations and declared data relationships. They cannot infer study design, discover all external leakage, or guarantee independent replication on another population. Provenance records should use controlled data references and avoid copying patient rows or identifiers into general logs. Dataset hashes identify snapshots; they are not anonymization.

Third party packages and native extensions

Future package work should start with local packages containing .aner modules, an Aner.toml manifest, exports, tests, and documentation. Resolve dependencies before execution and produce Aner.lock with exact versions/content hashes. Local directory dependencies allow useful library development before a public registry exists. Interpreter and compiler should consume the same resolved package graph.

Pure Aner packages can define models, losses, optimizers, metrics, and transforms using public tensor operations. Native extensions are a separate path for new kernels, hardware integrations, and existing C/C++ libraries. Design a versioned C interface with opaque handles, fixed width fields, explicit ownership, and status/error results; do not expose STL containers or C++ exceptions across the public boundary. Keep the interface experimental until tensor and error semantics settle.

A custom operation must declare its name/version, shape and dtype rules, device support, memory/aliasing behavior, completion semantics, and gradient rule or explicit nondifferentiable status. Include a reference implementation or reference vectors, numerical tolerances, gradient tests, and determinism tests. First party native backends should meet the same conformance requirements.

Use DLPack for future compatible tensor exchange and Arrow's C Data Interface for columnar exchange. These are interchange contracts, not neural network engines; zero copy is conditional on compatible layout, device, ownership, and synchronization. Arrow's C interface is in process sharing, not a cluster transport.

Package manifests and checksums do not sandbox native code or prove that it is trustworthy. Native extensions run with host process privileges. Record native artifacts and platform requirements, avoid hidden installation time execution, and treat process isolation/OS sandboxing as separate future security work. A public registry, signing/trust policy, and project/package licenses need decisions before distribution.

Implementation order and acceptance gates

MilestoneDeliverableAcceptance gate
1 Numeric and library foundationFloat32/Int32 semantics, contextual literals, scoped inference, local modules, dense typed CPU tensorsDefined conversion/overflow/shape rules; reference matmul and reductions; existing scalar behavior preserved or explicitly revised
2 CPU neural networkReverse mode differentiation, dense layer, loss, SGD, explicit RNG, experiment manifest and checkpointFinite difference gradient checks; deterministic synthetic training; interrupted/resumed run matches uninterrupted run within the declared strict environment
3 Optimized CPU and Apple GPUBLAS provider and a small Metal backend implementing the same operation setForward/backward agreement with CPU references within dtype tolerances; unsupported precision fails; synchronized performance measurements include transfers
4 NVIDIA GPU and native SDKCUDA backend, experimental C operation registration, third party package exampleIndependent package adds an operation without compiler edits; documented gradient/device contracts and backend comparison tests
5 Distributed executionTwo worker data parallelism, collectives, launcher and coordinated checkpointCorrect global gradient weighting; declared RNG/reduction behavior; checkpoint/restart test; same supported model API

Milestones 1 and 2 should also exercise bounded synthetic tabular batches, declared schemas, and train/test transformation boundaries so that dataset usability progresses with neural network execution. Benchmark representative training steps and data pipelines, including peak memory and synchronization, rather than treating factorial speed as a proxy for ML throughput. Run supported CPU tests on Windows, macOS, and Ubuntu; accelerator support is validated separately on actual hardware.

Implemented subsets: rank two Float64 CPU tensors; built in neural, visual, dataset catalogue, external CSV, classical ML, metric, and neutral tensor modules; reverse mode differentiation and SGD; small dense network examples; fitted KNN classification and seeded Lloyd K means; explicit CSV schemas, lazy scans, bounded summaries, and limited table collection. Classical algorithms have independent bounded work and validation contracts, detached outputs, and type directed prediction. No fitted pipeline, model serialization, or third party registration is implemented. Next prioritize data only checkpoints with repeatable resume, general local modules and parameter collections, and Float32 before accelerator work. CNN/RNN layers follow the shared tensor and gradient contracts. The milestone table remains a broader plan, not a claim that every item in stages 1 to 2 is complete.

Aner research & design · Proposal