Browse the handbook
Language

Language reference

Variables, functions, static types, control flow, classes, and the current language contracts.

This reference covers the scalar core and built in aner.nn, aner.viz, aner.dataset, aner.tensor, aner.metrics, aner.ml, and aner.data modules supported by aner check and aner run. Sections marked as proposals describe additional numerical syntax and large data features that are not implemented.

The examples use an installed aner command. See installation and release availability to choose a package for your system, and VS Code setup for editing and notebooks. Match the documentation and editor extension to the release you install.

Source and execution

Programs are UTF-8 text in .aner files. This is the canonical source extension for interpreted execution and future native compilation. A UTF-8 byte order mark is allowed at the beginning. Any editor can create them; the Aner executable checks and executes them.

The notebook document extension is .anernb (Interactive Aner Notebook), distinct from plain .aner source. VS Code extension 0.1.16 provides Aner — notebook, which follows the document’s execution mode: shared native values, imports, and functions by default, or independent executions when explicitly configured. A remembered Aner kernel selection does not override the document’s mode. No source wrappers or previous cell replay are used. Markdown and images belong to the document, not Aner syntax. A Jupyter kernel remains future work. See the notebook guide for setup, session lifecycle, and limits.

text
aner check program.aner
aner run program.aner
aner run program.aner --max-steps 2000000
aner --help
aner --version

check parses and statically checks the entire file without executing it, then prints OK. run performs the same checks before interpreting the program. It prints only the program's output on success. Diagnostic failures return a nonzero process status and identify a source location when applicable. A file may contain direct top level statements, or define fn main() -> Unit. Mixing executable top level statements with main is rejected. Helper functions remain file level declarations in either form. In a direct script, statements execute in source order; a function can access typed global bindings, but reading one before its initializer executes produces R1001. A file containing only helper functions and no statements or main has no entry point and is rejected by check and run.

aner build reports that native compilation is unavailable. Ahead of time native compilation is a later milestone. Native persistent notebook execution is available through aner session --stdio, a framed editor transport; an interactive terminal REPL and JIT remain separate possibilities.

Data types

TypeMeaningExample
Int64Signed 64 bit integer with checked arithmeticlet count: Int64 = 42;
Float64IEEE 754 binary64 numberlet weight: Float64 = 72.5;
BoolBoolean, without numeric truthinesslet included: Bool = true;
StringImmutable UTF-8 textlet label: String = "sample";
UnitNo meaningful returned value; written ()let nothing: Unit = ();
TensorDense rank two CPU Float64 values; requires one of the numerical/data imports listed belowlet x: Tensor = [[1.0, 2.0]];
DatasetImmutable source metadata and selected rowslet flowers = dataset.iris();
DatasetSplitImmutable training/test viewslet parts = dataset.split(flowers, 0.2, 2026);
StandardizerFixed training derived column statisticslet scaler = dataset.fit_standardizer(x_train);
KNNClassifierImmutable fitted nearest neighbor classifier; requires aner.mllet model = ml.fit_knn(x, labels, 5);
KMeansModelImmutable fitted clustering model; requires aner.mllet model = ml.fit_kmeans(x, 3, 100, 0.000001, 42);
DataSchemaImmutable external table schema; requires aner.datalet s = data.schema();
DataScanDeferred external CSV planlet q = data.scan_csv("data.csv", s, "NA");
DataReportCompleted selected row summarylet r = data.summarize(q, 1048576, 256);
DataTableCollected typed columns with validity informationlet t = data.collect(q, 1000, 1048576, 256);
NeuralNetworkImmutable sequential neural configurationlet net = nn.sequential(seed: 42);
ClassifierPipelineNeural classifier and preprocessing configurationlet plan = ml.classifier(net, standardize: true);
FittedClassifierTrained immutable classifier snapshotlet fitted = plan.fit(training, epochs: 2000, rate: 0.1);
EvaluationReportDetached classifier evaluation and provenancelet report = fitted.evaluate(testing);

Declarations may omit their type annotation: let weight = 72.5;. Variables keep their declared or inferred type. Function parameters require explicit types. A return annotation may be omitted when the body determines a compatible result type; explicit annotations remain supported and are required for value return inference cycles. There is no implicit conversion between integers and floating point numbers: write Float64(count). This built in conversion accepts one Int64; large values can round to binary64 precision.

Decimal integer literals have type Int64. The largest positive literal is 9223372036854775807. The minimum signed value is written -9223372036854775808: the integer must be the next token after the minus (whitespace and comments may separate them). An out of range positive magnitude is not made valid by placing it inside parentheses before negation.

Floating point literals require digits before a decimal point, digits after that point if present, and digits after an exponent if present. Examples: 15.0, 1e3, 2.5e-4. A decimal point or exponent is required to distinguish them from integer literals. .5 and 1. are invalid. Out of range literals, including overflow and underflow, are rejected; nonfinite values may still result from arithmetic. Unary plus is not supported.

Keywords, names, and strings

The reserved words are:

text
fn let var if else while return true false import class null self public private

Names are case sensitive. Identifiers start with an ASCII letter or underscore and continue with ASCII letters, digits, or underscores. Unicode text is allowed in strings. // begins a comment that ends at the next line break. Block comments are not supported.

Strings use double quotes and support \", \\, \n, \r, and \t escapes. Unescaped line breaks, raw ASCII control characters below U+0020 (including tabs), and unknown escapes are errors. Interpolation, raw strings, concatenation, and character indexing are deferred. String currently supports equality and printing.

Int64, Float64, Bool, String, Unit, Tensor, Dataset, DatasetSplit, Standardizer, KNNClassifier, KMeansModel, DataSchema, DataScan, DataReport, DataTable, NeuralNetwork, ClassifierPipeline, FittedClassifier, and EvaluationReport are built in type names; print and the Float64 conversion are built ins. These names and the nn, viz, dataset, tensor, metrics, ml, and data namespaces cannot be redeclared. Vector, Matrix, dot, transpose, sum, mean, len, rows, and cols are reserved as unqualified names for later numerical work. Some operations are available now as qualified functions such as nn.matmul, nn.transpose, and nn.mean. Future data vocabulary such as table, join, and gpu is not reserved now. Domain operations should become ordinary typed APIs unless their semantics require a language construct.

Bindings, scopes, and functions

Braces delimit blocks; indentation does not define scope. Imports, declarations, assignments, returns, and expression statements can end with a semicolon, a line break after a completed expression, the end of the source, or a closing brace. Multiple simple statements on the same line require semicolons.

Expressions continue across line breaks when an operator or call delimiter continues the expression. For example, 1 + followed by 2 is one expression; a leading + 2 on the following line can also continue 1. Use a semicolon to make that boundary explicit. Calls and tensor literals may span lines. A return immediately followed by a line break is a bare Unit return; put a returned expression on the same line as return, or begin its parentheses there.

A direct script needs no function wrapper:

aner
var a = 10
let b = 20
c = 30
a = a + 5
print(a + b + c)

The following complete program form remains valid:

text
fn square(x: Float64) -> Float64 {
    return x * x;
}

fn main() -> Unit {
    let threshold = 18.5;
    var total: Int64 = 0;
    var i: Int64 = 0;
    while i < 3 {
        total = total + i;
        i = i + 1;
    }
    if total == 3 {
        print(square(threshold));
    } else {
        print("Unexpected total");
    }
}

let creates an immutable binding; var permits reassignment of the same type. Both require an initializer. A bare assignment such as a = 10 creates a mutable binding in the current scope if no binding of that name is visible. Otherwise it assigns to the visible binding, requiring matching type and mutability. This rule applies both to direct statements and function bodies. Parameters are immutable. Blocks have lexical scope: duplicate declarations in one scope are rejected, and an explicit declaration in an inner block can shadow an outer binding. Compound assignments such as += are not supported.

Expression statements must have type Unit in .aner scripts and ordinary function/block bodies. At the outermost level of a persistent session cell, scalar expressions such as a or 1 + 2 are also allowed and display their result. Composite values still require explicit accessors or summary functions; no automatic tensor/table expansion is performed.

Functions are resolved before their bodies, so forward calls and recursion work. Calls accept positional and named arguments with exactly matching types; positional arguments must precede named arguments. Each argument expression evaluates once in source order, even when named arguments reorder parameters. Unknown, duplicate, or missing arguments are errors. User function labels come from their declared parameter names; built in labels are part of the API contract. Calls may have a trailing comma; function parameter declarations do not. Functions are not values and cannot be nested inside other functions. A local variable may shadow a user function name, but calling that variable is an error. A non Unit function must return a matching value on every path under a conservative check: even an apparently infinite loop does not establish a guaranteed return. A Unit function may fall through, use return;, or return ().

A result annotation is optional: fn square(x: Float64) { return x * x } infers Float64, while fn announce() { print("Aner") } infers Unit. Returned values must determine a compatible static type, and non Unit results still require a return on every path. Inference does not convert numbers or make parameter types dynamic. Value return inference cycles need an explicit annotation such as -> Int64; retaining annotations is also useful for public API contracts. The expression return at a line break is still a Unit return, not a value from the next line.

Return inference uses exact types except that references to the same class, that class's nullable type, and null can merge to the nullable class type. Mixed incompatible results, including Unit mixed with a value, report E1002; missing non Unit return paths report E1004. Bodies with no expression return, or only bare return/literal return (), establish Unit before checking their calls, so recursive Unit methods can omit annotations. A cycle of value return dependencies needs an explicit result annotation to break it (E1009). The inference dependency chain is bounded to 128 calls (E1007).

An inferred body is checked when its result is first needed. Any globals used by that body must already be known at that point. An explicit result annotation allows body checking to be deferred, but does not make a runtime read before a global's initialization valid. Generic methods are inferred for each concrete specialization.

if and while conditions require Bool. else is optional; else if is supported. break, continue, for, general user defined modules, value structs, and inheritance are outside the current core. User defined reference classes and their methods are supported. The seven included modules listed at the top can currently be imported. print is permitted in ordinary functions; no purity or checked effect guarantee is claimed.

Instance methods and concise model APIs

Built in values expose typed instance calls and chains. For example:

aner
import aner.dataset
import aner.nn
import aner.ml

let split = dataset.iris().split(test: 0.2, seed: 2026)
let network = nn.sequential(seed: 42)
    .dense(inputs: 4, outputs: 8)
    .tanh()
    .dense(inputs: 8, outputs: 3)
let pipeline = ml.classifier(network, standardize: true)
let fitted = pipeline.fit(split.train(), epochs: 2000, rate: 0.1)
print(fitted.evaluate(split.test()).summary())

A receiver is evaluated once, before that call's explicit arguments. Method names are resolved using its static type. A leading . may continue an expression on the following line; use a semicolon when a statement must end. These methods call native Aner implementations and preserve ordinary checker diagnostics and argument evaluation order. They are not dynamic attribute lookup or source translation.

NeuralNetwork holds an immutable configuration; layer methods return another configuration. ClassifierPipeline.fit creates a new fitted value and learns preprocessing only from its training input. FittedClassifier.evaluate reuses those fixed statistics and returns an EvaluationReport. The same fitted value can evaluate multiple compatible dataset views. See Aner Neural for supported layers, numerical semantics, provenance, and resource limits.

The built in object style interface is separate from the user defined reference classes below. Inheritance, value structs, traits/interfaces, tuple destructuring, and general list literals remain unimplemented. Tensor aliases retain their documented shared gradient behavior; a let binding does not promise deep immutability of every native value.

User defined reference classes

aner
class Node {
    let value: Int64
    var next: Node?

    fn get() {
        return value
    }
}

let node = Node(value: 10, next: null)
print(node.get())

Classes are nominal reference types with explicitly typed members. Fields are private by default and methods public by default; private and public explicitly override either default. Only a method of the owning concrete class may read/write its private fields or call its private methods. Synthesized constructors are public and require all declared fields, including private ones; no field defaults or custom constructors are provided. let fields cannot be reassigned; var fields can. An immutable binding to an object does not freeze that object's mutable fields. Copies alias the same object, and compatible object references compare by identity. Class methods receive an implicit typed receiver; writing self in the parameter list is optional. Their input parameters stay explicitly typed, while the return type can be inferred. Inside methods, unqualified fields and method calls resolve against the current object; locals and parameters may shadow fields, with self.field available to disambiguate. Names are scoped to the class; class declarations may refer to one another.

Node? permits Node or null. A nullable reference must use checked .unwrap() before field or method access; a preceding null check does not narrow its static type. Nullable syntax here applies to user defined class references, not general scalar missingness. Unwrapping null or exceeding the object heap's limits produces R2801.

Inside a method, bare value lookup prefers locals/parameters, then fields of the current object, then global bindings. For ordinary unqualified calls, a local binding or field of that name blocks the call; otherwise the current object's methods precede global functions. print and Float64 retain their built in unqualified call meaning; a same named method must be selected explicitly, such as self.print(). Explicit member calls retain the usual static type and visibility checks.

The managed heap allows cycles and performs iterative tracing from live globals after each successful session cell. It does not collect during a cell. Programs/sessions permit at most 256 class declarations and concrete specializations combined, with 64 fields per class and 1,024 functions/method templates/specialized methods combined. Outstanding allocations are limited to 16,384 objects and 262,144 fields, with the existing 128 call runtime depth limit. These are not process memory limits. Scripts release their heap at termination. Classes persist in sessions and cannot be redefined without a restart; static errors roll back new definitions, while runtime errors clear all session state. See the complete OOP tutorial for examples and ownership details.

Generic classes

aner
class Node<T> {
    let value: T
    var next: Node<T>?
    fn get() { return value }
}
let integer = Node<Int64>(value: 10, next: null)
let text = Node<String>(value: "sample", next: null)
print(integer.get())
print(text.get())

This example defines a separate generic Node; it cannot coexist with a nongeneric class of the same name. Constructor type arguments are explicit, with no inference from values. Type arguments may use supported scalar/native types, user defined classes, nullable class references, and nested generic classes. Unit is not an accepted type argument; null is a value, not a type. Each concrete specialization has its own nominal type. Fields are private unless declared public; the constructor still accepts all field values.

Aner substitutes class type parameters through fields, method parameters/results, and constructor references. It checks every method body when a class is specialized, including methods the program never calls. Generic functions, independently generic methods, dynamic Any, traits, and type constraints are not implemented. Int32 remains unavailable; the supported integer type is Int64.

Generic expansion is bounded by 262,144 cloned class/function/field/parameter/type/expression/statement nodes cumulatively, a type nesting depth of 128, at most 32 parameters/arguments in a type list, and 16 KiB per canonical concrete type name. Classes and methods also count toward the shared declaration limits above. Failed session checks roll back new specializations and their budget use. These safeguards are not a process memory quota. See generic classes and examples.

Built in argument labels

These are the accepted labels, in parameter order. Every parameter is required. Positional calls remain valid; names can be reordered after any leading positional arguments. A typed instance call supplies the first parameter through its receiver, so flowers.split(test: 0.2, seed: 2026) corresponds to dataset.split(data: flowers, test: 0.2, seed: 2026). A function's presence in this table does not add instance methods to unsupported receiver types; the checker still selects an exact signature.

Built in functionParameter labels
Float64value
data.booleantable, row, column
data.collectscan, max_rows, max_bytes, batch_rows
data.colstable
data.columnschema, name, type, nullable
data.column_nametable, index
data.describereport
data.explainscan
data.includescan, column
data.integertable, row, column
data.is_missingtable, row, column
data.meanreport, column
data.missingreport, column
data.numbertable, row, column
data.peak_bytesreport
data.rowsvalue
data.scan_csvpath, schema, missing
data.scanned_rowsreport
data.schemaNone
data.selectscan, column
data.summarizescan, max_bytes, batch_rows
data.texttable, row, column
data.to_tensortable, max_bytes
data.validreport, column
data.where_eqscan, column, value
data.where_gescan, column, value
data.where_presentscan, column
dataset.centerscaler
dataset.citationdata
dataset.class_namedata, index
dataset.classesdata
dataset.colsdata
dataset.countNone
dataset.descriptiondata
dataset.feature_namedata, index
dataset.feature_unitdata, index
dataset.featuresdata
dataset.fit_standardizerfeatures
dataset.idindex
dataset.irisNone
dataset.licensedata
dataset.loadname
dataset.namedata
dataset.one_hotdata
dataset.row_idsdata
dataset.rowsdata
dataset.scalescaler
dataset.sourcedata
dataset.splitdata, test, seed
dataset.targetsdata
dataset.testsplit
dataset.trainsplit
dataset.transformscaler, features
dataset.versiondata
dataset.wineNone
metrics.accuracypredicted, actual
metrics.adjusted_rand_indexpredicted, actual
metrics.confusion_matrixpredicted, actual, classes
metrics.mean_squared_errorpredicted, actual
ml.accuracyreport
ml.centermodel
ml.centersmodel
ml.classesmodel
ml.classifiernetwork, standardize
ml.convergedmodel
ml.evaluatemodel, testing
ml.fitpipeline, training, epochs, rate
ml.fit_kmeansfeatures, clusters, max_iterations, tolerance, seed
ml.fit_knnfeatures, labels, k
ml.inertiamodel
ml.initial_lossmodel
ml.iterationsmodel
ml.labelsmodel
ml.lossreport
ml.neighbor_distancesmodel, input
ml.neighbor_indicesmodel, input
ml.predictmodel, input
ml.predict_probamodel, input
ml.probabilitiesmodel, input
ml.row_idsreport
ml.rowsreport
ml.scalemodel
ml.summaryvalue
ml.training_accuracymodel
ml.training_lossmodel
ml.training_row_idsmodel
nn.accuracyscores, labels
nn.adda, b
nn.add_biasvalue, bias
nn.argmaxvalue
nn.backwardloss
nn.colsvalue
nn.cross_entropylogits, targets
nn.densenetwork, inputs, outputs
nn.detachvalue
nn.gradvalue
nn.itemvalue
nn.matmula, b
nn.meanvalue
nn.mseprediction, target
nn.multiplya, b
nn.parametervalue
nn.randomrows, cols, seed, scale
nn.reluvalue
nn.rowsvalue
nn.scalevalue, factor
nn.sequentialseed
nn.sgdparameter, rate
nn.sigmoidvalue
nn.softmaxvalue
nn.subtracta, b
nn.tanhvalue
nn.transposevalue
nn.valuetensor, row, column
nn.zerosrows, cols
printvalue
tensor.argmaxvalue
tensor.colsvalue
tensor.itemvalue
tensor.randomrows, cols, seed, scale
tensor.rowsvalue
tensor.valuetensor, row, column
tensor.zerosrows, cols
viz.capturestep, loss

Persistent notebook sessions

Each notebook session retains native global values, imported included modules, and checked function/class definitions. Cells are parsed and checked as a whole against the existing typed state, then execute once in the order requested. Types do not change between cells. Imports belong before functions or statements within each cell and remain available afterward. Session cells reject main; use direct statements or change the notebook to independent program mode with Aner: Use Independent Programs. Save the notebook to retain that mode; Aner: Use Shared Variables selects persistent execution again. A mode change cancels old work and clears state without running or replaying cells.

An explicit top level let or var declaration may reinitialize a name from an earlier cell when its type and mutability match. Duplicate declarations inside one cell are errors. Bare assignment still cannot update let. Function and class names cannot be redefined; restart the session to change an existing definition. A runtime diagnostic in an earlier cell's function retains that definition's source.

Parse/check failure preserves the previous state and executes nothing. Execution failure clears the entire session. Restart, Stop, timeout, notebook closure, and extension reload also clear its state; no earlier cells are replayed. Output and external effects already produced are not rolled back. Saved notebook outputs are historical document content, not serialized variables or model checkpoints. See session behavior and limits.

Operators and evaluation

OperatorsAccepted operands and result
+, -, *, /Two Int64 or two Float64; same type result
%Two Int64; integer remainder
unary -One Int64 or Float64
==, !=Matching Int64, Float64, Bool, or String; Bool result
<, <=, >, >=Matching numeric types; Bool result
&&, `

Precedence from lowest to highest: ||; &&; equality; numeric ordering; + -; * / %; unary - !; function calls. Arithmetic operators group left to right. Parentheses explicitly group expressions. Comparison chains, including mixed comparisons such as a < b == true, require explicit parentheses. Write (a < b) == true, or use a < b && b < c for a range check.

Operands and call arguments evaluate left to right. && and || short circuit; their unneeded right operand is not evaluated, but it is still type checked. The entire program, including unused functions and unreachable branches, must be well typed before any output or other execution occurs.

Int64 overflow, division/remainder by zero, and minimum integer division/remainder by negative one are runtime errors. Integer division truncates toward zero; a nonzero remainder has the dividend's sign. For example, -7 / 3 is -2 and -7 % 3 is -1.

Float64 arithmetic follows IEEE 754: overflow and division by zero may produce infinities or NaNs. A NaN is a numeric value category, not a missing observation. Fast math transformations are disabled. Cross hardware bitwise equality is not promised.

Output and runtime limits

print(value) accepts one Int64, Float64, Bool, or String, writes a newline, and returns Unit. It rejects Unit and all composite values; use module accessors and tensor.item or tensor.value to inspect them. Booleans print as true or false. Strings print their contents without quotation marks. Finite floating point output includes a decimal point or exponent. For example, a value otherwise formatted as 15 prints as 15.0. Nonfinite values print as inf, -inf, or nan. This initial human readable format is not a dataset serialization format.

The CLI limits source files to 2 MiB. Invalid UTF-8 is rejected. The frontend has a combined nesting budget of 128 across statement, expression, and unary parser entries, plus a separate expression tree depth limit of 128. These internal counters do not promise 128 nested parentheses; excessive depth produces an E0001 diagnostic. The interpreter permits at most 128 active function calls and 256 combined active statement/expression evaluations. These runtime nesting limits produce R0002 diagnostics and remain in force when the step budget is raised. Execution defaults to a budget of 1,000,000 steps. --max-steps N changes the budget for run; use a positive integer. Steps count interpreter work, not seconds, and their exact number is not a stable language contract. These safeguards are not a memory quota or security sandbox.

Native sessions additionally bound each source cell to 1 MiB, persistent global bindings to 4,096, functions/methods to 1,024, classes to 256, and retained function/class source to 8 MiB. The editor imposes a 120 second per cell timeout; its native response payload is limited to 1 MiB, reserving 16 KiB for diagnostics and allowing 1008 KiB of stdout. These are resource bounds, not a total retained memory budget. See the notebook limits.

See hello, scalars, control flow, and notebook variables for runnable examples.

Aner Neural module and tensors

Write import aner.nn at the start of a file or cell, before all functions and executable statements. This exposes Tensor values and the nn namespace; successful session imports remain available in later cells. Unknown modules, duplicate imports within one source, and imports nested in blocks or following functions/statements are rejected. Other supported imports may appear alongside it, in any order. These are included modules; general package loading remains future work.

text
import aner.nn;
fn main() -> Unit {
    let x: Tensor = [[1.0, 2.0], [3.0, 4.0]];
    let y = nn.matmul(x, nn.transpose(x));
    print(nn.value(y, 0, 0)); // 5.0
}

Literals must have one or more nonempty rows of equal length. Each entry is an expression of type Float64, evaluated in row order from left to right. Integer entries require explicit conversion. Flat array literals, trailing commas, indexing syntax, and nested tensor elements are unsupported. Tensor currently means rank two and Float64 on the CPU; generic dtype/rank/device syntax is not implemented.

Tensor operations use functions, not scalar operators. Shapes are checked at runtime; check checks argument types and arity but does not execute kernels or infer matrix dimensions. nn.add_bias explicitly broadcasts a single bias row; other elementwise operations require exact shapes. Nonfinite tensor inputs, outputs, and gradient contributions are rejected, unlike scalar Float64 arithmetic. Tensor literals and neural kernels use R2001; the additional modules translate their own API failures as specified below. Memory/graph/operation limits and the full differentiation contract are in Aner Neural.

Ordinary Tensor assignments share immutable numeric storage. Differentiation gradient state is shared by aliases and is updated by nn.backward; Tensor is not a pure value with no hidden state. Optimizer calls return new parameter tensors; reassignment must use var. General effects and concurrency guarantees are not implemented.

The interpreter step budget counts tensor dispatch, not each kernel iteration; neural operations have separate fixed resource limits. Neither policy is a process wide memory quota or security sandbox.

Visual capture module

import aner.viz; exposes viz.capture(step: Int64, loss: Tensor) -> Unit; use it alongside aner.nn. When aner run FILE --debug-viz NEW_DIRECTORY is enabled, the call records a copy of the graph reachable from a scalar loss, including available gradients. It does not execute backward or update parameters. Without recording, capture is an operation with no effect after ordinary argument evaluation; full type checking still applies.

Use a strictly increasing nonnegative step counter. Shapes, counter range, capture limits, optional raw values, and report files are specified in Aner Visual Debugger. Capture failures use source diagnostic R2101; output directory or report writing failures use CLI diagnostic E9004. Ordinary program output remains on stdout, and debug report status goes to stderr. Reports are written after successful execution only.

Proposed additional numerical syntax: not implemented

Vector<Float64>, Matrix<Float64>, flat array literals, indexing operators, and @ remain a later design/implementation milestone. Rank two Tensor literals and native matrix multiplication are available through aner.nn. The linear algebra example is a proposal and currently fails to check.

Initial numerical containers should be immutable values with fixed rank and runtime dimensions. Candidate operations are equal shape addition/subtraction and elementwise multiplication, explicit Float64 scaling, matrix multiplication with @, checked zero based indexing, dot, transpose, sum, mean, len, rows, and cols. Implicit broadcasting is not proposed. Shape, allocation size, memory budget, empty container, and reduction rules must be specified before implementation.

The proposed example's mathematical results are a dot product of 32.0 and a matrix product of [[19.0, 22.0], [43.0, 50.0]]. These are expected results, not an execution report.

Proposed additional data features: not implemented

Typed tables and datasets need nullable fields, dates/timestamps, categorical values, measurement units, exact decimals where required, and schema conversion. Zero, empty strings, NaN, and missing values must remain distinct. Float64? and missing are candidates, not active syntax. Patient identifiers should remain text when leading zeros and exact formatting matter.

Large data operations need explicit schemas, predictable join cardinality, batch processing, inspectable query plans, and bounded previews. Dense matrices alone do not address this problem. The initial external CSV module below adds explicit typed columns and missing tokens. Clinical meaning, patient/time partitions, units, and richer missingness reasons still require separate contracts. Each feature should gain valid/invalid examples and a precise contract before implementation.

Built in dataset values

import aner.dataset; enables the offline teaching catalogue and immutable Dataset, DatasetSplit, and Standardizer values. It may appear in any order with the other built in imports, once per module in each source, before functions and executable statements. Dataset APIs expose Tensor independently; nn.* operations still require an explicit neural import. Composite dataset/scaler values have no direct equality or print operation. Runtime dataset errors use R2201 with source locations. See the dataset reference for the complete implemented API and current numeric classification scope.

Shared tensor, metrics, and fitted ML values

aner.tensor, aner.metrics, and aner.ml are included modules imported once per source before functions and executable statements. Like aner.nn, aner.dataset, and aner.data, each exposes the Tensor type and rank two literals. Qualified calls require their own module import: importing aner.ml does not enable tensor.*, nn.*, or metrics.* calls. aner.viz alone does not expose Tensor. All module namespaces and type names are reserved regardless of imports.

aner.tensor provides zeros, random, rows, cols, value, item, and row wise argmax. It currently uses the same native storage implementation as Aner Neural, through a public facade; it is not a second tensor type or a completed backend extraction.

aner.ml exposes KNNClassifier and KMeansModel. Aner code creates these only through their respective fit functions. They may be function parameters, return values, or inferred bindings; copies share immutable fitted state detached from autograd. ml.predict(model, features) selects a built in signature using the model's static type. This does not introduce user defined overloads or a universal model class. K means does not support ml.predict_proba; an unsupported signature fails before execution. Models cannot be printed or compared directly.

metrics.accuracy(predicted_ids, true_ids) takes matching N × 1 label tensors. Neural code first obtains IDs with nn.argmax(scores) or tensor.argmax(scores). The existing nn.accuracy(scores, true_ids) retains its score matrix interface. metrics.mean_squared_error returns a detached Float64 evaluation value; differentiable neural loss remains nn.mse. Confusion matrices and adjusted Rand index are also available.

ML API failures use R2301, metrics use R2401, and neutral tensor calls use R2501, with source locations. Errors in an argument expression retain that expression's diagnostic; for example a nonfinite tensor literal still reports R2001. Shapes and numerical limits are runtime checks, whereas names, imports, arity, and types are checked before any execution. See the full API and algorithm contracts.

External CSV data module

import aner.data; exposes DataSchema, DataScan, DataReport, DataTable, and the existing Tensor type/literal syntax. Calls in other namespaces still need their own imports. These new values are opaque immutable handles; Aner code creates them with module functions, may pass or return them through annotated functions, and cannot compare or print them directly. data.explain(scan) and data.describe(report) return inspectable text.

CSV schemas declare column type names as Strings: "String", "Int64", "Float64", and "Bool", plus a Bool indicating whether missing cells are permitted. This adds missingness inside native data columns, not nullable scalar syntax or a missing language keyword. A missing cell must be checked through data.is_missing; reading it through a typed scalar accessor fails. No implicit scalar conversions are added.

The two data.where_ge signatures accept Int64 or Float64 thresholds; execution validates that the named column has that exact schema type. The data.rows signatures accept DataReport or DataTable. These are statically selected built in overloads, not user defined overloading. Shapes, named column existence/types, CSV parsing, and allocation limits remain runtime checks.

Plan construction performs no file read. Executing a plan resolves relative paths against the process working directory, reopens the source, validates the ordered header and every source field before filtering, and processes selected values in bounded column batches. data.collect has separate selected row and managed allocation limits; exceeding either returns an error without a partial table. data.to_tensor separately checks missingness, numeric types, exact integer conversion range, cell count, and output payload bytes. The source may change between executions; plans are not immutable data snapshots.

Native module errors use source diagnostic R2601. Resource limits constrain managed buffers per execution, not all interpreter values or total process RAM. The complete implemented functions and CSV/memory contracts are in Aner Data. No join, sorting, group by, external schema file parser, Parquet, SQL, or general nullable scalar API is introduced by this milestone.

Aner handbook · Guides and API reference