Language architecture
The current checked interpreter and the proposed path toward native compilation and a shared runtime.
Working design, October 10, 2026.
Implementation language
Use C++ for the first implementation. Compared with C, standard containers, RAII resource management, variants, and typed abstractions reduce manual bookkeeping in parsers, syntax trees, and runtime values. Keep the code simple: explicit ownership, modest templates, compiler warnings, and a small dependency set. C remains appropriate for stable native interfaces later. This is an engineering recommendation, not a claim that C cannot implement Aner.
Aner defines its own source semantics, type system, and execution model. Its native runtime can use existing numerical libraries and compiler infrastructure while preserving those language contracts.
One language with two execution paths
.aner source from any editor
-> lexer -> parser -> name resolution -> static type checking
-> checked Aner syntax tree
|-> interpreter -> Aner runtime -> results (current)
|-> typed IR -> native lowering -> LLVM -> executable (future)Currently, aner run interprets checked scalar control flow and dispatches tensor, neural, classical ML, metrics, dataset catalogue, and external CSV operations into native C++ libraries. Later, aner build compiles it ahead of time. Both paths share numeric kernels, shape validation, error conventions, and tests. The native runtime must not depend on the interpreter to execute ordinary compiled instructions.
LLVM's frontend tutorial covers language frontends and object code generation. LLVM ORC supplies JIT infrastructure, but a JIT is optional and is not the first interpreter. Install a standalone LLVM SDK only when the compilation milestone needs it; Apple's installed Clang alone is not proof that the LLVM development libraries are available.
Keep source spans in every frontend representation so errors can identify the actual expression. The first interpreter executes the checked syntax tree, with resolved bindings and expression types populated by the checker. A lower level typed IR can follow when native lowering or numerical operations need it. Avoid designing a large custom IR before real operations need it.
Large datasets are a design requirement
Doctors should be able to ask understandable questions of data without manually coordinating many libraries or learning memory management tricks. Data scientists still need ordinary functions, precise errors, inspectable plans, and extensible computations.
The first external data milestone now loads an explicit CSV schema, selects columns, filters records, computes summaries, and produces a small result while processing native batches. Its exact current contract is in Aner Data. The current built in Dataset is an immutable view of a small numeric teaching dataset, documented in Aner datasets. External ingestion uses the separate aner.data API and its DataSchema, DataScan, DataReport, and DataTable types; it does not extend the teaching Dataset type. DataScan retains a plan, while DataTable retains an explicit collection. A counting native allocator bounds parser, batch, aggregation, and collection capacities, including reallocations; bounded plan metadata and other interpreter values remain outside that budget. Explicit collection also limits selected rows. Sorting and joins will need spill to disk or documented capacity limits; batching alone does not make every algorithm bounded memory.
Design contracts should cover missing values, date/time parsing, patient identifiers, duplicate keys, join cardinality, measurement units, and data quality summaries. For repeated observations, partitioning and summaries may need patient/group or temporal structure. These semantics require domain examples and runtime validation; a language cannot infer the medical meaning of arbitrary columns.
Consider Apache Arrow for future columnar interchange. Its language neutral format describes typed buffers and validity bitmaps. Its native interfaces could support columnar interchange in Aner; a complete database engine would require additional components. Database connectors and Parquet readers remain future integration choices.
Use synthetic fixtures for development. No patient data or database credentials are needed to establish the toolchain or first language examples. The language itself is not a clinical decision system.
Leave room for acceleration
Preserve vector, matrix, and dataset operations until lowering so execution can later select suitable native kernels. Put kernel selection behind a small runtime interface. Do not make a per cell interpreter callback the only representation of an array operation.
Begin with one CPU thread and predictable results. Later measure multicore CPU kernels and BLAS integration. GPU support requires device memory, transfer, kernel, and failure policies; CUDA alone cannot serve every Mac, Windows, and Linux machine. Choose supported hardware backends explicitly after profiling. Clusters and supercomputers additionally require partitioning, communication, scheduling, checkpointing, and deployment formats.
Hardware acceleration is not automatic portability. An OS supported language executable can still use a CPU fallback when a GPU backend is unavailable. Define numerical tolerances and reduction policies before asserting agreement across parallel engines. None of these accelerator dependencies is needed for the current environment.
VS Code and other editors
The command line tools are the product interface. An editor saves .aner files and invokes those tools. Use VS Code's C/C++ and CMake Tools extensions to develop Aner itself.
The local extension in editors/vscode associates .aner files with Aner and supplies syntax highlighting, bracket handling, comments, and snippets. Standard syntax scopes let light and dark themes choose their own colors. It is declarative: no extension runtime, language server, or semantic completion is required for these features.
Workspace tasks build Aner, check or run the current file or the hello program, and run tests. A problem matcher sends CLI diagnostics to the Problems panel. A future language server can provide live diagnostics and semantic completion across editors; Aner source debugging is also future work. The VS Code language extension overview distinguishes declarative and programmatic features. Any future Node.js/TypeScript extension tooling remains separate from Aner program execution.
Development stages and completion criteria
| Stage | Work | Completion criterion |
|---|---|---|
| E0 environment, complete on this Mac | C++ project, CMake presets, editor tasks, specifications | Native environment check builds and passes on this Mac; other platforms have reproducible setup instructions |
| E1 scalar language, implemented | Lexer, parser, types, functions, conditions, loops, diagnostics | Valid/invalid scalar programs check correctly; precise source locations |
| E2 interpreter, implemented | Execute the checked program and basic IO | aner run executes the scalar examples through Aner’s native runtime |
| E3 numerical core, first subset implemented | Rank two Float64 CPU tensors, shape checks, shared kernels, reverse mode differentiation, SGD | Small dense network trains from Aner; generic Vector/Matrix syntax and broader numerical APIs remain future work |
| E4 native compiler | LLVM lowering, runtime ABI, standalone executable | Same programs/results/errors as interpreter under the declared numeric contract |
| E5 useful data handling, first CSV subset implemented | Explicit schema and missing tokens, native batches, filters/projection, count/mean summaries, bounded collection | Synthetic input larger than the managed working budget is summarized correctly; joins, sorting, grouping, richer types and additional formats remain future work |
| E6 expansion | Safe joins, provenance, models, acceleration | Measured user benefit and explicit correctness/resource guarantees for each addition |
Windows, macOS, and Ubuntu are supported platform goals from the start. Use portable paths, standard C++, and CMake; avoid platform specific shell commands in the core. Establish a three platform build/test matrix when a repository host is selected. Local checks cover the host build, scalar interpreter, CPU neural and classical ML libraries, metrics, and teaching datasets; they do not verify the Windows or Ubuntu toolchains, GUI debugging, GPU support, or the future native Aner compiler.
Neural network architecture proposal
The neural network execution proposal develops the next numerical stages into typed tensors, automatic differentiation, CPU kernels, Apple/NVIDIA GPU backends, reproducible experiments, and third party packages. It remains a broader design proposal. The first implemented subset is documented in Aner Neural: a built in module, rank two Float64 CPU tensors, automatic differentiation, and SGD. General third party packages, checkpointing, and accelerator execution are not implemented. Tensor acceleration can be reached from the interpreter before all scalar code has an ahead of time compiler.
Portable training observation
The aner_viz C++ target observes the existing neural graph through bounded snapshots; the interpreter passes an optional recorder into execution. Recording is off by default. Numeric buffers and gradient state are copied for inspection, never replayed or updated by a renderer. Reports describe reachable executed tensor graphs, not every path through the source program.
aner.viz produces a versioned trace, terminal summary, and a standalone browser report. The HTML template is embedded at build time so copying the native executable does not break report generation. Runtime execution on a headless host uses Aner’s native runtime and needs no browser, HTTP server, or GUI toolkit. Larger tensors retain summaries and larger graphs report explicit omissions. A future hierarchy of model/layer/operation groups can improve those summaries without changing training semantics.
The proposed aner.plot library should share chart specifications and renderers while serving datasets and statistics independently of neural training. Keep source/model definitions authoritative; graphical model editing requires a specified source translation and reviewable diff. See the current capture contract and roadmap.
First shared ML boundary
The current builtin prototype includes aner_ml for immutable KNN/K means fitted state, aner_metrics for detached evaluation, and aner_tensor for neutral creation/inspection. The latter currently delegates to existing aner_nn tensor storage; the public separation precedes an internal backend extraction. The interpreter resolves the two ml.predict signatures by fitted model type before execution. General module loading, user defined fitted types, pipeline protocols, serialization, and algorithm specific packages remain design milestones. The ML architecture distinguishes those proposals from the implemented APIs.
The external data engine uses native column batches with per column validity bytes; missingness is separate from each stored value. Float64 Tensor remains a distinct dense model format with an explicit copy/conversion boundary. User callbacks per CSV cell are not part of the execution path. Future Arrow interchange is a design option, not a current dependency. CSV projection reduces retained columns while the strict full schema validator still parses and checks all source fields. Record/field limits and byte aware flushing supplement row count batch limits.