Typed CSV data
Declare a schema, inspect missingness, filter rows, and control managed execution buffers.
aner.data reads external CSV files through declared schemas and lazy scan plans. data.summarize executes a plan in bounded batches; data.collect explicitly materializes a table with row and workspace limits. This is the first external file data module. It is separate from aner.dataset, the bundled Iris/Wine teaching catalogue.
The implementation is native C++. This milestone provides strict typed CSV reading, projection, a small filter set, missing value counts, numeric means, typed table access, and explicit conversion to a Tensor. It does not yet provide joins, dates, automatic cleaning, type inference from files, Parquet, SQL, arbitrary aggregate expressions, or model training directly from a batch stream.
Run the example
Follow the one time installation guide to put aner on your PATH. Run from the folder containing the binary bundle's examples directory so examples/fixtures/visits.csv resolves correctly:
aner check examples/data_csv.aner
aner run examples/data_csv.anerThe same commands work on macOS, Linux, and Windows, including VS Code's terminal. VS Code also provides Aner: Run CSV example, or use the current file check/run tasks. Extension 0.1.16 highlights the module and types with the active light or dark theme.
The file visits.csv contains six synthetic records; they do not represent patients or clinical rules. Its columns are record_id, age, measurement, site, and eligible. The example declares record_id as String so identifiers such as 0001 retain leading zeros. This fixture demonstrates type and missing value behavior, not a medical analysis.
Declare a schema and a plan
import aner.data;
fn main() -> Unit {
var schema: DataSchema = data.schema();
schema = data.column(schema, "record_id", "String", false);
schema = data.column(schema, "age", "Int64", true);
schema = data.column(schema, "measurement", "Float64", true);
schema = data.column(schema, "site", "String", false);
schema = data.column(schema, "eligible", "Bool", false);
let source: DataScan = data.scan_csv("examples/fixtures/visits.csv", schema, "NA");
let adults = data.where_ge(source, "age", 18);
let selected = data.include(data.select(adults, "measurement"), "age");
print(data.explain(selected));
let report: DataReport = data.summarize(selected, 1048576, 2);
print(data.describe(report));
print(data.scanned_rows(report));
print(data.rows(report));
print(data.mean(report, "measurement"));
}data.schema() starts an empty schema. Each data.column returns a new schema containing the added column; assigning it to var schema updates that binding. A scan requires the complete source schema, including columns you will later exclude. Type strings are exactly Int64, Float64, String, or Bool. The final Bool argument declares whether that column may contain the configured missing token. Names must be unique within the schema. Column names are case sensitive and the CSV header must match their names and order exactly.
DataSchema, DataScan, DataReport, and DataTable are opaque immutable value types. Copies share immutable state. They work in bindings and function signatures but cannot be printed or compared directly; use data.explain, data.describe, and accessors. Import aner.data to use these types. It also exposes Tensor annotations and literals, but qualified calls such as tensor.rows still need their own import aner.tensor;.
Building or changing a plan does not open the file. data.explain describes the plan without reading rows. Each summarize or collect call opens the path and executes a new scan. A plan is not a cached result or a source snapshot: if the file changes between calls, results can differ. Relative paths resolve against the process working directory at execution time, not the .aner file's directory. aner check checks the program without reading its CSV data.
Strict CSV and explicit missing values
The reader expects comma separated UTF-8 text with a header record. NUL bytes are rejected in CSV fields and schema/path/filter metadata. An initial UTF-8 BOM is accepted. Quoted fields may contain commas, doubled quotes, and embedded newlines; quoting does not prevent numeric conversion according to the schema. Record endings are LF or CRLF. A bare CR record separator is rejected. A blank line is a record with one empty field, not a row silently skipped; it therefore fails a multi column schema.
Every source field is checked for CSV structure, UTF-8, declared type, and nullability before its row is filtered. This includes unprojected columns and rows that do not satisfy the filter. There is no automatic trimming, coercion between column types, filling, or dropping of invalid rows. A successful scan validates the complete source; an error stops execution rather than returning a partially valid report.
The missing token applies only when the complete field is unquoted and exactly equal to the token given to scan_csv. For the token NA:
| CSV field | Meaning |
|---|---|
NA | Missing; accepted only if the column is nullable |
"NA" | A present String value NA, or an invalid numeric/Bool value |
"" | A present empty String; invalid for numeric/Bool columns |
| Empty unquoted field | A present empty String with token NA; numeric/Bool parsing fails |
NA or NA | Present text including the space; it is not the missing token |
The configured token must fit an unquoted field: it cannot contain commas, quotes, CR, or LF. If the configured token is the empty string, an unquoted empty field is missing while "" remains a present empty String. Missing values are never substituted with zero, false, or empty text. There is no general nullable scalar type added by this API: tables expose data.is_missing, and reading a missing cell through a typed getter fails.
Int64 fields accept a minus sign and decimal digits, including leading zeros, within the Int64 range. Float64 fields accept representable finite decimal/exponent forms, including .5, 5., and 1e3; the entire field must parse. Leading plus signs, surrounding whitespace, hexadecimal forms, NaN, and infinity are not accepted numeric data. These CSV field rules are separate from Aner source literal syntax. Bool fields are exactly true or false. String columns preserve text rather than interpreting numeric looking identifiers.
Plan operations
| Function | Result and behavior |
|---|---|
data.schema() | Empty DataSchema |
data.column(schema: DataSchema, name: String, type: String, nullable: Bool) | New DataSchema with one column appended |
data.scan_csv(path: String, schema: DataSchema, missing_token: String) | Lazy DataScan, initially projecting all schema columns |
data.select(scan: DataScan, column: String) | New scan whose projection is exactly that one column |
data.include(scan: DataScan, column: String) | New scan with another column appended; including a column already projected fails |
data.where_ge(scan: DataScan, column: String, threshold: Int64) | New scan comparing an Int64 column with the given threshold |
data.where_ge(scan: DataScan, column: String, threshold: Float64) | New scan comparing a Float64 column with the given finite threshold |
data.where_eq(scan: DataScan, column: String, value: String) | New scan requiring exact equality in a String column |
data.where_present(scan: DataScan, column: String) | New scan requiring a present value |
data.explain(scan: DataScan) | String description of the plan |
Filters combine with AND. A missing value never passes where_ge or where_eq; where_present makes explicit exclusions useful before Tensor conversion. Filters can refer to schema columns omitted from the projection. Numeric threshold types must match the schema exactly: use 18 for Int64 age and 18.0 for a Float64 column. There is no implicit threshold conversion or string to number comparison. Projection preserves the requested output order and successful execution preserves source row order.
Bounded summaries
data.summarize(scan: DataScan, memory_bytes: Int64, batch_rows: Int64) -> DataReport reads through the file without retaining all selected rows as a table. batch_rows limits selected/projected rows in one batch; an internal byte target may flush earlier. The byte budget still applies even to a single row. This is currently an execution API, not a user visible batch iterator.
| Function | Meaning |
|---|---|
data.scanned_rows(report: DataReport) | Int64 number of source data records validated, excluding the header |
data.rows(report: DataReport) | Int64 number of rows passing every filter |
data.missing(report: DataReport, column: String) | Int64 count of missing values among selected rows in a projected column |
data.valid(report: DataReport, column: String) | Int64 count of present values among selected rows in a projected column |
data.mean(report: DataReport, column: String) | Float64 mean of present numeric values in a projected column |
data.peak_bytes(report: DataReport) | Int64 high water mark for the managed data workspace allocations described below |
data.describe(report: DataReport) | String summary with counts, numeric statistics, and workspace information |
For each projected column, valid + missing == rows. Report column lookups use the projected column names. Numeric means exclude missing values; a mean with no present values fails rather than inventing zero or returning an untyped null. Means are not defined for String or Bool columns. Int64 observations are converted to Float64 for stable aggregation, so the mean is approximate beyond the exact Float64 integer range; stored Int64 table cells remain exact. A scan selecting no rows can still return its counts.
For the example's age ≥ 18 filter, all six records are validated and four are selected. Of those four measurements, three are present (1.5, -2.0, 6.0) and one is missing. Their mean is 5.5 / 3, approximately 1.8333333333333333. The missing age record fails the comparison; it is still part of scanned_rows. The eligible field is validated but not used as a filter in this example. These distinctions avoid treating source row counts, filtered row counts, and observed numeric counts as the same quantity.
Collect and inspect a table
// Inside main(), after constructing selected above:
let complete = data.where_present(selected, "measurement");
let table: DataTable = data.collect(complete, 100, 1048576, 2);
print(data.rows(table));
print(data.column_name(table, 0));
print(data.number(table, 0, "measurement"));
print(data.integer(table, 0, "age"));data.collect(scan: DataScan, max_rows: Int64, memory_bytes: Int64, batch_rows: Int64) -> DataTable retains the selected projection in memory. Exceeding max_rows or the managed byte budget fails; max_rows is not a request to truncate the file or return a prefix. Successful collection still validates all source fields and rows. An early limit failure returns no partial table and does not validate the remaining records. The example's additional presence filter gives three rows with columns measurement, then age.
| Function | Result |
|---|---|
data.rows(table: DataTable), data.cols(table: DataTable) | Int64 table dimensions |
data.column_name(table: DataTable, index: Int64) | String projected column name at a zero based index |
data.is_missing(table: DataTable, row: Int64, column: String) | Bool missing flag |
data.number(table: DataTable, row: Int64, column: String) | Float64 from a Float64 column |
data.integer(table: DataTable, row: Int64, column: String) | Int64 from an Int64 column |
data.text(table: DataTable, row: Int64, column: String) | String from a String column |
data.boolean(table: DataTable, row: Int64, column: String) | Bool from a Bool column |
Rows are zero based; column arguments use projected names. Getters reject out of range rows, wrong column types, and missing cells. data.number does not silently convert an Int64 field; use data.integer. data.rows is a built in static overload for reports and tables, not general user defined overload syntax.
Explicit Tensor materialization
// Requires import aner.tensor for tensor.rows/cols/value:
let values: Tensor = data.to_tensor(table, 1048576);
print(tensor.rows(values));
print(tensor.cols(values));
print(tensor.value(values, 0, 0));data.to_tensor(table: DataTable, max_output_bytes: Int64) -> Tensor is a separate, explicit allocation. It produces a detached rank two Float64 Tensor with the table's row and column order. Every column must be Int64 or Float64, every cell must be present, and the table must contain at least one row. Int64 values must lie in the inclusive interval [−2^53, 2^53]; larger integers are rejected rather than silently rounded. Text identifiers and Bool columns are not accepted.
The output budget covers the Float64 numeric payload: rows × columns × 8 bytes. It excludes Tensor metadata and does not include or release the already materialized table. The shared Tensor limit of 1,000,000 cells also applies. A detached Tensor can then be passed to numerical or model APIs, with fitting and preprocessing choices still explicit. This conversion is not a streaming model training pipeline.
What the byte budget measures
memory_bytes is a managed execution workspace cap for one summarize or collect call. It counts allocations made through the data memory resource for CSV parsing, column batches, aggregation, and collected values, including reserved capacity and temporary old plus new capacity during reallocations. A batch row limit does not override the byte budget: a small number of large String fields can still exceed it. Reducing batch_rows may reduce batch workspace, but it cannot make an oversized record or a too large collected table fit.
The cap excludes bounded immutable plan/schema/handle metadata, allocator bookkeeping and rounding, standard stream and operating system buffers, the interpreter and other Aner values, and returned text such as explain, describe, or copies returned by data.text. It is not a global process memory or RSS limit. peak_bytes reports the managed high water mark, not total RAM use. Keeping multiple collected tables alive retains each table's storage; separate successful calls do not establish a combined process quota. Tensor conversion has its own output payload cap.
Limits and failures
| Limit | Current bound |
|---|---|
| Schema columns | 1 through 128 when scanning; construction can start empty |
| Column name | 1 through 128 UTF-8 bytes |
| CSV path | 1 through 4096 UTF-8 bytes |
| Missing token | 0 through 128 UTF-8 bytes |
| Filters per plan | At most 64 |
| String equality filter value | At most 256 KiB of UTF-8 text |
| Encoded CSV record | At most 1 MiB |
| Decoded CSV field | At most 256 KiB |
batch_rows | 1 through 65,536 |
memory_bytes | 1 through 1,073,741,824 bytes; very small limits can immediately fail allocation |
max_rows for collection | 0 through 1,000,000; zero permits only an empty result |
| Tensor output | At most 1,000,000 cells, within the separate payload byte budget |
Runtime data errors use R2601 at the Aner call site. CSV parse diagnostics also identify the path and one based logical record/column, with record 1 being the header; quoted newlines do not start a new record. Malformed CSV, header mismatches, unknown columns, invalid data, nullability violations, unsupported conversions, file failures, and exceeded limits fail explicitly. Static checking reports missing imports and wrong Aner argument types before execution; schema dependent checks occur when the relevant operation executes. These limits bound particular data operations, not every possible program resource.
The current API does not automatically persist file hashes, row identifiers, source licenses, or a dataset manifest for arbitrary external files. Keep source identity and applicable notices with your research data and results; a path alone does not pin content. General package readers, row provenance, richer statistics, group/time aware splitting, and external storage formats remain future work.