Research directions
Active design work

Data contracts before model code

Make schema, missingness, provenance, and resource decisions explicit without a long cleaning script.

The current CSV reader requires a declared schema and an exact missing token. It validates every source column before applying filters, preserves text identifiers, and reports missing values among selected rows. Projection reduces stored data, not the amount of CSV validation.

Managed execution budgets cover defined parser, batch, aggregation, and collection allocations. They are not a quota for total process memory, and training directly from a stream remains future work.

The next questions concern richer schemas, group/time semantics, columnar sources, query planning, and transparent cleaning decisions. Evaluation should include malformed records, reproducibility of transformations, peak memory boundaries, and representative research workflows.

Read the technical context

Help investigate this question.

Bring evidence, experience, or a workflow we should understand.

Collaborate