Research directions
Active design work

Reproducibility as part of the workflow

Can explicit partitions, fitted preprocessing, and source identity reduce avoidable research mistakes?

Aner already carries dataset citations, source hashes, row IDs, explicit seeds, and preprocessing fitted on training data in its teaching workflows. These are a foundation for reproducibility, not a guarantee that every experiment is reproducible.

The research question is whether making these choices visible reduces leakage and makes an analysis easier to audit. Future work includes group and temporal splits, fuller run manifests, model checkpoints, and defined determinism levels across hardware.

A useful evaluation would compare equivalent workflows, record both correctness and researcher effort, and distinguish repeatability on the same platform from numerical agreement across devices. No comparative study or performance result is claimed yet.

Read the technical context

Help investigate this question.

Bring evidence, experience, or a workflow we should understand.

Collaborate