Browse the handbook
Design

Notebook inspection design

A proposal to connect variables, data, network structure, and training history in one research workspace.

Design proposal, October 10, 2026. Add a coordinated inspector to Aner Notebook so a researcher can connect source code, variables, data, neural network structure, and training history. VS Code extension 0.1.16 provides native persistent sessions alongside independent complete program execution. Variables, imports, and functions can now remain available between cells. The inspector panels, pause/step debugging, and streaming charts described here remain proposals.

The user experience

Keep the notebook as the main surface, with an Aner Inspector beside it and training history beneath it. Every computed view identifies the execution and observation it represents. Selecting an older loss point selects the corresponding recorded state when available; it must never suggest that the running program has rewound.

ViewWhat the researcher can inspect
VariablesName, Aner type, policy permitted scalar value or bounded preview, tensor shape/dtype/device, scope, and state version
DataSchema, nullability, available row counts and missingness, source identity and attribution
NetworkLayer/operation structure, shapes, parameters, and available activation/gradient summaries
TrainingNamed metrics over steps or epochs, aggregation and data partition, selected observation
DebuggerCurrent source location, call stack, pause/continue/step controls, and local variables in the selected frame

Start the variable explorer as read only. Editing a variable through a panel can change an experiment without changing its source; a later editing facility should require a valid typed assignment and record it in the execution history. An inspector must not execute arbitrary user code just to display a value.

The connected experience is the design goal. Spyder already provides a variable explorer and debugger; JupyterLab provides debugging and variable inspection with compatible kernels. These individual features are established. Aner should be evaluated on whether the combined workflow reduces mistakes and effort for its users, rather than claiming a language level advantage from having panels.

A consistent state across views

A variable table, network capture, and metric point must refer to compatible observations. Define a versioned observation envelope with session generation, run ID, cell ID and source revision, event sequence, snapshot ID, frame/binding identity, parameter version, step, optional epoch/batch, and capture phase. Include runtime version, dtype/device, seeds, and input identities when actually known. Do not infer provenance from a pathname or a random seed alone.

Use explicit states such as Running: last observation, Paused: current frame, Historical snapshot, Finished, and State unavailable. Display the age or step of an observation while computation continues. If a metric was recorded without a matching network/variable capture, say so; do not substitute a nearby frame as though it were simultaneous.

Inspectors consume immutable copied summaries or versioned handles with bounded lifetimes. They must not hold an entire autograd graph alive merely to keep a panel open. A session restart invalidates live handles; saved observations remain historical data. A saved notebook is not a complete process checkpoint.

Int64 values must survive transport without JavaScript precision loss. Use a typed representation for large integers and nonfinite floating values instead of relying on unqualified JSON numbers. Diagnostics and unsupported inspection requests should produce explicit unavailable states rather than guessed values.

Neural networks and loss charts

Use an architecture view for model structure and a separate operation view for the actual executed tensor graph. Small models such as XOR can expand to neurons; large models should default to collapsed layers and shapes. Do not draw every trainable scalar as a neuron. Selecting a layer should reveal its parameter shapes, operation, available summaries, and source location where the runtime can supply one.

Aner already has viz.capture(step, loss) and bounded native tensor/gradient snapshots for offline reports. Current graph node IDs are local to a capture, and labels are not stable model identities. Before linking layers across training observations, add stable logical identities and versioned runtime node references. Generated or conditional graph structure must remain distinguishable from an authored model diagram.

Capture phases matter: a loss and its gradients belong to the forward/backward computation that produced them. The existing Iris loop captures after backward and before the SGD updates. Do not pair that loss with post update weights without labeling the difference. Gradients not yet computed, invalidated by failure, or omitted by a capture budget must be shown as unavailable.

Keep metric collection lightweight and separate from heavier graph snapshots. A metric record needs a name, value, step, optional epoch/batch, train/validation/test partition, reduction, sample/element count, aggregation weighting, and observation identity. Distinguish batch loss from an aggregate epoch loss; unequal batches must not become an unweighted mean of batch means. A display should not assume that a variable named epoch establishes those semantics.

The current Iris example has a training/test split and evaluates its held out test rows after training. It does not supply a validation curve. A future validation curve requires a separate validation partition and an explicit evaluation schedule; never silently use held out test results to guide model selection.

Native execution and debugging

Two capabilities are related but distinct:

  • A running independent program can emit metric and graph observations before it exits. This can improve the existing independent program mode without promising shared cell state.
  • The implemented persistent session retains notebook level variables, datasets, and models for later cells. It does not yet expose a variable explorer. Inspecting function locals while paused additionally requires call stack and lexical scope information.

Today standalone execute creates fresh state, while the native Session owns globals, imports, and function definitions across evaluations. Function locals live in frames that disappear after returning. A frame's storage can retain values whose lexical block has ended; an inspector must respect active scope rather than listing every occupied storage slot. The implemented session lifecycle distinguishes static failure from runtime reset. Observation identities must invalidate correctly on those resets, cancellation, or restart.

Extend the native session foundation with an observation/control interface used by the VS Code inspector, notebook controller, future Jupyter adapter, and terminal tooling. The current framed execution protocol transports source and text results; it does not yet provide structured variable or training observations. Use framed, versioned messages distinct from human stdout/stderr; do not recover structured state by parsing printed numbers. The producer must bound message size, buffering, and capture cost, with visible dropped/coalesced observation counts. Source location mappings must identify the cell revision that ran.

Pause and step require cooperative safe points in the interpreter and long native operations, plus explicit paused frame ownership. Cancellation and pause are different: the current notebook Stop terminates the child process. It cannot preserve that process's variables for later continuation. Future GPU/cluster operations may only stop or observe at backend defined boundaries; the UI must report the actual boundary reached.

Use VS Code's debugging integration for breakpoints, stepping, and call stacks when those native capabilities exist. Custom data/model panels should consume the same state identities, not implement a separate debugger. Remote execution keeps native data on its execution host and sends only requested bounded observations.

Large data and research controls

Opening an inspector must not automatically execute a lazy CSV scan, collect a large table, transfer a GPU tensor to CPU, or compute an expensive reduction. Begin with cheap metadata; request further work explicitly and make its cost/cancellation visible. A tensor's numeric payload size is different from retained graph storage and total process memory; label each measurement's scope.

Default to bounded summaries for datasets and tensor values. Raw records, detailed arrays, scalar strings/identifiers, and potentially identifying path metadata require explicit inspection/export choices under the workspace policy, with clear indication that notebook outputs can persist them. Summaries are not a guarantee of de identification. This follows the existing statistics first visual capture approach and is especially relevant to the intended medical research workflow.

Keep loss history and snapshots bounded independently. Downsampled chart displays should disclose that reduction, preserve meaningful anomalies, and expose missing records. Do not replace the current capture caps with unbounded epoch by epoch graph retention. No performance or memory overhead claim is valid until measured with observation disabled and enabled.

Implementation sequence and acceptance

  1. Connect native observations to notebooks. First expose the existing recorded loss/topology/gradient report beside a complete program, explicitly as historical captures. Then reuse those graph summaries and stream named loss observations during one complete program. Display training history and matching captured topology, including omissions and capture phase. Verify that captured values match native results, prints cannot masquerade as events, slow clients remain bounded, and inspection does not retain old training graphs. The existing session foundation does not remove the need for this observation bridge.
  2. Add a read only variable explorer to the persistent session. The runtime now retains bindings with explicit type, rerun, scope, error, interruption, and lifetime rules. Expose bounded summaries with stable identities and verify that runtime reset/restart invalidates them. Merely opening the panel must not scan data or materialize tensors.
  3. Add coordinated pause and step debugging. Expose source locations and stacks, implement cooperative safe points, and keep all current state views on one paused snapshot. Verify a historical selection cannot mutate or resume the running program, cancelled work stays cancelled, and editing source while paused does not silently remap a breakpoint to different code.
  4. Add reproducibility and extensible inspectors. Track changed source/input dependencies and stale results, provide fresh ordered execution, and define versioned inspection/rendering contracts for third party types. Validate standard text fallbacks for terminals and unsupported renderers.

The first practical target is an Iris or XOR training run with a loss curve and selectable recorded network state inside VS Code. The full Spyder like workspace can now grow on the implemented native session and its owned values. Structured observations still need their own bounded interface; text output is not a substitute for native inspection.

Aner research & design · Proposal