ML tooling / Real-time systems / Engineering notes
ML Training Inspector
I built a live training dashboard and verified its charts and epoch table against a reproducible synthetic PyTorch CPU run.
Inspect the repositoryReviewed 6 October 2026 · Local verification recorded 4 October · Implementation and notes published

The problem
A training loop can produce numbers without making its behavior understandable. I built a browser dashboard to expose loss, accuracy, gradients, class-level results, and run state while training progresses.
My contribution
I connected PyTorch training to FastAPI and WebSockets, built React charts and an accessible epoch table, and added manual stopping and checkpoint saving. Recent work corrects example-weighted loss aggregation, bounds per-client queues, handles disconnects and terminal outcomes, and adds a seeded, download-free CPU demonstration.
One decision: keep run ownership explicit
The backend deliberately supports one shared run in one worker. Conflicting starts receive HTTP 409. Bounded client queues protect memory, and reconnects recover status and the latest epoch metadata rather than promising missed chart history. Stopped and error states remain distinct from successful completion.
Verification and evidence scope
The published record reports 15 backend tests and two frontend tests, plus a successful production build. A real browser CPU run rendered charts and an epoch table whose values matched the demo. The seeded SimpleCNN run used two epochs, 160 synthetic training examples, and 80 validation examples. Model/optimizer loading reproduced validation metrics, but scheduler, random-state, and sampler state are not preserved, so resume is absent. The measured synthetic accuracy is not evidence of generalization.
Demo
Follow the published README to install the CPU demo dependencies and run python scripts/demo.py. For the browser walkthrough, start the backend with one worker and the frontend, choose the synthetic dataset, and run two SimpleCNN epochs. Inspect the charts and matching epoch table, then explore stop and snapshot controls. The capture on this page comes from the verified synthetic run. No public hosted service is linked.
Limits and next work
GPU runs, full CIFAR-10 training, ResNet9 training, containers, and cross-platform repeatability were not verified in that record. There is no authentication or per-user ownership. Heuristic gradient signals have no detection-quality evaluation; some near-zero gradients can be benign. Checkpoint resume and reconnect chart-history replay are not implemented.