Laboratory Notebook
Experiment Log
Every run gets a record: hypothesis, setup, variables, results, and an honest conclusion. Failures stay visible — negative results steer the roadmap as much as positive ones.
13 RECORDS
- HELIX-200M-F35BHELIX-200M foundation run (helix-200m-foundation-35b)A 209M-parameter H-E-R hybrid can be pretrained to competitive validation perplexity on a 40B-token Chinchilla-matched corpus entirely on one M3 Max.2026-07-18ACTIVE
- REMY-ROUTER-LIVERemy sub-agent model router: live multi-provider verificationspawn_agents can run concurrent sub-agents across independent providers with per-agent routing, fallback, and enforced cost budgets.2026-07-18SUCCESS
- HX-STAGE0-AUDITHibbieFormer-HX Stage-0 forensic inventoryThe 2025 backup contains verifiable checkpoints and results sufficient to anchor the HX research program.2026-07-09SUCCESS
- SLM10M-EVALSLM-10M zero-shot leaderboard evaluationA 10M-parameter GQA transformer trained on 25B curated tokens reaches non-trivial zero-shot accuracy on standard small-model benchmarks.2026-06-23SUCCESS
- MLXR-HOTSWAPmlx-recurrence v2 kernels hot-swapped into a live training runFused Metal scan kernels can replace the chunked MLX path mid-flight in a live SSM+GLA training run with no loss discontinuity and material memory/throughput gains.2026-06-10SUCCESS
- DC2-BENCH-82KD-CSIL-2 capability benchmark across training checkpointsTask capability (dialogue, math, code, factual) emerges measurably across pretraining checkpoints and is sensitive to decoding strategy.2026-04-24SUCCESS
- DC1-DPO-E8D-CSIL-1 DPO/LoRA post-training evaluation (epoch 8)DPO via LoRA can improve instruction adherence on a 56M GLA model without degrading base quality.2026-04-10PARTIAL
- SN3-WT2-EVALSynapNet V3 WikiText-2 evaluation (epoch 50)The 42.6M O(L) SSM+GLA hybrid with bio-mechanisms reaches usable language-modeling quality on WikiText.2026-03-22SUCCESS
- FT-V13C-KILLFieldTransformer v1.3c training run — killedAdding latent-diffusion and entropy-bonus regularizers to v1.3 improves convergence over v1.2 on the same corpus.2026-03-11FAILED
- HF-GOLDEN-RATIOHibbieFormer V3 'Golden Ratio': φ-regularized stream balanceRegularizing the ratio of the wavelet-stream and SDM-stream activation magnitudes toward the golden ratio (φ ≈ 1.618), with a PID-tuned loss weight, improves language-model training.2025-09-06INCONCLUSIVE
- HF-GOLDEN-DUALITYHibbieFormer 'Golden Duality' run — never launchedA smaller (512-dim, 6-layer) HibbieFormer with the same φ-ratio + temperature-duality regularizers trains more efficiently than the 143M V3.2025-09-06FAILED
- HF-TINYSTORIES-215HibbieFormer TinyStories run: 215 plasticity cyclesA recurrent bio-inspired core trained with cycling plasticity phases (hebbian → scaling → pruning) converges stably on narrative text.2025-08-05SUCCESS
- HF-CYCLE150-EVALHibbieFormer cycle-150 comprehensive behavioral evaluationThe cycle-150 checkpoint (step 15,000, train loss 1.0015) generates stable, domain-coherent text across diverse prompt categories.2025-08-05SUCCESS