Laboratory Notebook
Experiment Log
Every run gets a record: hypothesis, setup, variables, results, and an honest conclusion. Failures stay visible — negative results steer the roadmap as much as positive ones.
25 RECORDS
- HXO-EXP-0008HELIX-Ω matched identity-diagonal rescueIdentity-initialized learned post-diagonals recover at least half of the pure TensorAxis loss gap while preserving the edge-hardware resource envelope.2026-08-25FAILED
- HXO-EXP-0005HELIX-Ω 10M-class TensorAxis engineering smokeA same-width TensorAxis MLP can reduce parameters, arithmetic estimates, and memory on Apple Silicon while remaining numerically healthy enough for a real-data pilot.2026-08-24PARTIAL
- HXO-EXP-0007HELIX-Ω three-seed WikiText-2 TensorAxis pilotReplacing only the dense MLP projections with same-width one-stage TensorAxis operators improves the validation-loss-versus-resource frontier on edge hardware.2026-08-24FAILED
- HELIX-200M-FINAL-BAKEOFFFinal post-training bakeoff: DPO-100 selected, a challenger left unresolvedMirroring the 120M's proven two-stage recipe (adapt → identity-clean SFT), then adding a short DPO pass, will produce a coherent, identity-stable checkpoint where four direct attempts failed.2026-07-24PARTIAL
- HELIX-200M-SFT-ACTIVEhelix-200m-chat-tonight-v1 — post-DPO SFT round, in progressA short, low-LR chat-oriented SFT pass on top of the DPO-100 checkpoint will improve conversational behavior without disturbing the identity/voice DPO already bound in.2026-07-24ACTIVE
- MEM-SHARED-3KOne contract, three memory systems: Cortex vs SAM vs SAMEThree local persistent-memory architectures can only be ranked fairly under a single deterministic contract — one event generator, one answer key, one stale definition, one payload counter.2026-07-22SUCCESS
- HELIX-200M-SFT-FAILURESDirect SFT on the 209M base: four failuresThe annealed 209M base can go straight to supervised fine-tuning, as the post-anneal evaluation gate recommended.2026-07-19FAILED
- REMY-ROUTER-LIVERemy sub-agent model router: live multi-provider verificationspawn_agents can run concurrent sub-agents across independent providers with per-agent routing, fallback, and enforced cost budgets.2026-07-18SUCCESS
- HELIX-200M-C1X-RECOVERYThe WSD recovery anneal: 'the plateau was the schedule, not the model'The 209M foundation's validation plateau (39.1–40.3) can be broken by actually running the WSD decay schedule the run was designed for — the prior 'anneal' branch never decayed its learning rate at all.2026-07-17SUCCESS
- HX-STAGE0-AUDITHibbieFormer-HX Stage-0 forensic inventoryThe 2025 backup contains verifiable checkpoints and results sufficient to anchor the HX research program.2026-07-09SUCCESS
- HELIX-200M-F35BHELIX-200M foundation trunk (helix-200m-foundation-35b)A 209M-parameter H-E-R hybrid can be pretrained on a 40B-token Chinchilla-matched corpus entirely on one M3 Max, from scratch after the 127M generation was ruled parameter-bound.2026-07-06SUCCESS
- HELIX-120M-5B-EXTThe 127M 5B-token extension: three attempts, one parameter-bound verdictDoubling the 127M's token budget toward 5B total tokens will push validation perplexity meaningfully lower.2026-06-29FAILED
- SLM10M-EVALSLM-10M zero-shot leaderboard evaluationA 10M-parameter GQA transformer trained on 25B curated tokens reaches non-trivial zero-shot accuracy on standard small-model benchmarks.2026-06-23SUCCESS
- HELIX-120M-FOUNDATIONHELIX-120M foundation run, the in-sample scandal, and the corrected held-out resultScaling to 127M params on a 2.5B-token Chinchilla-matched corpus improves on the 56M generation, and the transformer tax measured at 56M should be checkable a second time.2026-06-22SUCCESS
- HELIX-56M-VS-TRANSFORMERHELIX-56M vs. a matched transformer: the first honest tax measurementA same-size, same-data, same-protocol dense transformer will reveal how much perplexity HELIX pays for its O(L)-inference design.2026-06-12SUCCESS
- HELIX-BAKEOFF-0611Architecture bakeoff: rotation vs. nopulse vs. pulse HSLThe original 'pulse' absolute-time-gated HSL recurrence is the right design for HELIX's oscillator-bank state layer.2026-06-11SUCCESS
- MLXR-HOTSWAPmlx-recurrence v2 kernels hot-swapped into a live training runFused Metal scan kernels can replace the chunked MLX path mid-flight in a live SSM+GLA training run with no loss discontinuity and material memory/throughput gains.2026-06-10SUCCESS
- DC2-BENCH-82KD-CSIL-2 capability benchmark across training checkpointsTask capability (dialogue, math, code, factual) emerges measurably across pretraining checkpoints and is sensitive to decoding strategy.2026-04-24SUCCESS
- DC1-DPO-E8D-CSIL-1 DPO/LoRA post-training evaluation (epoch 8)DPO via LoRA can improve instruction adherence on a 56M GLA model without degrading base quality.2026-04-10PARTIAL
- SN3-WT2-EVALSynapNet V3 WikiText-2 evaluation (epoch 50)The 42.6M O(L) SSM+GLA hybrid with bio-mechanisms reaches usable language-modeling quality on WikiText.2026-03-22SUCCESS
- FT-V13C-KILLFieldTransformer v1.3c training run — killedAdding latent-diffusion and entropy-bonus regularizers to v1.3 improves convergence over v1.2 on the same corpus.2026-03-11FAILED
- HF-GOLDEN-RATIOHibbieFormer V3 'Golden Ratio': φ-regularized stream balanceRegularizing the ratio of the wavelet-stream and SDM-stream activation magnitudes toward the golden ratio (φ ≈ 1.618), with a PID-tuned loss weight, improves language-model training.2025-09-06INCONCLUSIVE
- HF-GOLDEN-DUALITYHibbieFormer 'Golden Duality' run — never launchedA smaller (512-dim, 6-layer) HibbieFormer with the same φ-ratio + temperature-duality regularizers trains more efficiently than the 143M V3.2025-09-06FAILED
- HF-TINYSTORIES-215HibbieFormer TinyStories run: 215 plasticity cyclesA recurrent bio-inspired core trained with cycling plasticity phases (hebbian → scaling → pruning) converges stably on narrative text.2025-08-05SUCCESS
- HF-CYCLE150-EVALHibbieFormer cycle-150 comprehensive behavioral evaluationThe cycle-150 checkpoint (step 15,000, train loss 1.0015) generates stable, domain-coherent text across diverse prompt categories.2025-08-05SUCCESS