Experiment Record · 2026-06-22
HELIX-120M-FOUNDATION
HELIX-120M foundation run, the in-sample scandal, and the corrected held-out result
Hypothesis
Scaling to 127M params on a 2.5B-token Chinchilla-matched corpus improves on the 56M generation, and the transformer tax measured at 56M should be checkable a second time.
Setup
127,252,744-param HER×5 model, 90% FineWeb-Edu / 10% WikiText-103 (2.5B tokens), 152,588 steps, ~4,400 tok/s. A matched dense-transformer control trained on Colab under identical data, steps, and schedule.
Variables
- ◆Architecture (HELIX vs. matched transformer)
- ◆Validation methodology (in-sample vs. held-out)
Results
- →Headline val PPL 30.07 was later found to be in-sample: val_path and train_path pointed at the identical file in the run config.
- →The transformer control had the exact same bug (26.30 in-sample).
- →Clean held-out re-evaluation at the same 152,588 steps: HELIX 39.83 vs. transformer 33.86 — a 15.0% gap, up from ~8–9% at 56M.
Observations
- ·Every 'val' number produced before this discovery had to be discounted.
- ·Held-out validation became a standing rule for every run afterward; new held-out val sets were built in direct response.
Conclusion
The architecture tax widened at 127M (15.0% vs. ~8–9% at 56M), and the program's worst methodological failure produced its best institutional fix: no more in-sample validation, ever.
Next Experiment
Carry the held-out-only rule into the 209M generation; re-check the tax trend at 209M.