Skip to content
D-CSIL

Experiment Record · 2026-06-22

HELIX-120M-FOUNDATION

HELIX-120M foundation run, the in-sample scandal, and the corrected held-out result

SUCCESSREPLICATED

Hypothesis

Scaling to 127M params on a 2.5B-token Chinchilla-matched corpus improves on the 56M generation, and the transformer tax measured at 56M should be checkable a second time.

Setup

127,252,744-param HER×5 model, 90% FineWeb-Edu / 10% WikiText-103 (2.5B tokens), 152,588 steps, ~4,400 tok/s. A matched dense-transformer control trained on Colab under identical data, steps, and schedule.

Variables

  • Architecture (HELIX vs. matched transformer)
  • Validation methodology (in-sample vs. held-out)

Results

  • Headline val PPL 30.07 was later found to be in-sample: val_path and train_path pointed at the identical file in the run config.
  • The transformer control had the exact same bug (26.30 in-sample).
  • Clean held-out re-evaluation at the same 152,588 steps: HELIX 39.83 vs. transformer 33.86 — a 15.0% gap, up from ~8–9% at 56M.

Observations

  • ·Every 'val' number produced before this discovery had to be discounted.
  • ·Held-out validation became a standing rule for every run afterward; new held-out val sets were built in direct response.

Conclusion

The architecture tax widened at 127M (15.0% vs. ~8–9% at 56M), and the program's worst methodological failure produced its best institutional fix: no more in-sample validation, ever.

Next Experiment

Carry the held-out-only rule into the 209M generation; re-check the tax trend at 209M.