Skip to content
D-CSIL

Experiment Record · 2026-07-18

HELIX-200M-F35B

HELIX-200M foundation run (helix-200m-foundation-35b)

ACTIVEMEASURED

Hypothesis

A 209M-parameter H-E-R hybrid can be pretrained to competitive validation perplexity on a 40B-token Chinchilla-matched corpus entirely on one M3 Max.

Setup

28 layers, d_model 1024, MoE top-2 of 4. Corpus: 90% FineWeb-Edu 10BT + 10% WikiText-103, per-doc repeat cap 4, 13-gram decontamination. LR 3e-4 cosine, 2,000 warmup, effective batch 16 sequences, BF16.

Variables

  • Scale (209M vs prior 57M/127M)
  • Corpus size and mixture

Results

  • Step 120,140: train loss 2.824 (PPL 16.85), validation PPL 39.4.
  • Throughput ~2,900 tok/s sustained; checkpoints every 5k steps (~1.2GB NPZ).

Observations

  • ·Val PPL 105.2 → 39.4 between steps 20k and 120k; gap to train PPL widening — watch for the plateau.
  • ·23h 43m elapsed at step 120,140 per run log.

Conclusion

In progress. Targets on record: GPT-2 small 37.5 PPL, prior HELIX-56M 48.6 (already passed), transformer baseline 44.8.

Next Experiment

Complete the token budget; final val PPL decides the HELIX-350M go/no-go.