Experiment Record · 2026-07-18
HELIX-200M-F35B
HELIX-200M foundation run (helix-200m-foundation-35b)
ACTIVEMEASURED
Hypothesis
A 209M-parameter H-E-R hybrid can be pretrained to competitive validation perplexity on a 40B-token Chinchilla-matched corpus entirely on one M3 Max.
Setup
28 layers, d_model 1024, MoE top-2 of 4. Corpus: 90% FineWeb-Edu 10BT + 10% WikiText-103, per-doc repeat cap 4, 13-gram decontamination. LR 3e-4 cosine, 2,000 warmup, effective batch 16 sequences, BF16.
Variables
- ◆Scale (209M vs prior 57M/127M)
- ◆Corpus size and mixture
Results
- →Step 120,140: train loss 2.824 (PPL 16.85), validation PPL 39.4.
- →Throughput ~2,900 tok/s sustained; checkpoints every 5k steps (~1.2GB NPZ).
Observations
- ·Val PPL 105.2 → 39.4 between steps 20k and 120k; gap to train PPL widening — watch for the plateau.
- ·23h 43m elapsed at step 120,140 per run log.
Conclusion
In progress. Targets on record: GPT-2 small 37.5 PPL, prior HELIX-56M 48.6 (already passed), transformer baseline 44.8.
Next Experiment
Complete the token budget; final val PPL decides the HELIX-350M go/no-go.