Model Registry · HELIX family
HELIX-56M
H-E-R stack · HER×4 (12 layers) · d_model 512 · 8 heads · rotation-HSL (post-bakeoff) · 4 experts
Benchmarks
| Metric | Value | Notes |
|---|---|---|
| WT-103 test PPL (token / word) | 27.29 / 48.59 | docs/RESULTS_56M_WT103.md |
| Matched transformer control (57.36M, same protocol) | 25.10 / 44.82 | Colab A100, same data/steps — transformer leads by 0.084 nats (~8–9%) |
| Tokens/sec (train) | 8,000–9,000 | operator-reported |
| Transformer comparison run | Colab | HELIX_Transformer_Comparison_56m.ipynb on Drive, 2026-06 |
Known Limitations
- −The controlled, matched-protocol comparison shows HELIX trailing a same-size dense transformer by ~8–9% perplexity (0.084 nats test) — the first measured point of the architecture's O(L)-inference tax, later confirmed larger (15.0%) at 127M.
- −Loss-curve anchor points are approximate (full step log not in registry yet); PPL and throughput are as recorded/reported.
- −Base LM only; no instruction tuning on this card (see the chat-adapt/SFT lineage in the HELIX project page).
- −Not publicly released — served through a custom, Apple Silicon-only inference engine behind an invite-only gate. There is no public checkpoint, API, or download.