Skip to content
D-CSIL

Model Registry · HELIX family

HELIX-200M

Hybrid H-E-R stack · HER×5 (15 layers) · d_model 1024 · 16 heads · rotation-HSL (per-bank ω, r_max 256) + ElasticAttentionGate + ResonantExpertRouting (MoE top-2 of 4, grouped-GEMM dispatch)

STATUS TRAININGVAL PPL 27.5PROJECT: HELIX
HELIX-200M — TRAINING LOSS (DEMO DATA)120,140 STEPS · MIN 2.82
2.84.45.97.48.90k30k60k90k120k

Benchmarks

MetricValueNotes
Train PPL @ step 120,140 (trunk)16.85run log, pre-branch
Trainer val PPL @ step 244,000 (final)28.45up from 37.32 pre-decay @ step 219,000 — WSD decay alone bought ~9 PPL
Held-out val PPL, 10M-token slice (final)27.50val200m_first10M, 9,999,360 scored tokens; up from 35.54 pre-decay
Open-SLM-style benchmark avg (limit-100)41.1%PIQA 68 / ARC-E 35 / HellaSwag 34 / ArithMark 32 / ARC-C 26 — best of any HELIX generation
Facttrack rank-114/19up from 10/19 pre-decay; 18/19 top-5
Tokens/sec (train)~2,900trunk phase, sustained BF16

Known Limitations

  • The HELIX-vs-matched-transformer perplexity tax (8–9% at 56M, 15.0% at 127M, both matched protocol) has never been re-measured at 209M — the open question that decides whether the architecture's scaling case holds.
  • The 10M-token held-out validation slice has no confirmed n-gram decontamination proof against the 40B training stream; the 27.50 figure could be optimistically biased by leakage — unresolved.
  • DPO step 100 was selected as the checkpoint on 2026-07-24, but a competing candidate (creator-unlikelihood-v3) was fully benchmarked afterward on the model's most visible residual failure and never adjudicated head-to-head.
  • Persistent unsolved weaknesses across every HELIX generation: arbitrary-symbol recall (Au/gold, rank 1746–9438), digit arithmetic (~30%), creator hallucination, empty multi-turn outputs.
  • A further SFT round is in progress now — the DPO-100 selection above is not necessarily the final checkpoint.