Experiment Record · 2026-06-23
SLM10M-EVAL
SLM-10M zero-shot leaderboard evaluation
SUCCESSMEASURED
Hypothesis
A 10M-parameter GQA transformer trained on 25B curated tokens reaches non-trivial zero-shot accuracy on standard small-model benchmarks.
Setup
Zero-shot log-likelihood evaluation on HellaSwag, ARC-Easy/Challenge, PIQA, ArithMark-2.0, plus a 10-question local MC probe.
Variables
- ◆Benchmark suite
Results
- →Average 32.38%: HellaSwag 26.53, ARC-E 30.47, ARC-C 25.00, PIQA 50.92, ArithMark 24.32.
- →Local MC probe: 3/10 correct via log-likelihood ranking.
Observations
- ·Open-ended generation loops at this scale — ranking-based eval is the honest protocol, and the report says so.
Conclusion
Published with full scorecard to HuggingFace. The value is the recipe (data mixture, GQA+QK-norm at 10M), not the absolute scores.