Skip to content
D-CSIL

Experiment Record · 2026-06-23

SLM10M-EVAL

SLM-10M zero-shot leaderboard evaluation

SUCCESSMEASURED

Hypothesis

A 10M-parameter GQA transformer trained on 25B curated tokens reaches non-trivial zero-shot accuracy on standard small-model benchmarks.

Setup

Zero-shot log-likelihood evaluation on HellaSwag, ARC-Easy/Challenge, PIQA, ArithMark-2.0, plus a 10-question local MC probe.

Variables

  • Benchmark suite

Results

  • Average 32.38%: HellaSwag 26.53, ARC-E 30.47, ARC-C 25.00, PIQA 50.92, ArithMark 24.32.
  • Local MC probe: 3/10 correct via log-likelihood ranking.

Observations

  • ·Open-ended generation loops at this scale — ranking-based eval is the honest protocol, and the report says so.

Conclusion

Published with full scorecard to HuggingFace. The value is the recipe (data mixture, GQA+QK-norm at 10M), not the absolute scores.