Experiment Record · 2026-04-24
DC2-BENCH-82K
D-CSIL-2 capability benchmark across training checkpoints
SUCCESSMEASURED
Hypothesis
Task capability (dialogue, math, code, factual) emerges measurably across pretraining checkpoints and is sensitive to decoding strategy.
Setup
Fixed benchmark suite run at steps 20k, 40k, and 82.5k with greedy and greedy+repetition-penalty-1.2 decoding.
Variables
- ◆Checkpoint step
- ◆Decoding strategy
Results
- →Overall accuracy 15.2% → 23.4% → 26.9% (rep 1.2) across steps 20k/40k/82.5k.
- →Dialogue: 40% → 64% → 80%. Train PPL 104.3 → 77.3 → 58.0.
- →Repetition penalty helped everywhere except factual (2.0% → 0%).
Observations
- ·Dialogue capability emerged earliest and strongest; factual recall stayed near floor at this scale.
Conclusion
Capability emergence is measurable and category-uneven; decoding strategy shifts results by up to 8 points overall.