Skip to content
D-CSIL

Experiment Record · 2026-04-24

DC2-BENCH-82K

D-CSIL-2 capability benchmark across training checkpoints

SUCCESSMEASURED

Hypothesis

Task capability (dialogue, math, code, factual) emerges measurably across pretraining checkpoints and is sensitive to decoding strategy.

Setup

Fixed benchmark suite run at steps 20k, 40k, and 82.5k with greedy and greedy+repetition-penalty-1.2 decoding.

Variables

  • Checkpoint step
  • Decoding strategy

Results

  • Overall accuracy 15.2% → 23.4% → 26.9% (rep 1.2) across steps 20k/40k/82.5k.
  • Dialogue: 40% → 64% → 80%. Train PPL 104.3 → 77.3 → 58.0.
  • Repetition penalty helped everywhere except factual (2.0% → 0%).

Observations

  • ·Dialogue capability emerged earliest and strongest; factual recall stayed near floor at this scale.

Conclusion

Capability emergence is measurable and category-uneven; decoding strategy shifts results by up to 8 points overall.