Skip to content
D-CSIL

Experiment Record · 2026-04-10

DC1-DPO-E8

D-CSIL-1 DPO/LoRA post-training evaluation (epoch 8)

PARTIALMEASURED

Hypothesis

DPO via LoRA can improve instruction adherence on a 56M GLA model without degrading base quality.

Setup

64-prompt sampled inference report at epoch 8; DPO LoRA epochs 1–3 plus merged final checkpoint.

Variables

  • DPO epochs
  • Sampling strategy

Results

  • Prompt adherence 0.336 mean; exact-match 37.5% on factual subset (66.7% within factual-knowledge category).
  • Validation PPL 455.77 — substantially degraded base quality.
  • Degenerate responses 3.1%; repetition ratio 0.088; zero artifacts.

Observations

  • ·Dialogue was the strongest category (~50% adherence); reasoning weakest (20.8%).

Conclusion

Half-credit: adherence signal exists, but the PPL cost at 56M scale was severe. Recorded as a lesson.

Next Experiment

Revisit post-training at D-CSIL-3 scale with the harness from DC2.