Experiment Record · 2026-04-10
DC1-DPO-E8
D-CSIL-1 DPO/LoRA post-training evaluation (epoch 8)
PARTIALMEASURED
Hypothesis
DPO via LoRA can improve instruction adherence on a 56M GLA model without degrading base quality.
Setup
64-prompt sampled inference report at epoch 8; DPO LoRA epochs 1–3 plus merged final checkpoint.
Variables
- ◆DPO epochs
- ◆Sampling strategy
Results
- →Prompt adherence 0.336 mean; exact-match 37.5% on factual subset (66.7% within factual-knowledge category).
- →Validation PPL 455.77 — substantially degraded base quality.
- →Degenerate responses 3.1%; repetition ratio 0.088; zero artifacts.
Observations
- ·Dialogue was the strongest category (~50% adherence); reasoning weakest (20.8%).
Conclusion
Half-credit: adherence signal exists, but the PPL cost at 56M scale was severe. Recorded as a lesson.
Next Experiment
Revisit post-training at D-CSIL-3 scale with the harness from DC2.