Experiment Record · 2025-08-05
HF-CYCLE150-EVAL
HibbieFormer cycle-150 comprehensive behavioral evaluation
SUCCESSOBSERVED
Hypothesis
The cycle-150 checkpoint (step 15,000, train loss 1.0015) generates stable, domain-coherent text across diverse prompt categories.
Setup
64 prompts across 8 categories (story starters, simple questions, completions, factual, philosophical, creative, instruction, emotional/social) with optimized inference (pre-allocated tensors).
Variables
- ◆Prompt category
- ◆Sampling temperature
Results
- →64/64 completions without runtime errors, gibberish, or artifacts.
- →Optimized inference measured at 1.4× speedup.
Observations
- ·Model answers every prompt as a children's story — including factual questions. Domain lock-in is total, as expected from the corpus.
Conclusion
Generation is stable and coherent within-domain; there is no instruction-following or factual capability, and the report says so.