Skip to content
D-CSIL

Experiment Record · 2025-08-05

HF-CYCLE150-EVAL

HibbieFormer cycle-150 comprehensive behavioral evaluation

SUCCESSOBSERVED

Hypothesis

The cycle-150 checkpoint (step 15,000, train loss 1.0015) generates stable, domain-coherent text across diverse prompt categories.

Setup

64 prompts across 8 categories (story starters, simple questions, completions, factual, philosophical, creative, instruction, emotional/social) with optimized inference (pre-allocated tensors).

Variables

  • Prompt category
  • Sampling temperature

Results

  • 64/64 completions without runtime errors, gibberish, or artifacts.
  • Optimized inference measured at 1.4× speedup.

Observations

  • ·Model answers every prompt as a children's story — including factual questions. Domain lock-in is total, as expected from the corpus.

Conclusion

Generation is stable and coherent within-domain; there is no instruction-following or factual capability, and the report says so.