TL;DR
D-CSIL-3 pretrains a 235.7M selective-state model on a seven-stage, 4.747B-token corpus — FineWeb-Edu, Wikipedia, SlimPajama, OpenWebText, WikiText-103, then SFT — with every byte staged locally. Measured throughput: 1,270 tok/s. Estimated wall-clock: 43.26 days.
The pipeline
Seven stages, each a local .npy cache: 1.5B FineWeb-Edu, 500M Wikipedia, 2.0B SlimPajama (two stages), 500M OpenWebText, 119M WikiText-103, and a 128M SFT mix at the end. Chinchilla-matched at ~20 tokens per parameter.
Why local staging matters
Streaming from the network mid-run couples training stability to connectivity. Staging everything into memory-mapped caches on the SuperDock makes the 43-day estimate a function of compute alone — and makes the run resumable from any of the 32,552-step checkpoints without re-downloading anything.