Flint: compressing reasoning without breaking it
Small reasoning models spend most of their thinking tokens narrating, not computing. Flint fine-tunes a model on its own section-aware-compressed traces and gets better accuracy at 2-3x fewer tokens — this pre-registered study tests whether the recipe transfers to a new model family, scales with data, and can shed its behavioral costs.
Running init · run 1 · protocol rev 1
- ns-flint initialized: Flint: compressing reasoning without breaking it