Compact models
We study when one smaller model can preserve useful behavior across multiple tasks without hiding specialists behind a router.
We study when smaller models are genuinely better deployment choices—and build the systems needed to find out.
We study when one smaller model can preserve useful behavior across multiple tasks without hiding specialists behind a router.
We make method selection, provenance, manifests, and portable artifacts explicit enough to inspect and reproduce.
We combine quality, uncertainty, out-of-distribution behavior, systems measurements, and deployment economics in one release decision.
We generate latent finance worlds before rendering prompts so accounting invariants and held-out regimes remain machine-checkable.
A completed training job is not evidence that a model should ship. Our experiments preserve losing arms, record the first failed gate, and keep quality claims bounded to frozen data and declared assumptions.
Read the experiment