Whitepapers

Original research and methodology writeups. Pieces here are typically published externally first, then linked back once they're live.

How we actually evaluate an AI model before betting on it

A methodology writeup on testing language models for a real task, at real scale, rather than trusting a benchmark or a short demo — including a real case where short-excerpt testing gave a misleading answer and testing at full scale reversed it. Coming soon.