Research that has to survive contact with a model in production.
Most of what we publish is written because we needed the answer ourselves. Below is the shape of the programme: what we study, how the results get used, and where the current work is stuck.
- Programme areas
- 7From circuits to labour-market measurement.
- Papers published
- 140+Peer-reviewed and preprint, with code where it exists.
- Eval suites
- 60+Run before release, and re-run when a threshold moves.
- Disclosures
- 100%Every released model ships with a public system card.
A steerable circuit for answer withholding
We localise a mechanism that suppresses an answer under uncertainty, then show that intervening on it changes behaviour on held-out tasks — and report where the intervention fails.
Drift in multi-day agentic tasks
Instruction adherence degrades measurably past a horizon that our tooling had not previously instrumented. We describe the failure modes and the mitigations that reduced them.
Task-level exposure in an AI-using labour market
A methodology note and first release of our recurring index: which tasks are being delegated, with regional breakdowns and disclosed sample sizes.
Negative results in preference optimisation
Four attempts to reduce sycophancy through preference tuning, all of which moved the measured behaviour without moving the underlying tendency. Published because the field keeps repeating them.
Research is a group activity, and we would like more people in the room.
We fund external researchers, host visiting positions, and publish our evaluation tooling. If you work on something adjacent and want access, the fastest path is to tell us what you would do with it.