AI ProductAI research2026
Evaluation moved from a shared document into the build pipeline.
Frontier model lab

- full suite runtime
- 38 minfull suite runtime
- regressions caught before release in quarter one
- 6regressions caught before release in quarter one
- number the org now argues about
- 1number the org now argues about
The situation.
Model quality was argued in review meetings. There was no shared harness, so every team measured a different thing and regressions reached customers.
What we did
- Built one harness with versioned datasets and reproducible runs
- Added adversarial and long-context suites written with the research team
- Gated deploys on regression thresholds set per capability
- Exposed results as a dashboard the whole company could read