Work in a realistic environment
Agents get a realistic environment to work in, with project state, context, and access to the development tools they need.
We evaluate model experiments across the Supabase developer journey, from building and deploying to investigating and resolving production issues, with real project context.
| Agent | Total | ||||
|---|---|---|---|---|---|
| 95% | |||||
| 95% | |||||
| 95% | |||||
| 86% | |||||
| 82% |
Agents get a realistic environment to work in, with project state, context, and access to the development tools they need.
Agents take on a task from somewhere along the Supabase developer journey, from building and deploying to investigating and resolving issues.
Scoring draws on SQL checks, client calls made as real users, and the files the agent creates. When broader assessment is needed, an LLM judge reviews the result.