Evaluating agentsacross Supabase

We evaluate model experiments across the Supabase developer journey, from building and deploying to investigating and resolving production issues, with real project context.

AgentTotal
Claude Code / Opus 5 (high)95%
Claude Code / Sonnet 5 (high)95%
Codex / GPT-5.6 sol (medium)95%
OpenCode / Kimi K386%
Codex / GPT-5.4 mini (medium)82%

Work in a realistic environment

Agents get a realistic environment to work in, with project state, context, and access to the development tools they need.

Work through a task

Agents take on a task from somewhere along the Supabase developer journey, from building and deploying to investigating and resolving issues.

Results are evaluated

Scoring draws on SQL checks, client calls made as real users, and the files the agent creates. When broader assessment is needed, an LLM judge reviews the result.