LaunchDevelopersAugust 1, 2026
Supabase open-sources Evals benchmark scoring Claude Code, Codex, OpenCode

Supabase open-sourced Evals, its benchmark and framework for testing AI agents on real tasks like building a schema, debugging a failed Edge Function, and fixing a broken RLS policy. The scoring harness is built around the Model Context Protocol (MCP) for running agentic coding workloads.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Temporal 5x'd AI spend, doubled revenue; CEO says link unproven
- Noam Brown highlighted as key architect behind OpenAI's o1 reasoning models
- AI beats human driver on Abu Dhabi race track
- Scoble: AI glasses will be commodities; experience is the moat
- EU mandates labels for authentic-looking AI content starting August 2