AnalysisAI ModelsAugust 18, 2026

Paper: Pass@k misapplied in AI coding agent benchmarks

arXiv paper argues current pass@k implementations misapply the Chen et al. (2021) estimator by setting n to the number of unit tests in a single submission rather than independent rollout attempts, skewing reliability and security measurements of agentic code generation.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed