AnalysisAI ModelsSeptember 13, 2026

Study: AI agents can't yet do open-ended AI research

A Princeton-led multi-institution study used "shadow evaluation" — having agents answer research questions from high-quality unpublished papers — and found agents could handle engineering tasks but lacked the judgment and creativity to produce work at top ML conference caliber. Claude Opus 4.8 was among the models tested.

People · Peter Kirgis, Sayash Kapoor

1 source

More stories today

Open the live feed