Terminal-Bench-Science evaluates AI agents on research workflows

Terminal-Bench-Science is a new benchmark for evaluating AI agents on scientific research workflows. It was announced on August 28, 2026, and has gained attention on Hacker News.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- OpenAI launches Thailand AI startup accelerator with MHESI
- Engineers increasingly say 'I don't know, Claude wrote this'
- Open source caught up because it's open
- Tencent releases AI model it claims outperforms Z.AI, Moonshot
- OpenAI hires Meta executive to lead Southeast Asia, Australia