Perplexity introduces Web Search Benchmarks for grounding AI agents

Suite covers 6 benchmarks backed by 2,464,345 task evaluations, last run Aug 14, 2026. Includes WANDR, a 500-task benchmark for wide and deep research agents; every score links to configuration, costs and telemetry.
How this story unfolded
2 days · 0 reports · 4 community posts · from Aug 14
- Aug 14
- Aug 17
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills