Qwen Code releases full SWE-bench Verified + Terminal-Bench 2.0 run
Qwen Code's DSW EAS release runs full end-to-end benchmark: SWE-bench Verified 500 cases first, then Terminal-Bench 2.0 89 cases after successful SWE publication. Benchmark-Qwen-Ref v0.21.12.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Google AI's EnvHarness turns static agent benchmarks into adaptive training worlds
- AI agents that pass authentication can still drift, expose data, or get memory-poisoned
- Local video watermark remover released on CivitAI
- Video walks through implementing Kimi K3 from scratch in PyTorch
- SeedVR2 TensorRT Studio offers free open-source video upscaling