Kimi K3 nearly matches Opus 4.8 on simple coding, fails on complex tasks

Kimi K3 scored roughly 64 of 70 on simple GitHub-style coding tasks (within a rounding error of Opus 4.8), but on adversarial trap tasks Opus failed ~8% of the time versus 36% for K3. It's priced at $3 per million input tokens and $15 per million output tokens on OpenRouter.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- DeepSeek-V4 Flash-0731 offers 80% of GPT-5.6 Luna's performance at 1/6 cost
- Chinese AI Chipmakers Poised to Gain From Beijing’s Tech Push
- Panther CEO Jack Naglieri: Using AI to build in the open is a good pattern
- Ethan Mollick: 95% of Kaggle submissions use seed 42, solutions diverge
- Shared terminal dashboard captures Claude Code, Codex, and OpenClaw sessions