Kimi-K3 beats Opus 4.8 in user oneshot eval

A user ran 34 oneshot prompts through Kimi-K3, scoring the generated HTML, screenshots, and GIFs with Sonnet 4.6, finding it better than Opus 4.8. Results were posted to r/LocalLLaMA with a full breakdown on oneshotlm.com.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Law firms urged to reclaim AI sovereignty from frontier labs
- Thread ping-back graph as primitive multiagent AGI future
- OpenAI asks judge to toss Apple's trade secret lawsuit
- Five-parallel-agent AI sales team runs inside Claude Code
- Deep Eye AI pen-testing tool scans 45+ vulnerability types