AnalysisAI ModelsJuly 31, 2026
Kimi-K3 tops Opus 4.8 in oneshot prompt evals

A LocalLLaMA user ran Kimi-K3 through 34 oneshot prompts, scoring generated HTML, screenshots, and GIFs with Sonnet 4.6, and found it beat Opus 4.8 in the evals.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- DeepSeek model so powerful it runs on two sparks
- UniFace unifies face detection, recognition, tracking, gaze in Python
- Poolside ships official Laguna S 2.1 FP8 and NVFP4 weights
- Gemini 3.6 Flash reacts to ChatGPT discoveries
- Google kills Earth AI generator after one day