AnalysisAI ModelsJuly 20, 2026
CRAFT and related rubric methods for LLM evaluation in new papers
CRAFT provides a rubric-based framework to diagnose weak LLM capabilities and generate targeted fine-tuning data. Other papers explore evolving rubrics from a single query, cross-rubric generalization in essay scoring, and biases in LLM-as-judge settings. These works aim to improve the reliability and granularity of LLM evaluation.
1 source
More stories today
- AI tool locks ground line on drone videos for real estate
- Testing OpenClaw with 12 subagents for automated QA
- Codex's Sol shows improved intent understanding in QA
- Opus 5 generates painterly world with wind-reactive grass in HTML
- User shares Analog Horror Krea 2 LoRA for Stable Diffusion