Developer benchmarks LLMs on GPT-2 prompt engineering capabilities

The benchmark evaluates model intelligence by testing their ability to write prompt templates for GPT-2 across 395 farm-action classification tasks. The project uses a minimal test framework to score model performance on this specific proxy task.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI detectors face criticism over reliability and impact on trust
- Student uses $3 chip to run Claude Code for automated betting
- MiniMax H3 CLIP swap cuts VRAM from 15.7 GB to 4.5 GB
- Artist's AI-generated 'Found [You?]' footage project blends video and music
- Anthropic's Haiku 4.5 nears 12 months without an update