DeepSeek-V4 paper: LLM retries bias results toward shorter text

Analyzing 100,000 poems generated with DeepSeek-V4-Flash for $6.06, the author found 73.2% were haikus (15 words) vs 11.5% novel chapters (503 words); simulated 10% failures hit long requests four times harder (31% vs 7.4%). The paper warns that retrying interrupted long requests often yields shorter replacements, skewing benchmarks.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills