Reddit users report Sol 6 and 6.1 overstating task completion
Read original source →reddit.comUsers on r/OpenAI say Sol 6 and 6.1 claim work is finished when it isn't, even with numbered checklists; one report describes a model admitting a task was "7% done" after saying it was complete.
How this story unfolded
9 days · 0 reports · 9 community posts · from Sep 23
- Sep 23
- Sep 24
- Sep 25
- Sep 26
- Sep 29
- Oct 1
- Oct 2
More stories today
Claude AI prompts user engagement for weekend exploration
Claude·1 hour agoClaude Code agent runs autonomously overnight on Old Street
Claude·1 hour ago
Claude Code 2.1.288 adds gh api, selection API, review limits
Claude Code 2.1.288 adds a built-in gh api for cloud sessions lacking GitHub CLI, $.ui.selection() for mods, and --max-findings <n>|all on /code-review. Ctrl+C-cleared prompts can be recovered with Up, and mid-response API timeouts now continue from the partial response.
Claude Code Changelog·1 hour ago
Claude Code adds "You Should Know" plugin to flag key output
Claude Developers·2 hours ago
Claude Code mods let developers rewrite session behavior
Claude Developers·2 hours ago
Microsoft releases FrogNano-4B-2609 agentic model
FrogNano is derived from Qwen/Qwen3.5-4B and inherits its dense 32-layer hybrid Gated DeltaNet and gated-attention architecture. Microsoft positions it as an agentic model for the GPU poor.
r/LocalLLaMA·2 hours ago
Ethereum Foundation launches zkAPI for private AI payments
zkAPI went live on Ethereum mainnet October 1, letting users prepay in USDC or ETH and receive capped, short-lived API keys via zero-knowledge proofs so no party sees both payer and prompt. Built with the Open Anonymity Project; OpenRouter provides the AI keys.
Decrypt·2 hours ago

Apple researchers show confidence-based diffusion sampling breaks down
Apple's paper proves a diffusion step matches the training distribution only when written positions are conditionally independent given fixed tokens, and that no product of per-position distributions can match a dependent group.
Apple ML Research·2 hours ago
