Uncensored LLMs more optimistic than base models

A new arXiv paper shows that "uncensored" LLMs are measurably more optimistic in their outputs compared to their original base models. The finding suggests that removing safety constraints alters model behavior beyond simple refusal patterns.
2 sources
More stories today
OpenAI acknowledges and partially fixes GPT-6 Astra downgrade
A Reddit user documented a real quality downgrade in GPT-6 Astra, which OpenAI acknowledged and partially fixed. The poster argues downgrades happen more often than users notice and are anti-consumer.
r/ChatGPT·48 minutes ago
DeepSeek Fails the Rubik's Cube Test
Matthew Berman's video examines DeepSeek's performance on a Rubik's Cube test, where the model fails the task.
YouTube·1 hour ago
ARC-AGI-4 to target autonomous open-ended invention
ARC Prize says ARC-AGI-4 will benchmark autonomous open-ended innovation, keeping the effort open-source as a shared research target. The announcement notes humans still significantly outperform AI at open-ended tasks.
r/Singularity·2 hours agosmolbenchmark ranks 8GB-fit models by decode speed and tokens per joule
A Reddit user released smolbenchmark, a leaderboard for models that fit in 8GB of memory, ranked by decode speed, tokens per joule, and heat. It targets tablets and other low-power local hardware rather than GPU servers.
r/LocalLLaMA·2 hours ago
Altman: OpenAI IPO would be 'ill-advised' in 2026
Altman told Fortune's Alyson Shontell that going public now would be "ill-advised" given safety concerns, saying "not 2026, yeah. We've got a lot of stuff to do." OpenAI has already filed confidentially for an IPO.
TechCrunch·3 hours ago

Dwarkesh Patel video examines why AI models sound alike
Dwarkesh Patel's video explores the convergence of AI model outputs, arguing that models from different labs increasingly produce similar responses. No specific models, benchmarks, or figures are cited in the available source material.
YouTube·3 hours ago
Quartermaster plugin manages Claude Code setup and rules
Quartermaster installs codebase-mapper, builds the project map, adds starter live-rules, and picks stack plugins and permissions, with every install, file edit, and settings change awaiting user approval. It then watches sessions, proposes one change at a time, and rolls back changes that didn't help.
r/ClaudeAI·3 hours ago
Real-SWE benchmark tests frontier AI models on private enterprise codebases
Real-SWE evaluates model-and-harness combinations on tasks drawn from private production codebases licensed from real companies, covering billing, tax calculation, and customer migrations. Tasks are not public, so agents cannot rely on internet-available solutions.
r/LocalLLaMA·3 hours ago