smolbenchmark ranks 8GB-fit models by decode speed and tokens per joule

A Reddit user released smolbenchmark, a leaderboard for models that fit in 8GB of memory, ranked by decode speed, tokens per joule, and heat. It targets tablets and other low-power local hardware rather than GPU servers.
1 source
More stories today
Altman: OpenAI IPO would be 'ill-advised' in 2026
Altman told Fortune's Alyson Shontell that going public now would be "ill-advised" given safety concerns, saying "not 2026, yeah. We've got a lot of stuff to do." OpenAI has already filed confidentially for an IPO.
TechCrunch·1 hour ago

Quartermaster plugin manages Claude Code setup and rules
Quartermaster installs codebase-mapper, builds the project map, adds starter live-rules, and picks stack plugins and permissions, with every install, file edit, and settings change awaiting user approval. It then watches sessions, proposes one change at a time, and rolls back changes that didn't help.
r/ClaudeAI·2 hours ago
Real-SWE benchmark tests frontier AI models on private enterprise codebases
Real-SWE evaluates model-and-harness combinations on tasks drawn from private production codebases licensed from real companies, covering billing, tax calculation, and customer migrations. Tasks are not public, so agents cannot rely on internet-available solutions.
r/LocalLLaMA·2 hours agoThariq: Claude Code would have looked like AGI in 2018
Thariq·2 hours agoClaude Code 2.1.270 fixes git permission prompts
Patch 2.1.270 fixes a regression from 2.1.269 where read-only git commands in Bash unexpectedly asked for permission after long-running sessions. The prior 2.1.269 release added `claude plugin eval` for scored, reproducible plugin eval suites with JSON and HTML reports.
Claude Code Releases·2 hours agoMiniMax H3 video model demoed in standard T2V workflow
A Reddit r/StableDiffusion post shows output from MiniMax H3 generated with the standard text-to-video workflow. No benchmark scores, pricing, or release details were included in the post.
r/StableDiffusion·3 hours ago
Fly Language Model wires fruit fly connectome into frozen 1.2B LLM
FLM couples the complete MaleCNS v1.0 fruit fly connectome to a frozen LiquidAI LFM2.5-1.2B-Instruct backbone via an architecture called GPF. The developer's own controls show the connectome wiring does not improve results.
MarkTechPost·3 hours ago

Conversation explores why robotics is about to take off
Robert Scoble·3 hours ago