Sage Attention cuts Minimax generation time to under 2 minutes
Read original source →reddit.comA Reddit user reports rendering 5-second Minimax clips drops from 5 minutes to under 2 with Sage Attention, hitting 4.32 s/it on a 5070 Ti (64GB RAM, int8 convrot model). The setup uses KJ Nodes' Patch Sage Attention setting with no launch argument.
1 source
More stories today
Show HN: Jev Plays Pokémon Red
A hobby project runs an AI agent called Jev through Pokémon Red, streaming every decision and its odds in a live feed. The creator says Jev decides fast but not fast enough to play Doom, and credits a guide for knowing where to go next.
Hacker News·55 minutes ago
Medicare's WISeR AI prior-authorization pilot drew errors, denials
Documents obtained by the EFF via FOIA show the WISeR pilot, launched in January, was rushed and error-ridden, with one vendor warning CMS a working product by launch was unrealistic. It runs in New Jersey, Ohio, Oklahoma, Texas, Arizona and Washington through 2031.
Ars Technica·1 hour ago

Show HN: Agentic CUDA Kernel Optimizer
A C++ CUDA test harness hands kernels to LangGraph-based AI agents that run them, collect benchmarks, and profile via Nsight. Built by bertaye as a side project to learn LangGraph.
Hacker News·1 hour agoContrastive-LM releases CLM-8B, an open model that scores agent actions
CLM-8B is the first open model in a new class called Contrastive Language Models: it does not generate text, but scores candidate actions against the current state and returns probabilities. Tests show it runs up to 9x faster than Jev, the proprietary System One model it targets as its main baseline.
VentureBeat·1 hour ago

Anthropic engineer tests when to change model effort settings
Thariq·1 hour agoReddit user runs break-even math on buying vs renting an 8-GPU H200 server
A LocalLLaMA poster worked through the rent-vs-buy decision for an 8-GPU HGX H200 server, saying the break-even point landed somewhere they did not expect. The post shares the full calculation and invites corrections.
r/LocalLLaMA·1 hour agoGoogle's PageBreak agent finds 500+ bugs in its own apps
Google's Product Security team disclosed PageBreak, an internal Gemini-based agent that has uncovered more than 500 XSS vulnerabilities in Google's first-party web apps. It only reports a bug after confirming it with a working exploit against a live environment, giving it a near-zero false-positive rate.
Decrypt·2 hours ago

Claude Code adds wrap-up allowance at 5-hour limit
When Claude Code hits a plan's five-hour usage limit mid-response, it now keeps working briefly to reach a graceful stopping point instead of cutting off mid-edit. The extra time comes from a small, fixed allowance drawn from the user's weekly limit.
Claude Developers·2 hours ago