EXL3 quants run dense 30B model on 12GB VRAM

A user runs Muse Glimmer 30B EXL3-SC 3.00bpw H4 fully on a 12GB VRAM GPU at 100K context with Q8_0 KV cache, achieving ~30 tok/s. Claims it's only slightly worse than the official 17GB quant.
1 source
More stories today
Anthropic adds `ant apply` for infrastructure-as-code management of Claude agents
Claude Developers·1 hour ago
OpenAI offers banked resets for Astra access delays
Tibo·1 hour agoMiniMax M3 live session on agents, coding, reasoning
GMI Cloud·1 hour agoAI Models by email
Get an email when there's news on AI Models
No news that day, no email.
Bespoke Labs' Inkling model boosts coding across the board
Tinker API·1 hour ago
Reddit thread discusses benchmarks big labs avoid
A Reddit thread on r/LocalLLaMA with 111 points and 10 comments discusses benchmarks that major AI labs allegedly don't want publicized. The post has no additional details.
r/LocalLLaMA·1 hour ago
Crusoe raises over $3B at $30B valuation
Crusoe, a cloud-computing provider and data center developer working with OpenAI, Microsoft, and Meta, raised over $3 billion in a funding round valuing it at roughly $30 billion.
Bloomberg Technology·1 hour ago

Anthropic's governance structure draws scrutiny from ValueEdge chair
Nell Minow, Chair of ValueEdge Advisors, says Anthropic is moving in the opposite direction from SpaceX on governance, but argues its structure is still unusual and potentially problematic. She says the company is combining elements of public companies and nonprofits, including a public benefit.
Bloomberg Technology·1 hour ago
Emad Mostaque: Human-led math results will be surpassed by AI
Emad Mostaque·1 hour ago