Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM

A user tested Qwen3.8 27B dense at Q2 and Q3 against Qwen3.6 35B-A3B MoE on a 12GB RTX 5070 Ti Laptop using llama.cpp CUDA with 4k context and q8 KV. Results show which quantizations are usable on 12GB VRAM.
1 source
More stories today
Kinetix research team joins Runway for robotics push
Runway·1 hour ago
Hugging Face blog: Refusing the right subset of a topic, not the whole topic
The blog post argues that AI safety refusal should target specific harmful subsets of a topic rather than refusing the entire topic, aiming to preserve legitimate uses. It discusses methods for fine-grained refusal control.
Hugging Face Blog·1 hour ago
New legal tech VC fund GCVC launches, backed by 50+ general counsels
GCVC, a new legal tech VC fund, launched backed by over fifty sitting general counsels from companies like Salesforce, Circle, and ElevenLabs. Wilson Sonsini is the first law firm to invest. The fund targets early-stage Seed and Series A startups.
Artificial Lawyer·1 hour ago

AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
Google DeepMind launches AlphaGenome Atlas mapping 9 billion DNA variants
AlphaGenome Atlas maps the molecular effects of all 9 billion possible single-letter DNA changes in the human genome. It includes a 1PB dataset with an AlphaGenome Variant Impact (AVI) score for each variant, accessible via a web portal, Antigravity, and the AlphaGenome interface.
Google DeepMind·1 hour ago
Runway hits $200M annual recurring revenue
Runway AI Inc.'s annual recurring revenue reached $200 million in September, driven by demand for its image- and video-generating software used in marketing and ads.
Bloomberg Technology·1 hour ago
Alphabet's CapitalG backs AI chip startup Celero at $3B value
Celero Communications raised $275 million in a funding round valuing the networking chip startup at over $3 billion, with Alphabet's CapitalG participating. The investment targets demand for improving AI data center efficiency.
Bloomberg Technology·1 hour ago
Viral 'Ox Alpha' model revealed as Zai's GLM-5.3-Flash
The anonymous model that topped coding benchmarks is Zai's GLM-5.3-Flash, a 320B-parameter hybrid with 18B active per token. It served 42 trillion tokens in six days on OpenRouter and costs just $0.09 per task on GDPval AA v2.
DeepLearning.AI·1 hour ago
Claude's new /design skill transforms design workflows
A YouTube video from Ben AI demonstrates Claude's new /design skill, claiming it changes design workflows. The video offers free resources and promotes the creator's AI courses and accelerator program.
YouTube·1 hour ago