1.56TB MoE model tested on 6GB laptop yields extremely slow inference
A Reddit user tested a 1.56TB Mixture-of-Experts model (96 shards, 93 layers, 896 experts/layer, MXFP4) on a 6GB RTX 4050 laptop GPU, reporting extremely slow inference speed due to memory constraints.
1 source
More stories today
Reddit users probe motion-context degradation in video diffusion
A r/StableDiffusion thread examines motion-context degradation, a quality issue in video generation workflows. The poster notes the H3-director node claims a refine pass can fix it, but calls that refine a black box when used with low-level motion-context nodes.
r/StableDiffusion·2 hours ago
Agnes-3.0-Flash 33B multimodal model posts AA score of 36
Agnes-AI's Agnes-3.0-Flash is a 33B hybrid-attention decoder with a 262,144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding. Three of every four layers use a gated delta rule with per-layer state.
r/LocalLLaMA·3 hours ago
Retrospectively Reverse-Engineering Apple's Neural Engine
Blog post maps the M1 ANE's compute, datapath, scheduler, memory and execution model, three years after the author abandoned the driver project. The 16 compute cores target dense CNN tensor reductions; M5 (2025) folded ANE cores into GPU cores.
Hacker News·3 hours agoDrug-discovery AI prioritizes likely-to-fail tests, founder says
Robert Scoble·3 hours ago
Reddit user posts Krea2 refusal-behavior experiment
A r/StableDiffusion user describes finding an "extremely CLEAN approach" to Krea2's refusal behavior, framing it as a follow-up to earlier discussion of the model's refusals.
r/StableDiffusion·3 hours ago
Reddit user asks llama.cpp maintainers for hot expert reload on GPU
A r/LocalLLaMA post requests hot expert reload on GPU for llama.cpp, claiming decode-speed gains on MoE models with few active parameters. It cites Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash and GLM 5.3 Flash, and says 2x 3090 cards would scale further.
r/LocalLLaMA·4 hours agoQwen3.8-27B-Q4 runs at ~170k context on a 32 GB GPU without KV cache quantization
A LocalLLaMA user reports fitting roughly 170k tokens of Qwen3.8-27B-UD-Q4_K_XL on a 32 GB GPU with MTP and mmproj enabled, avoiding KV cache quantization entirely. They say even q8_0 KV cache quantization is noticeably worse, and Qwen3.8 burns through the context at xhigh effort.
r/LocalLLaMA·4 hours ago
Codex usage limits reset again
Kimmonismus·4 hours ago