llama.cpp flag tuning boosted generation 70% on 40GB VRAM eGPU laptop

Benchmarking llama.cpp flags on a 40GB VRAM laptop with TB4 eGPU raised generation speed from 16 to 27 t/s (+70%), prefill from 376 to 573, and usable context to the full 262k. User also filed a bug in llama around MTP; final command included Qwen3.8-27B-UD-Q6_K_XL.gguf with -c 262144.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- LangSmith adds Preview Builds to test agent changes before production
- Exa plugin gives ChatGPT Work and Codex access to 100B+ websites
- Ramp launches its own AI model router, called Router
- Apple’s AirPods Should Avoid Meta’s Mistakes
- Docker's Tushar Jain on AI-native runtime for agent autonomy