llama.cpp flag tuning boosts generation 70% on 40GB VRAM laptop

A user benchmarked llama.cpp flags on a 40GB VRAM laptop with TB4 eGPU, boosting generation from 16 to 27 t/s (+70%), prefill from 376 to 573, and usable context from 220k to full 262k. Filed a bug in llama around MTP.
1 source
Developers by email
Get an email when there's news on Developers
No news that day, no email.
More stories today
- Google's Android update adds Motion Assist, Guided Vision, and more
- Srinivas warns AI agents could self-train on on-demand GPUs
- Perplexity launches Portable Computer for NVIDIA DGX Spark
- Reddit users share AI side-hustle earnings
- Palo Alto Networks tops profit outlook on AI security demand