llama.cpp PR 26291 speeds up RPC model loading from 5min to 1:30

PR 26291 cuts llama.cpp RPC model loading from 4min54sec to 1min38sec (~300% faster) using the new GGML_RPC_LOAD_THREADS=12 variable, bench-marked with a 300GB model on RTX 4060 Ti systems with DDR4 and DDR5 memory.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI detectors face criticism over reliability and impact on trust
- Student uses $3 chip to run Claude Code for automated betting
- MiniMax H3 CLIP swap cuts VRAM from 15.7 GB to 4.5 GB
- Artist's AI-generated 'Found [You?]' footage project blends video and music
- Anthropic's Haiku 4.5 nears 12 months without an update