llama.cpp PR 26291 speeds up 300GB RPC model loads 300%

On build b10173, RPC loading of a 300GB model dropped from 4m54s to 1m38s with GGML_RPC_LOAD_THREADS=12 (RTX 4060 Ti, DDR4/DDR5). PR 26291 is close to ready; needs a docs change for the new variable.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Paper examines the limitations of current AI evaluation methods
- OnlyHuman filter list removes AI-generated SEO spam from search results
- Qwen tokenizes 330-line code into 1,609 tokens; Gemma needs 4,258
- LifeOS: open-source AI harness for personal growth and work
- MINIMAX video drops Indiana Jones into Mortal Kombat