AVX2: Speed up large batch size prompt processing of IQ models

Draft PR: IQ quants are "particularly slow" on CPU at large batch sizes, the workload seen in imatrix and perplexity runs. Speed is measured via PPL on Qwen3.6-27B and Qwen3.6-35B-A3B with an EPYC 9654 using 24 threads and --chunks 50.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- LangSmith adds Preview Builds to test agent changes before production
- Exa plugin gives ChatGPT Work and Codex access to 100B+ websites
- Ramp launches its own AI model router, called Router
- Apple’s AirPods Should Avoid Meta’s Mistakes
- Docker's Tushar Jain on AI-native runtime for agent autonomy