AnalysisDevelopersJuly 7, 2026
llama.cpp adds support for Hy3 model via PR

llama.cpp has a pull request adding support for the Hy3 model, with GGUF quantizations available. Early testing shows coherent output at ~10-11 t/s on a 5090 with Q2_K quantization.
1 source
More stories today
- Alibaba reportedly tests standalone Qwen Office product
- Google shares Gemini usage data: multimodal AI useful for manual labor
- South Korea outlines AI future with NVIDIA at AI Summit
- PicoAgents framework for multi-agent systems released
- HuggingHack local HuggingFace tool moves to GitHub