NVIDIA MPS on EC2 cuts ASR inference costs by 75%
Read original source →aws.amazon.com
AWS, NVIDIA, and Heidi detail how NVIDIA MPS on Amazon EC2 reduces automatic speech recognition (ASR) inference costs by 75% while meeting strict latency requirements. The post targets low GPU utilization per request.
1 source
More stories today
Local voice assistant runs Qwen3.5 4B with tool calling, no GPU
A Reddit builder assembled a local voice assistant on Qwen3.5 4B with skills/tool calling, using Voxtral for speech-to-text and Pocket for text-to-speech, running without a dedicated GPU.
r/LocalLLaMA·38 minutes ago
Sakana AI paper proposes MASS for recursive self-improvement
David Ha (hardmaru)·41 minutes agoReddit user calculates $20 Claude plan yields ~$1,300 in API usage
A Reddit user measured API costs across all session logs after driving a new $20 subscription to near 100% of its 5-hour limit, then extrapolated. The same method put the $200 tier at roughly $8,000 of usage, though the poster calls that earlier figure imprecise.
r/ClaudeAI·58 minutes ago
Amazon AGI Lab demos perception agents for visual verification
Emile Baizel and Shruti Arora of Amazon AGI Lab present two perception agent primitives: visual annotation, where a person selects an element, and a browser verification pass that checks whether a coding agent's change actually appeared.
YouTube·1 hour ago
Snyk's Javier Garza demos AI bill of materials workflow
Talk walks through a CLI scan that turns a repository into a searchable inventory of models, datasets and agents, then surfaces it in a visual dashboard. The inventory is framed as the anchor for assessing AI supply-chain risk.
YouTube·1 hour ago
FriendliAI engineer: inference provider choice decides open-weight model speed
FriendliAI founding engineer Yunmo Koo argues open-weight models like GLM 5.2 and MiniMax M3 now compete with the best proprietary models, and shows GLM 5.2 and Claude Opus 4.8 building similar outputs at different speeds.
YouTube·1 hour ago
Karpathy: LLM capability understanding gap is widening
Andrej Karpathy·2 hours agoAuroraIMG-6M: 6M-parameter text-to-image model released
AuroraIMG-6M is a 6M-parameter text-to-image model that produces recognizable images at 64x64 resolution. The model card includes samples, architecture and training details, and a quick start.
r/StableDiffusion·2 hours ago