NVIDIA AVO scores 100% on ARC-AGI-3, solving all 183 levels

NVIDIA's general-purpose coding agent AVO completed all 183 levels across all 25 public ARC-AGI-3 environments with no instructions, explicit rules, or stated goals. TechCrunch reports the same harness lifted Claude Opus 5 from 30% to 100% on the benchmark.
How this story unfolded
1 day · 3 reports · 9 community posts · from Aug 21
- Aug 21
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agentsdeveloper.nvidia.com
Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.thenewstack.io
Nvidia just showed that the harness, not the AI model, is now the real herotechcrunch.com
- Aug 22
More stories today
Relight node brings 3D lighting studio to MiniMax H3 in ComfyUI
A community-built ComfyUI node adds a relighting studio for MiniMax H3, letting users place up to three lights on a 3D dome around an image. Each light's type, intensity, and color are configurable, along with background and atmosphere settings.
r/ComfyUI·1 hour ago
Archival photos and pose control models recreate a first meeting in documentary Love…
Google DeepMind·2 hours ago
Meta fixes Meta AI prompts after invasive personal questions
Meta AI suggested "Who is the child passenger?" under a user's video, then surfaced her daughters' ages, home location, and a photo she says she deleted years ago. Spokesperson Dina El-Kassaby said the company "missed the mark" and that the prompt feature is fixed.
The Verge·2 hours ago

Reddit users discuss what still runs on 8GB VRAM
A LocalLLaMA thread asks whether small models remain viable on 8GB cards like the RTX 2050 for office work, embeddings, reranking, and chat. The poster says small models feel abandoned and asks for recent good ones.
r/LocalLLaMA·2 hours agoApple's SimpleDesign jointly models protein sequence and structure
Apple researchers trained SimpleDesign on over 2M sequence-structure pairs with a single-stage end-to-end objective, combining discrete cross-entropy for sequences and a regression objective for structures. The model skips the usual multi-stage autoencoder-plus-latent-generative pipeline and reports competitive results on co-design and unconditional sequence/structure generation benchmarks.
Apple ML Research·2 hours ago

Apple introduces CapQuiz, a reference-free video caption benchmark
CapQuiz scores captions by how well they answer human-verified multiple-choice questions, spanning 10 question types across 24 video domains. The companion CapF1 metric combines CapP (factuality) and CapR (coverage), and correlates better with human judgments than existing metrics.
Apple ML Research·2 hours ago

Apple researchers introduce DiscoSign for discourse-aware sign language translation
DiscoSign is an LLM-based framework for text-to-ASL gloss translation that handles spatial coreference, Question-Answer Clauses, and concept-gloss consistency instead of sentence-level translation. Apple says it is the first systematic framework for discourse-level text-to-sign gloss translation, with new evaluation metrics for each dimension.
Apple ML Research·2 hours ago

Meta holds European credit investor roadshow to fund AI buildout
Meta Platforms has been meeting bond investors in Europe for at least a week in a non-deal roadshow, as hyperscalers tap different global debt markets to fund the AI boom.
Bloomberg Technology·2 hours ago
