Zai releases GLM-5.2 open-weight model
GLM-5.2 improves coding and agentic task performance with enhanced long-horizon capabilities. The open-weight model is available on Hugging Face with free inference and an NVIDIA NVFP4 quantization.
Daily AI Briefing
The 49 stories that mattered in AI, curated and summarized from dozens of sources by AIBriefs.
GLM-5.2 improves coding and agentic task performance with enhanced long-horizon capabilities. The open-weight model is available on Hugging Face with free inference and an NVIDIA NVFP4 quantization.
DeepSeek is reportedly developing its own AI chip, according to sources familiar with the matter. The effort would make the AI lab less reliant on imported semiconductors.
The Visual Concept Inference from Sets (VCIS) method enables VLMs to infer shared concepts from example images and apply them to new inputs, overcoming limitations in reasoning from purely visual context. Apple's approach uses a novel architecture that learns concept representations directly from image sets without textual descriptions.
VulnHunter is an agentic AI tool that scans source code for exploitable vulnerabilities, maps attack paths, and proposes fixes before code ships. It was open-sourced by Capital One and built internally.
DeepSeek told prospective investors it is suspending its second fundraising round days after comments attributed to founder Liang Wenfeng about US-China AI competition went viral. A leaked investor meeting transcript shows the company prioritizes AGI research over consumer products and near-term revenue.
OpenAI CFO Sarah Friar introduced a practical AI scorecard to measure ROI via four metrics: useful work, cost per successful task, dependability, and return on compute. The framework aims to help organizations evaluate AI investments concretely.
NVIDIA published a blog post detailing RoboLab, its simulation benchmarking platform for evaluating robot foundation models. The guide covers real-world deployment challenges and best practices for testing general-purpose robot policies.
The ACT-2 model achieves a 99% success rate on zero-shot laundry tasks, handling various garments without prior training. The system represents a significant advance in robotic manipulation for household chores.
Altman took responsibility for missing user and revenue targets. OpenAI projects losses up to $14 billion for 2026 despite 800 million weekly active users.
OpenAI proposes a 'reverse federalism' approach, where state-level AI laws inform a unified national framework for safe and democratic AI governance. The blog post outlines principles for balancing innovation with public safety across jurisdictions.
Altman criticized Musk for pitching short-term space data centers to investors, a view most experts already share. Musk responded that his company plans to start flying them next year.
Prosper AI raised $30 million from Andreessen Horowitz to scale its AI platform that manages the entire patient journey. The platform is already working across 150,000 healthcare encounters.
Shashwat Goel presents methods for using language models to forecast world events, covering leakage-free retrieval, RL training, and the FutureSim system. The talk also evaluates frontier models on forecasting benchmarks.
Apple ML Research proposes a method to reduce computational costs in machine unlearning by leveraging low influence points. Unlike existing methods that treat all forget-set points equally, this approach differentiates based on influence, potentially lowering compute requirements.
Claude Blog publishes a guide for CISOs on managing agentic AI risks. The post argues that eliminating all risk is not feasible and provides strategies for security leaders.
Google Cloud incorporates key Wiz capabilities into an agentic defense platform to automate threat detection and remediation against AI-driven attacks. The platform aims to outpace attackers by using autonomous agents for security operations.
Traceforce provides visibility and control over AI apps like ChatGPT and Claude across all devices, including laptops, sandboxes, and VMs. The YC S26 startup was founded by Xia and Varun.
MUGEN is a comprehensive benchmark for evaluating multi-audio understanding in large audio-language models across speech, general audio, and music. Experiments reveal current LALMs still struggle with multi-audio tasks.
The article surveys techniques for adjusting how much reasoning a model performs, building on OpenAI's o1 and DeepSeek-R1. It explains the reinforcement learning with verifiable rewards (RLVR) approach used to train such reasoning models. Sebastian Raschka also highlights open questions in balancing reasoning depth and cost.
The reference implementation replaces RAG and embeddings with continuous LLM consolidation, treating memory as a running process rather than context dropped after each query. It runs on Gemini 3.1 Flash-Lite and is available in Google Cloud's generative-ai repository.
Y Combinator President Garry Tan discusses how AI-native companies achieve 400x productivity leverage, allowing lean teams to operate at scale. He and Eve Bouff capped off the AI Engineer conference's startup and design engineer audiences.
Sakana AI's 'Diffusing Blame' paper trains DALE-compliant dual-stream networks using error diffusion, reaching 96.7% on MNIST and 61.7% on CIFAR-10 without backpropagation. The method sidesteps the weight transport problem by avoiding exact transpose of forward weights.
Daniel Ajisafe presents a method for improving text-to-video diffusion models' adherence to spatial controls like bounding boxes. The approach uses minor adjustments to better capture user intent while preserving generation quality.
Biomni performs research tasks across diverse biomedical fields. The AI agent, described in Nature Medicine, could be a powerful research partner for scientists.
xAI has open-sourced its Grok Build project under the Apache 2.0 license. The source code is now available on GitHub for developers.
The blog post traces Claude Code's origins in Anthropic safety research and its evolution through contributions from builders and early users. Boris Cherny, Claude Code's lead engineer, notes this is the first public telling of the story.
ZTE debuted a lineup of co-designed smartphones with built-in AI services, joining Chinese manufacturers reimagining mobile devices. The move signals growing competition in AI-integrated hardware.
Index Ventures co-founder Neil Rimer predicts the historic AI wealth generated in Silicon Valley will have to be redistributed, voluntarily or involuntarily. He suggests the current concentration of gains is unsustainable.
TikTok begins testing an opt-in tool that scans for AI-generated likenesses and lets creators report them. Initially tested with some US creators, the tool aims to help protect creator identity.
Google has changed its Gemini usage quotas, potentially reducing the number of AI responses. The article explains how the new rates work and how to monitor your usage to avoid hitting limits.
The AMD Instinct MI350P is a new HBM PCIe AI accelerator that has been observed in various settings. It is expected to offer high memory bandwidth for AI workloads.
Rising AI demand for memory chips is causing supply constraints and price hikes in India's smartphone market. The slowdown underscores how the AI boom is reshaping consumer electronics pricing and corporate strategies.
90% of organizations have adopted at least one internal platform, reducing environment request times from days to hours. The rise of AI agents now demands that platform engineering serve environments at even faster, agent-compatible speeds.
Coherence Guard is designed to enable service robots to behave appropriately around people. The software stack targets general-purpose robots and humanoids to conduct useful tasks.
ZUNA1.1 is released under Apache 2.0, supporting variable-length inputs from 0.5 to 30 seconds across arbitrary channel layouts. It builds on ZUNA1 with improved flexibility for reconstruction, denoising, and upsampling of EEG data.
Get tomorrow's AI brief in your inbox