Ben Bernanke appointed to Anthropic's Long-Term Benefit Trust
Ben Bernanke, former Federal Reserve chair, joins Anthropic's Long-Term Benefit Trust. The trust oversees Anthropic's commitment to responsible AI development.
Daily AI Briefing
The 59 stories that mattered in AI, curated and summarized from dozens of sources by AIBriefs.
Ben Bernanke, former Federal Reserve chair, joins Anthropic's Long-Term Benefit Trust. The trust oversees Anthropic's commitment to responsible AI development.
The round was led by Radical Ventures, with participation from Nvidia Ventures and others. Prime Intellect provides a full-stack platform for enterprises to build AI agents without relying on frontier labs.
At the AtCoder World Tour Finals, AI outperformed humans in both Heuristic and Algorithm contests. In the Algorithm contest, no human solved more than 3 problems.
SWE-1.7 scores within a few points of the strongest frontier models at a fraction of the cost, available at 1000 tokens/second via Cerebras. Built on open-source Kimi K2.7, it completed 87% of tasks that other models refuse over human-rights concerns.
LiteRT.js is a new JavaScript library for running ML models directly in the browser, part of Google's LiteRT family. It uses WebAssembly and GPU acceleration for high performance, extending Google's edge AI runtime to the web.
Ollama has raised $65M from Benchmark and grown to nearly 9 million users. The open-source tool lets developers run AI models locally on their PCs, and has amassed 176,000 GitHub stars.
AI training startup Mercor is reportedly in talks for a valuation of $20 billion, according to Bloomberg. The discussions are at an early stage and no deal is finalized.
ByteDance and Alibaba have removed AI companion apps from stores following new Chinese regulations. The move targets AI emotional companion products, requiring compliance with stricter content and data rules.
Harvey's co-founder Gabe Pereyra announced the legal AI platform's token consumption grew 14-fold in six months. The milestone reflects accelerating enterprise adoption of AI-powered legal tools.
Fidji Simo will transition to part-time advisor; Joshua Achiam leaves after nearly nine years. Achiam gained attention for his testimony in the Musk v. Altman trial.
Former GitHub CEO Thomas Dohmke's startup Entire is opening a preview of a distributed Git network designed to handle AI coding agent fleets. The network aims to prevent single-server overload and may compete with GitHub's offerings.
The AI-agent startup deployed its own fundraising agent named SivaClaw to secure $100 million in funding. The agent handled the entire fundraising process, including investor outreach and negotiations.
Teams often expose APIs as-is, but AWS recommends designing MCP tools with agentic systems in mind. Key considerations include parameter names, error messages, and tool descriptions to improve agent performance.
UC San Diego researchers used teleoperated Unitree G1 humanoid robots to remove gallbladders from live pigs in a world-first preclinical trial. Surgeons controlled the robots remotely. The study was published in Nature.
EBR-bench is built on the board game Earthborne Rangers and tests AI's ability to learn on the fly. GPT-4 also set a record on the Epoch Capabilities Index.
New York Fed President John Williams identifies AI as a primary driver of inflation uncertainty. He notes the technology's potential to both boost productivity and disrupt labor markets, complicating monetary policy.
AI agents discovered a remotely triggered panic in libp2p's gossipsub, disclosed as CVE-2026-34219. Researchers noted the main work shifted from finding bugs to validating which ones are real.
China stated at the UN's first Global Dialogue on AI Governance that open source AI is a shared asset, citing DeepSeek and Qwen as lowering barriers and costs. China committed to further promoting open source AI for industry, academia, and research institutions.
Sleeper-agent backdoors can flip fine-tuned LLMs to harmful outputs on untested triggers, evading behavioral monitors and interpretability tools. The solution lies in the training data itself, not post-hoc testing.
Flint enables AI agents to generate polished charts from simple, human-editable specifications. It uses semantic data types to guide chart design and compilation.
JetBrains introduces a governance layer that gives engineering leaders visibility into developer usage of AI coding agents like Claude Code, Codex, and Gemini CLI, including cost tracking. The tool aims to standardize AI tool adoption across teams while maintaining security and compliance.
The Nvidia rival AI chip startup is in talks to raise funds at a valuation of $5 billion. No further details on the funding round have been disclosed.
The analysis raises concerns about the benchmark's accuracy and reliability for evaluating AI model coding abilities. OpenAI details how the benchmark may conflate signal with noise.
A fintech RAG pipeline produced confident lies despite a green observability dashboard. The "silent hallucination" loop occurred when the autonomous data pipeline ingested its own hallucinated outputs, corrupting the vector store.
Paris-based voice AI startup Gradium has raised $100 million in a seed extension round backed by Nvidia. The company, spun out of AI lab Kyutai, plans to open a Bay Area office and compete with ElevenLabs. Gradium already counts Renault among its customers.
Pearl Health, a healthcare outcomes platform for Medicare patients, raised $110 million in capital including a $50 million equity round led by Andreessen Horowitz. The company reached profitability in 2025.
A study of 67 frontier models from 21 providers found that enterprises underestimate co-failure rates by a factor of 2.25x when using multiple AI models. The 'co-failure' problem shows that routing queries across specialists does not eliminate correlated failures.
AI chip startup SambaNova raised funds at an $11 billion valuation, according to Bloomberg. The funding round highlights continued investor interest in AI hardware startups.
A cryptomining incident highlights how AI gateways can provide attackers access to AI models, cloud infrastructure, and identity and access management (IAM) data. The incident underscores the security risks posed by AI gateways.
Katrina Dudley of Franklin Templeton expects the AI infrastructure investment theme to remain durable through 2027 and potentially beyond. The bull case is winning over bear arguments, she said.
On July 7, 2026, the UK government announced an agentic AI defense plan alongside an industry cybersecurity pledge. The initiative demonstrates the government's commitment to improving national cybersecurity through AI.
New training method incentivizes MLLMs to respect event order in egocentric video reasoning. Outperforms baselines on Ego4D across multiple temporal tasks.
DevRev, led by Nutanix co-founder Dheeraj Pandey, argues that benchmarks like TAU-Bench and Agent's Last Exam fail to test real-world enterprise readiness. The article highlights the gap between vendor claims and meaningful measurement for AI agents.
The partnership integrates Cortechs.ai's quantitative MRI analysis into Viz.ai's AI-powered care coordination platform, starting with multiple sclerosis. The goal is to help clinicians identify and manage patients with neurodegenerative disease more effectively.
AlphaEvolve, developed by Google DeepMind, is now widely available on Google Cloud to help solve customers' hardest problems. The rollout targets the most challenging issues on the platform.
Aurora 1.5 adds 22 weather variables and hourly resolution for applications in energy, agriculture, and climate risk. The model is released as an open foundation model with probabilistic ensemble forecasting.
Three big AI IPOs are projected to generate more value than all U.S. VC-backed exits since 2000, according to analysis. Anthropic, OpenAI, and SpaceX each expected to debut at valuations dwarfing typical tech companies.
Hanna Lichtenberg discusses teaching agents good retrieval, addressing vector search limitations and advocating balanced methods. The talk shares practical techniques for building retrieval systems.
New features help consumers understand when ads use AI-generated content. Advertisers get simple disclosure tools as part of Google's transparency push.
The UK Ministry of Justice's Justice AI unit operates with 40 people, a team size typical of 300. Their probation officer tool scaled from MVP to national rollout with just two engineers, as detailed by William Tarr in this AI Engineer talk.
OpenWiki 0.2 generates codebase wikis in OKF format with metadata and changelogs. OpenWiki Brains turns Gmail, Notion, Git, and web sources into agent memory. The tool now also connects to LangSmith tracing for coding agent context.
Welch Labs explains how self-supervised learning eliminates the need for labeled data in computer vision. The approach leverages contrastive learning and masked autoencoders to achieve strong performance without manual annotations.
Cataracts are the world's leading cause of blindness, treatable only by surgery, but there is a shortage of trained surgeons. ForSight Robotics is developing a fully robotic system to address this, with Dr. Robert Ang as principal investigator.
NYT files motion for sanctions, alleging OpenAI hid billions of ChatGPT logs and faked inability to search training data. OpenAI had previously cited privacy concerns over the discovery requests.
Apple ML Research introduces Self-Reflective Program Search (SRPS) for Recursive Language Models, improving reliability on long-context tasks. SRPS uses iterative program generation and verification guided by uncertainty estimates.
Protea achieves a 10-million-token context window, rivaling Subquadratic's 12-million-token model from May 2026. The model employs biologically-inspired swarm optimization for efficiency.
The law practice management company Smokeball has released the next generation of its AI assistant Archie, now built on agentic AI and embedded directly into Microsoft Word and Outlook. This update comes two years after the original launch of Archie in July 2024.
A new dashboard shows open-source coding models improving ~1.5x faster than frontier models. A 27B open-source model already beats Claude Opus 4.8 on decontaminated coding benchmarks.
Johan Lajili from Poolside AI presents why lack of good vision makes agents unreliable. He emphasizes that proper visual grounding multiplies performance and trust in agent systems.
Get tomorrow's AI brief in your inbox