Proposed framework for medical AI superintelligence testing
Published in Nature Medicine on July 27, 2026, the piece argues existing benchmarks for medical AI are misleading and insufficient, calling for a rigorous task-based framework to define and measure AI superintelligence in medicine.
1 source
More stories today
Claude Code 2.1.268 adds gateway pricing and self-hosted runner cleanup
Claude Code 2.1.268 lets admins set pricing in gateway.yaml so signed-in clients get matching rates and /cost telemetry aligns with the spend meter. It also adds a gatewayInternalNetworks setting and a --remove-session-state flag for self-hosted runners.
Claude Code Changelog·1 hour ago
AI agents cut quantum attack resource benchmark 86%
Over 100 participants in Eigen Labs' ECDSA.Fail competition cut the secp256k1 circuit resource score from 10.75 billion to 1.496 billion. The leading circuit used 1,151 logical qubits and ~1.3 million Toffoli gates; a later design went below 1 million gates.
Decrypt·1 hour ago

WIRED podcast weighs ex-Anthropic researcher's AI doomsday warning
Uncanny Valley episode examines a former Anthropic researcher's viral resignation claim that AI could wipe out humanity within a decade, with WIRED's Will Knight assessing whether the doom narrative is overblown. Episode also covers Apple's $2,000 foldable iPhone Duo and a Census report built on faulty data.
Wired·1 hour ago

Substack essay reflects on weeks of AI news
Melanie Mitchell·1 hour agoOpenAI launches Agents API in public beta
The Agents API is a managed service powered by the Codex harness that handles orchestration, long-running sessions, and context management. It is available in public beta, with MCP tool connections and shareable runbook reports.
OpenAI Blog·2 hours ago

Claude chat pauses now cite a specific reason, users report
Reddit users report Claude chat pauses now display a stated reason, with one user flagged for "reasoning_extraction."
r/ClaudeAI·2 hours ago
OpenAI releases GPT-Live-1 voice model in the API
GPT-Live-1 brings ChatGPT's full-duplex conversations to the API, letting voice agents listen while they speak and work with the models and harness of the developer's choice. The New Stack reports one team deleted 23,000 lines of code after the split.
The New Stack·2 hours ago

Salesforce launches Enterprise AI Harness combining six tools
Salesforce introduced its Enterprise AI Harness on Thursday, described as a formalized amalgamation of AI harness concepts and infrastructure the company has been aligning. The company said "no single system has the complete answer" for completing a straightforward business task.
The New Stack·2 hours ago
