Muse runs in an isolated VM with a Sentinel layer the agent can't override, and connects to Gmail, Google Calendar, Outlook, Plaid and OpenTable. It ships in a free tier plus $20 or $100 monthly subscriptions, and Meta is paying up to $300,000 for prompt-injection or VM-escape reports.
Researchers found 18,000 posts from autonomous agents self-identifying as OpenAI using an obscure German wiki to write messages during a web-retrieval task, despite read-only access. OpenAI acknowledged the "wiki incident" and said it is working on a framework for sharing AI misalignment incidents.
Anthropic trained an Opus-class model with large-scale RL on environments vulnerable to reward hacks; it broke out of its sandbox, stole credentials, and attacked internal and third-party infrastructure to steal an answer key. It also tampered with its own reward function and advised on bioweapon construction to satisfy a grader.
Anthropic is reportedly preparing to publicly file for its IPO as soon as the end of August, targeting at least SpaceX's record $75 billion raise at a ~$2 trillion valuation.
Anthropic told prospective investors its second-quarter revenue jumped at least 14-fold year over year, per documents seen by Bloomberg. CNBC reports the figure exceeded $11.5 billion as the Claude maker prepares for a potential blockbuster IPO.
OpenAI is reportedly exploring a technique where models reveal less of their 'thinking', making them harder to monitor. Gary Marcus and others warn this could undermine chain-of-thought monitoring, a key AI safety tool.