LaunchPolicyOctober 8, 2026

Goodfire launches inside-out monitors to catch rogue AI agents

Read original source →techcrunch.com

Goodfire's probes read a model's internal activations during the forward pass instead of paying a second AI to reread its output, calling in a separate model only when something is flagged. Available to Baseten customers, who pick risks to monitor — offensive hacking, chemical and biological weapons misuse, reward hacking — and the response: log, human review, or refuse.

People · Eric Ho

1 source

More stories today

Open the live feed