AnalysisPolicySeptember 18, 2026

SynthID-Text watermarking can weaken LLM safety guardrails

New research finds Google's SynthID-Text watermarking can alter which tools a model invokes and whether it follows safety guardrails, especially under adversarial prompts. Anthropic has said future Claude models will use SynthID-Text to comply with a new EU law.

1 source

More stories today

Open the live feed