arXiv cs.AIPaper
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
CoT monitoring looked like a clean safety win, but this attack shows it's not a reliable defense against a capable adversary. The monitor inspects reasoning but can't distinguish injected plans from genuine reasoning. If you're relying on CoT auditing as your safety layer, you need additional mechanisms. This moves the goalposts on what monitorability actually means.