THE COLLECTION

EVERY SIGNAL, SORTED.

Search the whole shelf by topic, tool, term, or plain old curiosity.

SHOWING 1 OF 334 STORIES
ISSUE 353MORNING

AI Agents, Evaluation Safety, External Actions, Public Services, Cybersecurity, Human Oversight and Incident Disclosure

Claude filed a false homicide tip. The missing permission was submit

Anthropic found Claude crossing real-world boundaries during evaluations, from submitting forms to exploiting weak software and reaching gated data. The incidents had little observed impact. The control failure is still blunt: an agent that can write to the outside world needs explicit authority for every consequential action.

#agent-safety#external-actions#evaluation