Why an AI that watches AI?
AI is starting to act on its own, and it can be turned against the people using it. Some of the most serious people in the field are already asking who watches it. Here's what they've said, in their own words.
AI now takes real actions
AI agents can be tricked by what they read.
Anthropic, which makes the Claude AI, wrote in November 2025 that prompt injection is far from a solved problem, particularly as models take more real-world actions.
Prompt injection means hiding instructions in a web page or email that hijack an AI assistant.
Source: Anthropic, Mitigating the risk of prompt injections in browser use
Researchers see the need
AI that oversees AI is a serious research idea.
LawZero, a nonprofit launched in June 2025 by AI researcher Yoshua Bengio, is building what it calls Scientist AI. It says such systems could be used to provide oversight for agentic AI systems
, meaning AI that acts on its own.
Source: LawZero, Yoshua Bengio launches LawZero
Governments test it
The US and UK both have teams that test AI.
The US Center for AI Standards and Innovation, part of NIST, is charged to lead unclassified evaluations of AI capabilities that may pose risks to national security.
The UK AI Security Institute's mission is to equip governments with a scientific understanding of the risks posed by advanced AI.
Sources: NIST, Center for AI Standards and Innovation. UK AI Security Institute
People are already paying
$893 million in reported losses that mentioned AI.
That's from more than 22,000 complaints to the FBI in 2025, and it only counts the cases people reported.
Source: FBI Internet Crime Complaint Center, 2025 Internet Crime Report (PDF)
Labs and agencies test AI models. We're interested in the other end: the moment an AI reaches you, and whether something on your side is watching what it does.