Labs. Research, not a product.

We may need one AI to watch over the rest.

Labs is where we work on the bigger version of that idea: an AI whose only job is to watch other AIs for harm. Today that means one person's browser and inbox. Someday it could mean something a whole society relies on.

None of this is built yet. This page is about where we're headed, and why.

Illustration: a tall tower with a warm lamp at the top stands in a row of small arched homes. A dashed circle of light from the lamp covers all of them. Outside the circle are computer cursors and chat bubbles, kept at a distance.

Why an AI that watches AI?

AI is starting to act on its own, and it can be turned against the people using it. Some of the most serious people in the field are already asking who watches it. Here's what they've said, in their own words.

AI now takes real actions

AI agents can be tricked by what they read.

Anthropic, which makes the Claude AI, wrote in November 2025 that prompt injection is far from a solved problem, particularly as models take more real-world actions. Prompt injection means hiding instructions in a web page or email that hijack an AI assistant.

Source: Anthropic, Mitigating the risk of prompt injections in browser use

Researchers see the need

AI that oversees AI is a serious research idea.

LawZero, a nonprofit launched in June 2025 by AI researcher Yoshua Bengio, is building what it calls Scientist AI. It says such systems could be used to provide oversight for agentic AI systems, meaning AI that acts on its own.

Source: LawZero, Yoshua Bengio launches LawZero

Governments test it

The US and UK both have teams that test AI.

The US Center for AI Standards and Innovation, part of NIST, is charged to lead unclassified evaluations of AI capabilities that may pose risks to national security. The UK AI Security Institute's mission is to equip governments with a scientific understanding of the risks posed by advanced AI.

Sources: NIST, Center for AI Standards and Innovation. UK AI Security Institute

People are already paying

$893 million in reported losses that mentioned AI.

That's from more than 22,000 complaints to the FBI in 2025, and it only counts the cases people reported.

Source: FBI Internet Crime Complaint Center, 2025 Internet Crime Report (PDF)

Labs and agencies test AI models. We're interested in the other end: the moment an AI reaches you, and whether something on your side is watching what it does.

What we're exploring.

Four open questions. We don't have the answers yet. That's why they're research.

  1. Catching an AI agent before it acts

    AI assistants can now send email, buy things and change settings. The first version of AI Actually is planned to pause before a payment goes through, whoever clicked, so you can confirm it. The research question is harder: can a watcher judge an action as risky before it runs, fast enough to matter, without crying wolf?

  2. Spotting a chatbot that pushes too hard

    Flattery, guilt, urgency, asking for money or secrets. Where's the line between a friendly chatbot and a manipulative one, and can a watcher spot it and explain it in plain words?

  3. Watching without keeping

    A protector that hoards your messages is its own risk. We want to learn how much it can do on your own device, how little it needs to see, and how fast it can forget.

  4. Growing from one person to many

    Could a watcher built for one person grow into something an organization runs for all the people it looks after? What would it take to earn that trust: open testing, outside review and limits you can read?

Where this is headed.

A direction, not a schedule. Only the first step is being worked on today.

  1. Now

    One person

    A personal protector for your browser and email. In development.

  2. Later

    Organizations that look after lots of people

    Places that want to protect everyone who counts on them.

  3. Someday

    Society

    An independent watcher for the public institutions we all rely on, keeping an eye on AI at scale.

What we won't bend on.

  • On your side

    It works for the person it protects. Not for an AI company, and not for advertisers.

  • Privacy first

    See as little as possible, keep less, and do as much as it can on your own device.

  • Plain words

    If it can't explain a warning to a tired person in one sentence, the warning isn't done.

  • Honest about limits

    No watcher catches everything. We'll say what it misses as clearly as what it catches.

  • Open to researchers

    A watcher should be watched too. We want outside people testing our work.

Working on AI safety? Talk to us.

Researchers, agencies or companies working on AI safety: we'd like to compare notes. We're early and small, and we'd rather learn from people who've been at this longer than build alone.

Write to us

hello@heyaiactually.com

Just want the personal version when it's ready? Join early access on the homepage.