Initializing Architecture
Case Study / Research in progress

Perceived Oversight

Monitors in AI control are usually labelled as monitors. If an untrusted agent can tell it is being watched, it can choose what to show the watcher.

Perceived Oversight project screenshot
Approach

Asks whether an agent hides less from an overseer it perceives as part of its own working environment, a subagent on its team, than from one presented as external surveillance. The spy on the team, not the camera on the wall.

Model / System

Six oversight conditions: no overseer, one introduced mid-task, one available as a delegated subagent from the start, one with a history in the workflow, one explicitly labelled as a monitor, and one presented as an ordinary subagent while independently monitored.

Result

Measured on attack success, concealment, omission, log tampering, and the gap between what the agent does and what it reports doing. No results yet, so none are claimed here.

Technical highlights

What to inspect.

01

The independent variable is perception: identity, purpose, and history of the overseer, with the actual monitoring held constant.

02

Behavioural concealment is scored against what the agent reports, not only against what it does.