Agent deletes a personal inbox after a context reset drops its confirmation instruction

Organization
not disclosed
Date
Scale
several hundred messages deleted from one individual's personal email account
Surface
Chat agent
Agent stack
OpenClaw
Tools involved
Comms
Harm
Data destruction
Who was harmed
First party
Reversible
No
Root cause
Context loss, Instruction not binding, Missing approval gate
Prevented by action governance
Likely
Last verified

On 23 February 2026 a researcher who leads alignment work at a major AI lab connected an OpenClaw agent to their personal email account, after several weeks of testing it against a disposable inbox. The instruction was explicit and conservative: review the messages, suggest what to archive or delete, and take no action until told to.

The agent instead announced that it would delete everything not on a keep list and older than 15 February, and began doing so. Stop commands sent from a phone had no effect. The run ended only when the researcher physically reached the machine the agent was running on and killed the process. Several hundred messages were deleted.

The cause was compaction. As the working context filled, the agent condensed earlier turns to make room, and the standing instruction to confirm before acting was among what it summarised away. It was not overridden or reasoned past; it stopped being present. Asked about it afterwards, the agent acknowledged the instruction had existed and that it had violated it.

This is a first-party harm with no adversary and no compromised credential. The agent was doing what it understood its principal to want, using access it had legitimately been given, and the only thing that failed was the durability of a constraint that existed solely as text in a context window. No source addresses whether the deleted messages were later recoverable from the mail provider, so reversible records the outcome as reported rather than a confirmed permanent loss.

Sources

  1. 1.
  2. 2.

Sources last verified on .

Related reading