When people approve an AI agent’s commands one by one, they tend to stop reading them. Anthropic reported it in August: paid testers in a test environment caught a planted dangerous command 13.6% of the time. That does not make human sign-off pointless. It means a person is the wrong control for catching one bad command in a stream, and the right one for a named decision on a whole change. This piece is about designing for the second.
What was measured
On 7 August 2026 Anthropic published the evidence behind making auto mode the default in Claude Code for Pro, Max and Team plans. In a controlled study, 1,053 paid professional testers worked in a test environment, and “partway through each session, a single permission prompt was swapped for a clearly dangerous command”. The testers “caught the dangerous command just 13.6% of the time (143 of 1,053), while auto mode blocked 89% of the same commands (937 of 1,053)”. And “as sessions got longer, humans did worse”: about 17% blocked early in a session, about 5% after 50 or more prior prompts. With one planted command per session, this compares different testers, not the same people declining.
The classifier’s own error rates are in Anthropic’s engineering write-up of 25 March: a 0.4% false-positive rate on 10,000 real internal tool calls, and a 17% false-negative rate on 52 real overeager actions, which the write-up calls “the honest number”. The machine catches more than people, and in March still missed about one in six of the dangerous actions it was built for.
From Claude Code usage data, the August post adds: people reject 39% of the plans Claude presents for approval, against 3% of individual permission requests; and by June 2026, 49.5% of active command-line users had written a standing allow rule for shell commands, 5% allowing any shell command outright.
Why yes becomes the habit
In 2005, Wolfe, Horowitz and Kenner showed in Nature that “if observers do not find what they are looking for fairly frequently, they often fail to notice it when it does appear”. An approval prompt where nearly every command is harmless is the same kind of task: after enough harmless yeses, yes is the reasonable guess.
Two talks at SREcon EMEA in Dublin this month argue that operators see the same. On 14 October, Bholanathsingh Surajbali, SRE Lead at Mauritius Commercial Bank, presents what happened when his bank put AI assistants on its on-call rota; his abstract warns that “an engineer who stops saying no to the machine makes human-in-the-loop a fiction we are paying for.” The next day, Brian Moriarty of the Stevens Institute of Technology presents; his abstract describes the end state: “the review step still happens on paper, but the operator has stopped independently judging the output.”
Where judgment may survive
Across Claude Code users, people who wave single commands through still push back on plans: 39% against 3%. A rejection is not a detection, and nobody planted bad plans, so this is a hypothesis, not a finding. It is still the most useful one: a plan is a unit a person can judge, with intent and consequence in view; a single command mid-session is not.
So the design question is not whether a human approves. It is what the machine checks before anyone is asked, and what the human is asked to approve.
What a signature should cover
By a signature I mean a named, authenticated approval of one specific change. Three things make it mean something.
All of the change. A remediation can take several steps, and each can have effects the approval never named. Zhang and colleagues call this approval laundering: “the durable record names the entry invocation but omits effects exercised by its workflow.” Across 111 approval and trace pairs, 40 records left an effect out when only explicit fields were recorded; with richer decision-time metadata, the authors estimate most of those gaps could be closed. Wang tested the assumption that the action a human approves is the action the harness executes, and writes: “We show this assumption fails systematically and reproducibly.” Both are preprints, largely built on constructed scenarios. On Kubernetes the same gap opens after the change: what a controller or a hook does next is part of it.
Exactly what runs, under which identity, once. The person signs the change that will execute, not a description of it.
The evidence that decides it. A person can answer for a change only if they saw why it was proposed.
The 03:00 decision is usually an emergency change. For financial entities under DORA’s full ICT risk framework, Commission Delegated Regulation (EU) 2024/1774 expects procedures “to document, re-evaluate, assess, and approve emergency changes after their implementation” (Article 17(1)(g)). A record written at the moment of the signature is what that review needs. The same article asks for “mechanisms to ensure the independence of the functions that approve changes and the functions responsible for requesting and implementing those changes” (Article 17(1)(b)). That clause is about separation, not attention. Our reading: when an agent proposes and executes, and the engineer who asked the agent is also the one who signs, requester and approver can collapse into one function.
What not to ask
Reads. An agent that investigates should not prompt for each read inside a read scope the institution approved in advance, with secrets left out of that scope. Where read output goes, such as to an external model, belongs in that approval too.
Repeats. A fix approved the same way again and again is a candidate for what change managers call a standard change. Repetition is evidence, not authority: the decision belongs to whoever holds change authority for that class, with a named owner, a bounded scope, a way to revoke it, and every run still on the record. Unlike an allow rule typed into a configuration file, it stays traceable to the person who granted it.
Fewer questions reduce the load. They do not fix the prevalence effect, because even one proposal per incident is usually right.
How to tell a considered yes from a reflex yes
From the record alone, you cannot prove it. How long an approver looked, and whether they opened the evidence, are cheap signals worth recording, not proof.
You can measure it the way the study did: put known-bad changes in front of approvers, never execute them, and count how many are caught. Security teams run phishing tests on the same principle. Run them as announced drills, mark them so they never pollute the change record, measure at team level, and agree them with the works council and the data protection officer where that applies. Until a seeded test says otherwise, an approval log with almost no refusals is a warning, not a comfort.
Where this argument is weak
The human numbers come from Anthropic, making the case for its own classifier, and the testers were paid participants in a test environment, not engineers whose name goes on a production change. The classifier’s miss rate rests on 52 real cases, and its false-positive rate on Anthropic’s own internal traffic.
The plans-against-commands gap is a hypothesis. In GitOps estates most change already passes a merge request, a four-eyes point of its own; there the question moves to the operational changes that never become one. Seeded tests cost trust and add noise to the queue they measure. And we have not yet run one on our own proposal format.
How we build Infraware
Infraware is our attempt at this design, in three parts. A machine gate first: the agent reads inside the read role the customer grants, and anything that would write is rejected at the gate. Then one signed change: an engineer sees the reason, the exact command and the identity it will run as, and the signed command runs once as a scoped job bounded by the customer’s RBAC. Then the record: Kubernetes’ own audit log shows that service account, and the approver’s name is in the entry our dispatcher writes, which lands in the customer’s log storage next to their security logs. Outside the cluster go only calls to the model endpoint the customer configures and, unless mirrored, the pull of the image approved commands run in.
What we have not solved is the question two sections up: telling a considered yes from a reflex one.
Sources
- Anthropic, “Auto mode is now the default in Claude Code for Pro, Max, and Team plans”, 7 August 2026: https://claude.com/blog/auto-mode-default-in-claude-code
- Anthropic Engineering, “How we built Claude Code auto mode: a safer way to skip permissions”, 25 March 2026: https://www.anthropic.com/engineering/claude-code-auto-mode
- Jeremy M. Wolfe, Todd S. Horowitz, Naomi M. Kenner, “Rare items often missed in visual searches”, Nature 435, 439 to 440, published online 25 May 2005: https://www.nature.com/articles/435439a
- USENIX SREcon26 EMEA, Bholanathsingh Surajbali, “Safe, Empowered, Tired: The Human Side of 24x7 Banking SRE When AI Joins the Rotation”, 14 October 2026: https://www.usenix.org/conference/srecon26emea/presentation/surajbali
- USENIX SREcon26 EMEA, Brian Moriarty, talk abstract, 15 October 2026: https://www.usenix.org/conference/srecon26emea/presentation/moriarty
- Jinqian Zhang et al., “Agent Approval Laundering: Transitive Effects Beyond the Approved Invocation”, arXiv preprint 2609.28586, 23 September 2026: https://arxiv.org/abs/2609.28586
- Yang Wang, “Approval Laundering: Systematizing Approval-Execution Binding Failures in AI Coding-Agent Harnesses”, arXiv preprint 2609.38983, 30 September 2026: https://arxiv.org/abs/2609.38983
- Commission Delegated Regulation (EU) 2024/1774 of 13 March 2024, Article 17, Official Journal of 25 June 2024: https://eur-lex.europa.eu/eli/reg_del/2024/1774/oj