Skip to content
Infraware.dev

Operations

The case against unattended remediation

Every vendor in this category will tell you their agent is safe. Fewer will tell you what “safe” means once the agent can act without anyone watching. We think that gap is the whole argument, so here it is, plainly.

The pitch, and where we agree with it

The pitch is obvious, and on its own terms appealing: alerts happen at night, humans sleep, so let the machine close the loop end to end, diagnose, decide, execute, and wake nobody up. For read-only diagnosis, we agree; the investigation runs on its own the moment an alert fires, bounded by a read-only policy, and no one is woken to grant a read. For the part that changes something in your cluster, we don’t, and the reason isn’t caution for its own sake. Unattended remediation quietly merges two different kinds of risk that deserve two different owners.

Two mistakes, two sizes

A wrong diagnosis costs a wrong paragraph. Someone reads a report that says the wrong thing, and the cost is the minutes it takes to notice. A wrong remediation costs a wrong change, a command that ran, against production, under some identity, because a model decided it should. Those are not the same size of mistake. Treating them as one continuous slider from “the AI helps” to “the AI handles it” erases the exact distinction that should decide who signs off on what.

The question that gives away the design

There’s a second cost, quieter but just as real: when a system both decides and acts without a pause, the honest answer to “who did this?” becomes “the agent did.” That isn’t an answer an auditor, an insurer, or your own postmortem will accept, because it isn’t one, it names a process, not a person. The question every vendor in this category should have to answer, and the one we ask ourselves in every design review, is: under whose identity does it act, and who approved it? “The agent” is not a valid answer.

What we do instead

Every action we run executes on a grant, and the grant always has a name attached. Reads run on a grant that’s standing and read-only, a policy your team sets once, enforced by a deterministic allowlist and a parser that rejects what it can’t understand, so the investigation runs on its own, bounded, the moment an alert fires. Writes run on a grant given in the moment: a named engineer reads the evidence and approves, and the fix executes as that engineer’s own mapped identity. A policy is a standing signature, even for the part that runs unattended, the read path’s grant traces back to whoever set the policy, the same way a write’s grant traces to whoever clicked approve.

We think a third kind of grant is worth building: a policy your team signs, once, for the handful of fixes that are genuinely routine, so that when a matching alert fires, the fix runs without the pause, under the same identity binding and the same audit line, traced to whoever signed the policy. We haven’t built it. Shipped Until it ships, no surface of ours gets to claim otherwise: unattended runs simply skip the write step today, and record that they skipped it.

The pause is not a tax

None of this is an argument against automation. It’s an argument against collapsing two decisions with different blast radii into one “the AI did it end to end” motion, because the pause is not a tax on speed, it’s the mechanism that makes the audit trail mean something. If a fix is truly routine and safe enough to run unwatched, the honest version of that is a policy your team writes and signs, not a human quietly removed from the loop on the hope that the model’s judgment holds.

The full grant model, what runs on its own today, what needs a signature, and what a signed policy will look like once it ships, is laid out stage by stage on how it works. The parser and allowlist that make the read-only rung real today are documented on the ladder.

In the live demo, we make this point rather than argue it: partway through, we hand the investigation a write command on purpose, unattended, to see what happens. It gets refused, on stage, and that refusal is the whole thesis of this essay in about four seconds.

You've read the mechanism. Now watch it run.

More engineering notes →