A command a model proposes has to become a command a shell would actually run. Somewhere between those two things sits a decision: what happens when the model’s text doesn’t parse cleanly?
Most systems in this category answer that question by not asking it. A string comes out of the model, a string goes into a shell, and whatever the shell makes of it is what happens. We think that’s the wrong place to put the decision, so we built a parser instead of a prompt, and we set its default the opposite way from most software: when parsing fails, the answer is no.
The shape of the problem
A proposed command isn’t one token , it’s
quoting, chaining, redirection, and substitution, all sitting in a
string a language model wrote. kubectl logs "checkout-7f9 looks
harmless until you notice the quote never closes. kubectl get pods; rm -rf /tmp/cache looks
like one command until you notice the semicolon. A regular
expression can catch the mistakes you thought of. It cannot catch
the ones you didn’t , and a model’s failure modes
are, by construction, the ones nobody thought of.
So the investigation’s command parser doesn’t pattern-match for danger. It parses , the same job a shell ’s own grammar does , quoting, substitution, chaining, redirection, all of it, before anything is allowed to mean something. If the parser can’t build one complete, unambiguous command from the model’s text, there is no command. There is a rejection, on record, and the investigation moves on. Parse errors reject. They never approve.
Two gates, then a principal with nothing to spend
That’s gate one. Gate two is narrower still: even a command
the parser fully understands has to match one of a short list of
read-only prefixes , the kind of thing an on-call engineer
runs by hand to see what’s going on. kubectl get,
kubectl describe, kubectl logs, journalctl, and the rest of a list that fits on one
screen. Nothing on that list writes anything, and the
investigation’s own credentials couldn’t perform a
write if the list let one through , the read-only
ServiceAccount it runs under has no write verbs to spend. Three
separate mechanisms, stacked: the allowlist, the parser, and the
principal. Any one of them failing leaves the other two standing.
One distinction worth being precise about, because careful readers will find it anyway. Once the parser and the allowlist accept a proposed command, it runs in a shell on the investigation service’s own pod , bounded by that pod’s own permissions. Approved writes travel a different, narrower path entirely: a fixed argv to a Kubernetes Job under the approver’s mapped identity, and that dispatch layer never re-parses request text through a shell at all , the argv is the command, nothing rewritten, nothing appended. The path that can write is the narrowest one we have, on purpose.
# illustrative only, the real allowlist ships as config; see the docs
kubectl get pods -n checkout # allowed
kubectl logs deploy/checkout --since=10m # allowed
kubectl logs "checkout-7f9 # rejected, quote never closes
kubectl get pods; rm -rf /tmp # rejected, a second command, chained Why reject instead of repair
An obvious objection: why not just fix the malformed command instead of throwing it away? Strip the bad quote, drop the second clause, run the part that looks safe. We considered it and rejected the idea for the same reason we reject the command: a parser that guesses at intent is a parser that can be steered. A rejection costs nothing , the model gets a report of exactly why the attempt failed and can try again , so there is no upside to guessing on its behalf. Ambiguity is treated as hostility, not as an editing opportunity.
A rejected command doesn’t vanish, either. It lands in the root-cause report next to everything else the investigation tried, so whoever reviews the fix can see exactly what didn’t make it through, and why. The design assumes the model is sometimes wrong , a rejected parse costs a line in a report. It never costs a shell.
We didn’t write this policy into a system prompt, because a system prompt is a request, and a request can be argued with. We wrote it into the parser and the allowlist, inside the Rust workspace behind every release, because a compiler doesn’t negotiate. The full mechanism, stage by stage, is on how it works; the shipped allowlist and parser reference ships with the deployment , and the security page walks the enforcement.
The easiest way to see a parse error reject something is to watch us try to make it happen. In the live demo, we hand the investigation a write command on purpose, at the exact moment only reads are allowed, and you watch it get refused , no cleanup, no drama, just a stopped proposal and a line explaining why.