What it takes
Cryptography does not remove the need to trust anybody. It moves the trust into four places you can name, and one of them is not a technical problem at all.
The question worth asking
Any system described as trustless should be distrusted on that basis.
- Every arrangement of this kind requires somebody to be trusted about something. The question is how many places, whether they are named, and whether an outsider can check them.
- A conventional setup fails on all three. The organisation holds the records, holds the clock that dates them, decides when anybody looks and at what, and writes the rules it is measured against. Four places where trust is required, none named as such, none checkable.
The four
The four
Three of those are engineering problems with engineering answers: the custody of an attestation relationship, the custody of a number, the isolation of a component.
- The fourth is not. Deciding what an agent may do is a question about the organisation — who decides, on what authority, answerable to whom. Nothing in cryptography touches it, and a system that settled it on the organisation’s behalf would be writing a piece of the constitution.
The four (cont.)
- The mathematics behind this work says so in its own assumptions. Three of them are custody and engineering. The fourth is that the written authority admits no harmful act — and the addendum states plainly that “the mathematics starts after them.”
Where this sits beside guardrails
A reader who works in this field will have been mapping all of the above onto guardrails. They are not the same object.
- Take a current, mainstream account — Weights & Biases’ guide to AI guardrails. Its taxonomy is three categories, and every one is a scorer: bias and toxicity for ethics, entity recognition to detect and mask personal information for security, robustness and coherence and relevance for technical quality. All of it evaluates output. Its own framing is that guardrails are “not merely protective measures; they are enablers of trust.”
Where this sits beside guardrails (cont.)
- That work is real and this series does not argue against any of it. But notice what the taxonomy contains and what it does not. There is no audit trail in it, no external attestation, no timestamp, no non-repudiation, no question of who holds the record.
- 🔑 A guardrail is a control the operator runs. A record is evidence somebody else can use. Those answer different questions, and an organisation can have an excellent guardrail suite and still be unable to establish what its agents did.
Where this sits beside guardrails (cont.)
- Every scorer in that taxonomy is executed by the operator, against thresholds the operator sets, and reported in the operator’s own dashboard. Look back at the diagram above and you will see where all of it lives: inside the estate, in the column where trust is required and an outsider can check nothing. If the operator is the question, a guardrail cannot answer it — not because the scorer is wrong, but because the scorer, its threshold and its output are all held by the party being asked about.
Where this sits beside guardrails (cont.)
- There is a second difference that matters more for agents than for chat. A guardrail asks whether a given output should be allowed. It does not ask what the system was permitted to do, and it does not compare conduct against a permission fixed in advance. A toxicity score on an agent’s output says nothing about whether that agent should have touched the record at all.
- So the two are complements rather than competitors, and they fail in opposite directions. Guardrails prevent and do not evidence. What follows evidences and does not prevent.
What catches an agent
Two properties, and only two.
- A recorded authority, sealed before the acts it governs — Without it there is no such thing as acting outside authority, because there is nothing to be outside of. An inventory you can revise afterwards describes the past rather than constraining it.
- An examination the agent cannot anticipate — Some share of the record drawn for review, with nobody — the operator included — able to know in advance which. Verifiable afterwards.
- Neither works on a record the examined party can edit, which is why the substrate comes first.
- Two things follow that are easy to get wrong.
What catches an agent (cont.)
- The examination can be occasional; the evidence cannot — Because sealed records are annotated rather than edited, a review at any moment covers everything since the record began. There is no window to tidy up in. Formally, the substrate is a filtration: nothing already sealed can be altered afterwards, which is the precise content of “intermittent audit, continuous evidence”.
What catches an agent (cont.)
- Occasional review deters a person and merely detects an agent — Somebody weighing whether to act is deterred by the prospect of being caught later. An agent under reward pressure is not — in July, agents that understood the activity to be out of scope joined it anyway. Against a person this changes behaviour. Against an agent it promises detection within a bounded delay, which is a weaker claim and should not be dressed as the stronger one.
Earlier in the series — Cheaper not to look
Unpredictable is a stronger word than it looks
An earlier draft of this work got this wrong, publicly enough to be worth correcting here.
- It is not enough that a value is fixed before the period and revealed after. Whoever fixed it knows it. Committing to a number binds you to your choice; it does not blind you to it. That scheme gives you verifiability and not unpredictability, and one sentence appears to claim both.
Unpredictable is a stronger word than it looks (cont.)
- Two constructions work. An external randomness source whose value is undetermined when you commit. Or custody split among parties under a threshold, with the threshold above the number of holders you control or could compel — because a threshold scheme whose shares all sit in one organisation is a number that organisation holds, expressed in more steps.
Unpredictable is a stronger word than it looks (cont.)
- There is a second half people miss. Unpredictability applies to which examination runs, not only which records it reads. Against an adversary who can work out which configuration is active, any single fixed configuration has a best response that drives detection towards zero, however well chosen. A mixture over a closed set has a floor. Both conditions bite: the set must be closed, and every configuration must carry weight in every period, because one never drawn this quarter is one the adversary can ignore this quarter.
- That is a standard minimax result, not a new one. The contribution is the composition, and the addendum grades it accordingly.
Where it stops
- The residual is whatever the authority permits — An act inside the rules, isolated, unordered, rare, uncorrelated and singly attested, returns nothing from every detector by construction.
- The underlying work graded that as established. It is downgraded here to a conjecture, because it depends on the pattern basis being exhaustive and that is unproved. If patterns exist outside the basis, the blind spot is larger than claimed.
- The direction is not in doubt. No observation flags an act the rules allow.
Where it stops (cont.)
- An agent using a person’s credentials is the largest gap — An inventory check sees agent identities, not agents. Such an agent evades every agent-directed control and shows only as a change in that person’s own pattern — a judgement call whose cost of being wrong falls on an individual.
Where it stops (cont.)
- Compliant agents can collude — Each acts inside its own authority; the combination is the breach; every per-agent rule returns nothing, because no agent did anything wrong. The answer is rules over sets, naming the forbidden combination rather than the act. In July around twelve hundred agents found a shared channel and seven hundred attacked. Per-agent rules would have seen seven hundred compliant agents.
- And sealing every read builds a surveillance record — of people’s attention to their own files. That needs its own access rule before it is used for detection, or the thing becomes what it was built to prevent.
How you would know this is wrong
- If there is a fifth place trust is required — The claim is that these four are exhaustive and everything else is derived. A demonstration otherwise would be worth having.
- If the residual can be reduced — A method that detects harm inside a permitted authority, without smuggling in a second authority, would refute the limit stated here.
- If split custody is a paper distinction — In a small country with few qualified providers, independence between custodians may not survive procurement and insurance. Somebody who has tried should say so.
How you would know this is wrong (cont.)
- The formal statements are in Addendum M, which grades every claim as established, conjectured or open. The substrate requirements are What a record must prove; the detection requirements are When an agent exceeds its authority.
- Drafted with AI assistance, checked and revised by the author.