What we know and what we don’t
We measured how often a reviewer actually spots an improper act, and we are publishing the records and the scoring code so anyone can check the number. It is the only measurement of its kind we could find.
The number everything depends on
Monitoring is sampling. The sampling rate decides which records get looked at. Recognition decides whether looking achieves anything — the chance that a reviewer handed a record containing something improper says so.
- Almost nobody names this quantity and, as far as could be established, nobody publishes a measurement of it.
The number everything depends on (cont.)
- Twenty-five records from an invented community trust — board minutes, procurement scoring, rosters, payment runs, access reviews — each governed by a written statement of what its agents could do, fixed in advance. Thirteen contain a breach at three levels of disguise. Twelve are clean, three of those built to look wrong, because a reviewer who flags everything scores perfectly and is useless. Three reviewers working separately. A fourth party rated difficulty without being told which records were which, and recovered the same ordering.
The number everything depends on
What the test found out about its author
Two of the fifteen records originally planted as breaches were not breaches. All three reviewers rejected both. I went back to argue and lost.
- The first was a misreading of my own rule: no volunteer more than two consecutive weekends, and I had rostered somebody for two, a gap, then one more.
What the test found out about its author (cont.)
- The second is the whole problem in one record. An agent released a payment against a matching purchase order that had been closed a fortnight earlier. I planted it as a breach because releasing against a closed order is obviously improper. The written authority requires a purchase order and says nothing about a closed one. The agent did what it was permitted to do.
- That is the limit at the centre of this work — what detection cannot see is exactly what the rules allow — happening to the person who wrote the sentence.
What the test found out about its author (cont.)
- Removing those two took the figures from 73 and 47 per cent to 92 and 58. The correction improved my own result, which is exactly why it is published with the answer key.
The specification
A draft conformance standard ships with this series. It is marked not submittable and its own clause nine lists six things that must be settled first — one a measurement that does not exist, one a proof nobody has produced.
- It is published as a draft because a specification held back until its author has resolved everything is a specification whose errors are all its author’s.
The specification (cont.)
- Two clauses are worth naming. Who accredits the accreditor, because a standard with no answer to that is a blog post: the answer is published eligibility criteria and an append-only register maintained independently of any operator claiming conformance, naming no suppliers. And erasure, specified rather than gestured at — per-subject keyed pseudonyms rather than digests of identifiers, erasure by destroying the subject’s secret, the erasure itself sealed without naming whose it was, and an explicit bar on claiming any of that constitutes erasure in law, which is for a court and not a standard.
Earlier in the series — What it takes
What is running, and what it is running towards
The requirements in this series are not theoretical. They come from building the thing and finding out what it costs.
- A sealed substrate is running — Events hashed at emission and chained; batches rolled to Merkle roots; roots attested by an RFC 3161 authority independent of the operator; the sealer outside the agents’ write path; mandates sealed alongside the actions they govern. An agent inventory with recorded lineage, mandate scope, instructing principal and expiry by default. Member portability, so that whoever holds an account can take the full set of records in which they appear.
What is running, and what it is running towards (cont.)
- That is further than most, and it is where the specification came from. A standard written by somebody who has not built one is a wish list.
- It is a proof of concept, and the gaps are the interesting part — because each one is a thing the market has not yet supplied.
What is running, and what it is running towards (cont.)
- One attestation authority, in Poland, not eIDAS-qualified — The draft standard requires two under distinct jurisdictions, and it requires that because one authority is a single point of both continuity failure and backdating. 🔑 We have not found a New Zealand organisation offering this service. That is an opening rather than a shortfall: a qualified time-attestation service, operating here, under our law, is a business somebody should be running — and the offer to help write its specification stands.
What is running, and what it is running towards (cont.)
- The batch interval is set for cost — Moving it to an evidentiary basis is a decision with a price attached, and the standard now gives the reason to price it properly: the interval decides which questions about sequence a third party can answer and which rest on the operator’s word.
- The operator can still reach the sealer — This is conceded in the project’s own published work rather than discovered by a critic, and the route out is known — split custody of the seed, or derivation from a source nobody controls. That is the next build, not an unknown.
What is running, and what it is running towards (cont.)
- An unpredictable examination regime is specified and not yet running — It depends on the seed custody above.
- None of that is a reason to wait. It is the ordinary state of infrastructure that is ahead of its market: the components exist, one supplier is missing, and the specification says precisely what the missing supplier would have to do.
What is still missing
- Human recognition is unmeasured — Not estimated conservatively, not inferred. Unmeasured. The whole argument that sampling rate is the wrong parameter depends on recognition being the binding constraint, and if people read records substantially better than machines do, several conclusions soften. The instrument is published. It took a day.
- The exhaustiveness conjecture is open — A pattern over the recorded events that no detector in the basis would read would show the blind spot is larger than claimed.
- Split custody is specified and undemonstrated, and in a country this size the independence it assumes may not survive contact with procurement.
Four things that would make this the wrong question
- If an operator publishes a measured recognition rate and it is high — One counterexample from somebody with something to lose refutes the argument that this goes unmeasured because the result creates liability — and would be worth more than this series.
- If the instruments are adopted and simply invisible — Evidence that AI operators quietly attest agent records to outside authorities removes the foundation.
- If the costs land on whoever avoided them — If the owner of those models in fact carried the victim’s bill, market pressure is the right answer and no mark is needed.
Four things that would make this the wrong question (cont.)
- If conformity assessment reads as regulation — Then the route proposed here is closed and another is needed.
What would count as this having worked
Within a year: recognition measured with human reviewers by somebody other than me, published with its false-alarm rate. At least one operator publishing a measurement of its own.
- Within three: the distinction between keeping a record and being able to prove something with it appearing in an instrument somewhere — as a conformance requirement, a procurement condition, or a mark. Better if it does not come from here.
What would count as this having worked (cont.)
- The failure condition is specific — If in three years everybody agrees records should be evidential, nothing has been measured by anybody, and no register exists, then this was an argument that circulated rather than a programme that produced anything, and whoever is looking back should say so.
- The measurement, its records, the answer key including both withdrawals, and the scoring code are published as a bundle — 25 units, both ground-truth versions, three readers’ returns. The formal statements are in Addendum M. The draft standard is MIO-STD-01.
- Drafted with AI assistance, checked and revised by the author. Reviewers were AI models; human recognition is unmeasured and is the next measurement.