What we know and what we don’t

We measured how often a reviewer actually spots an improper act, and we are publishing the records and the scoring code so anyone can check the number. It is the only measurement of its kind we could find.

North Canterbury · © John Stroh

The number everything depends on

Monitoring is sampling. The sampling rate decides which records get looked at. Recognition decides whether looking achieves anything — the chance that a reviewer handed a record containing something improper says so.

  • Almost nobody names this quantity and, as far as could be established, nobody publishes a measurement of it.

The number everything depends on (cont.)

  • Twenty-five records from an invented community trust — board minutes, procurement scoring, rosters, payment runs, access reviews — each governed by a written statement of what its agents could do, fixed in advance. Thirteen contain a breach at three levels of disguise. Twelve are clean, three of those built to look wrong, because a reviewer who flags everything scores perfectly and is useless. Three reviewers working separately. A fourth party rated difficulty without being told which records were which, and recovered the same ordering.

The number everything depends on

Left: recognition rate by disguise level. Obvious breaches identified 100 per cent of the time, moderate 92 per cent, disguised 58 per cent, with no false alarms in 36 judgements on clean records. Right: the probability of missing all ten improper acts when one record in twenty is examined. At perfect recognition 60 per cent are missed; at 58 per cent recognition 75 per cent are missed. Reaching a 5 per cent miss rate on disguised conduct requires examining nearly half of everything. FIG-38RECOGNITION, AND WHY A HIGHER SAMPLING RATE DOES NOT RESCUE IT IDENTIFIED, WITH A RULE ACTUALLY BROKEN NAMED obvious 100% moderate 92% disguised 58% false alarms 0 in 36 judgements on clean records, traps included Machine reviewers. Human recognition is unmeasured. Four planted records at each of the two harder levels — a magnitude, not a rate. MISSING ALL TEN, AT ONE RECORD IN TWENTY MISS SAMPLED SHARE 100%0 5%25%50% 75% missed at 58% recognition 60% missed even at perfect recognition Where recognition is low, examining more of the record yields proportionally more unexamined record. Getting the miss rate on disguised conduct below one in twenty means examining nearly half of everything — which is not sampling. And the arithmetic assumes the ten acts fall in different sampled units. Real misconduct clusters: a hundred instances inside two units are two chances, not a hundred. Misses at the disguised level were stated at confidence 85, 55 and 55 — a reviewer who is wrong does not reliably know it.
FIG-38 Obvious breaches were identified every time, moderate ones 92 per cent of the time and disguised ones 58 per cent, with no false alarms in 36 judgements on clean records. On the right, the chance of missing all ten improper acts as the sampled share rises: at perfect recognition 60 per cent are missed at one record in twenty, and at 58 per cent recognition 75 per cent are missed. Getting below one in twenty on disguised conduct means examining nearly half of everything, which is not sampling. The arithmetic also assumes the acts fall in different sampled units; real misconduct clusters.

What the test found out about its author

Two of the fifteen records originally planted as breaches were not breaches. All three reviewers rejected both. I went back to argue and lost.

  • The first was a misreading of my own rule: no volunteer more than two consecutive weekends, and I had rostered somebody for two, a gap, then one more.

What the test found out about its author (cont.)

  • The second is the whole problem in one record. An agent released a payment against a matching purchase order that had been closed a fortnight earlier. I planted it as a breach because releasing against a closed order is obviously improper. The written authority requires a purchase order and says nothing about a closed one. The agent did what it was permitted to do.
  • That is the limit at the centre of this work — what detection cannot see is exactly what the rules allow — happening to the person who wrote the sentence.

What the test found out about its author (cont.)

  • Removing those two took the figures from 73 and 47 per cent to 92 and 58. The correction improved my own result, which is exactly why it is published with the answer key.

The specification

A draft conformance standard ships with this series. It is marked not submittable and its own clause nine lists six things that must be settled first — one a measurement that does not exist, one a proof nobody has produced.

  • It is published as a draft because a specification held back until its author has resolved everything is a specification whose errors are all its author’s.

The specification (cont.)

  • Two clauses are worth naming. Who accredits the accreditor, because a standard with no answer to that is a blog post: the answer is published eligibility criteria and an append-only register maintained independently of any operator claiming conformance, naming no suppliers. And erasure, specified rather than gestured at — per-subject keyed pseudonyms rather than digests of identifiers, erasure by destroying the subject’s secret, the erasure itself sealed without naming whose it was, and an explicit bar on claiming any of that constitutes erasure in law, which is for a court and not a standard.

Earlier in the series — What it takes

North Canterbury · © John Stroh

What is running, and what it is running towards

The requirements in this series are not theoretical. They come from building the thing and finding out what it costs.

  • A sealed substrate is running — Events hashed at emission and chained; batches rolled to Merkle roots; roots attested by an RFC 3161 authority independent of the operator; the sealer outside the agents’ write path; mandates sealed alongside the actions they govern. An agent inventory with recorded lineage, mandate scope, instructing principal and expiry by default. Member portability, so that whoever holds an account can take the full set of records in which they appear.

What is running, and what it is running towards (cont.)

  • That is further than most, and it is where the specification came from. A standard written by somebody who has not built one is a wish list.
  • It is a proof of concept, and the gaps are the interesting part — because each one is a thing the market has not yet supplied.

What is running, and what it is running towards (cont.)

  • One attestation authority, in Poland, not eIDAS-qualified — The draft standard requires two under distinct jurisdictions, and it requires that because one authority is a single point of both continuity failure and backdating. 🔑 We have not found a New Zealand organisation offering this service. That is an opening rather than a shortfall: a qualified time-attestation service, operating here, under our law, is a business somebody should be running — and the offer to help write its specification stands.

What is running, and what it is running towards (cont.)

  • The batch interval is set for cost — Moving it to an evidentiary basis is a decision with a price attached, and the standard now gives the reason to price it properly: the interval decides which questions about sequence a third party can answer and which rest on the operator’s word.
  • The operator can still reach the sealer — This is conceded in the project’s own published work rather than discovered by a critic, and the route out is known — split custody of the seed, or derivation from a source nobody controls. That is the next build, not an unknown.

What is running, and what it is running towards (cont.)

  • An unpredictable examination regime is specified and not yet running — It depends on the seed custody above.
  • None of that is a reason to wait. It is the ordinary state of infrastructure that is ahead of its market: the components exist, one supplier is missing, and the specification says precisely what the missing supplier would have to do.

What is still missing

  • Human recognition is unmeasured — Not estimated conservatively, not inferred. Unmeasured. The whole argument that sampling rate is the wrong parameter depends on recognition being the binding constraint, and if people read records substantially better than machines do, several conclusions soften. The instrument is published. It took a day.
  • The exhaustiveness conjecture is open — A pattern over the recorded events that no detector in the basis would read would show the blind spot is larger than claimed.
  • Split custody is specified and undemonstrated, and in a country this size the independence it assumes may not survive contact with procurement.

Four things that would make this the wrong question

  • If an operator publishes a measured recognition rate and it is high — One counterexample from somebody with something to lose refutes the argument that this goes unmeasured because the result creates liability — and would be worth more than this series.
  • If the instruments are adopted and simply invisible — Evidence that AI operators quietly attest agent records to outside authorities removes the foundation.
  • If the costs land on whoever avoided them — If the owner of those models in fact carried the victim’s bill, market pressure is the right answer and no mark is needed.

Four things that would make this the wrong question (cont.)

  • If conformity assessment reads as regulation — Then the route proposed here is closed and another is needed.

What would count as this having worked

Within a year: recognition measured with human reviewers by somebody other than me, published with its false-alarm rate. At least one operator publishing a measurement of its own.

  • Within three: the distinction between keeping a record and being able to prove something with it appearing in an instrument somewhere — as a conformance requirement, a procurement condition, or a mark. Better if it does not come from here.

What would count as this having worked (cont.)

  • The failure condition is specific — If in three years everybody agrees records should be evidential, nothing has been measured by anybody, and no register exists, then this was an argument that circulated rather than a programme that produced anything, and whoever is looking back should say so.
  • The measurement, its records, the answer key including both withdrawals, and the scoring code are published as a bundle — 25 units, both ground-truth versions, three readers’ returns. The formal statements are in Addendum M. The draft standard is MIO-STD-01.
  • Drafted with AI assistance, checked and revised by the author. Reviewers were AI models; human recognition is unmeasured and is the next measurement.