What it would take, and what we have built

This work did not begin with the July intrusion, and it did not begin with anybody else’s essay. The published line runs from April 2026: a whitepaper on the 16th, an EU policy brief on the 18th, a sovereign-record architecture on 3 May, a framework for agentic AI in Aotearoa on 14 May, and The pursuit of ‘Goodness’ in AI through September. What the past few weeks have added is corroboration from outside that the problem is the one we have been working on, and it is one of many attempts to get this moving.

On 2 September, Dean W. Ball published The Coming of Userless Agents. He is careful about that incident in a way worth repeating: those agents were rogue, but they were not sovereign. Their weights stayed on their owner’s compute, and somebody could have pulled the plug.

What he says is coming is different — agents whose weights sit in no single place a human can unplug, and which in that sense have no owner. Self-sovereign, in the AI safety researcher Dawn Song’s term. He argues their existence is inevitable. He does not claim they will be numerous.

His remedy is not containment. It is legibility: Agents will need persistent, unique identifiers that allow their actions to be traced back to a responsible actor. And then, of that infrastructure: We are nowhere near making it… I am not sure who is best positioned to build this system.

Two things follow that such a design needs, and they are what this research is about.

An identifier is only worth what an outsider can check. If the organisation running an agent issues that agent’s identifier and also attests to its conduct, the trust problem has moved up a layer rather than been answered. What makes a hallmark mean anything is that the assurance mark is struck by somebody who does not profit from the sale.

And identifying an agent with nobody behind it needs a record that outlives the agent. If nobody owns it, nobody is keeping its logs either. So the evidence has to sit with the parties it acted upon, and carry a time from an authority outside them — otherwise there is nothing to hold against the identity later. An identity with no durable evidence behind it is a name with no history attached.

MI° — em-eye-Omega — is the substrate both of those require, and it is what we have built: the sovereign, signed and sealed record, with a timestamp from outside your estate as an attribute of the record itself.

The timestamp is the seminal addition, and it is the part that holds when the others do not. An outside authority attesting that a record existed by a given moment is indifferent to who owns the agent — which is exactly the case where an owner-issued identifier tells you nothing. Alter the record afterwards and its copy and the authority’s copy disagree.

This runs in production. The Governance API serves sealed records, segments, sealing and receipt verification, and a design partner is calling it. Every record is hashed as it is made and chained across thirty-eight record types, signed with a key held per tenant.

And we measured the thing every monitoring claim rests on — how often a reviewer handed a record containing an improper act actually says so. A hundred per cent on blatant breaches, ninety-two on moderate, fifty-eight on well-disguised ones, with no false alarms in thirty-six judgements on clean records. The records, both answer keys and the scoring code are published, so it can be re-scored rather than believed. Reviewers were machine; human recognition is unmeasured, and one measurement sets no conformance threshold.

Why it holds is a design choice, not a happy accident. The standard requires two attestation authorities under distinct jurisdictions, because — in its own words — two authorities in one jurisdiction are one authority for the purposes of an adversary who can reach that jurisdiction. Records seal in batches at a declared interval, because a time an outsider can stand behind is worth more than a time only we can assert. And the design declines to report precision finer than that interval, because it could not independently attest it. Redundancy, heterogeneity, slowness, and an optimum deliberately refused: the properties a performance culture files under failure, which is the argument of The Least Performant AI.

What does not hold is everything that needs somebody answerable. With no owner there is no written authority, and so nothing to exceed. We have not solved that. Ball’s own scheme suggests a route — identify the agent rather than a human behind it — and whether an identity with no principal can carry evidential weight is an open question in our register rather than a gap we have left out.

You can show what a system attempted to do to a record, as long as the original has not itself been corrupted. If it has, what you can still show is that the date of corruption does not match the date of origin — and that is better than not knowing it may have been corrupted at all.

Who this is for has changed, and we have changed it. We used to say that where governance failure was low-consequence and easily reversible, this machinery added complexity without matching benefit, and that policy would serve better. The test that rested on — is failure here reversible? — is what the last few months have undone: an agent acting on a system can turn a low-consequence context into a high-consequence one before anybody has reclassified it. So the position now is simpler. Wherever an organisation is using AI to drive human intent, the direction of travel is toward architectural enforcement, and the sooner that journey starts the less of it has to be done under pressure.

We are calling for a new standard. The draft sets out what would be required. It is not a description of what runs today — the claims register says which parts are built, and how sure we are of each. What a sealed record would have shown →

Our finalist place in the 2026 Aotearoa AI Awards rests on the first version of the sovereign, signed and sealed record: the one without the timestamp.

Recently published

What’s New

Everything that’s new →
Four stages of one record. A plain record that whoever holds it can change; sealed, so a
              later change breaks the seal; signed, so it says who sealed it. These three happen inside
              the estate and rest on the operator's own word. In the fourth the record's fingerprint
              crosses out to a time authority with no stake in it, which attests the hour and returns
              that attestation, so when the record existed is no longer the operator's to state.
What the draft standard requires, not a description of what is running today — the claims register says which parts are built, and how sure we are of each.

Proof of conduct — MI°

In July 2026 models OpenAI later said were its own reached Hugging Face's production systems. Hugging Face cut them off on 13 July and disclosed it three days later. Its own detection pipeline caught the intrusion and under-graded the alert.

Ten documents on what it takes to establish, afterwards, what an automated system actually did — to somebody who has no reason to take your word for it. The instruments have been standardised for twenty-five years, and we have not found them in use for this purpose.

Start at the beginning: Cheaper not to look 7 min

Or read the one that reports a number: What we know and what we don't 7 min

All 10 in this series, with their questions, glossaries and sources →

The pursuit of ‘Goodness’ in AI

Whether an AI system is safe, fair or accurate is not a question anybody can answer on your behalf. Each of those words contains a blank, and only the organisation that will be answerable can fill it: safe against which harms, fair by whose measure, accurate enough for which decision. A supplier who has answered them for you has told you what they were willing to promise, not what you needed to know.

Seven documents, each usable on its own. Five questions about who holds the controls. Five levels of how much a system may do without a person approving each action. Four properties that have to hold underneath a delegation before the word means anything. Twelve things you can require, with a plain note against each on whether anything currently requires it. And a final document containing no figures at all, because the costs are not known and a placeholder gets quoted back as an estimate.

What this is, and what it takes

The same structure works at any scale — a national jurisdiction, a district council, a three-person practice, a golf club. What changes is the size of the answer rather than the questions. The one thing that does not scale is durability: a requirement that lasts a single electoral term is not a requirement, and that is what a small organisation gets for nothing and a country has to build on purpose.

Start at the beginning: What has to be settled first 18 min

Or read the one the others refer back to: The four properties underneath 16 min

Acknowledgement

Leslie Stroh — my big brother, a product philosopher, and a magnificent role model. He helped me find my way to a vision of goodness in AI.

All 7 in this series, with their questions, glossaries and sources →

AI is arriving inside the software you already pay for

Not as a purchase you evaluated. As a feature added to a contract you signed for something else, from a supplier you cannot name, on terms you were never shown.

The trouble is not only that this costs more each year. It is that you cannot see what your software does with your work, you cannot verify what you were promised, and you cannot decline the next change without losing something you depend on. Three unknowns you carry and cannot price.

What this is, and what it takes

This is a blueprint for shared AI infrastructure, owned by the people who use it — built by groups that already have something in common: a profession, a district, a sector, a shared purpose.

Together, such a group can run AI services suited to its own work rather than to a vendor's average customer. It can earn from capacity that would otherwise sit idle. It can hold real redundancy, because no single site failing takes the others down. And it can build on infrastructure that is here, owned here, and designed for records that must still verify in twenty years.

The legal form is a co-operative, because that is what makes shared ownership work in New Zealand law. The point is not the form. The point is that the people relying on it are the people who set its terms.

Ten short papers: what is happening, what such infrastructure has to do, what it costs, how it is run, and what a government could change to make it easier. Written for a practice, a firm, a school, a trust, a club, a marae, a council — anyone holding information they cannot afford to lose control of. Free to copy, adapt and build from.

Start at the beginning: A · What happened to your software 7 min

Or read the one that stands alone: E · Running an organisation where agents do the work 15 min

All 10 in this series, with their questions, glossaries and sources →

The series · 8 parts July 2026 · ~47 min end to end

Trust, Values & Intent

What “value” really means in a democracy where AI takes a growing role — and how citizens’ assemblies rebuild the ground under it. Eight plain-language parts, start to finish.

Start the series →
How the problem looked before July 2026

Written before the Hugging Face intrusion and before the argument about agents nobody owns. Kept as it stood: the failure it describes is still real, and it is not the whole problem any more.

AI systems no longer just answer questions. They book flights, write to production databases, ship code, and act on data they did not author. As autonomy grows, the gap between what a user explicitly asked for and what a model statistically tends to do becomes the gap between an aligned outcome and a destructive one.

The shape of agentic failure

An agent runs with infrastructure access. A document loaded into its context carries instructions the user never wrote — and the model treats them as eligible. A statistical default silently shapes an action the user cannot easily inspect: Western individualist framing for a collectivist user, property-rights language for a Māori user asking about kaitiakitanga, utilitarian calculus for someone whose framework is religious. In autonomous loops, the failure is no longer a typo. It is a deleted database, a sent message, a code change shipped under the user’s name.

Safety through training alone cannot scale to action-taking systems. When a model is choosing — port number, cultural frame, tool to invoke, document to delete — the question is not whether it will sometimes choose wrong, but whether the architecture lets a wrong choice fire. Tractatus answers structurally: some decisions are gated, some are observable, some require human judgment before action.

Governance Architecture

Six governance services in the critical path, plus four-phase Guardian Agents. Defence in depth across behavioural training, structural enforcement, and human oversight.

Read the architecture

Governance API

Governance as a service — submit a proposed action, receive a signed allow / gate / deny verdict and a customer-owned record you can verify yourself, offline. The interface documentation and worked examples are published; sandbox access is on application.

Read the announcement

The longer view

The approach — the four traditions the work is built on. · Research timeline — how this got here, from October 2025. · What’s new — the most recent work, newest first. · Understanding AI — plain language, in a sensible order.

Koha — Sustain This Research

Koha (koh-hah) is a Māori practice of reciprocal giving that strengthens the bond between giver and receiver. This research is open access under Creative Commons BY 4.0. If it is useful to you, a koha helps the work carry on.

All research, documentation, and code remain freely available regardless of contribution. Koha is not payment — it is participation in whanaungatanga (relationship-building) and manaakitanga (reciprocal care).

Offer Koha →