A proposal — and a mark, because the phrase already belongs to somebody else.
Ask four questions of any system, and the answers give you its degree.
That is the whole instrument. It returns a position, not a verdict: two of four, and here is which two.
Stated as a category it would be useless. Almost nothing in service today answers yes to all four — not the frontier models, not the open-weight ones, not most public-sector deployments, and not, if the questions are asked properly, systems built by people who would describe themselves as democratic. A standard nothing meets is a slogan. A scale everything can be placed on is something an agency can put in a tender.
And the thing the measure exists to keep out:
Democratic is not the same as democratically produced. A system built in a democracy, sold by firms headquartered in democracies, and exported to allied democracies can still answer no to all four questions. Whose flag it flies under is not the measure. Whether you can say no to it is.
Anyone who has met the phrase elsewhere has met it meaning something else.
In March 2025 OpenAI made a submission to the White House Office of Science and Technology Policy on the forthcoming AI Action Plan. It contains a section headed Advancing democratic AI, and another headed Export Controls: Exporting Democratic AI. It proposes country tiers — Tier I: countries that commit to democratic AI principles — and asks that the world be built on what it calls, in the same document, both “democratic rails” and “American rails”.
Those two phrases, used interchangeably by the same author in the same submission, are the whole point. In that usage, democratic is not a property a system has. It is a description of where it was made and whose terms it ships on.
It is not the only occupant. Google DeepMind published a system called Democratic AI in Nature Human Behaviour in 2022 — a pipeline in which reinforcement learning designs a redistribution mechanism that “humans prefer by majority”, and which “successfully won the majority vote”. Elsewhere the phrase means participatory input into a single model, or simply that a tool is cheap and widely available.
Four meanings, then. A supply chain with a flag on it. A majority-preference selector. A consultation exercise. A price point.
This essay uses the phrase in none of those senses. That is what the mark is for.
This is not an argument about America. That submission is quoted because it is the clearest written instance of the move — our systems are the democratic ones, here are the tiers — not because the move belongs to one country. It is available to any bloc with a model industry.
The deployment runs the other way from the rhetoric. The open-weight models an organisation reaches for in order to avoid depending on a US vendor are now substantially Chinese — DeepSeek, Qwen and their descendants. This essay’s authors run their own community system on a Chinese base model, and it fails the fourth condition below for precisely the same reason an American one would. The question is not whose flag is on the system. It is whether anyone subject to it can say no.
There is a reassuring argument in wide circulation, put most clearly by Ramez Naam: there will not be one AI to rule them all, the frontier keeps changing hands, open-weight models are closing the gap — so the future is plural, no single firm will own it, and that is good for human freedom.
On the question it answers, that is substantially right, and worth defending against incumbents who would like regulation to become a barrier to entry.
What plurality genuinely buys:
Nothing in this proposal takes any of it away.
But the argument has started doing work it cannot do. It answers a question about market structure and gets used as reassurance about knowledge, culture and public language. Those are different claims, and the second does not follow.
Counting models tells you about the market. It tells you nothing about whether they know different things.
Several firms can differ in price, interface, size, latency, safety style and benchmark score while still building within a narrow family of architectures, drawing on heavily overlapping data, tuning toward the same idea of a good answer, and competing on one evaluation ecosystem. Many suppliers of similar products are not many products.
Five findings, ranked by how much weight each can bear. One of them cuts against this argument.
1. Models make correlated errors — and not because they are the same. The strongest of the five, and the most misquoted.
Kim and colleagues (ICML 2025) ran two separate tests. Conditional on both models being wrong, how often did they pick the same wrong answer?
| Dataset | Models | Questions | Chance | Measured | Effect |
|---|---|---|---|---|---|
| HELM | 71 | 12,032 | 33% | 60% | 1.8× |
| HuggingFace leaderboard | 349 | 14,402 | 12.7% | 42.3% | 3.3× |
The 60% is the figure that circulates, and it circulates without its denominator. HELM questions have four options, so three are wrong and two independent guessers would agree a third of the time by chance. The finding is 1.8 times chance, not near-total agreement. The HuggingFace set has more options, which is why chance there is 12.7% — and why 42.3% is the stronger result at 3.3×. Quoting the bigger number is quoting the weaker finding.
Two things matter more than either headline:
If shared error were a symptom of an immature field, it would fall as systems improve. It does not.
2. Compute is concentrated beneath the plural layer. Stanford’s 2026 AI Index reports industry produced over 90% of notable models in 2025, global AI compute grew 3.3× a year since 2022, and Nvidia supplies over 60% of the accelerators behind that growth. The leading-edge chips come from a single fabricator.
3. Co-writing with a model reduces diversity between writers. Padmakumar and He (ICLR 2024): 38 experienced writers, 300 essays. A feedback-tuned model significantly reduced lexical and content diversity and made different authors’ essays more similar — and the effect traced to the model’s contributions, not the writers’. The base model produced no significant effect on this sample, which points at tuned assistants rather than at language models as such.
4. AI suggestions shift writing toward Western styles. Agarwal, Naaman and Vashistha (CHI 2025), 118 participants across India and the United States. Indian participants got less benefit despite relying on suggestions more; cross-cultural similarity rose from 0.48 to 0.54; a classifier’s ability to tell Indian from American authorship fell seven points. The authors note their participants were crowdworkers, likely more familiar with American norms than the general population — so the real gap may be wider.
5. The population-level cultural effect is not established. A study comparing roughly 30,000 English news articles from 2018 with 30,000 from 2024 found no decrease in lexical diversity — one measure went up — while AI-associated vocabulary rose significantly.
Its authors then disown their own instruments, saying their methods were probably unsuited to detecting homogenisation at corpus scale. That is a measurement failure rather than a result in either direction.
The mechanism is demonstrated in controlled settings. The population-scale effect is plausible, uneven, and not yet shown. Anyone telling you otherwise — in either direction — is ahead of the evidence.
Something is excluded or transformed at every stage: knowledge that is not digitised, crawlable, licensed or in a dominant language; knowledge whose custodians did not consent to extraction, or whose conditions of transmission were stripped in the taking; oral, relational and place-specific knowledge that fits badly into pipelines built to treat information as transferable text.
**A cascade like this can be drawn for any medium that mattered. Print publishing narrows harder at every stage, in fewer languages, admitting orders of magnitude fewer items — and nobody thinks it produced a monoculture of mind. Selection is not convergence. The argument rests on the measurements above, not on the diagram.
The practical failure is false corroboration. Ask three systems, get one answer, and it feels like triangulation. It is not, if they share overlapping training data, benchmark incentives, tuning norms, errors already copied across the web, and distillation relationships with each other.
Agreement among models is a reason to inspect sources, not a substitute for sources.
And the political version of the same point, which is why this is a proposal rather than an observation:
Authoritarian ordering can be distributed. A system that cannot be contested by the people it represents, refused by those it governs, or made to yield authority over their own knowledge is compatible with authoritarian politics even when several private companies operate it, in a democracy, under competition law. Markets do not supply consent, recourse, or the right to determine how your language and knowledge are represented. Those have to be built.
That is the argument for a standard, and it requires no assumption about anybody’s motives.
A government having an ideological position is not the problem. It was in a manifesto, it was argued about, people voted on it, and in three years they can vote it out. That bias runs through a chain of accountability ending in someone who can be removed. It is not a flaw in democracy — it is democracy. A public service enacting nobody’s politics would be a public service answering to nobody.
The vendor’s position has none of those properties. It was in no manifesto. Nobody voted on it. An election does not touch it: a new minister inherits the contract, and re-letting it takes longer than a term. Which is the part that should concern a permanent secretary more than a minister:
It survives the change of government.
An incoming administration wins a mandate to do things differently and inherits a delivery layer whose priors were set elsewhere, by people pursuing another purpose. Not sabotage. Defaults nobody re-examined. The democratic mechanism for changing direction runs straight through the one layer that does not change.
Ideology in enactment is not new. It has always been there — in discretion, in caseworker culture, in which precedent gets cited, in whose file gets read first on a Friday afternoon. What was missing was any way to see it. You cannot audit a thousand officials’ priors, and the gap between what Parliament legislated and what citizens actually experienced has been hard to measure at anything beyond sample scale.
A system leaves a trace. Inputs, the rule that fired, the threshold, a distribution of outcomes you can slice by region, language, income. For the first time the question is enactment doing what the statute said? can be answered across a whole caseload rather than sampled.
That is the shift worth wanting. Not removing judgement from government — making it visible, and therefore answerable. Policy stays contestable and political, exactly as it should be. Enactment becomes evidence-based, because it becomes inspectable.
And the two halves are one argument. You cannot audit priors you cannot see. You cannot see priors in a system you do not govern, cannot interrogate, cannot refuse, and whose evidence base belongs to somebody else. Authority over the delivery layer is not a separate concern from evidence-based government. It is the precondition for it.
The same technology is unusually good at making a political choice look like a fact.
“The system determined you are not eligible” sounds like arithmetic. It is a policy, with a threshold somebody chose, wearing the authority of a calculation — and far harder to appeal than a decision signed by a person, because there is nobody to argue with and nothing that resembles a judgement to challenge.
So AI does not push government toward evidence or toward ideology. It amplifies whichever one you build for. The four questions are the difference between the two:
Seven layers. Independence has to be built at each rather than hoped for — and the reason it is seven rather than one is that fixing a single layer while leaving the others alone produces the appearance of plurality without the substance.
This layer is furthest along, and not in the places usually credited with leadership on AI governance.
Māori data governance has gone further on this question than anywhere else I have found, and it has an instrument written for AI specifically rather than retrofitted to it.
Te Kāhui Raraunga, acting on behalf of the Data Iwi Leaders Group, publishes the Māori Data Governance Model (Tuia te korowai o Hine-Raraunga) and, extending it, a Māori AI Governance Framework. Its core statement is more direct than anything in this essay:
“AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data.”
Alongside it, the CARE Principles for Indigenous Data Governance, published by the Global Indigenous Data Alliance, set out Collective benefit, Authority to control, Responsibility and Ethics — written deliberately as a complement to the FAIR principles of open data, because findable and reusable are not the same as rightly held.
⚠️ A note on which framework, and why it matters here. The 2016–2018 principles of Te Mana Raraunga — the Māori Data Sovereignty Network, which remains active as an advocacy body — are widely cited, and Dr Karaitiana Taiuru’s September 2025 critical analysis found they do not address AI, model training, algorithmic discrimination or digital colonialism, having been scoped in 2016 to government statistics and research repositories. An AI-governance document citing them as the operative framework in 2026 would be a year behind the published critique. This essay follows Te Kāhui Raraunga’s AI framework for that reason.
These instruments remain, outside narrowly-defined public-service contexts, closer to stated principle than to demonstrated practice. That is a gap in implementation, not in the thinking — and the thinking is what this proposal is short of.
This is not an illustration. It answers the question the rest of the field is still circling: not what data may be used, but who decides.
The position is that knowledge cannot always be separated from its provenance, its conditions of use, the collective responsibilities attaching to it, and the authority required to transmit it. A pipeline that treats knowledge as universally extractable content — and successful paraphrase as preservation — commits a category error before it commits a factual one.
That is a claim about what a statement is. A model can preserve a proposition and delete who was entitled to say it, what obligations constrained them, and what would settle a disagreement about it. A similarity score reports that the meaning survived. The accountability that made it meaningful has not.
In Aotearoa this stands on constitutional ground rather than in an emerging debate: WAI 262 — a Treaty of Waitangi claim lodged in 1991 and the Waitangi Tribunal’s report Ko Aotearoa Tēnei (2011) — concerns taonga species, mātauranga Māori and intellectual property. It is not a data governance framework and is not cited here as one. It is why these questions have an answer with legal weight behind them here, and a discussion paper elsewhere.
And it is why the mark on this essay’s title matters. Indigenous data sovereignty is not a majoritarian claim. It exists in part because majorities overrode it. Any use of the word democratic that collapses into what most people voted for would file tino rangatiratanga under the very thing it stands against — which is precisely what the fourth question — retain authority — is there to prevent. Authority to control is not a preference to be aggregated, and a measure that scored it as one would be measuring the wrong thing.
Two things follow from that:
What the layer requires: community-governed corpora; consent and provenance controls that travel with the data; support for minority and Indigenous languages; and the use of instruments that already exist — datasheets for datasets, data statements, kaitiakitanga licences — rather than their reinvention.
| Layer | What it requires |
|---|---|
| Models | More than parameter variation: materially different corpora, objectives, languages and institutional purposes. Note the evidence — different architectures and providers alone did not decorrelate errors |
| Evaluation | Multiple frameworks including uncertainty disclosure, cultural appropriateness and provenance preservation. With the hazard named: a plural benchmark culture everyone adopts becomes a single benchmark culture again |
| Infrastructure | Local and regional compute; alternatives to a small number of chip and cloud chokepoints. Slowest to shift — the AI Index figures suggest a decade rather than a budget cycle |
| Institutions | Procurement that tests correlated failure; public records of model use; verification against independent sources |
| Interfaces | Retrieved text rendered visibly differently from generated text; visible uncertainty; friction before an unsupported claim is reproduced as fact |
| Public culture | Disclosure where AI materially shaped a public claim; continued recognition of first-person, local, nonstandard and unfinished speech |
And the answer is not to ask individuals to write oddly, refuse assistance, or perform authenticity through roughness. That converts an institutional problem into an aesthetic obligation imposed on ordinary people. It is unfair and it does not work.
Most of the above is addressed to states, standards bodies and frontier labs. One thing is not.
A public agency can avoid vendor lock-in by buying three systems and still reproduce one blind spot through all three. Multi-vendor procurement without diversity of data, evaluation and governance is redundancy in name only.
Make it a tender condition — call it the shared blind spot test:
It is a small piece of work set against the size of the contract, it follows from the strongest of the measurements above, and it tests the thing that matters rather than the thing that is easy to count.
It weakens if:
That fifth item is close to satisfied — the authors of the one naturalistic study available are the ones who disowned their instruments.
And the proposal itself, separately from the diagnosis, is wrong if:
The case against a single controlling system is sound. Competition lowers prices, constrains vendor power, widens access and makes regulatory capture harder. Open weights materially expand the room for local and sovereign deployment.
What does not survive is plurality as reassurance.
Many models are a defence against monopoly. They are not, by themselves, a defence against shared error, shared infrastructure, shared evaluative assumptions, a narrowing public register, or the loss of provenance when situated knowledge becomes generic answer text.
The question was never whether there will be one AI. There will be many.
The question is whether they leave room for more than one world — and whether the people living in those worlds have any say in the systems that speak for them.
That is what the mark is for.
Alongside: questions and answers · glossary · sources and evidence · slides · The Marks It Leaves