Democratic AI°

A proposal — and a mark, because the phrase already belongs to somebody else.

A dense meadow of wildflowers in many colours and forms growing together, none dominant.
Plurality of kind rather than of count. Many things in one place, none of them a version of the others.

The measure

Ask four questions of any system, and the answers give you its degree.

  1. Govern — can the people it decides about govern it? Not be consulted about it. Govern it.
  2. Contest — is there a route by which a decision can be challenged, and something actually changes?
  3. Refuse — is declining available, and survivable?
  4. Retain authority — does the knowledge it was built from stay under the control of those it belongs to, on their terms?

That is the whole instrument. It returns a position, not a verdict: two of four, and here is which two.

Stated as a category it would be useless. Almost nothing in service today answers yes to all four — not the frontier models, not the open-weight ones, not most public-sector deployments, and not, if the questions are asked properly, systems built by people who would describe themselves as democratic. A standard nothing meets is a slogan. A scale everything can be placed on is something an agency can put in a tender.

And the thing the measure exists to keep out:

Democratic is not the same as democratically produced. A system built in a democracy, sold by firms headquartered in democracies, and exported to allied democracies can still answer no to all four questions. Whose flag it flies under is not the measure. Whether you can say no to it is.

Anyone who has met the phrase elsewhere has met it meaning something else.

A note on the name, since it is not free ground

In March 2025 OpenAI made a submission to the White House Office of Science and Technology Policy on the forthcoming AI Action Plan. It contains a section headed Advancing democratic AI, and another headed Export Controls: Exporting Democratic AI. It proposes country tiers — Tier I: countries that commit to democratic AI principles — and asks that the world be built on what it calls, in the same document, both “democratic rails” and “American rails”.

Those two phrases, used interchangeably by the same author in the same submission, are the whole point. In that usage, democratic is not a property a system has. It is a description of where it was made and whose terms it ships on.

It is not the only occupant. Google DeepMind published a system called Democratic AI in Nature Human Behaviour in 2022 — a pipeline in which reinforcement learning designs a redistribution mechanism that “humans prefer by majority”, and which “successfully won the majority vote”. Elsewhere the phrase means participatory input into a single model, or simply that a tool is cheap and widely available.

Four meanings, then. A supply chain with a flag on it. A majority-preference selector. A consultation exercise. A price point.

This essay uses the phrase in none of those senses. That is what the mark is for.

This is not an argument about America. That submission is quoted because it is the clearest written instance of the move — our systems are the democratic ones, here are the tiers — not because the move belongs to one country. It is available to any bloc with a model industry.

The deployment runs the other way from the rhetoric. The open-weight models an organisation reaches for in order to avoid depending on a US vendor are now substantially Chinese — DeepSeek, Qwen and their descendants. This essay’s authors run their own community system on a Chinese base model, and it fails the fourth condition below for precisely the same reason an American one would. The question is not whose flag is on the system. It is whether anyone subject to it can say no.

A wide field of ripening wheat stretching to a distant treeline under a broad sky.
The word monoculture was borrowed from agriculture for a reason. A field like this is productive, uniform, and shares one vulnerability across every plant in it.

Why the number of models does not settle it

There is a reassuring argument in wide circulation, put most clearly by Ramez Naam: there will not be one AI to rule them all, the frontier keeps changing hands, open-weight models are closing the gap — so the future is plural, no single firm will own it, and that is good for human freedom.

On the question it answers, that is substantially right, and worth defending against incumbents who would like regulation to become a barrier to entry.

What plurality genuinely buys:

Nothing in this proposal takes any of it away.

But the argument has started doing work it cannot do. It answers a question about market structure and gets used as reassurance about knowledge, culture and public language. Those are different claims, and the second does not follow.

Counting models tells you about the market. It tells you nothing about whether they know different things.

Several firms can differ in price, interface, size, latency, safety style and benchmark score while still building within a narrow family of architectures, drawing on heavily overlapping data, tuning toward the same idea of a good answer, and competing on one evaluation ecosystem. Many suppliers of similar products are not many products.

What is actually measured

Five findings, ranked by how much weight each can bear. One of them cuts against this argument.

What is measured, and what is not Six findings, sorted into three tiers by how much weight each can bear. Measured and peer-reviewed: correlated errors across 349 and 71 models, Kim and colleagues at ICML 2025, at 1.8 to 3.3 times chance; co-writing reducing diversity between 38 writers, Padmakumar and He at ICLR 2024, and only for feedback-tuned models; and AI suggestions shifting 118 participants toward Western styles, Agarwal and colleagues at CHI 2025. Measured but not peer-reviewed: compute concentration, from Stanford's 2026 AI Index, which is an institutional annual report. Suggestive but awaiting replication: a preprint on shrinking linguistic diversity under LLM assistance. Not established: the population-level narrowing of public language — the one naturalistic corpus study found no decrease in lexical diversity, and its authors state their metrics were probably unsuited to detecting the effect at that scale. The mechanism is demonstrated in controlled settings; the population effect is plausible and not yet shown. MEASURED, PEER-REVIEWED Models make correlated errors HELM 71 + HuggingFace 349 models · ICML 2025 · 1.8× and 3.3× chance Co-writing reduces diversity between writers 38 writers, 300 essays · ICLR 2024 · tuned models only Suggestions shift writing toward Western styles 118 participants, India and US · CHI 2025 MEASURED, INSTITUTIONAL REPORT Compute is concentrated beneath the plural layer Stanford AI Index 2026 · not peer-reviewed SUGGESTIVE, AWAITING REPLICATION Linguistic variation shrinking under LLM assistance Four studies · PREPRINT · not peer-reviewed NOT ESTABLISHED Public language is narrowing at population scale News corpora 2018 vs 2024: no decrease — and the authors doubt their metrics Shown in the lab. At population scale: plausible, and not yet demonstrated.
Ranked by what each can bear. The strongest evidence concerns errors and infrastructure; the cultural claim that gets made loudest is the one with the least behind it.

1. Models make correlated errors — and not because they are the same. The strongest of the five, and the most misquoted.

The correlated-error finding, with its baseline Kim and colleagues, ICML 2025, ran two separate tests of how often two language models pick the same wrong answer when both are wrong. On the HELM dataset — 71 models, four options per question — chance agreement is 33 per cent and measured agreement was 60 per cent, an effect of 1.8 times chance. On the HuggingFace leaderboard dataset — 349 models, more options per question — chance agreement is 12.7 per cent and measured agreement was 42.3 per cent, an effect of 3.3 times chance. The 60 per cent is the figure that circulates, and it circulates without its denominator: it reads as near-total agreement only if the baseline is assumed to be near zero. The smaller headline number is the stronger finding. Two results from the same paper matter more than either figure: distinct architectures and distinct providers did not decorrelate the errors, and correlation rises as models become more accurate. WHEN TWO MODELS ARE BOTH WRONG, HOW OFTEN DO THEY PICK THE SAME WRONG ANSWER? HELM · 4 OPTIONS · 71 MODELS by chance: 33% measured: 60% 1.8× HUGGINGFACE · MORE OPTIONS · 349 MODELS by chance: 12.7% measured: 42.3% 3.3× WHAT THE HEADLINE NUMBER HIDES 60% reads as near-total agreement only if you assume the baseline is near zero. It is 33%. The smaller figure is the stronger finding. AND TWO THINGS THAT MATTER MORE • Distinct architectures and providers did NOT decorrelate the errors • Correlation rises as models get more accurate
Kim et al., ICML 2025 — two datasets, 71 and 349 models. The effect is real at 1.8× chance on HELM and 3.3× on the HuggingFace leaderboard. It is also routinely quoted without the denominator that makes it interpretable.

Kim and colleagues (ICML 2025) ran two separate tests. Conditional on both models being wrong, how often did they pick the same wrong answer?

Dataset Models Questions Chance Measured Effect
HELM 71 12,032 33% 60% 1.8×
HuggingFace leaderboard 349 14,402 12.7% 42.3% 3.3×

The 60% is the figure that circulates, and it circulates without its denominator. HELM questions have four options, so three are wrong and two independent guessers would agree a third of the time by chance. The finding is 1.8 times chance, not near-total agreement. The HuggingFace set has more options, which is why chance there is 12.7% — and why 42.3% is the stronger result at 3.3×. Quoting the bigger number is quoting the weaker finding.

Two things matter more than either headline:

If shared error were a symptom of an immature field, it would fall as systems improve. It does not.

2. Compute is concentrated beneath the plural layer. Stanford’s 2026 AI Index reports industry produced over 90% of notable models in 2025, global AI compute grew 3.3× a year since 2022, and Nvidia supplies over 60% of the accelerators behind that growth. The leading-edge chips come from a single fabricator.

Plural where you look, concentrated where you do not The visible layer of the AI market is genuinely plural: many assistants and interfaces from many firms. Beneath it, according to Stanford's 2026 AI Index, industry produced over 90 per cent of notable models in 2025, Nvidia accounts for over 60 per cent of global AI compute, and almost every leading AI chip is fabricated by a single company. Diversity at the layer people check sits on top of concentration at the layer that decides who can build anything at all. THE LAYER PEOPLE CHECK Assistants and interfaces Many firms, many products, real competition THE LAYERS BENEATH IT Notable models Over 90% produced by industry in 2025 Compute Nvidia: over 60% of the global total Fabrication Essentially one company Plurality where it is visible can sit on concentration where it is not — and the visible layer is the one people check.
Figures from Stanford's 2026 AI Index, reporting 2025 data. Competition at the point of sale rests on a markedly narrower base.

3. Co-writing with a model reduces diversity between writers. Padmakumar and He (ICLR 2024): 38 experienced writers, 300 essays. A feedback-tuned model significantly reduced lexical and content diversity and made different authors’ essays more similar — and the effect traced to the model’s contributions, not the writers’. The base model produced no significant effect on this sample, which points at tuned assistants rather than at language models as such.

4. AI suggestions shift writing toward Western styles. Agarwal, Naaman and Vashistha (CHI 2025), 118 participants across India and the United States. Indian participants got less benefit despite relying on suggestions more; cross-cultural similarity rose from 0.48 to 0.54; a classifier’s ability to tell Indian from American authorship fell seven points. The authors note their participants were crowdworkers, likely more familiar with American norms than the general population — so the real gap may be wider.

5. The population-level cultural effect is not established. A study comparing roughly 30,000 English news articles from 2018 with 30,000 from 2024 found no decrease in lexical diversity — one measure went up — while AI-associated vocabulary rose significantly.

Its authors then disown their own instruments, saying their methods were probably unsuited to detecting homogenisation at corpus scale. That is a measurement failure rather than a result in either direction.

The mechanism is demonstrated in controlled settings. The population-scale effect is plausible, uneven, and not yet shown. Anyone telling you otherwise — in either direction — is ahead of the evidence.

Where the selection happens

The selection cascade, and what returns to it Human plurality passes through a series of narrowing stages before it reaches a model's output: only some of it is written down, only some of that is crawlable or licensed, only some of that survives filtering, and preference tuning then rewards what is legible to a generic idea of helpfulness. Model output re-enters the corpus as people reuse it, which is why the shape is a loop rather than a simple funnel. The important caveat is that a diagram like this can be drawn for print publishing or broadcast, both of which narrowed harder at every stage in fewer languages, and neither produced a monoculture of mind. Selection alone therefore proves nothing. What matters is whether this cascade narrows along dimensions the earlier ones did not, and how fast — which is a question for measurement rather than for a diagram. WHAT GETS THROUGH WHAT DROPS OUT Human plurality Written down at all oral, relational Digitised archives, letters Crawlable, licensed restricted, unconsented Survives filtering minority languages Preference-tuned what resists summary Model output reuse re-enters the corpus WHAT THIS DIAGRAM DOES NOT PROVE The same shape can be drawn for print publishing, which narrowed harder at every stage, in fewer languages, admitting far fewer items. It produced no monoculture of mind. Selection is not convergence — which is why the argument has to rest on measurement, not on this.
Every stage excludes something, and output re-enters at the top. But the same cascade describes print and broadcast, both narrower still. The diagram states the mechanism; it cannot carry the conclusion.

Something is excluded or transformed at every stage: knowledge that is not digitised, crawlable, licensed or in a dominant language; knowledge whose custodians did not consent to extraction, or whose conditions of transmission were stripped in the taking; oral, relational and place-specific knowledge that fits badly into pipelines built to treat information as transferable text.

**A cascade like this can be drawn for any medium that mattered. Print publishing narrows harder at every stage, in fewer languages, admitting orders of magnitude fewer items — and nobody thinks it produced a monoculture of mind. Selection is not convergence. The argument rests on the measurements above, not on the diagram.

What the count cannot tell you

The practical failure is false corroboration. Ask three systems, get one answer, and it feels like triangulation. It is not, if they share overlapping training data, benchmark incentives, tuning norms, errors already copied across the web, and distillation relationships with each other.

Agreement among models is a reason to inspect sources, not a substitute for sources.

And the political version of the same point, which is why this is a proposal rather than an observation:

Authoritarian ordering can be distributed. A system that cannot be contested by the people it represents, refused by those it governs, or made to yield authority over their own knowledge is compatible with authoritarian politics even when several private companies operate it, in a democracy, under competition law. Markets do not supply consent, recourse, or the right to determine how your language and knowledge are represented. Those have to be built.

That is the argument for a standard, and it requires no assumption about anybody’s motives.

Two kinds of bias, and only one of them is legitimate

A government having an ideological position is not the problem. It was in a manifesto, it was argued about, people voted on it, and in three years they can vote it out. That bias runs through a chain of accountability ending in someone who can be removed. It is not a flaw in democracy — it is democracy. A public service enacting nobody’s politics would be a public service answering to nobody.

The vendor’s position has none of those properties. It was in no manifesto. Nobody voted on it. An election does not touch it: a new minister inherits the contract, and re-letting it takes longer than a term. Which is the part that should concern a permanent secretary more than a minister:

It survives the change of government.

An incoming administration wins a mandate to do things differently and inherits a delivery layer whose priors were set elsewhere, by people pursuing another purpose. Not sabotage. Defaults nobody re-examined. The democratic mechanism for changing direction runs straight through the one layer that does not change.

The opportunity, which is larger than the risk

Ideology in enactment is not new. It has always been there — in discretion, in caseworker culture, in which precedent gets cited, in whose file gets read first on a Friday afternoon. What was missing was any way to see it. You cannot audit a thousand officials’ priors, and the gap between what Parliament legislated and what citizens actually experienced has been hard to measure at anything beyond sample scale.

A system leaves a trace. Inputs, the rule that fired, the threshold, a distribution of outcomes you can slice by region, language, income. For the first time the question is enactment doing what the statute said? can be answered across a whole caseload rather than sampled.

That is the shift worth wanting. Not removing judgement from government — making it visible, and therefore answerable. Policy stays contestable and political, exactly as it should be. Enactment becomes evidence-based, because it becomes inspectable.

And the two halves are one argument. You cannot audit priors you cannot see. You cannot see priors in a system you do not govern, cannot interrogate, cannot refuse, and whose evidence base belongs to somebody else. Authority over the delivery layer is not a separate concern from evidence-based government. It is the precondition for it.

Which way it goes is a design decision

The same technology is unusually good at making a political choice look like a fact.

“The system determined you are not eligible” sounds like arithmetic. It is a policy, with a threshold somebody chose, wearing the authority of a calculation — and far harder to appeal than a decision signed by a person, because there is nobody to argue with and nothing that resembles a judgement to challenge.

So AI does not push government toward evidence or toward ideology. It amplifies whichever one you build for. The four questions are the difference between the two:

The specification

Seven layers. Independence has to be built at each rather than hoped for — and the reason it is seven rather than one is that fixing a single layer while leaving the others alone produces the appearance of plurality without the substance.

Data — where the thinking is furthest ahead

This layer is furthest along, and not in the places usually credited with leadership on AI governance.

Māori data governance has gone further on this question than anywhere else I have found, and it has an instrument written for AI specifically rather than retrofitted to it.

Te Kāhui Raraunga, acting on behalf of the Data Iwi Leaders Group, publishes the Māori Data Governance Model (Tuia te korowai o Hine-Raraunga) and, extending it, a Māori AI Governance Framework. Its core statement is more direct than anything in this essay:

“AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data.”

Alongside it, the CARE Principles for Indigenous Data Governance, published by the Global Indigenous Data Alliance, set out Collective benefit, Authority to control, Responsibility and Ethics — written deliberately as a complement to the FAIR principles of open data, because findable and reusable are not the same as rightly held.

⚠️ A note on which framework, and why it matters here. The 2016–2018 principles of Te Mana Raraunga — the Māori Data Sovereignty Network, which remains active as an advocacy body — are widely cited, and Dr Karaitiana Taiuru’s September 2025 critical analysis found they do not address AI, model training, algorithmic discrimination or digital colonialism, having been scoped in 2016 to government statistics and research repositories. An AI-governance document citing them as the operative framework in 2026 would be a year behind the published critique. This essay follows Te Kāhui Raraunga’s AI framework for that reason.

These instruments remain, outside narrowly-defined public-service contexts, closer to stated principle than to demonstrated practice. That is a gap in implementation, not in the thinking — and the thinking is what this proposal is short of.

This is not an illustration. It answers the question the rest of the field is still circling: not what data may be used, but who decides.

The position is that knowledge cannot always be separated from its provenance, its conditions of use, the collective responsibilities attaching to it, and the authority required to transmit it. A pipeline that treats knowledge as universally extractable content — and successful paraphrase as preservation — commits a category error before it commits a factual one.

That is a claim about what a statement is. A model can preserve a proposition and delete who was entitled to say it, what obligations constrained them, and what would settle a disagreement about it. A similarity score reports that the meaning survived. The accountability that made it meaningful has not.

In Aotearoa this stands on constitutional ground rather than in an emerging debate: WAI 262 — a Treaty of Waitangi claim lodged in 1991 and the Waitangi Tribunal’s report Ko Aotearoa Tēnei (2011) — concerns taonga species, mātauranga Māori and intellectual property. It is not a data governance framework and is not cited here as one. It is why these questions have an answer with legal weight behind them here, and a discussion paper elsewhere.

And it is why the mark on this essay’s title matters. Indigenous data sovereignty is not a majoritarian claim. It exists in part because majorities overrode it. Any use of the word democratic that collapses into what most people voted for would file tino rangatiratanga under the very thing it stands against — which is precisely what the fourth question — retain authority — is there to prevent. Authority to control is not a preference to be aggregated, and a measure that scored it as one would be measuring the wrong thing.

Two things follow from that:

  1. This proposal follows those frameworks. It does not extend them, speak for them, or claim their endorsement. They state the point more precisely than this essay does.
  2. Their concern is extraction and consent, which is a different failure from convergence. A hundred genuinely different models, each trained without consent on restricted knowledge, commits that error a hundred times over. Plurality alone does not fix it — which is why a proposal about plurality alone would be insufficient, and why the data layer comes first here.

What the layer requires: community-governed corpora; consent and provenance controls that travel with the data; support for minority and Indigenous languages; and the use of instruments that already exist — datasheets for datasets, data statements, kaitiakitanga licences — rather than their reinvention.

The other six

Layer What it requires
Models More than parameter variation: materially different corpora, objectives, languages and institutional purposes. Note the evidence — different architectures and providers alone did not decorrelate errors
Evaluation Multiple frameworks including uncertainty disclosure, cultural appropriateness and provenance preservation. With the hazard named: a plural benchmark culture everyone adopts becomes a single benchmark culture again
Infrastructure Local and regional compute; alternatives to a small number of chip and cloud chokepoints. Slowest to shift — the AI Index figures suggest a decade rather than a budget cycle
Institutions Procurement that tests correlated failure; public records of model use; verification against independent sources
Interfaces Retrieved text rendered visibly differently from generated text; visible uncertainty; friction before an unsupported claim is reproduced as fact
Public culture Disclosure where AI materially shaped a public claim; continued recognition of first-person, local, nonstandard and unfinished speech

And the answer is not to ask individuals to write oddly, refuse assistance, or perform authenticity through roughness. That converts an institutional problem into an aesthetic obligation imposed on ordinary people. It is unfair and it does not work.

What you can do on Monday

Most of the above is addressed to states, standards bodies and frontier labs. One thing is not.

Three vendors, one blind spot A public agency buys three AI systems from three different vendors to avoid lock-in, and believes it has redundancy. But because the systems share overlapping training data, similar tuning objectives, a common evaluation culture and distillation relationships with one another, their errors overlap far more than chance predicts — so the same blind spot reaches the decision through all three routes. Multi-vendor procurement without diversity of data, evaluation and governance is redundancy in name only. The test that would reveal this is to give each candidate system the same held-out set of hard cases from the actual domain, record which items each one gets wrong rather than only how many, and measure whether the errors overlap more than chance would predict. WHAT THE AGENCY THINKS IT BOUGHT Vendor A Vendor B Vendor C Three suppliers. No lock-in. Redundancy. WHAT THEY SHARE overlapping web-derived training data · similar tuning goals one benchmark culture · distillation relationships · the same chips Measured: distinct architectures and providers did not decorrelate errors. One blind spot, three times over THE TENDER CLAUSE THAT WOULD CATCH IT Same held-out cases. Record WHICH items fail, not how many.
Buying three systems is not buying three judgements. The test is cheap: give each the same hard cases from your own domain, record which items each gets wrong, and check whether the failures overlap more than chance predicts.

A public agency can avoid vendor lock-in by buying three systems and still reproduce one blind spot through all three. Multi-vendor procurement without diversity of data, evaluation and governance is redundancy in name only.

Make it a tender condition — call it the shared blind spot test:

  1. Give every candidate system the same held-out set of hard cases from your real domain
  2. Record which items each one gets wrong, not just how many
  3. Check whether those failures overlap more than chance predicts
  4. Require decorrelation, not branding
  5. Re-run it at renewal, because correlation rises as systems improve

It is a small piece of work set against the size of the contract, it follows from the strongest of the measurements above, and it tests the thing that matters rather than the thing that is easy to count.

What would show this proposal wrong

It weakens if:

That fifth item is close to satisfied — the authors of the one naturalistic study available are the ones who disowned their instruments.

And the proposal itself, separately from the diagnosis, is wrong if:

The boundaries of the claim

A single cabbage tree standing on a ridge against snow-covered mountains and open sky.
One thing, distinctly itself, in a landscape large enough to hold it.

What survives

The case against a single controlling system is sound. Competition lowers prices, constrains vendor power, widens access and makes regulatory capture harder. Open weights materially expand the room for local and sovereign deployment.

What does not survive is plurality as reassurance.

Many models are a defence against monopoly. They are not, by themselves, a defence against shared error, shared infrastructure, shared evaluative assumptions, a narrowing public register, or the loss of provenance when situated knowledge becomes generic answer text.

The question was never whether there will be one AI. There will be many.

The question is whether they leave room for more than one world — and whether the people living in those worlds have any say in the systems that speak for them.

That is what the mark is for.

Alongside: questions and answers · glossary · sources and evidence · slides · The Marks It Leaves