Democratic AI°

Democratic AI° — a proposal

An AI system is Democratic AI° when the people and communities it affects can govern it, contest it, refuse it, and retain authority over the knowledge it is built from.

The mark ° is a degree sign — because the four conditions below measure a position on a scale, not a badge a system either has or lacks.

It is there because the phrase already belongs to somebody else.

The measure — four questions

  1. Govern — can the people it decides about govern it? Not be consulted. Govern it.
  2. Contest — can a decision be challenged, and something actually change?
  3. Refuse — is declining available, and survivable?
  4. Retain authority — does the knowledge it was built from stay with those it belongs to?

It returns a position, not a verdict: two of four, and here is which two.

Almost nothing scores four

Not the frontier models. Not the open-weight ones. Not most public-sector deployments.

Not ours either — we score three.

A standard nothing meets is a slogan. A scale everything can be placed on is something an agency can put in a tender.

Whose phrase this is

OpenAI, to the White House, March 2025. Two section headings:

Advancing democratic AI

Export Controls: Exporting Democratic AI

With country tiers. Tier I: countries that commit to democratic AI principles.

The same submission asks for the world to be built on “democratic rails” — and on “American rails”.

Used interchangeably. That is the whole point.

And this is not an argument about America

The move — our systems are the democratic ones, here are the tiers — is available to anyone with a model industry and a government that wants leverage.

Meanwhile the open-weight models most organisations actually deploy — the ones you reach for to avoid depending on a US vendor — are increasingly Chinese.

We run our own community system on a Chinese base model. It fails the fourth question for exactly the same reason an American one would.

The question is not whose flag is on the system. It is whether anyone subject to it can say no.

Many models are a defence against one kind of concentration

They are not the defence they get used as.

  • Competition constrains vendor power
  • Prices fall, access widens
  • Open weights allow local deployment and inspection
  • No single firm decides who may use these systems

All true. None of it is about knowledge.

The move that does not follow

From a claim about market structure

to reassurance about knowledge, culture and public language.

Counting models tells you about the market. It tells you nothing about whether they know different things.

Many suppliers of similar products are not many products

Firms can differ in price, interface, size, latency, safety style and benchmark score while still:

  • building within a narrow family of architectures
  • drawing on heavily overlapping data
  • tuning toward the same idea of a good answer
  • competing on one evaluation ecosystem

Four firms selling near-identical cars are not diversity of transport.

What gets through, and what drops out

The selection cascade, and what returns to it Human plurality passes through a series of narrowing stages before it reaches a model's output: only some of it is written down, only some of that is crawlable or licensed, only some of that survives filtering, and preference tuning then rewards what is legible to a generic idea of helpfulness. Model output re-enters the corpus as people reuse it, which is why the shape is a loop rather than a simple funnel. The important caveat is that a diagram like this can be drawn for print publishing or broadcast, both of which narrowed harder at every stage in fewer languages, and neither produced a monoculture of mind. Selection alone therefore proves nothing. What matters is whether this cascade narrows along dimensions the earlier ones did not, and how fast — which is a question for measurement rather than for a diagram. WHAT GETS THROUGH WHAT DROPS OUT Human plurality Written down at all oral, relational Digitised archives, letters Crawlable, licensed restricted, unconsented Survives filtering minority languages Preference-tuned what resists summary Model output reuse re-enters the corpus WHAT THIS DIAGRAM DOES NOT PROVE The same shape can be drawn for print publishing, which narrowed harder at every stage, in fewer languages, admitting far fewer items. It produced no monoculture of mind. Selection is not convergence — which is why the argument has to rest on measurement, not on this.
Every stage excludes something, and output re-enters at the top. But the same cascade describes print and broadcast, both narrower still. The diagram states the mechanism; it cannot carry the conclusion.

But that diagram proves less than it looks

Draw it for print publishing:

literacy → a publisher → an acquiring editor → house style → a distributor

Narrower at every stage. Fewer languages. Orders of magnitude fewer items.

Nobody thinks twentieth-century print produced a monoculture of mind.

So selection is not convergence. The argument has to rest on measurement.

What is actually measured

What is measured, and what is not Six findings, sorted into three tiers by how much weight each can bear. Measured and peer-reviewed: correlated errors across 349 and 71 models, Kim and colleagues at ICML 2025, at 1.8 to 3.3 times chance; co-writing reducing diversity between 38 writers, Padmakumar and He at ICLR 2024, and only for feedback-tuned models; and AI suggestions shifting 118 participants toward Western styles, Agarwal and colleagues at CHI 2025. Measured but not peer-reviewed: compute concentration, from Stanford's 2026 AI Index, which is an institutional annual report. Suggestive but awaiting replication: a preprint on shrinking linguistic diversity under LLM assistance. Not established: the population-level narrowing of public language — the one naturalistic corpus study found no decrease in lexical diversity, and its authors state their metrics were probably unsuited to detecting the effect at that scale. The mechanism is demonstrated in controlled settings; the population effect is plausible and not yet shown. MEASURED, PEER-REVIEWED Models make correlated errors HELM 71 + HuggingFace 349 models · ICML 2025 · 1.8× and 3.3× chance Co-writing reduces diversity between writers 38 writers, 300 essays · ICLR 2024 · tuned models only Suggestions shift writing toward Western styles 118 participants, India and US · CHI 2025 MEASURED, INSTITUTIONAL REPORT Compute is concentrated beneath the plural layer Stanford AI Index 2026 · not peer-reviewed SUGGESTIVE, AWAITING REPLICATION Linguistic variation shrinking under LLM assistance Four studies · PREPRINT · not peer-reviewed NOT ESTABLISHED Public language is narrowing at population scale News corpora 2018 vs 2024: no decrease — and the authors doubt their metrics Shown in the lab. At population scale: plausible, and not yet demonstrated.
Ranked by what each can bear. The strongest evidence concerns errors and infrastructure; the cultural claim that gets made loudest is the one with the least behind it.

The strongest finding, and the most misquoted

The correlated-error finding, with its baseline Kim and colleagues, ICML 2025, ran two separate tests of how often two language models pick the same wrong answer when both are wrong. On the HELM dataset — 71 models, four options per question — chance agreement is 33 per cent and measured agreement was 60 per cent, an effect of 1.8 times chance. On the HuggingFace leaderboard dataset — 349 models, more options per question — chance agreement is 12.7 per cent and measured agreement was 42.3 per cent, an effect of 3.3 times chance. The 60 per cent is the figure that circulates, and it circulates without its denominator: it reads as near-total agreement only if the baseline is assumed to be near zero. The smaller headline number is the stronger finding. Two results from the same paper matter more than either figure: distinct architectures and distinct providers did not decorrelate the errors, and correlation rises as models become more accurate. WHEN TWO MODELS ARE BOTH WRONG, HOW OFTEN DO THEY PICK THE SAME WRONG ANSWER? HELM · 4 OPTIONS · 71 MODELS by chance: 33% measured: 60% 1.8× HUGGINGFACE · MORE OPTIONS · 349 MODELS by chance: 12.7% measured: 42.3% 3.3× WHAT THE HEADLINE NUMBER HIDES 60% reads as near-total agreement only if you assume the baseline is near zero. It is 33%. The smaller figure is the stronger finding. AND TWO THINGS THAT MATTER MORE • Distinct architectures and providers did NOT decorrelate the errors • Correlation rises as models get more accurate
Kim et al., ICML 2025 — two datasets, 71 and 349 models. The effect is real at 1.8× chance on HELM and 3.3× on the HuggingFace leaderboard. It is also routinely quoted without the denominator that makes it interpretable.

Two things that matter more than the headline

Distinct architectures and distinct providers did not decorrelate the errors.

Correlation rises as models get more accurate.

If shared error were a symptom of immaturity, it would fall as systems improve.

It does not.

Plural where you look

Plural where you look, concentrated where you do not The visible layer of the AI market is genuinely plural: many assistants and interfaces from many firms. Beneath it, according to Stanford's 2026 AI Index, industry produced over 90 per cent of notable models in 2025, Nvidia accounts for over 60 per cent of global AI compute, and almost every leading AI chip is fabricated by a single company. Diversity at the layer people check sits on top of concentration at the layer that decides who can build anything at all. THE LAYER PEOPLE CHECK Assistants and interfaces Many firms, many products, real competition THE LAYERS BENEATH IT Notable models Over 90% produced by industry in 2025 Compute Nvidia: over 60% of the global total Fabrication Essentially one company Plurality where it is visible can sit on concentration where it is not — and the visible layer is the one people check.
Figures from Stanford's 2026 AI Index, reporting 2025 data. Competition at the point of sale rests on a markedly narrower base.

The population effect is not established

The one naturalistic study — 30,000 news articles from 2018 against 30,000 from 2024:

  • lexical diversity did not fall — one measure rose
  • AI-associated vocabulary did rise, significantly

And the authors disown their own instruments:

“we suspect the lexical diversity methods we applied are inappropriate for revealing a loss of lexical diversity on the scale of a very large text corpus”

So the claim has to be conditional

The mechanism is demonstrated in controlled settings.

The population-scale effect is plausible, uneven, and not yet shown.

Anyone telling you otherwise — in either direction — is ahead of the evidence.

Several models are not several sources

Ask three systems. They agree. That feels like triangulation.

It is not, if they share:

  • overlapping web-derived material
  • common benchmark incentives
  • related tuning norms
  • errors already copied across the public web
  • distillation relationships with one another

Agreement among models is a reason to inspect sources, not a substitute for sources.

Two kinds of bias

A government having one is not the problem.

It was in a manifesto. People voted on it. In three years they can vote it out.

That is democracy working.

The vendor’s was in no manifesto

Nobody voted on it. It cannot be removed at an election.

It survives the change of government.

An incoming administration wins a mandate to do things differently — and inherits a delivery layer whose priors were set elsewhere. Not sabotage. Defaults nobody re-examined.

The opportunity is bigger than the risk

Ideology in enactment was always there — discretion, caseworker culture, whose file gets read first on a Friday afternoon.

What was missing was any way to see it.

A system leaves a trace. So for the first time:

Is enactment doing what the statute said?

has an answer you can compute rather than infer.

But only if you can see the priors

You cannot see the priors in a system you do not govern, cannot interrogate, cannot refuse, and whose evidence base belongs to someone else.

Authority over the delivery layer is the precondition for evidence-based government — not a separate concern from it.

Get it wrong and it does the opposite

“The system determined you are not eligible” sounds like arithmetic.

It is a policy. With a threshold somebody chose. Wearing the authority of a calculation.

Far harder to appeal than a decision signed by a person — because there is nobody to argue with, and nothing that looks like a judgement to challenge.

What this does to public language

One register becomes the standard by which speech is judged.

Then two things happen at once:

  • A weakly supported claim can be made to sound finished
  • Someone speaking from direct experience can sound less authoritative

The surface markers of competence are now cheaper than the research that ought to underwrite them.

Compression removes provenance before content

A model can keep the proposition and delete:

  • who is speaking, and what they inherited
  • what obligations constrain them
  • what place and history make it meaningful
  • whether the knowledge was offered, entrusted — or taken

A similarity score reports that the meaning survived. The accountability has not.

And this part is not ours to claim

Māori data governance states it more precisely than we do — and has an instrument written for AI rather than retrofitted to it:

  • Te Kāhui Raraunga — Māori Data Governance Model, and the Māori AI Governance Framework
  • CARE Principles — Global Indigenous Data Alliance
  • WAI 262 — a Treaty claim and Tribunal report. Not a data governance framework, and not cited as one

“AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data.”

— Māori AI Governance Framework

Two consequences:

  1. What counts as restricted is not a procurement decision
  2. Plurality does not fix this. A hundred different models trained without consent commit the error a hundred times

The best case against this argument

Convergence may be a phase. Differentiation becomes valuable as easy gains run out. Open weights may be exactly what lets different systems emerge.

There is evidence: cultural alignment improves with a better language mixture.

Two things stop it settling the question:

  • Specialisation is not plurality of mind
  • Timing — the convergent phase is the one entering schools and public services now

What would falsify this

  • marker sets stop generalising across vendors
  • differently-trained models diverge in register and error structure, not just capability
  • leaderboards fragment into incommensurable regimes
  • multi-model systems show low correlated error against independent ground truth
  • corpora show no homogenisation, measured with instruments their authors stand behind
  • community-governed systems produce durable measurable difference

One of these is close to satisfied.

The test worth running on Monday

Three vendors, one blind spot A public agency buys three AI systems from three different vendors to avoid lock-in, and believes it has redundancy. But because the systems share overlapping training data, similar tuning objectives, a common evaluation culture and distillation relationships with one another, their errors overlap far more than chance predicts — so the same blind spot reaches the decision through all three routes. Multi-vendor procurement without diversity of data, evaluation and governance is redundancy in name only. The test that would reveal this is to give each candidate system the same held-out set of hard cases from the actual domain, record which items each one gets wrong rather than only how many, and measure whether the errors overlap more than chance would predict. WHAT THE AGENCY THINKS IT BOUGHT Vendor A Vendor B Vendor C Three suppliers. No lock-in. Redundancy. WHAT THEY SHARE overlapping web-derived training data · similar tuning goals one benchmark culture · distillation relationships · the same chips Measured: distinct architectures and providers did not decorrelate errors. One blind spot, three times over THE TENDER CLAUSE THAT WOULD CATCH IT Same held-out cases. Record WHICH items fail, not how many.
Buying three systems is not buying three judgements. The test is cheap: give each the same hard cases from your own domain, record which items each gets wrong, and check whether the failures overlap more than chance predicts.

Five steps, one afternoon

  1. Same held-out hard cases from your real domain, to every candidate
  2. Record which items each gets wrong, not how many
  3. Measure whether failures overlap more than chance predicts
  4. Make decorrelation a tender condition, not branding
  5. Re-run at renewal — correlation rises as systems improve

Three vendors is not three judgements.

And not this

The answer is not to ask individuals to write oddly, refuse assistance, or perform authenticity through roughness.

That converts an institutional problem into an aesthetic obligation imposed on ordinary people.

It is unfair, and it does not work.

What survives

Many models are a defence against monopoly.

They are not, by themselves, a defence against:

  • shared error — measured
  • shared infrastructure — measured
  • shared evaluative assumptions — visible
  • a narrowing public register — plausible, not shown
  • lost provenance when situated knowledge becomes generic answer text

The question was never whether there will be one AI

There will be many.

The question is whether they leave room for more than one world.

agenticgovernance.digital/democratic-ai