Democratic AI°
Democratic AI° — a proposal
An AI system is Democratic AI° when the people and communities it affects can govern it, contest it, refuse it, and retain authority over the knowledge it is built from.
The mark ° is a degree sign — because the four conditions below measure a position on a scale, not a badge a system either has or lacks.
It is there because the phrase already belongs to somebody else.
The measure — four questions
- Govern — can the people it decides about govern it? Not be consulted. Govern it.
- Contest — can a decision be challenged, and something actually change?
- Refuse — is declining available, and survivable?
- Retain authority — does the knowledge it was built from stay with those it belongs to?
It returns a position, not a verdict: two of four, and here is which two.
Almost nothing scores four
Not the frontier models. Not the open-weight ones. Not most public-sector deployments.
Not ours either — we score three.
A standard nothing meets is a slogan. A scale everything can be placed on is something an agency can put in a tender.
Whose phrase this is
OpenAI, to the White House, March 2025. Two section headings:
Advancing democratic AI
Export Controls: Exporting Democratic AI
With country tiers. Tier I: countries that commit to democratic AI principles.
The same submission asks for the world to be built on “democratic rails” — and on “American rails”.
Used interchangeably. That is the whole point.
And this is not an argument about America
The move — our systems are the democratic ones, here are the tiers — is available to anyone with a model industry and a government that wants leverage.
Meanwhile the open-weight models most organisations actually deploy — the ones you reach for to avoid depending on a US vendor — are increasingly Chinese.
We run our own community system on a Chinese base model. It fails the fourth question for exactly the same reason an American one would.
The question is not whose flag is on the system. It is whether anyone subject to it can say no.
Many models are a defence against one kind of concentration
They are not the defence they get used as.
- Competition constrains vendor power
- Prices fall, access widens
- Open weights allow local deployment and inspection
- No single firm decides who may use these systems
All true. None of it is about knowledge.
The move that does not follow
From a claim about market structure
to reassurance about knowledge, culture and public language.
Counting models tells you about the market. It tells you nothing about whether they know different things.
Many suppliers of similar products are not many products
Firms can differ in price, interface, size, latency, safety style and benchmark score while still:
- building within a narrow family of architectures
- drawing on heavily overlapping data
- tuning toward the same idea of a good answer
- competing on one evaluation ecosystem
Four firms selling near-identical cars are not diversity of transport.
What gets through, and what drops out
But that diagram proves less than it looks
Draw it for print publishing:
literacy → a publisher → an acquiring editor → house style → a distributor
Narrower at every stage. Fewer languages. Orders of magnitude fewer items.
Nobody thinks twentieth-century print produced a monoculture of mind.
So selection is not convergence. The argument has to rest on measurement.
What is actually measured
The strongest finding, and the most misquoted
Two things that matter more than the headline
Distinct architectures and distinct providers did not decorrelate the errors.
Correlation rises as models get more accurate.
If shared error were a symptom of immaturity, it would fall as systems improve.
It does not.
Plural where you look
The population effect is not established
The one naturalistic study — 30,000 news articles from 2018 against 30,000 from 2024:
- lexical diversity did not fall — one measure rose
- AI-associated vocabulary did rise, significantly
And the authors disown their own instruments:
“we suspect the lexical diversity methods we applied are inappropriate for revealing a loss of lexical diversity on the scale of a very large text corpus”
So the claim has to be conditional
The mechanism is demonstrated in controlled settings.
The population-scale effect is plausible, uneven, and not yet shown.
Anyone telling you otherwise — in either direction — is ahead of the evidence.
Several models are not several sources
Ask three systems. They agree. That feels like triangulation.
It is not, if they share:
- overlapping web-derived material
- common benchmark incentives
- related tuning norms
- errors already copied across the public web
- distillation relationships with one another
Agreement among models is a reason to inspect sources, not a substitute for sources.
Two kinds of bias
A government having one is not the problem.
It was in a manifesto. People voted on it. In three years they can vote it out.
That is democracy working.
The vendor’s was in no manifesto
Nobody voted on it. It cannot be removed at an election.
It survives the change of government.
An incoming administration wins a mandate to do things differently — and inherits a delivery layer whose priors were set elsewhere. Not sabotage. Defaults nobody re-examined.
The opportunity is bigger than the risk
Ideology in enactment was always there — discretion, caseworker culture, whose file gets read first on a Friday afternoon.
What was missing was any way to see it.
A system leaves a trace. So for the first time:
Is enactment doing what the statute said?
has an answer you can compute rather than infer.
But only if you can see the priors
You cannot see the priors in a system you do not govern, cannot interrogate, cannot refuse, and whose evidence base belongs to someone else.
Authority over the delivery layer is the precondition for evidence-based government — not a separate concern from it.
Get it wrong and it does the opposite
“The system determined you are not eligible” sounds like arithmetic.
It is a policy. With a threshold somebody chose. Wearing the authority of a calculation.
Far harder to appeal than a decision signed by a person — because there is nobody to argue with, and nothing that looks like a judgement to challenge.
What this does to public language
One register becomes the standard by which speech is judged.
Then two things happen at once:
- A weakly supported claim can be made to sound finished
- Someone speaking from direct experience can sound less authoritative
The surface markers of competence are now cheaper than the research that ought to underwrite them.
Compression removes provenance before content
A model can keep the proposition and delete:
- who is speaking, and what they inherited
- what obligations constrain them
- what place and history make it meaningful
- whether the knowledge was offered, entrusted — or taken
A similarity score reports that the meaning survived. The accountability has not.
And this part is not ours to claim
Māori data governance states it more precisely than we do — and has an instrument written for AI rather than retrofitted to it:
- Te Kāhui Raraunga — Māori Data Governance Model, and the Māori AI Governance Framework
- CARE Principles — Global Indigenous Data Alliance
- WAI 262 — a Treaty claim and Tribunal report. Not a data governance framework, and not cited as one
“AI systems must not be implemented in Aotearoa without fully realising Māori authority over Māori data.”
— Māori AI Governance Framework
Two consequences:
- What counts as restricted is not a procurement decision
- Plurality does not fix this. A hundred different models trained without consent commit the error a hundred times
The best case against this argument
Convergence may be a phase. Differentiation becomes valuable as easy gains run out. Open weights may be exactly what lets different systems emerge.
There is evidence: cultural alignment improves with a better language mixture.
Two things stop it settling the question:
- Specialisation is not plurality of mind
- Timing — the convergent phase is the one entering schools and public services now
What would falsify this
- marker sets stop generalising across vendors
- differently-trained models diverge in register and error structure, not just capability
- leaderboards fragment into incommensurable regimes
- multi-model systems show low correlated error against independent ground truth
- corpora show no homogenisation, measured with instruments their authors stand behind
- community-governed systems produce durable measurable difference
One of these is close to satisfied.
The test worth running on Monday
Five steps, one afternoon
- Same held-out hard cases from your real domain, to every candidate
- Record which items each gets wrong, not how many
- Measure whether failures overlap more than chance predicts
- Make decorrelation a tender condition, not branding
- Re-run at renewal — correlation rises as systems improve
Three vendors is not three judgements.
And not this
The answer is not to ask individuals to write oddly, refuse assistance, or perform authenticity through roughness.
That converts an institutional problem into an aesthetic obligation imposed on ordinary people.
It is unfair, and it does not work.
What survives
Many models are a defence against monopoly.
They are not, by themselves, a defence against:
- shared error — measured
- shared infrastructure — measured
- shared evaluative assumptions — visible
- a narrowing public register — plausible, not shown
- lost provenance when situated knowledge becomes generic answer text
The question was never whether there will be one AI
There will be many.
The question is whether they leave room for more than one world.
agenticgovernance.digital/democratic-ai