The Marks It Leaves
How much of what you read was built with a machine
Not a rhetorical question. It has an answer.
- LinkedIn long-form — roughly two in five, written entirely by a machine
- New web articles — about half
- Medium, Quora — near two-fifths
- Substack — about a fifth
- Reddit — 2.5% to 11.6%, and replies are 98% human
- Facebook · Instagram — not counted
- WhatsApp · Messenger — cannot be counted from outside
Drafted with AI assistance across several systems. Declared, because this piece argues for declaring it.
Every one of those is the open internet
Look at what that list has in common: an outsider can read all of it.
- Facebook — no measurement found
- Instagram — no measurement found
None found, not none existing — one search, and it could not reach every source. - WhatsApp — cannot be measured from outside
The table is a map of what can be crawled. It has been standing in for a map of where people write.
The closest thing to a Facebook figure
DiResta & Goldstein, peer-reviewed 2024. One unlabelled AI image among the ten most-viewed Facebook posts of Q3 2023 — 40 million views.
But: 125 Pages, found by hand. The authors’ own words —
The Pages we studied are not necessarily reflective of how unlabeled AI-generated images are used on Facebook as a whole.
A real finding about reach. Not a percentage. Any percentage citing it was invented in transit.
What Meta marks — and what it does not
Labelling since February 2024. Every scope statement names the same three things:
image · video · audio
- Generated text is named in none of them
- Enforcement reporting: 26 policy categories, none for synthetic media
- Only published number: 590,000 — requests its own generator refused, to make images of named politicians
On the platforms where most people write, the provenance system does not cover writing.
The law splits at the same seam
EU AI Act, Article 50 — applying from 2 August 2026
| Who | Duty | Text covered? |
|---|---|---|
| Built the model — Art 50(2) | Mark output machine-readable | Yes — “audio, image, video or text” |
| Runs the platform — Art 50(4) | Disclose deep fakes | Barely — “image, audio or video”; text only for public-interest publishing, and not where there was human review or editorial control and a person holds editorial responsibility |
DSA Art 35(1)(k) — one of the risk-mitigation measures, not a flat mandate: image, audio or video. Text absent.
Text is named where the model is made. Not where the platform carries it.
Meta’s own answer, July 2026
As AI makes it easier to generate content, profiles, and messages, we want to provide a way for you to know there is a real person on the other side of a profile.
Facebook Verified — a badge for “a real person — someone who completed selfie verification”.
Those two categories are the ones its labelling never covered.
Marking the machine, and now verifying the person.
A declaration costs nothing. A selfie check costs a name.
WhatsApp is not uncounted. It is uncountable from outside.
End-to-end encryption: no researcher, no regulator, no third party can survey it.
That is the property it is supposed to have.
- Only public groups a researcher joined can be studied — a self-selected sample
- A stated percentage of WhatsApp traffic could not have been produced
- Better detectors do not help. The missing thing is access, not technique
Which is the whole argument, arriving from the other side: inside an encrypted conversation, somebody saying what they did is not the best signal. It is the only one.
The figure to stop quoting
“Experts estimate that as much as 90% of online content may be synthetically generated by 2026.”
Europol, 2022. Follow the footnote: one trade paperback, published 2020.
No measurement was ever performed. The claim in the book was about video.
A claim about pictures, laundered into a statistic about prose.
Machine writing does leave marks
Four kinds, and they are not equally worth having.
| Kind | What it is | Worth |
|---|---|---|
| Provenance | Deliberately embedded — watermarks, credentials | Strongest. Unreadable by you |
| Language | The words themselves | Measured. Trivial to remove |
| Discourse | How the thing is organised | Dearest to remove. Least tested |
| Transfer | Traces of how it travelled | Strong about the route, weak about the author |
Watermarking became real this month
- Google SynthID — running on Gemini since 2024, published in Nature
- Anthropic — announced for Claude, mid-August 2026, applied worldwide
- OpenAI — built one, measured 99.9% internally, did not ship it
What forced it: European law, applying from 2 August 2026.
And you cannot read any of them
The Act requires the mark to be machine-readable. It does not require anyone but the vendor to be able to read it.
- Google’s detector — a waitlist, for Google’s own content
- Anthropic’s — announced, unpublished
- OpenAI’s — nothing deployed
The mark exists. The key does not travel with it.
The strongest measurement used no detector at all
15.1 million biomedical abstracts. Count how often a word appears; compare against how often it should.
- delves — 28× its expected rate
- underscores — nearly 14×
At least 13.5% of 2024 abstracts had been through a model.
“Our analysis is performed on the corpus level and cannot identify individual abstracts.”
Everything cheap to spot is cheap to erase
| Marker | What removes it |
|---|---|
| Excess vocabulary | find and replace |
| Em-dashes | one instruction: 9.09 → 0.19 |
| Markdown scaffolding | one instruction |
| Statistical detection | one paraphrase: 70.3% → 4.6% |
| Vendor watermark | repeat it: 99.3% → 9.7% |
| Noun-heavy density | real rewriting — and never tested |
A sign-list works best on whoever was not hiding anything.
Which is generally the person doing nothing wrong
Why none of it is proof — 1
Signs that share a cause do not add up.
Writes at length · writes formally · repeats themselves · will not concede · sounds angry
That looks like five findings.
It is one person arguing hard, seen from five angles.
Why none of it is proof — 2
Rarity beats accuracy.
If machine-written posts are a modest share of the posts you actually read, a test that catches most of them still returns mostly innocent people.
There are so many more innocent people to catch.
This does not improve much with a better test.
Why none of it is proof — 3
The signs land hardest on the wrong people.
Seven detectors, run over essays by writers using English as a second language:
- 61% flagged as machine-written
- 5% for US eighth-graders
- Rewritten to sound native: under 12%
- Native prose simplified: 57%
The authorship never changed. Only the register did.
What the detectors were measuring
Not machine authorship.
How restricted the writer’s English is.
A sign-list that does not say so becomes a weapon aimed at people writing in a second language.
When several marks appear at once
The tempting move: count the layers, treat four layers as four findings.
It fails. One ordinary cause routinely produces marks in several layers:
- a house style guide — vocabulary, register, template
- a scheduling tool — formatting and paste damage
- a translation workflow — three layers, no AI involved
- accessibility guidance — short sentences, headings, stripped hedging
So ask a better question
Before a cluster means anything, go looking for the one pipeline that would produce all of it.
If one fits, you have found the explanation — and the cluster is one observation wearing several coats.
What a sign is for
Not a verdict. A reason to be careful.
- Verify a claim rather than passing it on
- Look for the source behind a confident summary
- Notice the difference between sounding authoritative and demonstrating expertise
- Stop treating polish and completeness as evidence of reliability
What changes is how you read — not what you conclude about the writer.
Which leaves the thing that does settle it
Someone saying so.
A declaration costs nothing. Needs no key. Survives a paraphrase. Misfires on nobody. Requires no authority to sit in judgement over what is authentic.
And its absence means nothing — most posts made without a machine will carry no note saying so.
Plurality is not the cure it reads as
Many makers means no single company owns the answer. Good, and worth defending.
It says nothing about sameness.
A few hundred words identify machine writing across vendors. If the plurality ran deep, they would mark one maker.
An average over a wider base is still an average.
The harm is not deception
Most machine-written text does not lie.
It sounds like everything else — and crowds out the range of ways a thing can be said.
The odd construction. The paragraph that runs hot. The point made badly by someone who has thought about it for twenty years.
Those carry information about who is speaking. A smooth register deletes them and sounds more reliable for having done so.
This deck, measured against its own argument
The essay runs at 10.2 em-dashes per thousand words.
Above the model that drafted it. More than three times the human baseline it cites.
That human range ran from 0.3 to 17.
You cannot tell those two explanations apart from the page. Neither can any detector.
Say so. Take the answer. Argue with the argument.
It is a low standard.
It is also higher than anything else on offer.
agenticgovernance.digital/machine-marks