The Marks It Leaves

How much of what you read was built with a machine

Not a rhetorical question. It has an answer.

  • LinkedIn long-form — roughly two in five, written entirely by a machine
  • New web articles — about half
  • Medium, Quora — near two-fifths
  • Substack — about a fifth
  • Reddit — 2.5% to 11.6%, and replies are 98% human
  • Facebook · Instagramnot counted
  • WhatsApp · Messengercannot be counted from outside

Drafted with AI assistance across several systems. Declared, because this piece argues for declaring it.

Every one of those is the open internet

Look at what that list has in common: an outsider can read all of it.

  • Facebook — no measurement found
  • Instagram — no measurement found

None found, not none existing — one search, and it could not reach every source. - WhatsAppcannot be measured from outside

The table is a map of what can be crawled. It has been standing in for a map of where people write.

The closest thing to a Facebook figure

DiResta & Goldstein, peer-reviewed 2024. One unlabelled AI image among the ten most-viewed Facebook posts of Q3 2023 — 40 million views.

But: 125 Pages, found by hand. The authors’ own words —

The Pages we studied are not necessarily reflective of how unlabeled AI-generated images are used on Facebook as a whole.

A real finding about reach. Not a percentage. Any percentage citing it was invented in transit.

What Meta marks — and what it does not

Labelling since February 2024. Every scope statement names the same three things:

image · video · audio

  • Generated text is named in none of them
  • Enforcement reporting: 26 policy categories, none for synthetic media
  • Only published number: 590,000 — requests its own generator refused, to make images of named politicians

On the platforms where most people write, the provenance system does not cover writing.

The law splits at the same seam

EU AI Act, Article 50 — applying from 2 August 2026

Who Duty Text covered?
Built the model — Art 50(2) Mark output machine-readable Yes — “audio, image, video or text”
Runs the platform — Art 50(4) Disclose deep fakes Barely — “image, audio or video”; text only for public-interest publishing, and not where there was human review or editorial control and a person holds editorial responsibility

DSA Art 35(1)(k) — one of the risk-mitigation measures, not a flat mandate: image, audio or video. Text absent.

Text is named where the model is made. Not where the platform carries it.

Meta’s own answer, July 2026

As AI makes it easier to generate content, profiles, and messages, we want to provide a way for you to know there is a real person on the other side of a profile.

Facebook Verified — a badge for “a real person — someone who completed selfie verification”.

Those two categories are the ones its labelling never covered.

Marking the machine, and now verifying the person.

A declaration costs nothing. A selfie check costs a name.

WhatsApp is not uncounted. It is uncountable from outside.

End-to-end encryption: no researcher, no regulator, no third party can survey it.

That is the property it is supposed to have.

  • Only public groups a researcher joined can be studied — a self-selected sample
  • A stated percentage of WhatsApp traffic could not have been produced
  • Better detectors do not help. The missing thing is access, not technique

Which is the whole argument, arriving from the other side: inside an encrypted conversation, somebody saying what they did is not the best signal. It is the only one.

The figure to stop quoting

“Experts estimate that as much as 90% of online content may be synthetically generated by 2026.”

Europol, 2022. Follow the footnote: one trade paperback, published 2020.

No measurement was ever performed. The claim in the book was about video.

A claim about pictures, laundered into a statistic about prose.

Machine writing does leave marks

Four kinds, and they are not equally worth having.

Kind What it is Worth
Provenance Deliberately embedded — watermarks, credentials Strongest. Unreadable by you
Language The words themselves Measured. Trivial to remove
Discourse How the thing is organised Dearest to remove. Least tested
Transfer Traces of how it travelled Strong about the route, weak about the author

Watermarking became real this month

  • Google SynthID — running on Gemini since 2024, published in Nature
  • Anthropic — announced for Claude, mid-August 2026, applied worldwide
  • OpenAI — built one, measured 99.9% internally, did not ship it

What forced it: European law, applying from 2 August 2026.

And you cannot read any of them

The Act requires the mark to be machine-readable. It does not require anyone but the vendor to be able to read it.

  • Google’s detector — a waitlist, for Google’s own content
  • Anthropic’s — announced, unpublished
  • OpenAI’s — nothing deployed

The mark exists. The key does not travel with it.

The strongest measurement used no detector at all

15.1 million biomedical abstracts. Count how often a word appears; compare against how often it should.

  • delves28× its expected rate
  • underscores — nearly 14×

At least 13.5% of 2024 abstracts had been through a model.

“Our analysis is performed on the corpus level and cannot identify individual abstracts.”

Everything cheap to spot is cheap to erase

Marker What removes it
Excess vocabulary find and replace
Em-dashes one instruction: 9.09 → 0.19
Markdown scaffolding one instruction
Statistical detection one paraphrase: 70.3% → 4.6%
Vendor watermark repeat it: 99.3% → 9.7%
Noun-heavy density real rewriting — and never tested

A sign-list works best on whoever was not hiding anything.

Which is generally the person doing nothing wrong

Why none of it is proof — 1

Signs that share a cause do not add up.

Writes at length · writes formally · repeats themselves · will not concede · sounds angry

That looks like five findings.

It is one person arguing hard, seen from five angles.

Why none of it is proof — 2

Rarity beats accuracy.

If machine-written posts are a modest share of the posts you actually read, a test that catches most of them still returns mostly innocent people.

There are so many more innocent people to catch.

This does not improve much with a better test.

Why none of it is proof — 3

The signs land hardest on the wrong people.

Seven detectors, run over essays by writers using English as a second language:

  • 61% flagged as machine-written
  • 5% for US eighth-graders
  • Rewritten to sound native: under 12%
  • Native prose simplified: 57%

The authorship never changed. Only the register did.

What the detectors were measuring

Not machine authorship.

How restricted the writer’s English is.

A sign-list that does not say so becomes a weapon aimed at people writing in a second language.

When several marks appear at once

The tempting move: count the layers, treat four layers as four findings.

It fails. One ordinary cause routinely produces marks in several layers:

  • a house style guide — vocabulary, register, template
  • a scheduling tool — formatting and paste damage
  • a translation workflow — three layers, no AI involved
  • accessibility guidance — short sentences, headings, stripped hedging

So ask a better question

Before a cluster means anything, go looking for the one pipeline that would produce all of it.

If one fits, you have found the explanation — and the cluster is one observation wearing several coats.

What a sign is for

Not a verdict. A reason to be careful.

  • Verify a claim rather than passing it on
  • Look for the source behind a confident summary
  • Notice the difference between sounding authoritative and demonstrating expertise
  • Stop treating polish and completeness as evidence of reliability

What changes is how you read — not what you conclude about the writer.

Which leaves the thing that does settle it

Someone saying so.

A declaration costs nothing. Needs no key. Survives a paraphrase. Misfires on nobody. Requires no authority to sit in judgement over what is authentic.

And its absence means nothing — most posts made without a machine will carry no note saying so.

Plurality is not the cure it reads as

Many makers means no single company owns the answer. Good, and worth defending.

It says nothing about sameness.

A few hundred words identify machine writing across vendors. If the plurality ran deep, they would mark one maker.

An average over a wider base is still an average.

The harm is not deception

Most machine-written text does not lie.

It sounds like everything else — and crowds out the range of ways a thing can be said.

The odd construction. The paragraph that runs hot. The point made badly by someone who has thought about it for twenty years.

Those carry information about who is speaking. A smooth register deletes them and sounds more reliable for having done so.

This deck, measured against its own argument

The essay runs at 10.2 em-dashes per thousand words.

Above the model that drafted it. More than three times the human baseline it cites.

That human range ran from 0.3 to 17.

You cannot tell those two explanations apart from the page. Neither can any detector.

Say so. Take the answer. Argue with the argument.

It is a low standard.

It is also higher than anything else on offer.

agenticgovernance.digital/machine-marks