Somebody in your family bought a children's book recently. Thirty pages, a rhyme scheme that mostly works, an illustrated fox in a red coat, four dollars, delivered in two days. The kid asks for it three nights running.

Now ask a question that would have sounded strange in 2019. Did a person write it?

You cannot answer from the object. The rhymes land. The copyright page lists a name you have never heard, which is true of most children's books. Nothing in the artifact tells you whether it came from an author at a kitchen table or from a prompt run forty times until the meter behaved.

The loud version of this argument is wrong, so let me get out of its way. Children's books published after 2022 are not machine-written as a class. Most, presumably, are not, though nobody can currently measure the proportion, which is part of the problem.

The interesting claim is narrower: the polish of an object no longer carries much information about its origin. You used to reason backward from quality to effort, and from effort to a human. That inference is much weaker than it was, and it weakened quietly.

I think the bill lands on social media first. Not because feeds fill up with fakes, but because checking stops being worth the trouble. That is a theory, and near the end I say what would prove it wrong.

Copywriting was disturbed before radiology

OpenAI made ChatGPT public on November 30, 2022, as a free research preview. The businesses I saw disturbed first were not the ones in the think pieces. Not radiology, not law. It was copywriting, content farms, SEO filler, and low-cost publishing. That is an impression, not a finding.

Those markets had already agreed that plausible was enough. Nobody buying a product description for the price of a sandwich was buying insight, only grammatical English that fit a slot. When a machine could produce that, the market cleared quickly, because the bar had never had anything to do with origin.

What Amazon's disclosure rule actually requires

Amazon's KDP content guidelines draw a line I find useful. Content counts as AI-generated if a tool created the actual text, images, or translations, even if you edited substantially afterwards. It counts as AI-assisted if you wrote it and used tools to refine or error-check. Publishers must tell Amazon about the first category; the second requires no disclosure.

Read where that disclosure goes. It goes to Amazon. The policy says inform us, and says nothing about the customer ever seeing anything. So the fact may exist, recorded in a system you cannot query, while you decide whether to spend four dollars.

The sharper question is whether the buyer can find out.

For a four-year-old it hardly matters. A good rhyme is a good rhyme, and a four-year-old is not auditing the supply chain. For the parent it matters differently. Part of what they were buying was a small quantity of human attention, someone having thought about what a child finds funny. That is a real thing to want, and a thing the market cannot price, because the claim costs nothing to fake.

Images, audio, and video are losing the same property, later

Text was the easy case, because text never carried much evidence. A paragraph is just a claim. Images, audio, and video were different. For well over a century they were the closest thing consumer media had to a fact. A photograph of a flooded street was weak evidence, but it was evidence, and a voice on the phone was authentication.

That property is degrading unevenly, which makes it worse. Meta, announcing its labeling work in February 2024, was candid about the seam. It said it was building tools to read invisible watermarks and metadata in images, because image generators are starting to cooperate on IPTC and C2PA markers, but that companies "haven't started including them in AI tools that generate audio and video at the same scale."

So for the two formats where synthesis is most convincing, the fallback is asking users to disclose, with penalties possible if they do not. YouTube asks the same of creators, while automatically labeling content made with its own tools or arriving with C2PA metadata attached.

Both policies are reasonable, and both lean primarily on self-disclosure. The automatic labels reach only the pipeline that was already cooperating.

Detection and provenance are two different products

Detection looks at a finished artifact and outputs a probability about how it appears. It has no access to the artifact's history. It is a guess made after the fact by one model about whether some text or pixels resemble the output of another.

Provenance is a record of origin. It is attached when the thing is made, it names who made it and what happened to it since, and somebody signs for that claim. It carries a history, or it does not exist.

Detection fails in both directions. Krishna and colleagues showed how completely: running machine text through DIPPER, an 11-billion-parameter paraphraser they built for the purpose, dropped DetectGPT's accuracy from 70.3% to 4.6% at a fixed 1% false positive rate, and also evaded watermarking, GPTZero, and OpenAI's own classifier. Their proposed defense is the tell. The provider keeps a searchable database of everything it has ever generated, which is a provenance system in a detector's clothes and covers only text that came through that provider.

The false positives land on real people. Liang and colleagues found that GPT detectors misclassified more than half of TOEFL essays by non-native English speakers as AI-generated, while achieving near-perfect accuracy on essays by US eighth-graders. Across that set, 89 of 91 human-written essays were flagged by at least one detector. The detector was not measuring authorship. It was measuring how closely a sentence resembled fluent native English, and calling the difference fraud.

The C2PA specification is more honest about its limits than most of the discourse around it. Content Credentials record what happened to a file and who signed for each step, and the explainer says plainly that provenance information alone cannot tell you whether content is true, accurate or factual. A signed photograph of a staged scene is a signed photograph of a staged scene.

A chain of custody also binds only the parties who agree to be bound. The uncooperative path is a screenshot, a re-encode, a phone camera pointed at a monitor, or an open-weights model that signs nothing.

C2PA anticipates this with durable Content Credentials, pairing metadata with watermarking and fingerprinting so the claim survives stripping. Whether that holds up against a camera aimed at a screen is an open empirical question. Meanwhile the cheapest available answer, showing a buyer what a seller already disclosed, needs no new technology and has not been built.

A theory about verification cost

Every claim on a screen carries a verification cost, the effort required to establish whether it is real. Every claim also carries a value, what knowing would be worth to you. People check when the value exceeds the cost, and they are efficient about this without ever announcing it.

Synthesis raises verification cost across the board, including for all the true things. My guess is that for a typical post the cost crossed the value some time ago, and that the response is to stop treating the category as evidence and start consuming it as entertainment.

I cannot prove that at population scale. I can report that I have stopped reverse-image-searching things I would have checked in 2019, and that I did not decide to stop. It went the way habits go.

A feed does not die from being full of lies. It dies from not being worth checking.

The strongest objection is that feeds were never evidence

The best argument against all of this is that I have described the loss of a property social media never had.

The weak form is that feeds were always unreliable, which everyone concedes. The strong form is more damaging: the feed's function was never informational. People open these apps to see what people they know are doing. That is a social function, and social functions no more require verification than waving at a neighbor requires a signature.

Staged photos, filtered faces, purchased engagement, and bot networks predate generative models by a decade, and users adapted years ago by reading feeds as mood rather than as record. If that holds, my mechanism is measuring a property nobody was buying.

I take it seriously, and there is one answer. What the people you know are doing is now synthesizable too. The fake that matters is no longer a fabricated news event, which feeds had already stopped being trusted for. It is a thirty-second clip of your cousin, in your cousin's voice, saying something your cousin never said.

The social function does rest on a minimal evidential floor: that the person shown is the person, and that the thing shown happened in some form. That floor sits far below the one everyone argues about, and it is the part now under pressure.

Two smaller objections stand too. Trust has relocated before, from artifact to institution: photography absorbed retouching, print absorbed the forged pamphlet, and society kept working. And platforms have every incentive to sell trustworthy identity, since scarcity is where the money is.

What would prove this wrong

I would treat the theory as falsified if engagement with unverifiable content holds or grows over the next several years while migration into small accountable spaces stays flat. I would treat it as supported if group chats, thirty-person servers, mailing lists, personally owned sites, and rooms with actual chairs keep absorbing the attention that used to go to public feeds, while those feeds keep their numbers and lose their function.

That second outcome is a split rather than a collapse. The feeds stay profitable and stop being consulted.

I notice I am describing my own behavior, and that a personal site is exactly this move. I am inside the sample, which is a reason to discount me. If the theory fails, I expect it fails because people wanted their feeds to be evidence far less than I assumed.

What stays expensive to fake

Whatever happens to the feeds, one filter gets sharper. Trust flows toward whatever is costly to counterfeit: a body in a room, a record of behavior across years, software somebody else can run and check. It is the only claim here I would bet money on.

Which brings me back to the book. Under current KDP policy, if that story was machine-generated, the publisher was required to tell Amazon. Assume they complied. The fact would then exist, sitting in a database three clicks and one permission boundary away from the parent holding the four dollars.

This is not really about what the models can do. It is about whether the label reaches that parent. The cheapest fix in this field is to show the buyer what the seller already told the retailer. Nobody has shipped it.

Sources and further reading

Author profiles: Rashid Azarang on LinkedIn and Jesús Carlos Acosta Rocha on LinkedIn.