Helmur Helmur
All essays
#101 July 25, 2026

I Analyzed All 693 Comments on Substack's AI Detection Post. The Real Story Is the Comments Nobody Liked.

The unanimity is real. The finding that made me publish is at zero likes.

ai-detection disclosure writing-with-ai substack pangram
A vintage spectrometer as a metaphor for AI detection tools: a scientific instrument that measures with precision, but only what falls within its spectrum. Pointed at the wrong layer of a writer's work, it reads keystrokes without seeing the thinking underneath.
Detection measures keystrokes, not thinking.

Drafted with AI. Refined with care. Full disclosure at the end, and yes, you should scan this piece.

Last week Substack gave every reader a scan button. The launch came via a founding essay from CEO Chris Best titled “Against Claudefishing,” his coinage for the con of passing off machine writing as human connection. Point the scanner at any post and Pangram, Substack’s new AI detection partner, estimates how much of the text was written by a human. Writers got tools too: pre-publish draft scanning, per-post opt-out, and a native “How I make this” statement that appears whenever someone scans your work.

Then Substack published a follow-up titled “How writers are reacting to Substack’s AI transparency tools.” It quoted a warm collection of voices embracing the transparency and drafting their statements.

I analyzed all 693 comments underneath that post. Every single one. The comments tell a different story than the post, and a more interesting one than the backlash does.

The method, so you can check my work

I pulled the complete comment section: 693 comments from 490 unique authors, carrying 1,693 total likes. The pipeline, me plus AI as disclosed at the end, processed every comment with three or more likes in full and the rest in compressed form, then classified stance and themes across the whole corpus. I personally read the entire top of the thread and every comment quoted or credited in this piece. Where I cite a number below, it comes from that dataset, not from vibes. Where a claim is an anecdote from a commenter, I say so. Anecdotes are evidence of experience, not proof of error rates.

One honest caveat about the sample. This corpus is the reaction post’s thread, which naturally attracts people provoked by the framing of the reaction post itself. The original announcement, “Against Claudefishing,” carries a separate and larger thread of over 1,600 comments. That corpus is the control group, and it is the subject of a follow-up analysis. What you are reading here is the complete record of one thread, not a survey of the whole platform.

Finding one: the top of the thread is not divided. It is unanimous.

Substack’s post frames the reaction as an honest mix. The community’s own voting says otherwise.

The twenty most-liked comments are all critical of the tool, the rollout, or the post itself. Not eighteen of twenty. Twenty of twenty. Together they hold 971 likes, which is 57 percent of every like in the thread. Widen the lens to all 36 comments with ten or more likes and the count of supportive comments is still zero.

The most-liked comment in the entire section, from Frederick Woodruff at 192 likes, is three sentences: he tried Pangram on writing he knew was fully human, it flagged him twice, and he asked the only question that matters. Now what?

Finding two: the objection is a pile of experiments, not a mood.

Here is what surprised me about the criticism. The loudest voices aren’t venting. They’re running experiments.

I count 27 firsthand experiment reports in the thread, comments where a writer ran their own work, or a control text, through a detector and reported the result. Most reported wrong or inconsistent outcomes. The standouts:

Travis Culliton ran the same essay repeatedly and mostly got flagged as AI. Then he swapped one colon for a semicolon. The tool declared the essay 100 percent human. He fixed the semicolon and it flipped back to 100 percent AI. One punctuation mark, a hundred-point swing.

Emma Al fed identical text to nine different detectors and got scores from 1 percent AI to 100 percent AI, with Pangram at the top of the range. When independent systems disagree that dramatically about the same input, the disagreement is itself the finding.

Then there are the time travelers. Thirteen commenters reported that texts written before generative AI existed were flagged as AI, including blog posts from 2013, an early-90s book a commenter ran through Pangram specifically, and award-winning poems. (Two other tests, both on the King James Bible, used unnamed detectors, so file those under genre color rather than a Pangram indictment.)

And the errors run both directions. One writer fed a fully AI-written paragraph in and got 100 percent human back. A detector that fails open and closed does not protect readers from slop or writers from suspicion. It manufactures noise in both markets.

Casey Bohn, buried two replies deep, compressed all of this into one line that deserves to outlive the thread: “If you write like you have some sense, it’s going to say AI did it.”

In fairness, the thread has exactly one sustained technical defender. Paddy Alton argued across nine patient comments that Pangram detects statistical word-sequence patterns rather than surface tells, and cited the company’s claimed false positive rate of one in ten thousand. He engaged critics respectfully and cited sources. He earned almost no likes for it. Hold that thought, because the like economy is about to become the whole point.

Finding three: the most honest comments in the thread got zero likes.

This is the finding nobody is talking about, and it’s what made me publish.

Of the 693 comments, 435 have zero likes. The median comment in this thread was read by the algorithm and nobody else. And buried in that silent majority is a distinct cohort: writers describing, openly and in detail, how they actually work with AI.

One writer spends an hour or two crafting and interrogating a prompt before generating anything. Another spent a week training a model on her own voice. A third drafts in Dutch, uses AI to translate into English, and hand-verifies every technical term. Someone else starts with fifteen minutes of voice transcription before any model touches the piece. Others describe AI as an editor, a sparring partner, a translator, a secretary. Most say they reject the majority of what the model suggests and are willing to tell you which parts they kept.

These are the people doing exactly what Substack’s transparency policy asks for. They are the “How I make this” statement made flesh. And the thread’s engagement economy renders them invisible. The anger at the top earned hundreds of likes per comment. The disclosure in the middle earned an average of roughly nothing.

Both of these things are true at once: Substack curated a rosier picture than its community holds, and the community’s own voting curated an angrier picture than its members hold. The post hid the critics. The likes hid the practitioners. The actual distribution, when you read all of it, is a hostile plurality, a small supportive minority of maybe eight to ten percent, and a large, thoughtful, almost entirely unliked middle arguing about what authorship even means now.

Finding four: the critique with the highest stakes has the lowest visibility.

Twenty-six comments in the thread come from disabled, neurodivergent, dyslexic, and non-native-English writers, and they all make versions of the same point: detection structurally penalizes anyone who writes with assistance.

One had used assistive tech for dyslexia for decades. Autistic writers described using AI to bridge communication formats. Non-native English speakers pointed to the documented tendency of detectors to flag ESL writing as machine-generated. One commenter tested formal academic citation style on its own; it scored 100 percent AI.

Those 26 comments earned 33 likes combined. Two percent of the thread’s attention for the critique most likely to end up in a courtroom or a compliance review. If this technology migrates to platforms where writing is tied to employment, and it will, this is the fault line.

Finding five: the tax has already started.

Scattered through the low-engagement comments are the behavioral changes. A novelist is delaying publication because he cannot predict what the scan will do to his launch. A retired journalist has decided not to start a Substack. Some writers refuse, on principle, to inject typos into clean prose to pass a humanity test. Others now run every draft through Pangram before publishing, defensively, the way you check your fly before a presentation.

One commenter named the trap precisely: now that per-post opt-out exists, disabling the scan itself reads as an admission. Another linked the sharpest concept in the whole discussion, the verification tax: the ambient, unpaid labor of proving provenance after the fact, levied on exactly the people who did the work.

What I think this means

The comments revealed a dispute about disclosure. But underneath it was a more important question: what exactly are we trying to detect?

I have been putting a line at the bottom of every essay here since early on: Drafted with AI. Refined with care. I did not do it because I predicted a scan button. I did it because I concluded that the only durable version of trust between a writer and a reader is the one the writer volunteers.

The thread confirms it from both directions. The thing that stops me every time I look at this data: detection measures keystrokes, not thinking. It cannot see the five hundred conversations behind an argument, the years of operating experience behind a framework, or the four editorial passes between a draft and a published piece. It can see whether your punctuation is too clean. That is not transparency. That is a spectrometer pointed at the wrong layer.

Anyone who has looked at Yelp knows this shape. Most reviews are five stars or one star, because those are the reviews people are motivated to write. The vast middle, the meal that was fine, the visit that was pleasant, the coffee that was just coffee, never gets logged. Restaurant owners running actual businesses know the middle is where their revenue lives, but the Yelp page shows them the extremes and asks them to make decisions from that data. They end up chasing the outliers on both ends and losing the customers they already had. Substack’s comment threads do the same thing to a writer, and to whoever at Substack is trying to figure out what the community actually thinks.

Substack’s founding document already admits all of this. In “Against Claudefishing,” Best concedes plainly that Pangram can only tell you whether AI touched the text. It cannot tell you whether real human care went into the work, and it cannot distinguish AI as a writing tool from AI as a research source. The CEO wrote the limitation into the launch essay. The comment thread is what happens when the product ships anyway and the caveat stays in the fine print. The semicolon flip isn’t a surprise failure. It’s the founding document’s disclaimer, made visible.

Substack got one big thing right, and it is hiding in their policy sentence: the use of AI is not necessarily the problem, but a lack of transparency around it definitely is. The scan button undermines that principle. The “How I make this” statement fulfills it. One tool guesses at your process from the outside and gets semicolons wrong. The other lets you state your process from the inside and be held to it. Monica Hebert, whose response to the launch drew more than a hundred replies, put the whole asymmetry in eleven words: “The presence of AI does not prove the absence of a human.”

So my advice, for whatever it is worth, splits by audience. If you write here: write your statement. Make it specific. Not “I sometimes use AI” but the actual shape of your pipeline, what is yours, what is assisted, what you would want a reader to know before they trust you. Publish it before anyone scans you, because after the scan it reads as defense, and before the scan it reads as character.

And if you read here: when a scan says 100 percent AI, treat it as the opening of a question, not the closing of one. The thread is full of people who did the hard human work and got a machine’s shrug in return. The statement next to the score is where the real information lives.

The scan button will reach other platforms. The writers who come through intact will not be the ones who hid the tools or the ones who raged against them. They will be the ones who told you how they work before you thought to ask.

How I made this

The ideas, the analysis, the methodology, and the conclusions in this piece are mine. I pulled the full comment dataset, directed the classification, checked the counts, and made every editorial judgment about what mattered and why. To be precise about the verbs, because precision is the point of this essay: AI did the exhaustive reading of all 693 comments; I read the entire top of the thread and every comment quoted here myself, and I chose the word “analyzed” over “read” in the title deliberately. AI also helped draft prose that I then rewrote, cut, and rearranged through multiple passes. If a scan flags this piece as AI-assisted, that is accurate and expected. It is also, I would argue, beside the point, which is roughly 2,000 words long and sitting above this line.

Comments credited in this piece, with thanks: Frederick Woodruff, Travis Culliton, Emma Al, Casey Bohn, and Paddy Alton, who deserves particular credit for defending an unpopular position with patience and sources in a thread that rewarded neither.

Drafted with AI. Refined with care.

New essays land here first.

One argument at a time, on the operating system underneath AI-augmented work. Get each new piece by email.