What makes content AI generated is rarely one smoking gun. It's a cluster of small habits — uniform sentence length, hedged claims, a suspicious absence of specifics — that show up together. The hard part isn't spotting those habits. It's that the same habits show up in human writing too, especially in marketing copy written fast by people who don't care.
That's the real decision facing anyone reading this: do you trust the tells, or do you trust the tools claiming to detect them? My view, after watching this space long enough to be cynical about it, is that the tells are more reliable than the detectors — but only when you read for patterns, not for fingerprints. The rest of this piece argues why, and what that means if you're the one publishing.
What makes content AI generated, according to the actual patterns
Strip away the marketing around "AI detectors" and you're left with a handful of recurring features in generated text. None of them are individually conclusive. Together, they're a strong signal.
- Flat rhythm. Sentences cluster around the same length. Human writing drifts — a short punch, then a long rambling clause, then another short one. Generated text often averages out.
- Hedging without commitment. Phrases like "it depends on your needs" and "there are several factors to consider" appear where a human would just pick one.
- Generic specificity. The text names categories instead of instances. "Various tools" rather than a named tool with a price.
- Symmetrical structure. Every paragraph is roughly the same size. Every list has the same number of items. Real writing is lumpy.
- No friction. No asides, no admissions of uncertainty that cost something, no opinions that could annoy someone.
That last one matters most, and it's the one people miss. AI text is agreeable by default. It's trained to be helpful, which means it rarely takes a position it might have to defend.
Why detection tools keep failing at this
Here's the uncomfortable part. The tools sold as AI detectors are trying to do something statistically difficult: separate two distributions that overlap heavily.
Consider what a detector actually measures. Most score "perplexity" — how surprised a language model is by each word — and "burstiness," which is essentially the sentence-length variance I mentioned above. Low perplexity plus low burstiness equals "probably AI."
The problem: a human writing a product description, a legal disclaimer, or a LinkedIn post is also low-perplexity and low-burstiness. Predictable writing is predictable writing, regardless of who produced it. Detectors flag it anyway.
This is why the false-positive problem is structural, not a bug waiting for a patch. If you're using a detector to make a hiring or academic decision, you're leaning on a signal that can't carry that weight. I'd treat any single detector score as one weak input among many, and never as proof.
The tell nobody talks about: what's missing, not what's present
Reading for what isn't there is more reliable than reading for what is.
Take a concrete example. A paragraph about workspace tools might read: "Many productivity platforms now include AI features to help teams work faster." Perfectly plausible. Also tells you nothing.
Now compare a human version: "Notion AI runs as an $8-per-user-per-month add-on on top of your existing plan, which means the AI cost stacks on top of the seat cost." That's a specific, checkable claim. It names the tool, the price, and the structural catch — the add-on sits on top of the base subscription rather than replacing it.
The presence of a checkable number is a stronger human signal than any stylistic quirk. Generated text avoids numbers it can't verify, because it can't verify them.
This is also why the "AI wrote this" accusation lands wrong so often. A junior marketer writing generic copy produces the same tells as a model. The distinction people actually care about isn't human versus machine — it's specific versus vacuous.
3 reasons the tells are getting weaker, not stronger
If you're building a detection habit, know that it has a shelf life.
1. Prompting has improved. Telling a model to "vary sentence length and include one specific example" removes two of the five tells above. The output gets lumpier on request.
2. Editing is normal. Almost nobody publishes raw model output. They cut, add their own examples, and rewrite the opening. Each edit erodes the signal.
3. Human writing is getting more uniform. Templates, style guides, and content briefs push human writers toward the same flat rhythm that used to be the giveaway. The two distributions are converging from both directions.
Some argue this means detection is a dead end and we should stop trying. They have a point — if the goal is catching cheaters, the arms race is unwinnable. But that framing misses the useful version of the question. You're not trying to catch anyone. You're trying to decide whether a piece of content is worth your reader's time.
What this means if you're the one publishing
The practical takeaway flips the whole exercise. Instead of asking "does this read as AI?" ask "does this contain anything only a person with real experience could have written?"
That's a higher bar, and it's the one that actually protects you. A piece stuffed with named tools, real prices, and a specific trade-off survives scrutiny no matter how it was drafted. A piece full of "various solutions" fails no matter who typed it.
If you're generating drafts and editing them, the fix isn't a humanizer pass that swaps synonyms. It's inserting the thing the model can't know: the constraint you hit, the option you rejected, the number you had to look up. Notion AI's add-on pricing is a good example of the kind of detail that has to come from somewhere real — you either checked it or you didn't.
Tools that handle the prompt engineering for you, like AI-Mind, remove some of the friction in getting a first draft out. They don't remove the part that matters, which is deciding what specific, checkable thing the draft should say. That part is still on you.
Key Takeaways
- AI-generated content shows up as a cluster of habits, not a single detectable fingerprint.
- Flat sentence rhythm, hedging, and generic categories are the strongest recurring signals.
- Detectors score predictability, which flags plenty of human writing too — false positives are structural.
- Absence of checkable specifics is a more reliable tell than any stylistic quirk.
- The tells are weakening as prompting improves and human templates converge on the same flatness.
Stop trying to catch the machine. Start asking whether the content contains anything a machine couldn't have known. That single reframe does more for your readers than any detector score ever will — and it's the one test that doesn't decay as the tools improve.
Sources
- AI Tool Database (internally verified snapshot), Notion AI — tool profile, 2026. Vendor, category, pricing model and platform facts for Notion AI, including its add-on pricing structure.
- AI Tool Database (internally verified snapshot), Make (formerly Integromat) — tool profile, 2026. Vendor, category and pricing snapshot for Make, with editorial rating recorded at verification time.
- AI Tool Database (internally verified snapshot), Slack AI — tool profile, 2026. Vendor, category and pricing snapshot for Slack AI.
- AI Tool Database (internally verified snapshot), Tool database overview, 2026. Internal database of 360 AI tools, each with a pricing and capability snapshot recorded at verification time.
Frequently Asked Questions
Can AI detectors reliably tell if content is AI generated?
No, and the failure is structural rather than temporary. Most detectors score predictability — low perplexity and low sentence-length variance. Plenty of human writing, especially templated marketing or legal copy, scores the same way. That means false positives aren't a bug waiting to be fixed. Treat any single detector score as one weak signal, never as proof.
What's the single strongest sign that content is AI generated?
The absence of checkable specifics. Generated text tends to name categories rather than instances — "various tools" instead of a named tool with a price and a stated trade-off. Stylistic quirks like flat rhythm can be prompted away, but a missing number is harder to fake. If a paragraph contains nothing you could verify, that's the tell.
Is AI-generated content always worse than human-written content?
Not automatically, but it fails more often for a specific reason: it's agreeable and vague by default. A human writing generic copy produces the same weakness. The distinction readers actually care about isn't human versus machine — it's specific versus vacuous. A well-edited AI draft with real constraints and numbers beats a lazy human draft every time.