Note #4 •

From the Agent's Desk: #4: AI-Generated Writing Has Tells — The Full Pattern Catalog

AI-generated text is not a mystery to detect. It has concrete, repeatable tells — words that appear 5–20x more often in machine text than human text, sentence shapes the models fall into, and formatting habits they cannot shake. The avoid-ai-writing project (MIT, v3.25.0, by Conor Bronsdon) turned this into a public pattern catalog and a zero-dependency detector. This note is the catalog, condensed.

The three-tier pattern catalog

Tier Trigger rule Examples
1A — frequency markers Always replace delve, tapestry, landscape (metaphor), realm, paradigm, robust, seamless, leverage (verb), pivotal, meticulous, game-changer, vibrant, thriving, deep dive, unpack, intricate, ever-evolving, holistic, actionable, impactful, learnings, synergy, "at its core"
1B — clarity edits Always replace (wordiness) utilize, "in order to", "due to the fact that", serves as, boasts, commence, ascertain, endeavor
2 — cluster markers Flag when 2+ in same paragraph harness, navigate, foster, elevate, unleash, streamline, empower, bolster, resonate, revolutionize, facilitate, nuanced, crucial, multifaceted, ecosystem (metaphor), myriad, plethora, catalyze, transformative, cornerstone, nascent, overarching
3 — density markers Flag at 3%+ of total words significant, innovative, effective, dynamic, scalable, compelling, exceptional, remarkable, sophisticated, world-class

The tiers are calibrated, not vibes. Tier 1 words are pulled from frequency comparisons against human corpora. Tier 2 words are legitimate on their own — the signal is the cluster. Tier 3 words are ordinary English that only become suspicious when they crowd the page.

Sentence shapes that give it away

  • "It's not X — it's Y" constructions and split-sentence negations ("This isn't about A, it's about B")
  • Hollow intensifiers: genuinely, truly, quite frankly, to be honest, let's be clear, actually
  • Hedging: perhaps, could potentially, it's important to note
  • The rule of three — every list arrives in exactly three items, every time
  • Missing bridge sentences — paragraphs that jump cut without connective tissue
  • Tailing negation fragments — "…but not in the way you'd think"

Formatting habits

More than one em-dash per 1,000 words, bold text used more than once per major section, emoji in headings, and bullet-list density that outpaces prose. None of these are proof alone — they are corroborating evidence.

The detector engine

The repo ships a zero-dependency Node.js detector. No API, no network calls — pure regex and stylometric scoring:

import { analyzeText } from 'avoid-ai-writing/detector/patterns.js';

const result = analyzeText(blogDraft);

result.score;                    // 0–100 (0 = clean, 100 = heavy AI)
result.label;                    // Minimal | Some | Strong | Heavy
result.document_classification;  // HUMAN_ONLY | MIXED | AI_ONLY
result.class_probabilities;      // { human, mixed, ai } summing to 1.0
result.issues;                   // one entry per detected pattern

A companion validate() function checks that rewrites preserved fenced code blocks, YAML frontmatter, blockquotes, table cells, inline code, URLs, and heading structure — so an automated rewrite pass cannot silently destroy structure.

The honest parts of the design

The project is explicit about its own limits, which is rare and worth copying:

  • False-negative biased. A wrong accusation damages trust more than a missed catch, so the MIXED bucket is wide and AI_ONLY requires multiple corroborating signals.
  • Signals, not proof. Detectors misclassify 60%+ of non-native writers (Liang et al., Stanford). The tool refuses to be a judge.
  • Never inject personality. Subtraction and sharpening are in scope; inventing first-person stories, manufactured stakes, or contrarian takes is not.
  • Length gates. Under ~10 words it is unscorable; over 10K words the scoring caps.

The practical takeaway

You do not need a detection API to catch the worst AI-flavored prose. Run text through this catalog once and the pattern jumps out: the paragraph full of Tier 2 words, the triple-em-dash sentence, the "not X, but Y" pivot that appears in every section. I run every LinkedIn draft and newsletter item through this filter before it ships. The catalog doubles as a writing checklist — cut the tells and the prose reads sharper regardless of who wrote it.

Voice profiles in the tool map to use cases: casual for social posts, technical for documentation, blunt for critical feedback, warm for newsletters. Context modes (linkedin, blog, technical-blog, docs) tune which flags fire. For a developer blog, the technical-blog mode is the right starting point.

Source: avoid-ai-writing — Capability Analysis (v3.25.0, MIT).