# Why AI writing keeps saying 'it's not X, it's Y'

> RLHF mode collapse, markdown training data, and reward bias each explain part of AI's 'not X, but Y' tic. None explains all of it, and the evidence disagrees.

Canonical: https://brandonlazovic.dev/articles/llm-negative-parallelism-tic/  
Author: Brandon Lazovic  
Published: 2026-07-21

## The short version

- AI chatbots overuse 'it's not X, it's Y' (negative parallelism) at a measurably higher rate than human writers, and researchers have at least three competing, unreconciled theories for why.
- A Washington Post analysis of 328,744 ChatGPT messages found the construction in 6 percent of July 2025 chats, and the same model's em dash usage jumped from under 10 percent to over 50 percent of responses in a year.
- RLHF measurably narrows a model's output diversity, but the paper proving that says it cannot explain the mechanism, and a newer study finds the em dash tendency already present in base models before any RLHF step runs.
- Stripping these tics from copy is restructuring, not substitution: swapping em dashes for semicolons just trades one repeated shape for another that reads just as generated.

Ask ChatGPT, Claude, or Gemini to describe almost anything, and there is a decent chance the answer arrives in the shape "it's not X, it's Y." Researchers at the AI detection company Pangram estimate that "Not just X but Y" sentences appear roughly three times as often in AI writing as they do in human writing, according to founding engineer Elyas Masrour. [1] A Washington Post analysis of 328,744 ChatGPT messages found the construction, known as negative parallelism or contrastive phrasing, in 6 percent of chats by July 2025, while the same model's em dash usage jumped from under 10 percent to over 50 percent of responses in a single year. [2] At least three technical explanations exist for why models lean on this pattern so heavily, and they do not fully agree with each other. For anyone editing AI assisted copy, that unresolved mechanism matters less than the fix, because the tic is structural and structural problems have a structural repair.

> Pangram estimates "not just X but Y" appears roughly three times as often in AI writing as human writing; a Washington Post analysis of 328,744 ChatGPT messages found it in 6 percent of July 2025 chats.

## Why does every AI chatbot keep writing "it's not X, it's Y"?

AI chatbots overuse the "it's not X, it's Y" construction, known as negative parallelism or contrastive phrasing, because it has become their default move for adding apparent nuance to a description. It shows up across every major lab's models, from OpenAI to Anthropic to Google, rather than as one company's house style. Pangram's Masrour told The Atlantic that ChatGPT, Claude, Gemini, and several open source models all lean on it, to varying degrees. [1]

The scale is what separates this from an ordinary stylistic quirk. Barron's found the phrase's appearance in corporate communications more than quadrupled between 2023 and 2025. [1] The Washington Post's data puts a hard number behind that trend: across 37,929 publicly archived ChatGPT conversations from May 2024 to July 2025, the phrase turned up in 6 percent of July 2025 responses, and em dash usage climbed from under 10 percent to over 50 percent over the same window. [2] Both figures come from the same dataset and the same model, GPT-4o. The counts and the date range are stated plainly enough for anyone to check.

OpenAI has acknowledged the problem in its own words. Laurentia Romaniuk, the company's product manager for model behavior, told The Atlantic she prefers the term "contrastive phrasing" to the more common "negative parallelism," and admitted ChatGPT overuses it to the point of feeling formulaic. Her team, she said, is "working on ways to broaden the chatbot's repertoire." [1] That admission, from inside the company that built the model, is the closest thing to a settled fact here: a known, named, unresolved defect.

## Is this construction actually new, or did AI just turn an old rhetorical trick into a tic?

Negative parallelism, the "it's not X, it's Y" construction, predates chatbots by centuries. Classical rhetoric already had names for versions of it, antithesis and metalinguistic negation, and English literature is full of examples. What changed with AI is frequency: a device writers once reached for occasionally, for emphasis, now appears so often that readers recognize it as a tell rather than a flourish.

The Atlantic's own examples make the lineage plain. Shakespeare wrote "The fault, dear Brutus, is not in our stars, but in ourselves" in Julius Caesar. Vince Lombardi popularized "Winning isn't everything; it's the only thing" in the 1960s. A 1990s DiGiorno ad line ran "It's not delivery. It's DiGiorno." All three are genuine rhetorical devices doing real work, used sparingly, in contexts built around a single memorable line. [1]

What broke the scarcity is training scale. A pattern rare enough in human writing to feel fresh each time becomes a statistical habit once a model has ingested enough of it and generates text one likely word at a time. Oremus notes that before ChatGPT, the pattern was obscure enough that it lacked an agreed name at all. It has one now, which is itself a sign of how visible the frequency has become.

## What actually causes AI to overuse negative parallelism and em dashes?

Three separate technical explanations exist for why AI models overuse these tics, and the newest one complicates the most repeated theory. Reinforcement learning from human feedback (RLHF) is confirmed to narrow a model's output diversity, but the paper establishing that finding says explicitly it cannot explain the mechanism. [3] A newer, unreplicated study on em dashes specifically finds the tendency already present in base models, before any RLHF step runs at all. [5]

Start with what is confirmed. Kirk et al., in a peer-reviewed style paper measuring the effects of RLHF against supervised fine-tuning alone, found that RLHF measurably reduces both syntactic diversity (distinct word sequences) and semantic diversity (embedding-based sentence similarity) in model outputs, while logical or NLI-based diversity was not significantly affected. [3] That is a real, demonstrated narrowing effect. What the same paper does not do is explain why RLHF narrows outputs this way. The authors state this as an explicit limitation, and they specifically rule out the KL-penalty coefficient (the term that keeps a fine-tuned model from drifting too far from its starting point) as a full explanation on its own. [3] RLHF collapses variety. Nobody has published a mechanism for why.

For negative parallelism specifically, the Atlantic's reporting lays out a plausible, compounding story rather than a single cause. Tuhin Chakrabarty, a computer science professor at Stony Brook who studies AI writing, told Oremus it is plausible that human RLHF reviewers rated "it's not X, it's Y" responses highly because the construction gives an impression of nuance and insight, the model appears to reason its way from a weaker descriptor to a stronger one. [1] A second, independent theory concerns token prediction itself: because models generate one chunk of text at a time, balancing likely wording against reward, starting a description with what something is not is a safer, more probable continuation than committing immediately to a direct characterization. Chakrabarty also points to a feedback loop: newer models increasingly train on text that earlier AI models generated, so a tic already present in the training pool gets reinforced rather than diluted. In Chakrabarty's words, describing that loop: "There's already negative parallelism in the text, and then AI is preferencing negative parallelism," to the point that a model trained this way struggles to write without the pattern at all. [1] These four factors, training-data presence, RLHF reward, token-prediction dynamics, and self-reinforcing training loops, plausibly compound rather than compete.

Em dashes are where the theories genuinely disagree. Sean Goedecke, analyzing punctuation frequency across model generations, argues em dash overuse is a training-data era effect: digitized 19th century books, where em dash usage peaked around 0.35 percent of characters versus roughly 0.25 to 0.275 percent in current model output, entered training pipelines around the same window that GPT-3.5 gave way to GPT-4o, which is exactly when em dash rates jumped roughly tenfold. Goedecke explicitly tested and rejected three alternative explanations: tokenization efficiency, next-token-prediction optimization, and an RLHF dialect-preference theory, the last one falsified by Nigerian English's em dash rate, at roughly 0.022 percent, being lower rather than higher. [4] A separate, unreplicated study called "The Last Fingerprint" tested 12 models across 5 providers, suppressing markdown formatting in prompts and comparing base models against their instruction-tuned, RLHF'd counterparts. It found elevated em dash usage already present in base models, before any RLHF step runs, which points toward training-data composition (markdown-heavy technical documentation, specifically) rather than RLHF rater preference as the primary driver. [5] The McGill Office for Science and Society, reviewing similar evidence, lands closer to a plain frequency-and-clarity explanation: the construction is common in training data and rewarded by a training objective focused on clarity, and the office notes that human content moderators typically do not alter a model's linguistic style during review, a detail that further weakens the rater-preference theory as the primary cause for em dashes specifically. [6]

Put plainly: nobody has one confirmed causal chain. RLHF narrows output diversity, confirmed in general. Negative parallelism has a plausible compounding story. Em dashes have two training-data theories that partly disagree, plus a newer experiment that weakens, without eliminating, the RLHF-rater explanation. Three real mechanisms, real evidence, no study that adjudicates all of them at once.

## Why do these tics matter beyond annoyance, are they a detection signal?

AI's negative-parallelism and em-dash tics matter beyond annoyance because they are exactly what AI-detection tools are built to key on, and Pangram's own data shows that persistence is doing the detection work. Masrour told The Atlantic that although the specific markers of AI writing keep shifting, AI text is not getting any harder for Pangram's software to detect, and he pointed to the stubborn persistence of constructions like negative parallelism as one reason why. [1]

My read: copy loaded with negative parallelism and heavy em dash use reads as a legible pattern to any tool built to flag stylometric anomalies, low sentence-length variation, repeated rhetorical shapes, whether or not that tool publishes its detection features. Some SEO practitioners speculate Google's own quality systems reward or penalize similar uniformity signals, but that mechanism is unconfirmed by Google and I found no primary source for it. Treat it as speculation, not fact.

A second wrinkle is worth flagging precisely. A German study on spontaneous human speech found AI-associated word choices, terms like "delve" and "showcase," turning up more often across roughly 740,000 hours of recorded conversation. [7] That is a lexical finding, specific word choices migrating into human speech, not evidence that syntactic tics like negative parallelism are migrating the same way. Claiming rhetorical-structure drift from a vocabulary-drift study would be exactly the overstatement this article is trying to avoid.

## How do you actually strip these tics from real copy at scale?

You strip negative parallelism and em-dash overuse by restructuring sentences rather than substituting one marker for another, because the tell is usually the repeated shape rather than the specific punctuation mark. I run a fixed set of rules on every deck, talk track, and article that passes through multi-critic QA before it ships, and each rule exists because a mechanical shortcut created a new, equally visible tell the first time I tried it.

The first rule is the one this article is itself proof of: kill the negative-parallelism drumbeat. A single factual contrast, stating what something runs on and what it does not, is fine. The repeated "it's not X, it's Y" rhythm is not. The fix is to rewrite toward the direct, positive claim and cut the negated half unless the contrast genuinely carries information the reader needs.

The second rule is an em dash sweep done by restructuring, never by substitution. An independent clause becomes a period. An appositive that itself contains commas gets wrapped in parentheses instead. A label-and-value or setup-and-explanation pair takes a colon. Swapping em dashes for semicolons as a reflex creates its own tell, a repeated semicolon cadence that reads exactly as generated as the em dashes did, and I have made that mistake and had to redo the pass.

Third, run a cadence check after any sweep. No more than about two sentences in a row should share the same shape. Vary openers on purpose: a declarative sentence, then a subordinate clause, then a prepositional opener, then a direct instruction. A document where every sentence follows the same "clause, then restatement" pivot reads as formula-generated in exactly the way em dash-heavy titles do, even with zero em dashes left in it.

Fourth, cut scaffolding that announces itself instead of saying the thing: "the key point here," "it's worth noting that," "here's the striking part." Fifth, drop dramatic openers and soften false absolutes: an "estimated" is more honest than a confident-sounding claim the evidence does not support. Sixth, strip sales and urgency language, "no longer optional," "game-changer," "table stakes," in favor of stating the number and its source. Seventh, and this is the rule the other six depend on: vary confidence deliberately. Uniform certainty across every sentence is itself a tell, because real practitioners hedge where the evidence is actually thin, which is most of section three above.

## What should writers and publishers do about a mechanism nobody can fully explain yet?

Writers and publishers should treat negative parallelism and em dash overuse as a known, recognizable pattern worth editing out on sight, without waiting for the labs to fix it at the source. OpenAI has said it is working on broadening the model's repertoire, and Romaniuk suggested users lean on custom instructions in the meantime, but the underlying construction has shown no sign of disappearing on its own. [1] Compare that persistence to shorter-lived tics such as ChatGPT's brief obsession with goblins and gremlins, which OpenAI retired outright last fall once it traced the habit to the model's "nerdy" personality setting. [1]

The practical stakes are twofold. For publishers, these tics are now a stylistic signal that both detection tools and readers have learned to spot, which carries a real, if modest, credibility cost. For anyone editing AI-assisted drafts, the honest position accepts that the causal chain is still contested among researchers, while treating the fix as fully available today: the pattern is diagnosable by eye, and removing it is a matter of restructuring sentences, not waiting on a training fix that may or may not arrive. A well-documented symptom with a genuinely unsettled cause is unusual enough to name plainly rather than resolve with false confidence.

## Sources

1. The Atlantic: The Most Famous AI Writing Tic Is Also the Most Mysterious (Oremus, 2026) — https://www.theatlantic.com/technology/2026/07/ai-chatbot-writing-tic-negative-parallelism/687892/
2. Washington Post analysis of ChatGPT writing style, via Boston Globe — https://www.bostonglobe.com/2025/11/13/business/chatgpt-writing-style-clues/
3. Kirk et al., Understanding the Effects of RLHF on LLM Generalisation and Diversity (arXiv 2310.06452) — https://arxiv.org/html/2310.06452v2
4. Sean Goedecke: Why do AI models use so many em dashes? — https://www.seangoedecke.com/em-dashes/
5. The Last Fingerprint: How Markdown Training Shapes LLM Prose (arXiv 2603.27006) — https://arxiv.org/pdf/2603.27006
6. McGill Office for Science and Society: Why Did LLMs Steal Our Em Dashes? — https://www.mcgill.ca/oss/article/critical-thinking-student-contributors-technology/why-did-llms-steal-our-em-dashes
7. German study on AI word-choice drift in human speech (arXiv 2409.01754) — https://arxiv.org/abs/2409.01754
