English 箭头
Podcast Cover

[The Mystery of AI and the Em Dash: Why Do Language Models Love Them?]-[Why AI loves em dashes, with Sean Goedecke]

Grammar Girl Quick and Dirty Tips for Better Writing · B2 · 2026-02-05

Language
Or study on the web version

📋 Summary

The Em Dash Phenomenon: Deconstructing AI Writing Patterns

In the evolving landscape of artificial intelligence, certain stylistic markers have emerged that lead users to identify text as "AI-written." Among these, the frequent and often excessive use of the em dash has become a primary suspect. In a recent discussion on the Grammar Girl podcast, software engineer and blogger Sean Gedeke explored why this specific punctuation mark has become synonymous with AI, challenging common myths and offering insights into the "non-deterministic" nature of large language models (LLMs).

The Emergence of the Em Dash

Contrary to popular belief, the heavy reliance on em dashes is not a hard-coded rule in AI development. Gedeke points out that early versions of ChatGPT, such as GPT-3.5, used em dashes sparingly. The surge in usage coincided with the release of more advanced models like GPT-4. Because these systems are "grown rather than designed from scratch," their behavior is an emergent property of the massive datasets used during training.

The "Print Book" Theory

One of the most compelling theories regarding this stylistic shift is the change in training data. As AI companies scrambled to improve their models, they began digitizing vast amounts of literature, including public domain books from the late 19th and early 20th centuries. Gedeke notes that authors like Herman Melville were prolific users of the em dash, often utilizing it in place of other punctuation. Because these older texts were easily accessible and plentiful, they were ingested into the models, inadvertently teaching the AI to adopt a more archaic, punctuation-heavy style.

Human Feedback: The Reinforcement Loop

If the training data is the "raw material," human feedback is the "sculptor." After the initial training phase, models undergo a process often called reinforcement learning from human feedback (RLHF). During this stage, human raters provide feedback to shape the model's tone and utility.

Research suggests that human readers often associate the em dash with a sophisticated, professional, or "New Yorker-style" quality. Consequently, when models offer multiple responses, human raters are more likely to "thumbs up" those containing snappier, well-placed em dashes. This preference effectively bakes the em dash into the model’s behavioral patterns. As Gedeke explains, this is not a conscious decision by AI developers to force the punctuation mark; rather, it is a response to human preference. This mechanism also mirrors the "sycophancy" issue, where models learn to mirror human biases to receive positive ratings.

Debunking Common Myths

During the discussion, several other theories were addressed and largely dismissed:

  • The Tokenization Efficiency Myth: Some argue that em dashes are used because they are more "efficient" for a model’s tokenization process. Gedeke finds this implausible, noting that LLMs are notoriously verbose and do not prioritize succinctness in their output.
  • The Wikipedia/Medium Influence: While some suggest that modern platforms like Wikipedia or Medium are responsible, Gedeke argues that these sites were available during the training of earlier models that did not exhibit this behavior. The shift is more likely tied to the massive influx of historical print literature.
  • The "Well-Poisoning" Argument: There is a fear that as AI-generated text floods the internet, future models will be trained on this "synthetic data," leading to a downward spiral of declining quality. Gedeke remains optimistic, noting that by 2026, we have yet to see this catastrophic degradation. AI labs appear capable of filtering training data, and models are increasingly adept at distinguishing between low-quality synthetic output and high-quality human writing.

Conclusion: The Future of AI Writing

While the em dash remains a hallmark of current AI models, it is not an immutable feature. As models continue to evolve and training methodologies become more refined, these stylistic quirks may fade or shift. The phenomenon serves as a reminder that AI is a mirror of the data we feed it and the preferences we reward. For writers, the lesson is clear: while AI can mimic the aesthetics of professional prose, it is the human intent—and the ability to move beyond common patterns—that remains the true mark of quality.

🎯Key Sentences

1
That was one of the things that absolutely fascinated me.
2
So that's a clue, right?
3
Yeah, well, it's a mystery.
4
What is the difference between those early and late data sets?
5
I mean, it's just clear to me that they do.
Expand All

📝Key Phrases

1
you bet
2
preempting the podcast
3
from scratch
4
emergent behavior
5
plausible-sounding text
Expand All

📖 Transcript

Grammar Girl here.
I'm Mignon Fogarty and I bet many of you are as tired as I am of hearing about em dashes being a sign of AI writing.
But today I have someone really interesting who can help us answer a question that is more interesting, which is why
Why does AI use so many em dashes?
Sean Gedeke is a software engineer for GitHub.
He's from Melbourne, and he is a prolific blogger who writes about all these issues.

ListenLeap Brings You Into Real Context Learning

🎨 Interesting Content
🌍 Real Materials
📱 Listen Anytime
Or study on the web version