Grammar Girl here.
I'm Mignon Fogarty and I bet many of you are as tired as I am of hearing about em dashes being a sign of AI writing.
But today I have someone really interesting who can help us answer a question that is more interesting, which is why
Why does AI use so many em dashes?
Sean Gedeke is a software engineer for GitHub.
He's from Melbourne, and he is a prolific blogger who writes about all these issues.
Sean, welcome to the Grammar Girl podcast.
Hi, Mignon.
Thanks for having me.
You bet.
So I saw your blog post and I was instantly fascinated, because it's this thing that people in my world all the writers are talking about and debating whether we should leave em dashes out of our writing because everyone thinks it's a sign of AI.
And I have always loved the em dash.
I use it a lot.
So I didn't even really notice that it was being used so much.
But obviously now it's a thing.
I always felt like so much normal writing, so much good writing has em dashes.
Like, isn't that just why AI is doing it?
Because it's in human writing.
Well, the short answer and maybe I'm preempting the podcast entirely with this is yes, that's basically why it is.
But there are interesting questions around why it emerged.
Because the initial versions of AI ChatGPT-3 didn't use em dashes and the latest versions of the models use em dashes less.
That was one of the things that absolutely fascinated me.
So 3.5 barely used em dashes at all, right?
Yeah.
Yeah, that's right.
It certainly was not enough for people to talk about it as a sign of AI.
It might have used the odd one here or there, but it wasn't the kind of extreme overuse we saw out of later models like 40 and 40.
Yeah.
So that's a clue, right?
Yeah, well, it's a mystery.
I think one of the most interesting things about this is that I have as we'll get to, I think some theories about why this is happening.
But nobody knows, like nobody actually knows, not even the people who sort of build the models know, because it's so non-deterministic, this process of like constructing an AI model because it's trained or grown rather than kind of designed from scratch.
It's really like emergent behavior, this use of em dashes.
Yeah, why don't for the audience?
You know a lot of writers.
Why don't you sort of give a really quick overview of how these models are made and why it's non-deterministic?
Yeah, sure.
So a normal computer program.
Somebody sits down and they kind of go through step by step everything that can happen in the program.
And that's what becomes the program.
Whereas an AI model is more like.
Imagine like a series of like knobs and dials the size of a football stadium.
And if you tune all those knobs and dials to precisely the right values, what you end up with is something like ChatGPT.
But there's no human being going and tuning all those knobs and dials.
What happens instead is a bunch of training data gets automatically shoveled through the whole thing and automatically turns the dials to various things.
So it's how the model behaves is really a function of like what kind of things it was exposed to as it learnt how to speak English and learnt how to produce plausible-sounding text.
Models that are trained on different things can behave very differently.
Not even the people who trained them know exactly how that works or exactly what kinds of things they ought to train the models on in advance.
One reason why GPT-3 is different is because it was so early.
It was trained on fundamentally different things than later models.
They just didn't have access to the kind of data sets that were used later on.
And I think that's partially why we see more MDashes later on.
Yeah.
What is the difference between those early and late data sets?
Yeah.
Well, at the start, I think people weren't even sure that this process of building a large language model would even work.
So there wasn't this huge effort to get a lot of high quality data.
It was really just trained on the internet on variations of what people call the pile, which is this like huge data set of sort of conversations and articles and blog posts and comments scraped from the internet.
And that's what kind of formed these initial AI models.
So it was really biased towards kind of short-form content and it was biased towards short-form content from the last 20 years, basically.
Then, once ChatGPT took off and once it became clear that there was frankly, an enormous amount of money in these tools, companies started scrambling to find more data and better data.
They started doing things like – I think this is where the MDash comes from – digitizing books, and digitizing huge amounts of print books which would have contained more em dashes than the kind of writing we see today.
Why do you think print books contain more em dashes?
Well, I don't know why.
I mean, it's just clear to me that they do.
When I say older books, I'm talking about kind of the early 1900s, the late 1800s.
And I think if you read books from that era, it's clear that, like There are certain styles that are not present today, such as the almost German-like capitalization of words for emphasis or the use of em dashes,
My favorite book is Moby Dick by Herman Melville and that has, I think, 20 em dashes per page the whole way through.
It just has a ridiculous amount of em dashes, almost in place of many other punctuation marks.
You wouldn't see a book written like that today.
It would just read as kind of hopelessly archaic.
Right.
He loved them.
In the late 1800s.
I think you cited a study that does show that the use of the MDASH really peaked in that time.
And so those would all be public domain books now, right?
And so they were easily accessible to the people who wanted to train models.
Is that correct?
Yeah, I'd say that's the gist of my theory that there was this influx of public domain books from the late 1800s, early 1900s that just sort of brought the em dash in to language models somewhere in between the launch of GPT-35 and the launch of GPT-4.
Right.
So semicolons actually peaked during that same time, and I never hear anyone complaining that ChatGPT uses too many semicolons.
So what do you think is going on there with those two different punctuation marks?
Yeah, well, that's a good question.
And it touches upon kind of probably the biggest problem with my theory, which is if this is what's causing the rise of em dashes, why don't Why doesn't ChatGPT write more like Herman Melville?
Why doesn't it write in a kind of archaic style?
And to answer that, I think we've got to kind of talk about the other half of building language models, which is after the training phase, after you've fed all this data through it.
It goes through what's called a – sometimes called a reinforcement learning phase, or a human feedback phase, or many other things.
But it's the process of turning what is like called in the industry a base model into an actual, usable assistant model that wants to help you.
And that's kind of something you can rely on to talk to.
And the way that works, or at least the way it's –
Labs are very, very secretive about this process, but certainly the way that it used to work is there would be hundreds of human beings who would talk to the model and thumbs up or thumbs down responses.
That's what would turn it from this really powerful but not very useful tool into something that actually behaves like a human being.
You have the preferences of these hundreds of people get baked into the model.
Incidentally, this is why ChatGPT, certainly in its GPT-4 days, used to use words like delve a lot, which are not super common in American or Australian English but happen to be super common in Nigerian English, which is where a lot of the reinforcement learning through human feedback was done.
Because OpenAI needed to pay hundreds of people who were quite literate and the combination of low average wage and high literacy, a lot of that work ended up being in Nigeria.
So a lot of peculiarities of Nigerian English got baked into the models early on, with interesting consequences.
Does Nigerian English use more em dashes than American or British English?
I thought it might, and that would have been a really satisfying explanation for this, but unfortunately no.
I did a bit of analysis on public domain Nigerian English text and they actually use fewer em dashes.
No, you can't use that as an explanation, unfortunately.
But the The point I'm making as to why we don't see kind of these archaic constructions and why we don't see a lot of semicolon use is that when human beings are reading and rating language model outputs, I think they often prefer em dashes to other punctuations.
So they would.
They would thumbs down a response that sounded like Herman Melville, but they would.
But they're probably more likely to thumbs up a response with like a snappy em dash at the end.
You know, it reads a little bit more like professional.
I think people associate it with a kind of a magazine like effect like, like a New Yorker style kind of thing.
Right.
Yeah.
I used to be a journalism professor and, you know, so I'm very familiar with journalistic writing.
And I think that's why I'm so comfortable and use em dashes a lot.
It's just a sign of that kind of writing which you're saying then people sort of associate with high quality.
So when they're doing that thumbs up or thumbs down, they're more likely to maybe pick one that has an em dash.
I mean, and I have to say the feedback I get.
People don't like semicolons that much.
And so I can imagine people thumbs downing something with a lot of semicolons.
That kind of makes sense to me.
It breaks my heart.
I'm a semicolon lover.
I use semicolons all the time.
I have to edit them out constantly.
I know.
They're nice.
Sometimes, when you're using ChatGPT or the other models, it gives you two responses and asks you which one you like better.
Are you giving them free reinforcement learning when you do that?
While it's possible, I think the short answer is no.
I think what you're doing when you do that is you're helping them decide which version of a new model they're going to launch with.
So you're kind of giving them feedback on their products, but I don't think that's being fed directly into training.
Incidentally, I don't know if you recall, but the ChachiBT sycophancy scandal where there was like –
Two weeks where ChatGPT would tell you yes to whatever you asked it.
You could tell that an actor was sending secret messages to you through watching their TV show and they wanted you to come and visit.
And ChatGPT would be like, wow, you're so perceptive for noticing those secret messages.
You should definitely go and do that.
Buy that ticket.
Yeah, buy that ticket.
And with other sort of less funny consequences.
But that was due to or at least what OpenAI said, was that it was due to an overuse of direct human feedback, sort of in that style people thumbs up and thumbs down in messages that when you ask people to do that, they overwhelmingly prefer more sycophantic responses, just like they tend they seem to prefer em dashes.
So it's a little bit dangerous, kind of just going off what people like at the individual response level.
You need to be a little bit careful.
I've heard people say that when they get those which one do you like better, that that's a sign that a new model is coming soon.
Do you think that's true?
Yeah, probably.
Yeah.
Interesting.
So there were some other theories that we should probably talk about that you dismissed pretty much.
But one is that em dashes are favored because they use fewer tokens.
So they're more efficient.
Maybe a little bit of context here.
Language models don't think or talk in words, and they don't think or talk in letters.
This causes a lot of interesting behavior that's hard to understand until you know about tokens.
Language models don't The basic unit of language for a language model is kind of like a word fragment.
So, for instance, the word like semicolon would not be represented probably as a single unit in a language model brain.
It would be represented as the combination of semi or colon or, even more confusingly, it might be the combination of sem and then ic and then o-l-o-n.
Basically as the model gets trained or before the model gets trained.
Basically, it learns ways to break words up that are like, quite efficient in representing kind of common words.
And often that will mean that it pulls out like English prefixes.
But it'll also do it in a way that's a little bit less intuitive.
This is one reason why the classic question to stump a language model used to be can you count the number of R's in the word strawberry?
It doesn't think in letters, so it was actually quite hard for it to do that.
It doesn't think of strawberry as having R's in it.
It thinks of strawberry as being made up of letters. straw and then ber and then ry you know it's it's sort of three letters from the from from from the model's perspective but coming back to m dashes the the theory here is that like maybe m dashes are kind of preferenced because of this tokenization thing the core way language models think might prefer m dashes because they're efficient to express as tokens Is this correct?
I mean, maybe it's really hard to disprove stuff like this, but it just doesn't seem particularly plausible to me.
Language models do all kinds of things that aren't efficient.
As you know, if you've ever talked to one, they go on and on and on.
They don't mind kind of rambling around the point.
So this idea that they're using em dashes because em dashes are so succinct?
It would be more compelling to me if I saw other evidence of language models trying to be succinct and succeeding at it.
Mm-hmm.
And what about sort of large collections of data like Medium or Wikipedia?
It seems like some people say well, because it was trained on that and those sites use a lot of em dashes that could have shuffled the whole thing so that it's favoring them.
Why is that not a great argument?
I think something like that argument is broadly correct.
It's just that it's not Wikipedia.
It's the scanned books from the late 1800s, early 1900s.
So I am in favor of the kind of training data style explanation.
But I don't think Wikipedia is a particularly compelling example, just because I think – they would have trained on it for GPT-3.
It was around, it was available.
It would have been right at their fingertips when they were trying to train the early models.
And if that's where the M-dashes were coming from, then that's where it would have been.
When I posted my blog post about this a while ago and there was a lot of discussion, I saw some people claiming that they were responsible for the M-dashes in.
In particular, I think it was the CEO of mediumcom said that he felt responsible for mdashes because the people who built mediumcom were typography nerds and they made sure that mdashes were kind of – double dashes were automatically converted to mdashes, for instance in medium.
So it rendered more mdashes.
I don't find that category of explanation compelling at all because – I mean, Google Docs does that.
Yeah.
Word does that.
Yeah.
The question is not why ChatGPT uses specifically the em dash punctuation mark.
It's why it uses em dash, the grammatical unit, like why it does em dash things.
It could do them with a single dash if it wanted, and it would still be as puzzling.
Yeah.
You know, I have to say, as a professional editor, I do love that I'm seeing more proper em dashes these days, because you know, people who aren't professional writers will often use the hyphen in place of a dash and that drives editors crazy and we always change it.
And ChatGPT, at least, it might overuse it, but it uses the right one.
Well, I apologize because I know I've done that in my blog, so you probably encountered that in your research.
I
I apologize for that.
Oh, I didn't notice.
That's funny.
I was so captivated by your argument.
I've actually heard people doing the opposite of deliberately using the short dash, the short incorrect dash, instead of the em dash, so that they can do em dash things without it being immediately read as chat GPT.
I've heard that now too.
I think it's so sad that people are changing their writing style to try to appear less like they're writing like ai like to see people changing their natural writing style because of ai makes me sad and frustrated we shouldn't have to do that but i understand why people do too but i see people even talking about putting in typos so their work looks more human like Oh, please don't.
We should all still be good writers.
So you referenced some work from, I hope I get this name right, Maria Sakharova.
And she talked about how the fact that you know there's so much more AI writing appearing on the Internet.
And it's going to have maybe these hallmarks of AI writing as people talk about with Delve or the MDash.
And then that's going to get sucked up into the new training data.
And it's going to become this downward spiral where all these hallmarks get amplified even more in future models.
But it sounded like you were saying...
ChatGPT-5 doesn't maybe use as many em dashes.
So what do you think is going on?
That seems like a reasonable argument, but what might be going on there?
Yeah, this is sometimes called the well-poisoning argument or the synthetic data argument.
It's gone by a couple of different names.
People were talking about this a lot in the early-ish days of language models.
I remember in 2023, early 2024, people were taking this very seriously.
I think, by 2026.
It's pretty clear, certainly from where I'm sitting, that that hasn't happened and that it would have happened already if it was going to happen.
We've had like three or four generations of language models trained, since there was a ton of LLM produced data on the internet.
If it was possible, like if you could not avoid having an LLM poisoned by this, then all current LLMs would be poisoned.
And I think it's pretty clear just that they're not, that just hasn't happened.
Why hasn't that happened?
It's probably some combination of AI labs putting in a lot of effort and being successful at filtering out examples of this from their training data.
Absolutely, they've got the capability of doing that.
But it could also be simply that it might just not be that big a deal.
Language models are trained on all kinds of information right now, some of it very low quality, some of it very high quality.
Part of the training process is the model learning to distinguish the one from the other.
So it's possible that you could train a model on like 75 of sort of synthetic data, like AI generated data.
And, as long as it was properly guided, it would classify that as something to avoid rather than to copy and would end up okay.
You know, just yesterday I was reading about this phenomenon where AI comes up with the same name for people in professions.
So if you ask it to generate a name for a software developer, it will almost always name that person Marcus Chan, and a sci-fi protagonist is very often a Lara Voss or some slight variation of that.
Have you heard about this phenomenon?
And if so, does that play into this anyway?
How does that happen?
Oh, that's fascinating.
I know you haven't written about this, but No, I hadn't heard those examples before, but I am familiar with the phenomenon.
Language models are non-deterministic and they're hard to predict, but they aren't random.
When you have all the big AI labs training in basically the same way on basically the same data which they are.
There are lots of little subtle differences in how they're doing it, but all the people at these labs live in San Francisco.
They all know each other.
They all talk.
They all go to the same parties.
Fundamentally, they're doing very similar things.
I think what we're seeing is that that kind of leads to very similar results even at the level of specific questions like what you name characters.
It's just the space of all possible language models is a little bit more defined, I think, than we might have expected from the start.
And there are these attractors in the space that are more specific than we expected at the start.
And it's just...
If you do the same things, you're going to get the same outcome.
Why these particular examples happen, I could speculate, but I don't know.
It's some pattern in the training data that's getting kind of focused on.
Yeah.
The thing that made me think of it is because I went to Amazon and I just searched for a Laura Voss and I could see there were, you know, multiple self-published novels that had that same protagonist.
And so you know people were saying well, this is how you can tell the book was written, most likely with AI, because it's picking this name that is so common, coming out of AI.
And it made me wonder then, if well, if those books were then used in future training, like we were talking about the AI for the MDASH it's like is that going to like reinforce the idea of these common names appearing again and again?
And so maybe if the companies are aware of it, they can just filter it out so it doesn't happen.
Or maybe it's such a low rate that it wouldn't matter.
But it seems like if it's not happening with the MDASH, then maybe it wouldn't happen with these unusually common names too.
I don't know.
What do you think?
I would not expect this to be like a huge problem for new generation models.
I would not expect it to become more common that models name sci-fi protagonists the same thing.
I would expect that to kind of fade away just by inference from like the previous examples of like common model behavior that have been kind of trained away.
I mean, the em dash is kind of one of those things.
Like the reason it hasn't been fully trained away, I think is because people like it.
Like it's They could eliminate the em dash if they wanted.
I think it just does so well when they do kind of human feedback testing that they've kept it in.
Yeah.
Sam Altman gave an interview where he said they put it in because people liked it.
And then people were debating whether that was an intentional thing or maybe it was just.
He was saying they saw that people liked it in the human feedback part.
I certainly don't think they intentionally put em dashes in the model.
I don't think it works like that.
To the extent that it was intentional.
It was noticing an existing behavior and choosing to reward it and keep it in rather than try and get rid of it.
Well, speaking of signs of AI writing, in the bonus segment we're going to talk about AI detectors and the flaws in them.
We're going to talk about something that might be a little controversial.
Sean suggests that people should write for AI, which will be an interesting discussion.
And then we're going to get his book recommendations.
So, Sean Gidecki, thank you for being here for the main segment of the Grammar Girl podcast.
What's your blog?
Where can people find you?
Thank you so much, Mignon.
People can find me at my blog, which is www.SeanGedeke.com.
That's G-O-E-D-E-C-K-E.
And that's basically it.
I don't have a social media presence.
I just write on my website.
Wow.
Wow.
Old school.
Very much.
Well, thank you so much for being here.
Thanks so much.