Hammer Girl here.
I'm Mignon Fogarty.
Today, I'm here with Abigail Eisenstadt, a science writer at Science magazine.
I'm really interested in speaking with her today because They did a study on AI writing that was much more expansive than what I see most people doing.
You know I see people give Chachapit a try once or twice and you know form their opinion about it.
But as good science people do, they did a methodical test.
And we're going to hear all about it today what they found that it can and can't do, and what it means for the future.
Abigail, thanks so much for being here.
I'm happy to be.
Thank you for inviting me.
Yeah.
So can you describe for people sort of what the science writers of science did?
Yeah.
So we are a press package team and we send out summaries to reporters on a newsletter each Sunday.
And what we put in that newsletter is 250 to 350 words of a news brief, essentially.
And so these are summaries that go out.
There's usually five on average summaries.
And what we wanted to do is we wanted to see if ChatGPT Plus could emulate our specific style.
And so our style is pretty typical for a news pyramid, whereas typically for a pyramid, you would expect the most important sentence to be actually at the bottom.
So, most important sentence at the top then we do the background, then we do the methods, then we do a conclusion.
And we wanted to see if ChatGPT could emulate that essentially.
And as well as science writers, we have to be very careful about the word choices we use, because we don't want to say groundbreaking a new, you know, first of its kind, study.
All science is built on the shoulders of other science.
So you can't truly say anything is novel.
You have to maintain context in whatever you do.
So each week we had ChatGPT analyze two papers that the writers on my team, including myself, would nominate.
And we compared its summaries to our summaries.
So we had already written our summaries on it as well.
And then we would look at the differences between the two to see if there were any hallucinations, any claims that we felt were a little too aggressive.
Would it be something that we would be confident sending out to reporters that we believed would not break their trust?
They could trust our credibility as a representing research from our institution.
And also if it even followed the style that we use.
And then, over the course of that experiment, what we found is it did a good job transcribing studies.
So it could summarize in sort of a lay person's abstract, but it didn't really translate those studies.
It was missing the element of contextualization.
It loved to use groundbreaking, for example, and it didn't quite provide the narrative within, like this study that exists within a field of other research.
And ChatGPT couldn't really do that part, which is so critical for science writing.
Can you explain a little bit more what you mean by the difference between transcribing and translating?
Like, maybe give a concrete example for people?
I would describe translating as when you take a sentence and you don't just restate the sentence with a different word for each.
You know each word.
You also... explain what the sentence means.
So transcribing is just resummarizing.
Translating is really broaching that understanding element to it.
And again, I would just call it adding context for the sentence, providing the sentence beforehand and the sentence after.
So you have a setup.
Sure.
And so how many, you did this for a lot of weeks.
You did a lot of different tests.
How many did you do?
Yeah, we ran this for a year.
So, you know, give or take.
I would say like the papers that we studied were over 60.
So I would call around 64 papers.
And we definitely did this for at least 40 to 50 weeks.
So it was a pretty wide range of content for a pretty long time.
But, you know, there are some human biases involved because we're writers.
So we also had to be cognizant of that, which is why we spent such a long time doing the project to try, to you know, balance human bias human, the human side, with you know a little bit of a longitudinal side.
And one thing I think is important to mention is this was actually in AI terms.
This was quite a while ago.
So the time frame.
You know, one thing I was as I was reading your report you just referred to ChachiPT Plus.
And I was like, well, which model did they use?
And it turned out there were multiple models because this field advances so quickly.
So can you talk a little bit about sort of the time frame this happened in and how, how maybe you saw things change over that time?
Yeah, that's a great topic, too, because I think when we started, it was different. around December, 2023, we had this idea and that was kind of before prompt engineering had percolated into the public conscious and I wasn't a developer and no one on my team is.
So our prompts were pretty.
I would call them pretty basic in terms of what you would ask a human to do, rather than specific for what you would ask an AI to do, knowing what we know now.
And so we started with I would say just generic ChatGPT+.
We had a coworker who would put the summaries in, and we just tried to maintain the same model.
It was all 4.0, I believe, if that's how you say it.
As you can see, I haven't interfaced with the AI itself that much.
But we tried to account for model diversity by keeping those prompts the same.
But those prompts stayed the same, and so they did not really evolve according to prompt engineering education as it entered.
I guess the zeitgeist over 2024.
Yeah, that was going to be.
One of my questions actually is as the models change, did you change your prompts to sort of adapt to how the models are changing?
No.
At that point, we I mean, you know, there was human biases involved.
There were the paper.
Topics were changing every week because we're covering, you know, at the press package.
We cover studies from science immunology, science magazine, science advances.
So this is archaeology, neurology.
And, you know, that's hard to standardize as well.
So we kept them very much exactly the same.
And we kept a very we had a very general summary.
Just like write me a paper in layman's language, layperson's language on this one study.
Write me a precise summary on this study as you would for a high school reader.
And then write me X, Y, Z, our specific outlining process prompt.
And we never changed that.
So I think they were all written to how you would instruct a human who had a bit more creativity and ability to interpret the prompt.
Did you give it examples of human written summaries?
So you would say like write a summary kind of like this, and give it like three examples, or something like that.
Yeah, so we didn't give an examples, but we would say write it akin to a news study or a news story that you might find in a magazine.
But again, the examples component hadn't really been on our radar when we started in January of 2024.
I will note, too, our coworker was experimenting with those things.
But at that point, we didn't want to, you know, those are outlier data.
And it was interesting what it could do.
The main point I want to stress is that this is a very selective study with a lot of factors going in it.
I would call it, well, my supervisor calls it sandboxing.
And I think that that's a great term because, you know, we did what we could with what we had.
And I think the results are pretty representative of how, what a lot of people are finding.
But also there are caveats.
So I'm glad we can get into those.
Yeah.
And so I think it ended.
Your study ended right about the time when, I think, the first reasoning models became available.
So it sounds like that you didn't do it with any of the reasoning models.
Is that right?
That is correct.
Yeah.
So how did you nominate the... How did you choose the articles to nominate for AI treatment?
Yeah.
So we kind of used a process similar to how we... select our summaries.
So the first qualification is it had to be something that we had already written a summary on, because we have to be able to compare it to.
And then we had a series of reasons why it could be nominated.
So there's technical jargon.
Is this something we really want to see the AI delve into and see if it can unpack it accurately?
Are there human subjects involved?
Because when you're doing medical coverage, you want to make sure you treat any study with human subjects with sensitivity.
There are phrases we will avoid.
So you never want to say this study involving disease participants.
You always want to say participants with XYZ disease instead, because it's a matter of personhood.
We also had, is it just a fun topic?
Like, are we just writing about mammoths for some reason?
Is it a particularly controversial topic? topic.
We had a study come out on Facebook and election biases.
So can the AI handle that and make sure to represent nuances that we would be very careful in discussing?
And so those were the qualifying factors.
And I had each writer complete a survey at the end during their assessment, so that we could see the breakdown of what type of topics were nominated.
But again, there was so much diversity.
And also each study has diversity in how it's conducted.
So the AI is tackling... variety of different methods and approaches.
So it's very hard to standardize, you know?
Yeah.
And so, as you discovered things, did you give it rules like you know, always describe you know a person with this instead of you know like We did not.
What I would say happened over the course of the experiment is we too learned how to evaluate more.
There were certain things like okay, if it's going to say people with X disease versus you, I'm going to flag that once and then ignore it.
And we'll just assume that that's a continuity we've already noticed.
Um, there was a point where we finally changed the prompt a little bit to say please stop saying groundbreaking, particularly because it just it was.
You know it's something that you can't tune out when you read it and you have to flag it as inaccurate every single time.
So that was.
That was the one intervention I would say that we did.
But yeah So, not even subconsciously, I think, just in general, we all kind of unilaterally agreed that this was not.
You never wanted to see that again.
Certain things were just let it go, yeah.
I'm curious is this something that you, as the team of writers, decided that you were really curious about and wanted to do?
Or is it something that management came to you and said, well, we want to see if it can do this.
Will you please test it?
So it was a little bit of column A and column B, I think.
When ChatGPT Plus first came on the scene, there was a lot of debate about whether the LLM would replace writing jobs.
And so I would say that a lot of the writers on my team and myself had differing views, or not even differing views.
Just you know.
What does this mean for us?
I've always been of the stance that, you know, like any event, the industrial revolution or whatever, we will evolve and our skills will be suited to use a tool eventually, in theory.
And I think Perhaps I was blasé enough about it that my supervisors noticed, and so they asked me to run it, stressing that nobody would be fired.
It was just a way to, you know, stay on top of the landscape and whatnot.
And because, again, I was like, well, sure.
Because I wasn't too concerned about it.
And I think it was a really good experiment to do, because I actually have more respect for these tools afterwards.
I think knowing limitations allows you to respect something more.
Yeah, absolutely.
And how did you evaluate the output?
So I tried my best, without a scientific degree, to create a survey where we had a few qualitative and a few quantitative answers.
So what I did was I would say, did this emulate?
Was this compelling?
Was this a compelling summary, yes or no?
Was this a summary that is feasibly able to blend into the rest of our press package?
Because if it's not, we couldn't use it.
Can it stand alone and go into the press package?
On average, all of the scores that came from the course of the year were all in between two and three.
So there's decimal points in there, but I can't remember.
Out of five?
Yeah out of five, which is like not too bad, but does indicate there's a requisite for editing or prequisite.
The qualitative was just I would ask the writers for... you know, their thoughts.
And those are in the paper, but more as sentences, because I'm not going to, you know, be like.
Certain phrases keep emerging because again, we're writers.
We all have our own, you know, schticks that we lean into.
So there was no way to quantify those.
Yeah.
Yeah.
So you so the.
It was the writer.
The writers in your group were the one evaluating the output and comparing it.
It sounds like you're saying it was pretty steadily like between two and three.
There weren't some that were one and some that were five.
Yeah.
So we had different formats of papers.
That's another thing I should mention, which is why it was kind of hard to adapt to the prompts we had.
We would put in policy forums, which are something that science magazine specifically publishes, or we would have reviews or research resources, which are things that the sibling journals use.
And so all of those, it kind of performed differently on.
It did very well.
On reviews, I will say but again, that aligns with the transcription.
You know it's good at resummarizing something and a review.
It's kind of a translation.
So if you transcribe a translation, you do pretty good.
In terms of research articles, there were a few where I was just.
I received some feedback in a survey from my colleague Walter, who was like are you sure you put the right paper in this?
Because it's so off base.
So that was startling.
And I had one where it started talking about bacteria in the brain, which was also a little concerning.
But...
On the whole, yeah.
They were pretty consistently at the level where they required a lot of editing to improve, but I wouldn't dismiss them off the bat.
But the amount of editing that it would take to improve, I would not bother with.
Right.
Yeah.
In some ways, I'm not surprised that you didn't find that it was helpful to your group, because you are probably some of the best science writers in the country, in the world.
You know, I think that one thing that people who've tested um Chachupati and the like have found is that, you know, for people who are already really good at what they do, it doesn't help them that much.
They find that like, people at sort of the lower skill levels get more of a benefit than people at the higher skill levels.
And I wonder, if you feel like you know for a less skilled group, if something like you did might be helpful.
So I think that it depends on audience a lot and kind of the context you have.
That comes in to it for sure.
We actually had a study published at Science Advances which is my eye report on Science Advances studies specifically.
It's our open access journal.
And it came out talking about how abstract language has changed from LLM use.
And I don't know if that's necessarily for the worst because, if you think about it, who is the audience for these papers?
You know, often research papers involve a lot of passive voice.
If LLMs can reduce that audience, that's pretty great, because then it kind of percolates more into society that we don't need to use as much passive voice as there always is.
Um.
So yeah, I think it depends.
If you are writing for an audience of your peers and you're a researcher, you know it's a good idea to have a little bit of more colloquial influence in there.
If you're writing as a science writer, you pretty much have your style down pat.
Like you, have to think about all these different factors and our audience is specifically our audience at the press package as reporters.
So they're pretty sharp.
And then as a reporter, you have to walk the delicate line of your audience being the public.
And then you have to be very careful about what you say, because every word matters and you don't want to misrepresent something, but you also want people to keep reading your study.
So I think it does provide a useful tool for people who are not writers, in so much as it can influence the styles in a way that changes the general trend for society, not in so much as each paper it tackles gets better.
Yeah, sort of a general flattening of style.
So, not long after you released your results, OpenAI released a study where they did something exponentially bigger because, of course, everything they do is huge.
And they looked at more than 1,300 tasks across 44 industries.
And they had experts come in and write the prompts and then evaluate the output and everything using the newest models available.
And some of those were science writing tasks.
Now, they didn't break it out. task by task.
So even by science writing task.
But I wonder if you saw that and if you had any thoughts about that after, like finishing up your study and then seeing that, and if you've had any thoughts on like compare, compare and contrast or anything like that.
Yeah, I had time to review that.
And I think overall I have heard good things about Claude and that seemed to be one of the main takeaways of the study in the weeks since putting out our own report,
And, you know, again, I wasn't super aware of Claude when we started the experiment.
But I will note also the overall results for the study, the OpenAI study specifically, where Claude was graded as better or equal to human writers 475.
So again, that kind of brings back the issue of return on investment on editing.
That's not super convincing to me.
Right.
And they didn't break it out.
So we don't know for their study either if it was there were ones and fives or if it was all 40, if they were all half as good.
Yeah.
Yeah.
So, and then I also noticed that at the bottom of page three, which is this is very incremental, but there was a note too that it's hard to you know see if the summaries are fully blind, because a lot of these different models have their own quirks.
So yeah.
Perhaps there was an ability of reviewers to see if it was human or not.
You know, or just assume, because there's a lot of M-dashes, as the myth goes.
Maybe they assume it's Chachupati.
So you know, I think there were flaws in that study as well, but I do think the results were promising.
Also, they noted that there was a heavy cost in that study, and that doesn't seem super surprising.
Right.
I think the cost was in paying all the experts because they they had experts with an average of 14 years experience in each industry.
And then they they spent I think they spent an extraordinary amount of time writing the prompts.
The prompt underwent five rounds of human review before they were tested.
So the thing about that study is it was as good as it possibly could be today.
I think is what that study says.
So if you're writing the best prompts possible by the experts in the world, then for writing it was about, And again evaluated by these experts, it was is good or better than the human output about half the time for writing and about 75 of the time for editing.
So those were numbers to be aware of and that I'll be keeping my eye on, because those are pretty good numbers actually.
I think they're good numbers.
I think, again, it depends on the field.
I don't think for my specific field they're good numbers.
I think they have great potential to be useful. for other fields.
Yeah.
And I think it depends on how you count the time that then you have to spend cleaning it up for the things that don't work and and you know whether that and the time going into creating the prompts
I mean, if you can use the prompt over and, over and over again, then maybe that cost gets amortized across all your work.
But if you're having to do that for every task individually, then it's, it's too much, you know?
Yeah.
I can't imagine the level of expertise and and training that I I'm very impressed by the human reviewers and prompt engineers involved on this project as well.
Yeah.
Yeah.
It's a lot to keep an eye on in the future with the new models.
Do you, I mean, do you have any interest in testing it again?
I mean, you've done this huge test.
I can see you being like, we just, we're done with it.
We don't want to do it anymore.
We we're good at our job.
It's not saving us any time, but like then the models keep improving.
So do you think that you'll test it again or are you done?
I think, It's difficult to say.
We're definitely curious about different models, but to that point they are evolving so quickly that I hesitate to do another longitudinal study because things might change so drastically again.
We have thought about asking reporters if they'd be interested in a shorter-term study where they evaluate our summaries and the LLM summaries side-by-side.
Unclear about what model that would be either.
But again, reporters times are pretty limited, especially today.
And so I don't know about the buy in with that.
Also, when you're selecting reporters, do we go for whoever wants to join?
Do we pick, you know, non-technical outlets?
And by technical, I mean like not wired because they already know a lot about AI.
So it's kind of it would be very difficult, I think, to standardize, in the same way that this study was, or sandbox study was, standardized.
So I don't know.
It's difficult to say.
We're definitely curious.
And I don't think we're done staying on top of it.
But I don't know if we would launch something this grand at scale.
Yeah.
No, I mean, I commend you for doing it in the first place this way.
I mean, it's so nice to see it all laid out like that.
I had a follower on social media who is a science writer who asked me to ask you since you spent so much time with these models, is there anything you think that they are good for?
Yeah, I actually think that they might have a lot of potential for social media.
And I would be very curious at how they do at distilling our summaries for social media.
In that same vein, I have heard and I am very curious about.
I've heard of writers putting their summaries or just their pieces I don't know if they're summaries to say but putting them into the models and asking what are the main points.
And I'd be kind of curious if I could get something from that.
You know, like if I put in a study on biomolecular engineering and I ask it what is the news here?
Maybe it can tell me if I'm actually presenting that correctly.
And I guess that would bring out more of the analytical side that we saw was lacking in this study.
I keep calling it study, but that was lacking in our experiment as well.
Yeah.
Yeah.
Well, Abigail Eisenstadt from Science Magazine, science writer, thank you so much for doing the tests and for talking about them with us here today.
Thank you for having me.
It was great to think more about it and, you know, discuss it with you.
Yeah.
So interesting.