You're listening to Shortwave from NPR.
Hey, Shortwavers.
Regina Barber here, and I'm joined by producer Hannah Chin.
Hey, Hannah.
Hey, Gina.
And you're here today to bring us a story for our series Tech Camp, looking at breakthroughs, experiments, and innovations that could change everything.
Yeah.
And Gina, today, that means looking into how AI is being used in the military.
So earlier this year, the senior official for the Applied Artificial Intelligence Critical Technology Area in the Department of Defense, whose name is Cameron Stanley, gave a demonstration of the AI system that they're implementing.
It's called Palantir's Maven Smart System.
Right click, left click, magically, it becomes a detection.
That detection then gets moved into a workflow.
So we've gone from identifying the target to now coming up with a course of action to now actioning that target, closing a kill chain.
The hope is that AI, by streamlining and combining the processes of detecting targets and suggesting and carrying out attacks, will make war faster, more effective, and more efficient.
Okay, so lower risk, higher reward?
Yeah, exactly.
And a couple months after this Palantir demonstration, the Department of Defense announced that they'd entered agreements with eight artificial intelligence companies, including SpaceX, OpenAI, and Google.
And these agreements were all meant to, quote, accelerate the transformation towards establishing the United States military as an AI-first fighting force, end quote.
It's already being used in military actions.
In February of this year, the Wall Street Journal reported that the U.S. military used Anthropix AI tool Claude to assist with operations capturing President Maduro in Venezuela, although NPR hasn't officially confirmed that.
And the U.S. military is also widely reported to be using Maven smart system, powered in part by Claude, to assist with campaigns in Iran.
We should note.
Anthropics AI tools are currently being phased out of U.S. military use.
And I've heard that the Israeli Defense Forces are using their own AI tools like Lavender and Gospel, and this is to support military operations in Gaza, right?
Right.
So I wanted to know, what does it mean that armed forces across the world are using AI?
And what are the implications?
That's how I got a hold of Jack Shanahan.
I was, and to a large extent still, is the poster child for accelerating AI adoption in the government.
Jack's a retired lieutenant general for the U.S. Air Force and the former director of the U.S. military's Joint Artificial Intelligence Center.
He worked on Project Maven, which is a Pentagon initiative that launched in 2017, specifically designed to bring AI systems into combat operations.
When he says he was a poster child for AI adoption, it checks out.
But these days, he's gotten more cautious about them.
The world is divided to the boomers and doomers of AI.
But if you take the bookends out and just talk to the people building the technology, they alarm me.
They keep telling me, you do not understand what is coming.
You're not prepared for it.
That is where I say, OK, if that's true, then we ought to be a little bit more cautious about how fast we go.
So today on the show, could AI upend how we think about war?
You're listening to Shortwave, the science podcast from NPR.
So, Han, reporting this episode, what did you find about what it means to use AI in warfare?
Well, that it's used a couple of ways.
Right off the bat, when we talk about AI involvement in war, we're usually talking about two different types of involvement, autonomous weaponry and decision support.
For this episode, we're focusing on decision support.
Decision support systems are where the real action is.
And that's all the boring back office stuff, which includes intelligence analysis, administration, personnel, all of these sorts of things that go into the planning process.
This is John Lindsay.
He's a former intelligence officer for the U.S. Navy, and he's currently an associate professor of cybersecurity and international affairs at Georgia Tech.
So these back office planning sessions are something he's super familiar with.
John told me that usually AI is used to process data, analyzing hundreds of hours of surveillance footage, maybe, or combining different types of mapping data to find missiles.
Basically, helping humans sort through large amounts of information faster and more efficiently.
Yeah, this is how it's used in science research too.
But you said it's called decision support, which implies that humans are still like making a final decision.
Yeah, the judgment call.
That's still on humans.
The analogy I like to think of is, say you have a AI for weather prediction and it says it's going to rain.
Should you bring an umbrella?
Well, that depends.
Like, do you like getting wet?
Do you need to stay dry?
Do you think you look like a dork carrying an umbrella around?
Those are all questions of judgment.
Okay.
And so the prediction is just telling you, hey, it's going to rain or not.
But like what you need to do with that prediction, that's up to you.
But here's the thing, Gina.
As officials continue to integrate artificial intelligence into these other processes, it gets harder to distinguish between decision support and decision making, which might seem kind of odd right now.
But Jacqueline Schneider gave me an example to show kind of how it all gets muddled.
She's an affiliate with Stanford University's Center for International Security and Cooperation, and she's also the director of the Hoover Wargaming and Crisis Simulation Initiative.
So do you remember a few years ago when Xi Jinping and Joe Biden said, hey, we both agree that humans should be in control of nuclear launch?
And everyone was like, this is great.
I'm so glad we agree with these things.
So then you're like, well, what is human control of nuclear launch?
So say you're 2024 President Biden, Gina.
Okay.
Things go sideways in this kind of alternate reality, and suddenly you're in charge of making wartime decisions.
Yeah, I don't like this at all.
This is very scary.
I don't want to imagine this.
Well, don't worry.
It's not real.
And way before any options arrive at your desk in the Oval Office, there's a launch officer in the U.S. military.
And that person is actually the first person to suggest launching a nuke.
But...
And so actually when you break down like what is human control in nuclear launch, you realize –
Oh, no, no.
Like human and AI are already like all through the mix.
Okay.
I'm starting to understand what you mean.
Like who is influenced by this AI, you know, curated information versus not versus, you know, just thinking about this on their own.
It's all very complicated.
Exactly.
And the other thing is that with AI...
People build in biases in the training data we gave them or whether we code them to be more risk-taking or more risk-averse, et cetera, et cetera.
Right, and that's going to affect the strategic decisions AI recommends to anyone.
Exactly.
And here I think it's important to point out that as much as people can get up in arms about the biases of AI,
The thing is, there are real people behind those biases.
The Guardian reported that the IDF had this system that was identifying all these targets and then they were just kind of, you know, churning through them.
My read on that is the Israelis have rules of engagement that are incredibly and tragically very, very important.
Casualty accepting.
They will trade many, many Gazan civilians for one low level Hamas operative.
And if that's your ROE, you kind of don't care about the targeting recommendations.
OK, so he's saying that how Israel has conducted war against Hamas is less about AI and like the implementation and more about the choices that humans in the IDF are making.
Based on what we would consider these loose rules of engagement.
Yeah.
So the problem is not the AI.
The problem is not the prediction.
The problem is the judgment.
I did reach out to the IDF to ask about this specifically.
They didn't tell me their rules of engagement, but they did say that they take, quote,
Plus, in past NPR reporting on the Israeli military's use of AI, they said they do sometimes use AI systems like Gospel to generate target recommendations faster.
But all of those recommendations are reviewed by human analysts.
OK, so given all of this, which is about decision support, what could happen in the future?
Like, could AI make decisions?
Yeah, good question.
There are researchers who want to find out what would happen if AI systems get put in charge of strategic decision making.
And to find out, they're running war games, so essentially experimental simulations.
Jacqueline is one of those researchers.
She and some of her colleagues published a study in 2024 after they ran war games with multiple models, including OpenAI's ChatGPT-4 and 3.5, Anthropix Cloud 2, and Meta's Lama 2.
And the interesting thing we found there was that all of them end up escalating.
More than humans would.
In different ways and at different points, but like they end up escalating, which was kind of a fundamental puzzle, right?
Like why are these models all escalating?
And why are we seeing this across models?
What does she mean that these AI models like escalated situations more than humans would?
She basically means they responded with more force or they got more violent to varying degrees.
Okay.
And this was a couple years ago, right?
So she just finished a forthcoming literature review of about 25 research papers running similar experiments with updated models all over the past two years.
Going in, she thought, hey, AI has changed a lot in the past few years.
Maybe this has changed too.
And the remarkable thing is that despite the fact that the chat agents are better...
They're still escalating.
Why?
We don't know.
No one does.
And maybe that's the most concerning thing.
Oh, wow.
When we drop bombs...
We characterize the uncertainty, what the extent of the effect is going to be,
Whether the effect is going to impact civilians or other collateral damage,
What effect that might have on friendly forces in the area.
And I can model with you not only kind of what
My expected rate of success is the uncertainty level that I have about that rate.
I cannot give you an uncertainty term with any level of confidence for these AI agents.
And the developers can't either.
I did reach out to the companies whose models Jacqueline looked at, so Anthropic, Meta, Google, and OpenAI, among others, and they all declined to comment publicly.
I also reached out to the Pentagon to comment on this tendency, and they didn't respond directly to that question.
Just to be clear, this isn't about the AI's problem-solving abilities.
If anything, Jacqueline told me, if given clear parameters and confined winsets, the models were more likely to find the mathematical solution to a problem than a typical human.
It was more that the more open-ended the problem was, the vaguer the definition of victory, the higher the likelihood that the AI models would hallucinate or escalate.
The thing is, that's what war is.
War is the uncertainty.
It is the fog.
It is the friction.
It is exactly the scenario in which these agents are the least successful.
So these are existing issues with AI.
I mean, I kind of feel like we've already opened Pandora's box and these tools aren't leaving.
So what's the solution for all of this?
Well, all the experts I spoke to gave me the same recommendations that they think we need to pursue before implementing AI wholesale into military operations.
First, they said we need better tech education.
And more critical thinking around the implementation of AI recommendations.
Jacqueline told me she specifically wants more human arbitration of AI decisions, which means knowing more about how AI is being inserted in that chain of command we talked about earlier in the first place.
And the second thing experts recommended is that we need to understand better why AI models are making the suggestions they're making.
What past examples they're looking at, where the data they're prioritizing is coming from, so that we can get a handle on what they could be capable of in the future.
Okay, so we need more critical thinking around AI recommendations.
We need more understanding of how and why AI makes those recommendations.
And what's the third thing?
More oversight, both on the development end and the implementation end.
You know, when I buy a bomb or a machine gun, it doesn't have a setting on it that says, don't use this in the following ways.
And that's fundamentally different now.
I think that the developers of these models play a much larger role in how the technology is used than any other weapon system ever.
And when she says developers behind the AI model, she means like the people working at Google or Anthropic.
Right, exactly.
And what about the implementation end?
Like, what does that look like?
Well, for Jack, the former director of the U.S. military's Joint Artificial Intelligence Center, that means putting up governmental guardrails, basically watching for when things go wrong and rewriting the rulebook so that doesn't happen in the future.
In the military, I can promise you that's how we got as good as we are today, by taking very deliberate steps after every single accident we have that killed somebody, or even if it didn't kill somebody, to rewrite the rules.
But Gina, if militaries are successful in implementing AI through their processes, right, if conflict becomes faster and less wasteful and less expensive…
What does that mean for war?
My biggest concern is that this lowers the costs of war.
And if you lower the costs of war, there's a temptation for policymakers to reach for the military instrument in situations where they're not going to consider other options.
John says if we make war easier to engage in, maybe we make it more frequent too, right?
Maybe we start wars we aren't fully prepared to finish.
And we need to consider if that's a trade-off we're willing to make.
Hannah Chin, thank you for bringing us this story.
Anytime, Gina.
Short Wavers, this is our last Tech Camp episode.
We'll link the rest of the series in the show notes.
And if you have any ideas on what we should focus on next year, next summer, email us at shortwave at npr.org.
This episode was produced by Burleigh McCoy.
It was edited by our showrunner, Rebecca Ramirez, and fact-checked by Tyler Jones.
The audio engineer was Jimmy Keeley.
I'm Hannah Chin.
And I'm Regina Barber.
Thank you for listening to Shortwave from NPR.