Imagine a moment in the not-too-distant future when thousands of AI agents are making decisions on our behalf.
What could possibly go wrong?
McKinsey partner, Rich Eisenberg, says leaders need to start managing the risk of these agents' stat to get the benefits of AI safely.
These are not set it and forget it.
They need to be consistently monitored and tuned and tweaked and sometimes fired.
And it has to be clear who are the owners of these agents and who's accountable for their performance.
And I think that's become the hardest question in this new operating model construct, where the answer is it's not always the technology team anymore.
This is the McKinsey Podcast, where we help you make sense out of the world's toughest business challenges.
I'm your host for today, Lucia Rahilly.
Rich Eisenberg, welcome to the McKinsey Podcast.
Thank you.
Happy to be here.
Today we're going to talk about one of the hottest issues of the AI moment, and that is AI and trust.
So tell us data security, cybersecurity we've been talking about these as a priority for leaders for years.
How does agentic AI specifically up the ante?
A way to think about it that I like is you know, people sometimes think about Agentech as just a better chatbot.
In reality, it's delegated agency.
So you're giving these software programs agency.
They can plan, they can call tools, they can execute workflows.
And in order to manage this, your entire operating model has to change.
You know, Agentech AI is not only content generation.
It's decisioning and action at machine speed.
So the questions often shift from is the model accurate to who is accountable when the system acts.
Your own governance, you know, must define scope, inventory, ownership, and make it auditable.
So this transformation to agentic is agency isn't a feature.
It's a transfer of decision rights.
And in the article you cite research saying that 80 of organizations have encountered risky behaviors from AI agents.
Help us understand a little bit what's at stake.
What's an example of this kind of risky behavior?
I'll give you two recent well-cited examples.
Anthropic has published some research where they've been doing some interesting work of can you make an agent behave poorly in an enterprise environment?
And they do it across multiple LLMs.
And one example, they inserted some emails from senior executives that the agent had access to read.
And the discussion was about shutting down the agent.
The agent then independently started to mine this senior executive's personal emails, discovered he was having an extramarital affair and independently started submitting blackmail emails to the senior executive to not shut the agent down.
Second example will be there was a customer service agent, which is a common use case.
We're replacing customer service intake with agentic technology.
They pose as a customer and are repeatedly asking the agent, are you a computer or a human?
And an agent was so insistent it was a human.
It threatened to show up at the customer's front door wearing a blue coat and a red tie.
A flaw in one agent can propagate downstream and massively amplify the impact, right?
The agent risk isn't just about wrong answers.
It's wrong answers at scale.
And the thing executives need to have in mind is the scariest failures are the ones you can't reconstruct because you didn't log the workflow.
It's amazing to hear these examples.
To be clear, they were simulated to stress test agentic AI.
And Anthropic says no human was hurt in the process.
Leaders are under huge pressure to deliver ROI quickly from AI.
And these are obviously dramatic examples.
But how do you see leaders balancing that pressure, striking the balance between the really ambitious productivity goals that are being bandied about and then the potentially less sexy concerns like governance and risk mitigation?
What we find is most people, through brute force, can get one or two use cases up and running because they get the six magical smart people in the room and they have a debate.
But that doesn't scale.
And if you think about, after most companies have completed an AI transformation five, 10 years from now, they'll have thousands of agents running around their enterprise.
So the way to really split the difference of how do I enable innovation and allow my firm to capture the tangible benefits of lowering the cost of innovation, accelerating time to market, transforming your cost basis of operations and getting more efficient and productive?
And the winning pattern is really archetypes plus tiered approvals plus monitoring, right?
What we find is most organizations suffer from fragmented assessments, inconsistent steps and multiple committees reviewing the exact same issues.
What you have to do is really accelerate safety by using agents to do a lot of these things.
Checking completeness, surfing missing evidence, drafting review summaries and then humans just approving the decision.
Innovation only scales when governance becomes a repeatable product, not a bespoke committee debate.
A common mistake I see, or a failure mode, is executives assume it's covered without upgrading inventory.
Identity management observability.
They just assume we know how to protect sensitive data, so this is no different.
But the difference with agentic AI is you can't govern what you can't see.
If you don't inventory it and identity bind it, you're not scaling agents, you're scaling unknown risk.
We've talked before about implementing AI with speed and safety.
How does that taxonomy change with agentic?
What are some examples of new risk drivers?
If you think about levels of autonomy, a very low autonomous agent could be a co-pilot, a knowledge agent, where the risk is really around accuracy.
Is it accurate?
Is it citing the work that it came from?
Do we have hallucination controls?
That's far different than when you get to a semi-autonomous agent maybe, for example, in a procurement department, one that's approving invoices and looking for inconsistencies.
And it's got still within bounded rules.
But now you're really needing to put some more human in the loop to manage the inappropriate approvals, because there's financial risk attached.
Now you get to a fully autonomous agent.
That's maybe managing your public cloud and technology infrastructure and making system changes and performing upgrades and optimizing performance.
Then it's about are the agents doing what they've been assigned to do within the boundaries they were designed for, and not making independent decisions?
So the risk taxonomy has to be updated around these different risk outputs, around accuracy bias harm, as well as cybersecurity risk and cross-agent containment.
Picture those same three agents.
You've got an agent in procurement approving invoices, you've got an agent managing your Cloud infrastructure and you've got an agent in a call center talking to customers.
But they've all been trained on the same data and they've all been trained on the same internal knowledge.
So now one single data poisoning attack can have ripple fast failures across operations, finance and customers.
And that's something people are just starting to get a handle on and figure out how do they adjust their risk taxonomy.
How do they do intake and triage and enforcement?
It's amazing.
OK, so suppose I'm a leader and I am looking to embark on deploying agentic within my organization.
I'm thinking about how to mitigate these novel risks that AI-enabled autonomous decision-making represents.
How flexible do leaders need to be to ensure compliance, given the potential for all these changes in the offing, presumably?
I will try and simplify a very complex answer.
A lot of this is geospecific based on where you operate.
I'd say the most specific regulations, which are actually more helpful.
A good example is the EU AI act, right?
They break down and classify different agentic AI use cases into risk levers, with clear accountability on what you are allowed to do and what you're not, and what you have to prove.
You know US regulations.
Take a bit of a more framework approach and let you decide how you manage it.
In Asia, there's even a different set of regulations.
But they all have a very similar grounding.
And it's really about being able to do risk-based triage and prove enforcement and automate evidence.
Agentic risk management fails when guardrails are opt-in and anyone can go around them, right?
That's what leads to shadow agents and inconsistent standards, right?
You know, if I were to have to think of a soundbite, you know it'd be the best.
Guardrail is one you can't bypass.
And that's a lot of what the regulations get to.
How do you know the bias controls are there?
How do you know the factual groundness checks are passing?
How do you know that you're not doing harm against a consumer protection standard to a protected class?
These all have to be monitored in real time so that you can deploy these agents without surprises.
So complex.
So what else should leaders be asking themselves?
To assess their own readiness to adopt agentic at any kind of meaningful scale at least?
The goal for leaders isn't to slow innovation.
It's to make safe scaling repeatable.
What it means is that the leaders can't think of this as just another technology upgrade.
This is an operating model change.
You need clarity on decision rights, accountability, escalation paths, controls.
If you don't redesign those, you're not leading a transformation.
You're hoping the system behaves.
And that is not a defensible posture talking to your board or regulators.
And we were joking before this podcast started, about dystopian outcomes and so forth, and that's extreme.
But the truth is that folks are pretty jittery, generally speaking, about deploying a Gentic.
And as you've said, trust is vital.
It's not just a feature.
What kinds of communications should CEOs and tech leaders adopt to maintain trust, both externally with their customer base and internally with their own talent?
I will say one core question is, are the systems working the way you intend?
You have to look behind uptime and focus on three things.
Outcomes, behavior, and control.
Did the agent achieve the intended outcome without unintended side effects?
You have to have an answer for that in every transaction.
Does it behave consistently under edge cases or stress?
And can we reconstruct every decision and action end to end?
Design for trust first, speed second.
Start with bounded autonomy, but make sure you're keeping humans accountable for high impact decisions and scale only when the monitoring shows the system behaves predictably.
The customers don't care that it's AI.
They care that it's safe fair, fixable when something goes wrong.
For a tech leader, I'd tell them, make sure your teams convince you they've earned autonomy.
Don't just grant it because agents know how to do it.
Do you see organizations having achieved this already or is it an aspiration?
I think it's very much an aspiration, with some of the leaders showing us what works versus what doesn't.
That technology is so new and changing so rapidly, but it still needs to be steeped in traditional technology risk management frameworks.
You have to have an inventory of things.
You have to have a mapping of who is responsible for the outcomes.
You have to have a mapping of who sets the thresholds on control performance and whose job is it to take action when a threshold is exceeded.
Those principles don't change.
What changes is how you do this and the impact to your workforce.
Do organizations have the skills they need to make sure AI is deployed with appropriate safety and security, or is this a work in progress as well?
It's a work in progress.
A challenge folks are having is this is so new.
You can't post a job for someone that's been doing this with 10 years experience because nobody has.
And the few people that have gotten real expertise in this are being incentivized never to leave and change jobs.
So we've learned things, you know.
A good first use case is people started adopting development tools for developers.
You know, how can I get more efficient and faster in development?
Can I replace mid-level developers with some agentic actions?
What we learned is you fundamentally have to change your operating model.
The same people that are used to executing tasks when they now become the governor of the system.
That's a different skill set.
They need to be trained in kill switches, prompt engineering, and it might be a different type of person that has the rigor to actually play the human in the loop role.
I think a lot of firms just assume people are just going to change these roles and self-learn.
It has to be very purposeful.
And the folks that have been able to move the fastest are are those that have really invested in internal.
You know, call them company name university of how do we get everybody upskilled to an AI native employee.
I understand that you're talking about reimagining the process from end to end, but where do you stand on starting with a pilot or how to approach this, given the scale of the change that needs to happen across organizations?
So what we've seen is pilots should be bounded to six to 12 weeks with a clear business case.
You may decide, well, it's ready and the business case has been proven.
Let's execute and invest.
You may decide the technology is not there yet.
We need to wait a little bit before the hypothesis on the business case can be proven.
But you want to have an environment that encourages organized experimentation with the fewest roadblocks possible.
But you need to have them bounded and attached to a business case so that you can actually figure out what bets you're going to make and actually measure if they're paying off or not.
And how do organizations, particularly big global organizations, maintain line of sight into various agentic activity that's going on?
So the term I like to use is what's required is centralized governance, but federated execution.
You think back to one of the first topics we were talking about, right?
The people are used to this multiple committees with people opining.
That doesn't get you to any decision.
And what we often find is nothing gets denied.
So you end up in this, everything got approved and you're not clear where the value's coming from.
Now some folks have very mature tech governance and it's more of a bolt-on to get it to work in that forum.
Some folks build a completely separate standalone end-to-end AI governance.
That doesn't get away from the principle of federated execution.
Let the businesses decide their own use cases.
Let them operate the platform.
Let them hire talent and let them build things, while the central tech teams manage the platform and the automated guardrails.
But you have to have end-to-end visibility around what you're doing, with strong business cases that you can measure, or else it very quickly becomes technology experimentation for the benefit of technology experimentation and smart tech.
People like doing cool things.
But if it's not mapped to the business and what they're trying to achieve, it gets very, very dangerous very, very quickly.
Suppose I have undertaken agentic activity in my organization.
Talk to us a little bit about what has to be in place to get agent-to-agent protocol in line.
I'll give you an answer on two dimensions.
One on the technical dimension.
The industry is quickly forming around standards with how to securely allow agents to talk to each other.
And the 800-pound gorilla is if an agent is going to be accessing tools or data, it should be going through an MCP gateway.
And MCP is short for model context protocol.
Think of it as a USB-C port for AI, right?
It's a control point where agents can access tools, access data, but with controls and policy-based enforcement in a very standardized, pattern-based way.
Now there's some other standard protocols around agent-to-agent communications, but the important thing is, you cannot allow open communication.
There has to be, within a platform, an AI gateway and an MCP gateway to govern and control all of this agent-to-agent communication.
That's the technical answer.
Now, the people-based answer is the onboarding.
A lot of these agents are designed to be reusable.
So I'd say it's not about the agent itself.
It's about the workflow and the use case.
You may have a data analyst agent named Dana, and maybe Barbara uses Dana to access some highly sensitive data.
But Joe uses Dana, the data analyst agent, to access non-sensitive data.
Well, that agent has the same agent, two very different purposes.
That needs two very different policies.
And you have to allow for how does it get situational access that's temporary.
And to do that, that's back to the people side of it.
The ones designing the use cases have to, in a centralized way, be the ultimate authority on what is the use case supposed to do and work with the control owners to put the guardrails around that bounds in place.
That is so interesting.
So just so I'm clear, you know, when you were describing not having open communication between agents and having the MCP gateway, like what happens if there is open communication.
Does agent to agent communication go haywire?
Here's a great analogy that I like to use.
You know, people think these AI agents are supercomputers so smart.
I liken their intelligence to a human two-year-old.
It's literally a toddler.
And if anybody's had a toddler, if you're in an upstairs hallway with wood floors and steep stairs and you tell the toddler to run down the hall and stop at the top of the stairs, they might.
They'll generally go in that direction.
They may also pull some crayons out of their pocket and draw on the wall on the way and probably trip and face plant down the stairs and hurt themselves or somebody else.
Without a baby gate in place and if you're not there, how do you know they did what you told them to do in the way you wanted them to do it?
Because these agents are capable of independent reasoning and problem solving.
So if you just let it run wild without an enforceable control policy, which is what the purpose of forcing these communications through an AI gateway and an MCP gateway, that's where you apply controls.
The opposite model is you write a bunch of documents saying here's the requirements and standards, and every business owner then has to figure out their own way to meet them.
And it's impossible to manage that at scale, to test it at scale, to provide assurance at scale.
Is it harder to manage for risks introduced by humans or by agents, in your view?
I would say that it's easier to manage humans only because we have an enormous past history of human behavior to rely on.
We don't yet have that for agents.
So I'd say that's why it's a little trickier, even though humans can be more cunning.
They also move slower.
They don't work 24 hours a day, seven days a week, and they're easier to monitor.
Agents work 24 by seven, work at the speed of computers and you can't monitor them with a team of four people doing targeted assessments.
You have to monitor them with their same technology.
You know, I think the EHR and CHRO example is I would not advocate.
The HR department needs to get involved with managing agents.
But each agent does need an owner.
These are not set it and forget it.
They need to be consistently monitored and tuned and tweaked and sometimes fired.
And it has to be clear who are the owners of these agents and who's accountable for their performance.
And I think that's become the hardest question in this new operating model construct, where the answer is it's not always the technology team anymore.
Wow.
And what's the contingency plan?
Like what happens if an agent does go amiss in some way, or two agents you know start getting into it and making some unfortunate decisions for the business.
What next?
You could look at it on a scheme of what's the worst thing that can happen.
Go watch the original Terminator movie.
That probably sums it up.
Yeah.
I would say the challenge is these agents reason, and they work very fast.
And if they're directly talking, colluding with each other, the spiral risk of massive problems at scale is going to be important to manage.
And that's why the purpose of you have to think about the kill switches.
And under what circumstances do you have logs that would automate pulling one?
They can't be manual decisions.
So far, we've been talking primarily about agents that exist in virtual space, meaning that our interactions with them are synthetic interactions.
But in reality, when agents are embodied in robot form, that's on the horizon, right?
And maybe nearer than we've previously thought.
So how might the rise of humanoid robots affect security protocols?
Think about it this way.
We have a lot of autonomous self-driving cars on the road today, right?
It's not that different.
They all run on an operating system.
They're starting to introduce agentic features into it so they can adapt and learn and optimize and fix themselves.
But if you think about, they're all software.
It's like an agent is a software program.
That's what's running on these cars as well.
And they all go back to a central location.
You know, One rogue agent, you know, this is a little hyperbole, right?
Figure in the very near future, has the intelligence to do it, takes over that operating system and suddenly has control of all cars on the road that have self-driving features.
And in the wrong hands, that can be enormously destructive.
So, when you think about the future and robotics, this idea of guardrails that are enforceable and ensuring that only decisions that have been designed and anticipated are what's happening.
And if something does go wrong, what is the rollback plan?
And what about, you know, it's not like agentic is the end of this.
We've got AGI, we've got quantum, which has cybersecurity implications, security implications.
Anything leaders should be thinking about now to anticipate those new categories of risk or resilience that might emerge as technology continues to advance.
Yeah, absolutely.
And I talk to a lot of boards, right?
And this is a common conversation, right?
And here's my assertion.
Boards don't need to be technical.
They need to be precise, right?
I always encourage boards.
We'll use Agentic AI as the example, but any big tech transformation.
Ask five questions to your leaders and demand precise answers, right.
For Agentic AI, it's, do we have a complete inventory of agents and owners?
Yes or no.
How is autonomy tiered by risk?
There should be an answer that has five or six segments.
Do agents have verified identities and least privilege access?
Again, yes or no.
And then more importantly, the last two should be, can we reconstruct decisions end to end?
Yes or no.
And then lastly, do we have a real rollback plan if something goes wrong?
If their leadership can't answer those five questions with precision, they don't yet have agentic risk under control.
And that's how boards and top management should think about it.
Good AI governance is not about knowing the model or being super technical.
It's about the ability to prove control.
I want to go back to one point that you made earlier.
I'm interested in your take on sovereign AI and national security in data.
That seems to be emerging as a hot topic.
Yeah.
I don't want to scare people, right?
But, you know, a sovereign AI is really about who's in charge when AI makes decisions.
Right.
We talk about sovereign AI.
We're really talking about control.
Who has control over the data, the models, the infrastructure and decision making.
And it's the idea that governments and enterprises should not be outsourcing their most critical AI capabilities to opaque, foreign or uncontrollable platforms.
Once the AI systems become agentic and have this ability to act, learn and make decisions.
Sovereignty is a risk issue, not just a policy debate.
And leaders are realizing that whoever controls the AI stack ultimately controls the outcomes.
And this is a whole new, when you think about third-party risk and, you know, nth-party risk, right?
Agentech AI has finally turned sovereignty from an abstract concept into an operational one.
You know.
If an AI agent is acting on your behalf, leaders are going to need confidence in where does it run, whose laws apply, who can inspect it and who can shut it down.
You can't govern what you don't control, and you can't control what isn't sovereign.
Rich, this is a huge issue, super complex, super high stakes.
If there were just one or two things top line, you would want CEOs or other leaders to take away.
What would that be?
I would say you take away that the future is not humans versus AI.
It's humans with AI.
As AI systems move from generating ideas to taking action, the real differentiator won't be who adopts it the fastest.
It will absolutely be who governs it the best.
You know, agentic, sovereign, secure AI is not about slowing innovation.
It's about earning the right to scale it with trust, accountability, and control.
Rich Eisenberg, thanks so much for joining us.
This is fascinating.
I'm happy to be able to talk about something I spend most of my weekends reading about, writing about and tinkering with, and hope some other like-minded folks find it interesting.
Great to have you.
Thank you.
Thanks so much for listening to the McKinsey Podcast.
I'm Lucia Rahilly.
And I'm Roberta Fasaro.
Find us on McKinsey.com.
We'll have a transcript of this episode up shortly.
And download the McKinsey Insights app, where you can find this podcast and other helpful content updated daily.
If you enjoyed the show, we'd love for you to leave a rating and a review.
We'll see you in two weeks.