Welcome to Goldman Sachs Exchanges.
I'm Alison Nathan, and I'm here with George Lee, the co-head of the Goldman Sachs Global Institute.
Together, we're hosting a series of episodes exploring the rise of AI and everything it could mean for companies, investors and economies.
George, good to see you again.
Great to see you, Alison.
Good to be here.
So George, we've had several conversations about how AI is shaping the economic and business landscape.
But today we want to actually get a little bit more under the hood and talk about the technology itself and, in particular, the role that data will play in enabling or possibly stalling its progress.
We have a terrific guest to dive into these issues with us Nima Raphael, the Chief Data Officer and Head of Data Engineering here at Goldman Sachs.
Nima, welcome to AI Exchanges.
Thanks for having me.
Excited to be here.
So, before we dig into the topic at hand, first just tell us a little bit about how you got here, your career journey and what your current role here at Goldman Sachs entails.
Yeah, this is 20 plus years for me at Goldman.
I started right out of college studied computer science, and During that I sort of realized that I wanted to apply technology to a domain that I had not known before.
And so really the finance was sort of like this black box to me.
I came to Goldman, met amazing people, started as an analyst here.
Software engineering, you know typing at the keyboard, writing code.
And then really the data thing came, I'd say five years in, global financial crisis, 2008.
Lehman Brothers collapses and a group of technologists called CoreStats at the time was going around the firm saying hey, we have to figure out what our exposure to Lehman is.
We have to figure out our liquidity profile.
See what's going on at Goldman.
And the way they had structured it was to try to get all of the data from the front office, middle office and back office together in one place to sort of figure out the end-to-end exposure to Lehman in a very technology and data heavy way.
And that was sort of like the genesis of, I'd say, my data journey here.
That project actually was super interesting because we had heard other banks and other financial institutions actually have to go into their filing cabinets to dig out their ISDAs that were signed with Lehman to figure out what their contracts were.
We luckily had a lot of our data sort of corralled in one place.
And actually that database we built it was called Copter ended up becoming this place.
Not only that, people realized like the power of data not just being sort of like an exhaust but actually an enabler for the business.
And then not only did people recognize okay, this database sort of saved the firm in some interesting way.
But then when we gave that same data to traders salespeople, strats on the desk, quants on the desks people started coming up with new innovative ways to use that data for helping our clients and just running the firm a lot more efficiently.
And so it became this sort of launching pad for people outside of technology to say hmm, data maybe can be a powerful concept here, that's great.
And so in your role as chief data officer, now you, you oversee all of that.
Other things you've done in your career have been involved with what we used to call, in the dim, dark past, machine learning.
The point is, you know, ai's been around, we've used it at the firm.
It's been, you know, broadly proliferated.
But The rise of generative AI has garnered so much attention.
Is it fundamentally different than the journey we've been on or is it just an extension of the continuum of good, old-fashioned AI?
Of both, because it feels like some sort of step change function from the historical you know, i always talk about the first 50, 60 years of computer science sitting down and humans have to code rules to tell the computer what to do, and we talk about determinism.
Like the rules were deterministic, like if you push this button, please do this, or if you type these keys, please do that, and so there was this really fundamental shift, I guess, in machine learning in general, which is like learn by example instead of learn by rules.
In some ways, the generative AI stuff is just a continuation of learn by example.
But I don't think people naturally saw it go from hey I could learn maybe how to predict some patterns, to now the computer could create anything.
And so there's a little bit of that continuum like hey, if we just feed the machine more and more examples, more and more data, it could start learning things is probably the path of continuum.
But the sort of step change was like oh well, but can we feed it and create images, create audio, create images, language?
And so I think that's sort of the novel step change in the generative part.
And you illustrated something I think is very fundamental in terms of company culture in this shift, which is we're used to deterministic computing.
For a given input, the outputs are correct, repeatable, and traceable.
We're no longer in that sphere.
As you pointed out, these are probabilistic machines.
Something emerges from it that you can't trace and is often right, but not always.
Talk about the mindset difference inside an organization, of getting business users in particular, to be comfortable with that.
I'd say a little bit in finance.
I mean, it was always stochastic in that way anyways.
So there was always a little bit of, okay, like the world is non-deterministic.
And so prices are non-deterministic.
The markets are non-deterministic.
Economies are non-deterministic.
So I think there was maybe a willingness to sort of understand that here in the finance world.
But I agree.
I think when non-engineers sit at a computer, they sort of want a thing to be a repeatable pattern.
That's how we build workflows here, that's how we build client insights or anything we do here to help our clients.
So I think it's really about teaching people this isn't just some magic crystal ball, right?
What it's really doing is taking a lot of examples and giving you an extrapolation from those examples.
Let me just have some thought to that though, because we've had a lot of conversations on this podcast about the ultimate potential of the technology.
There's so much hype around it.
We're having another, I think, leg up in the hype in the last month or two here.
Given what you know about the technology, do you think it's overhyped or maybe even underhyped?
Yeah.
So as George knows, I'm always a little bit of a skeptic of new technology.
Historically, we've talked a lot about blockchain and things like that.
And that was supposed to revolutionize.
And is the next thing going to revolutionize?
And look, I think from an AI perspective, it's obvious that it's real.
It's here to stay.
There is absolutely a hype to it.
But also when you go on your phone and you ask Claude Gemini GPT, take a picture and you ask like what is this?
Or you ask, give me some research on a topic I'm curious about.
And you get great answers and you research more.
It's definitely, definitely real in the sort of consumer world, I think.
I think where the hype, I don't know, I would say it's slightly differently than hype.
I'd say the potential, I think, in the enterprise is still to be seen.
I think there's some really slam dunk use cases we've seen, right?
Agent coding, for example, is the thing that sort of flipped my brain from.
This might be vaporware to like wow, this is really real.
When I sat down at the computer and I was coding with an agent and it was helping me with problems that I've never been able to solve before, I was like wow, this is incredibly powerful as a superhuman ability, amplifying my abilities.
So I think there's definitely real there.
I think from the enterprise perspective, the thing to be seen is where can people harness their data and their enterprise data and the proprietary data?
They have to make some differentiation in the enterprise space.
That's the to be seen part.
But we're only a couple of years into this newer generation of these models.
Do you foresee a future where we actually do, though, run out of data?
I mean, we're early here, but is that ahead?
I would frame it a different way.
We've already run out of data.
We've already run out of data.
When you read about the new models, the undertone of what people say and you've seen this in like the deep seek moment and things like that is like Everyone wonders how did they do that with less money?
And one of the big hypotheses is they trained against another model, right?
And so it already incorporated the previous thing.
I think the real interesting thing is going to be how previous models then shape what the next iteration of the world is going to look like in this way.
So let me reframe my question, which is more that do you think this is going to restrain the potential of the technology?
No, I don't think so.
The explosive nature of the synthetic data and the fact that now the computer could generate infinite amount of more data.
Again, I think there'll be a sort of a cursor of what people call like AI slop versus maybe more insightful data.
But I don't think it's going to be a massive constraint only because a lot of trapped enterprise data still has not been harnessed.
And I think you see that in the work that we're doing at Goldman, for example.
We want to help our salespeople, our traders, our quants, our PMs, to sort of again, get that superhuman capability, that information synthesis capability, being able to help with their hypothesis.
And there's still a lot of data here at Goldman that can be used for that.
So I think from a consumer world model, I think it's interesting.
We've definitely in the synthetic sort of explosion of data.
But from an enterprise perspective, I think there's still a lot of juice, I'd say, to be squeezed in that.
Yeah, I would echo that.
I think these machines have become an enormous part distance in their quality, and they've done it largely in the back of publicly available and synthetically generated data.
The amount of data that lives behind firewalls, trapped inside corporate repositories that's highly salient to garnering business value, that has yet to be unlocked.
It's the work that Nima is doing here.
Then there are also other horizons.
Think about all the video data in the world.
Think about spinning up virtual environments, where you're creating a platform for virtual robots to generate their own data about understanding the world.
I think while we've exhausted one pool of data, there are many others to go attack.
I think I'll get a little philosophical out of my realm, but I think what might be interesting is people might think there might be a creative plateau.
I mean, if all of the data is synthetically generated, right then, like How much human data could then be incorporated?
New human data, new human intellect, new human creativity.
I think that'll be an interesting thing to watch from a philosophical perspective.
For sure.
You know, one or two of our prior guests have made the observation germane to this discussion, that the quality of outputs changes from these models.
Particularly in enterprise settings is highly dependent on the quality of the data that you're sourcing and referencing inside the business.
Do you agree with that?
Kind of goes to this, what's the value of these behind the firewall data stores?
Maybe just illustrate a little bit of that how we can make models smarter with our own proprietary data stores.
Yeah, I think first again, you got to remember what this thing is doing, what this machine is doing, right?
Whatever data patterns you are feeding this machine is what it's going to learn and what it's going to extrapolate from.
And so I think, from an enterprise value perspective cleaning your data, normalizing it, having the semantics of that data well understood, how it links to other pieces of data all of this stuff is what's going to allow enterprises to level up from what we think the consumers get to what enterprise value could be created.
So what are some things that Goldman is doing to unlock that value?
Yeah look, I don't think people had always thought of data as sort of like this thing that could Give more insight to the world.
I mean, it's always historically been thought of as like business exhaust in some way.
Right.
Like a trader executes a trade.
They're sort of like, OK, I'm done now.
I'm just managing the risk.
But there's a whole machine behind that about what happens after that, all the workflows that happen after that and before that.
And so the real challenges are getting that disparate data into some place where you could organize it in a sane way and then normalize it in ways where the data is correct when you ask it a question.
It's linked to the other facts of the world.
And so all of these challenges are really, you know.
That's why the role of data engineering was even created.
People are like, we need a practice of engineering that's like software for data.
And so just like people write code in a specific way and there's specific architectures and engineering practices to that, is the same in data.
You have to sort of understand what the data actually means.
You have to understand, are these two concepts the same?
Are they linked differently?
And so really, the challenge is understanding the data, understanding the business context of the data and then being able to normalize it in a way that makes sense for the business to consume it.
Can you actually use the models to help you organize the data?
Is there some synergy happening there?
Definitely.
Like people have built software agents, people have built engineering agents to do like this cleansing this normalization, this linking.
So absolutely in the same way, where we're seeing software being created by these agents, there's also a feedback loop of data cleansing and normalization and wrangling too.
It's a good insight.
Nima, we often close these interviews by asking our guests how they might actually use AI themselves.
Either.
We talked a lot about how you were using it in the office, but even personally.
What do you find is the most interesting and helpful usage?
Yeah, you're asking a tech nerd.
So obviously I'm thinking of like the tech nerd, the coding answer is like the base case.
But I also have a three and a half year old son and he's in his Y phase, which is awesome.
But I'm just like I run out of, run out of like the turtles of the why and so like.
Actually a lot of the times he's like what is that, why is that?
And and actually bouncing ideas off of him with the agents, i think is really cool and i think it's been, it's been powerful for me.
So now he asks me questions.
I ask the ai questions.
We learn together about questions he's curious about.
So i love that And as he gets older he's going to be able to ask himself.
So when my teenagers ask me questions, I say, look it up.
Use AI.
Let me Google that for you.
Exactly.
Though part of Nima's genius is he gets to free ride on the knowledge acquisition of his son.
That's right.
Exactly.
And be involved.
So I love that.
Exactly.
Well, thanks very much, Nima.
That was a fascinating conversation.
Thanks.
Thanks for having me.
I mean, George, Nima had so much insight.
You know, what really stood out to you the most about the conversation?
Well, as I predicted, you know, grounded, objective, thoughtful.
I agree it was a great discussion.
You know, people sometimes peer past the data problem.
But, as Nima, I think, illustrated well, it really lies at the heart of bringing value from these systems in business.
So onward and upward.
As always, thanks for the conversation, George.
Always great talking to you.
Thank you.
Thank you, Nima.
Thank you both.
It was awesome to be here.
This episode of Exchanges was recorded on September 25th, 2025.
I'm Allison Nathan.
This material may contain forward-looking statements.
Past performance is not indicative of future results.
Neither Goldman Sachs nor any of its affiliates make any representations or warranties, expressed or implied, as to the accuracy or completeness of the statements or information contained herein, and disclaim any liability whatsoever for reliance on such information for any purpose.
Each name of a third-party organization mentioned is the property of the company to which it relates, is used here strictly for informational and identification purposes only and is not used to imply any ownership or license rights between any such company and Goldman Sachs.
Copyright 2025 Goldman Sachs.
All rights reserved.