Imagine building an engine with 54 billion parts.
Now imagine each piece is the size of a gnat's eyelash.
That gives you some idea of the scale Jonah Alban works at.
Jonah is the co-lead of GPU engineering at NVIDIA.
The engines he builds are GPUs. Without these chips, your favorite computer games and special effects movies would look pretty lame.
GPUs also power scientific simulations of everything from a Mars lander to a protein spike on the coronavirus.
And these days, they do much of the heavy lifting for the latest and greatest form of computing, AI.
Welcome to the AI Podcast. I'm Rick Merritt, a staff writer at NVIDIA.
And welcome, Jonah. Hi, Rick. Thank you.
Some of our listeners may not know you, so can you describe what a leader of NVIDIA's GPU engineering team does?
Well, you know, I guess you try to figure out what the future should look like, I guess.
You know, GPUs fundamentally are you know, accelerated computing machines, right?
That that's how we think about them. And so, you know, to do accelerated computing well, a lot of it is, is what problems do you want to go solve? for which people and how to solve those, right?
It's an exciting field because you're always thinking about the problem end to end, right?
You're not just running a small little piece of the code, you're looking at the whole problem that customer is trying to solve and figuring out, you know, how to, how to solve that whole problem.
So it means you've got to talk to. Customers, you've got to talk to the hardware team, the software team, the systems team, everybody's got to work together. to find the right answer.
And that's most of our lives every day, and the team is working on that activity.
It finally comes to fruition when we have a new product like A100 out in the market and can get to show it off to everybody, which I guess is where we're at today.
So it's been quite a journey. It's very satisfying to have gotten there.
And you work at a different time scale than most people in their jobs.
I mean, you're looking at, it takes... This A100 you just released, you worked on that for what, three, four years?
Yeah, I'd say probably at least that amount of time.
There's a time when you're thinking about it and you haven't started working on it yet. know in terms of development and then of course the development time as well and then uh getting it to production so yeah it's definitely it's it's a long journey you got it you got to again sort of try to think a little ahead into what you think the perfect future could be, and then you go try to build that.
But yeah, as you said, these are really complicated devices.
We of course go to a fab to build them, but it's actually mind boggling to me that it's even possible for a fab to build something. even one of this kind of device and actually have it be functional.
But amazingly enough, they're able to do that.
And then, of course, the challenge for us is to do all the work to figure out how to how to build something that complicated to solve important problems for people.
And tell us just a little bit about what that's like.
So, I mean, if you started, for instance, today, you're going to design the next generation ship.
And you know it's going to take three years or so before it's actually out.
How do you think about where AI is going to be in three years?
I mean, most people have trouble thinking beyond 18 months.
You can't know, right? Everybody has to guess.
One of the important things for us is trying to always be on the leading edge of where things are now.
NVIDIA does a lot of our own work in AI, and that's really important to me, to talk to the people inside of the company that are doing work in AI, and understand sort of what their pain points are now and what their dreams are they're not able to achieve.
Certainly, within one company, you can obviously just talk about anything you want to because there's no secrets inside of it. side of a company.
So, so we, you know, we rely a lot on that and we rely on, on good conversations with customers also, and then looking at where the researchers are going.
Yeah. So you've been doing this work for like 20 some years now.
And I'm curious, I mean, you started off as a as a GPU engineer when GPUs were all about graphics, computer graphics.
And somewhere along the line, AI happened in a big way.
I mean, it's been going on for a long time, but it really became a commercial reality somewhere down that road for you.
So do you remember when AI hit your radar screen and what was that like?
Yeah, so definitely, I think I have very, very fond memories of that even today.
It's fun to think about how far we've gone since then.
Before AI, of course, we had our GPU computing revolution that we built up and brought to market with CUDA.
And we had a vision then or a dream that when we put GPUs out in the world, these great parallel computers that And we put them in every GPU, the capability of every GPU we released that somebody out there in the world would find these GPUs and would use them for some new problem that we didn't even know about at the time.
That was probably over 10 years ago now, 15 years ago that we did that.
And then AI, I think was really, you know, this, you know, great problem that emerged that sort of I think proved that that dream had been a good dream to have had.
We never talked to Alex Krzyzewski personally right before he started doing this work with GPUs.
One of the founders of AI, commercial AI.
Yes, exactly. Yeah, and one of the creators of the famous AlexNet But people in this field started to use GPUs.
They started to see that they give them this amazing boost in compute, which is really what the field needed.
And they started making great developments with it.
Personally, I'd say I got involved when we started seeing more customer engagement, particularly Google was one of the very early, innovators in this area and I got to meet Alex who was at Google at the time and they were doing work and we're looking for even more acceleration in the future.
We had some really great conversations with them and I read the papers. the early AlexNet papers and other papers in the field.
And it was just very exciting at the time, I'm not an AI, researcher myself, but just sort of seeing how people have been able to harness this power of compute, of matrices to mysteriously produce intelligent output was, uh, was very impressive to see what they were able to do.
Turning a pile of matrix math into recognizing images of cats, that's pretty magical.
Yeah, there was no human that was ever in there telling it to do something. have a lot of experience with algorithms where you try to sit down and think yourself about how to make a computer do something.
This was sort of setting the computer up to learn on it by itself how to do something, which was a really sort of amazing leap forward.
Very exciting for us to see how to make it better.
Take us back for just a moment to, if you can remember, what it was like when you maybe first met Alex and what you thought this was going to be then compared to what you think now.
I always thought it was cool. I would say that it's gone so far since then.
I think that the applications back then were interesting, but just The fact that so many other people sort of also picked it up and applied it for lots of different things, it's been amazing to see how far it's gone.
When I first was looking at this, of course, one of the things I didn't realize, I didn't think about how it could possibly affect our core gaming business.
But then in the last couple of years with deep learning, super sampling, other efforts, we've now found that this is actually a great application for gaming as well, which makes sense since The original applications were imaging-based, but it just wasn't something that was sort of initially in my dreams when it started out.
Yeah, it's like it's going everywhere. Yeah, so since that time, It is going everywhere.
And in the semiconductor world now, there's tons of startups.
There's big companies talking about we've got to make GPUs to accelerate AI training and inference.
It's become a superheated, this sector you're in of building chip accelerators, building the engines for AI has become superheated.
What's that been like for you? So in some respects, it's actually been a little nostalgic.
You mentioned I've been at NVIDIA for 23 years.
When I was first at NVIDIA, there were something like 30 plus companies doing graphics accelerators for PCs for 3D gaming.
And it was a crazy time, right? There's all kinds of folks out there. doing interesting different things.
And, you know, definitely we felt the pressure all the time back then to, to do our best work and to, come out with products that our customers would love.
This has felt a lot close to that time in many respects.
There's definitely a lot of people in this space that are all looking to see how to make the next great thing.
So I think for me, I think back to that time and what NVIDIA did and how we thought to be successful back then.
You know, which is focus on doing your best work.
Don't be too proud of your past. You got to think about the future.
Our intention, of course, is to be able to look back 10 years from now and be proud of the work we've done and the work that we are doing, being a leader in the industry.
You've always got to be thinking about what you can do to keep raising the bar.
Because engineers like a hard problem. Yes.
Yeah. And I think this has been a really exciting problem.
It's great to have a problem like this that's so complicated.
I think one of the things I'd say from In the beginning, it's actually been more complicated, more interesting than it originally appeared to be.
Just looking at the original AlexNet papers, those networks were not... super complicated.
And since then, things have gotten much more interesting.
And I think every generation we realize there are some new challenge or new nuance that we hadn't seen before.
So it's a very exciting computer science problem to go solve.
Now here's one of the interesting twists of history to me, of technology history, I guess.
Just at the time when there's this whole new method of computing, AI comes up, We're also coming to the end of the traditional methods for making chips that accelerate things.
I don't want to get into the deep goo here, but Moore's Law, because of just the dynamics wavelengths of light that we can use to make the lines on silicon is coming to an expected end, or at least it's slowing down.
And so it's becoming more expensive to make chips go faster.
Sometimes it's hard to make chips go faster at all.
That's just a historical reality for the whole industry.
So right when we need more performance, it's getting harder to get.
What's this new reality? How would you say this new reality is like for you and for the industry?
So I think if anything, it makes the job that people in my field have even more important. if Moore's law was like it was 20 years ago and you can just take an old, old idea and put it in a new process technology, and then it gets a bunch faster than, than,
You know, you, maybe you don't necessarily need to be so clever, right.
You know, with Moore's law slowing, I think that that puts pressure on the whole computer industry to find. smarter ways in architecture and computer architecture and system architecture and data center architecture to make applications go faster.
That's true not just of AI, but of every other application that's important in computing.
So I feel like this is a great opportunity and a great challenge for the folks in this business of building chips to do their best work to make up for Moore's law, right?
To make our own Moore's law is one thing we like to say internally.
Making your own Moore's Law. Okay. Any examples of that?
Well, Moore's Law has a certain pace of how much better things are supposed to get.
And so It doesn't just have to be the fabric that makes it better.
It could be your own ideas that make it better.
So if we can make chips 10 times or 20 times better or, you know, however we wherever we can pull off, then that's that's making our own war's law.
Well, kudos here. You've just come off from what's being billed as the biggest generational improvement to date. in a new NVIDIA GPU architecture with the A100.
Can you, for an audience of people that may be listening in that aren't chip architects, What should they be reading in the tea leaves of the A100?
What does it say about your team's thinking about where AI is going?
AI is going in many different diverse directions.
That's one of the most exciting things about it.
And A100, I think, reflects that, right?
We thought about everything from The very advanced researcher who's working on a very speculative network and making sure we have the flexibility and the usability for them all the way to deployment of a network that you want to run inferencing on.
You're actually running the network in production for customers and being able to take this giant chip and chop it up into a bunch of little chips that makes it more easily usable for, for running lots of different little networks for people in production.
So I think AI is not just one thing. It's a ton of different problems all stuck together into one.
And with A100, we tried to imagine an architecture that could span that. that breadth of challenge in AI.
Cool. I wonder if there's any good behind the scenes story, like if we were doing the B-roll or the blooper roll even. of the making of AI and we had cameras following you for these three years or whatever.
Are there any interesting stories to tell?
I don't know if there are happy, funny, crazy deadline stories or whatever about – behind the scenes with the A100.
Anything come to mind? If you think about just the fact of this ship existing, that was a daunting thing to think about in its own right internally for the team.
You know, going to seven nanometer, This was our first 7 nanometer chip, over 50 billion transistors, It was an amazing challenge for the team to think about.
Would the fab be able to build this chip?
Would we be able to to build the chip that we had in mind to fit.
One of the challenges of a reticle limited chip is that You can't make the chip any bigger, right?
So you better be pretty sure of what you're building because if the chip gets bigger in the reticle, then you just have to stop and start over again.
How would you define reticle? Yeah, you can think of the way chips are built basically is that the fab effectively is shooting pictures onto a piece of silicon.
And there's a picture frame size. And they can't make a picture bigger than the picture frame.
So the reticle is the fab's term for the size of that picture frame.
So you can only make pictures that are a certain size.
And you're going right to the edge of the frame with this a 100.
Yeah. And so we, you know, we, we did, we, we wanted to make sure we put everything that we could imagine into, into making a great chip for, uh, for our customers, you know, leave nothing back.
And so Reddick Limited is the way to do that.
Everybody's pulling out all the stops for this because it is this whole new form of computing which is being born and we see people doing whole whole wafers as a chip and a dozen little chiplets glued together and So it becomes a weird thing of, well, how do you measure this stuff?
I mean, the workloads are changing. As you say, they're getting more diverse and the chips that are running them are getting more diverse.
So you know we've had some success. I think in the past couple years with the ML Perth group coming up with some training and some benchmarks to measure this.
How are we as an industry doing and coming to grips with giving kind of apples to apples measurements to people who need to know?
So I think we're doing well due to a huge amount of work from a bunch of people in the industry, a bunch of companies that came together to go to go build MLPerf.
I think that None of us really knew what we were getting into.
It turns out that AI is probably the hardest thing that anybody's ever tried to benchmark.
We were benchmarking at a system level. It's a huge scale problem.
It's a problem that starts at a very high level.
It's not just a binary you're executing, it's a Python program.
Now, it's been a couple years now that we've been working on it, and I think that we have come a very long way.
So you can now run these workloads that are based on real customer workloads and see the performance of different systems and different scales of systems solving those problems for people. our goal was to make sure this measured workloads that would matter to customers and weren't some synthetic type things.
When you talked earlier about matrix math or linear algebra as being like the basis of the workloads behind AI, And so naturally, a lot of the chips and the NVIDIA chips too, running them are accelerating that matrix math.
And I'm wondering, as these workloads become more diverse and they become larger and the applications become different, Do you think we're going to see moving in different areas.
I mean, today the chips are mainly accelerators for that kind of math.
Is it going to move elsewhere, do you think?
So I think there's interesting diversity of approaches already today, actually.
You have some folks that are looking at having lots of very small processors, other folks big processors that are you know there's.
Individually less less flexible but- just have a lot of throughput per processor.
So there's a lot of different ideas out there.
Obviously, we have our own ideas about sort of the right place to be in that spectrum.
But AI is not just linear algebra. It's a lot of other computations that have to go around that.
And so there is an interesting challenge to figure out what the right architecture is.
I think our mindset is that we do believe that being able to handle diversity is very important.
AI workloads are already quite diverse and every year some new interesting thing comes out.
And going forward, I think in terms of where processors should go, it should be related to what those workloads look like.
Look at the workloads people are running and workloads people want to run. and try to build processors that are going to be able to do a great job with those workloads.
You mentioned in the past, NVIDIA created this environment CUDA so it could run any job that needed parallel programming or acceleration.
To some extent, you got lucky that this huge job of AI came along when you had put that infrastructure in place.
I wonder looking back, are there any cool and unexpected things that you found over your career that people were doing with GPUs, whether they're for AI or not? that surprised you or were interesting?
So I think the first thing that I remember seeing that gave me a sense that GPUs were really out there being used for things that we didn't tell a customer to do was when I read a paper that someone had used to use a GPU to simulate how the human nose worked.
And that's a really interesting problem, but I knew that was a problem that there was no salesperson from Nvidia that had ever called up that researcher to try to sell a GPU to him for that.
He must have done all on his own. And so that has stuck with me as sort of the, you know, the first time I was like, okay, you know, This isn't just for the three problems that were listed on our to-do list to do that people are now finding their own problems to go solve with GPUs.
And what was he trying to do with those?
Just figure out how the nose smells stuff.
Inside the nose, there's cells in there that are sensitive to chemicals in the air, and he was trying to simulate how all those cells together in the nose ended up sending a signal to your brain to say that it smelled something.
And he was using a GPU for it. Okay, so our time's kind of coming to the end, but I wanted to ask you if you had, you know, maybe advice to young people in the audience.
What would you say to somebody who's thinking about, yeah, this GPU engineering stuff sounds interesting.
Maybe I want to get involved in it. What would you tell them?
The first thing I'd say is figure out what you love to do and do it.
This is the thing that I... realized a long time ago I love to do and I still very much love to do today so I'm very happy with the choices that I made but you know it really had to be founded on caring a lot about something, finding it very interesting to go pursue.
You know, if you are considering this area, I certainly would encourage you to do so.
It is a fascinating area. Computers in general are influential over so many of the things that are going on in the world today.
And it's become this great system level problem.
I would say hardware, software, Every chip today is a very complicated system, lots of software and hardware working together, and it's just a really intellectually interesting challenge as well as a fun challenge for the people you get to work with to go try to solve those problems.
And how about you, Jonah? If you could just do it all over again, starting today, what would you choose?
I'm pretty happy with life. I don't know if it's serendipity or something, but...
You know, I had the good fortune to meet Jensen a long time ago and to feel inspired by what he was planning to do when he was much younger man than he is now or even that I am now and and, you know, get to work with a great bunch of people.
So it's been a really great privilege to have been able to work in this field for so long and to have this still amazingly interesting new challenges coming up to work on.
Well, it's been a great privilege talking to you too.
So thanks for taking some time out to be on the AI podcast, John.
Thank you. Thank you. ¶¶