Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. As I record this, Nvidia researchers are gearing up to present 19 accepted papers and posters. seven of them speaking sessions at the annual Computer Vision and Pattern Recognition Conference this June in Salt Lake City, Utah. joining us today to discuss some of what's being presented at CVPR and to no doubt give us his unique perspective on the world of deep learning and AI in general. is one of the pillars of the computer science world.
Dr. William Daly, chief scientist here at NVIDIA.
Dr. Dahle joined NVIDIA in 2009 after chairing the computer science department at Stanford. which came after stints at MIT and Caltech.
Honestly, I could spend the which include numerous awards and fellowships, co-founding two companies, and having his name on over 250 papers, 150 patents, and yes, four textbooks.
But instead of me reading laundry lists of accolades, let's talk computer vision and AI.
Dr. Daly, thank you so much for making the time to join the podcast.
Hey, you're very welcome. So we're going to talk about the work NVIDIA Research is presenting at CVPR.
But before we do, maybe you can set the stage for the listeners.
You've been working on neural networks for 30 years and then some, going back to your days at Caltech.
Now you're chief scientist in NVIDIA. And as I walked in the door to do the session today, I noticed the big IMAI sign out in front of the headquarters.
So kind of broadly speaking, what's changed?
What's the same? And where's the field headed?
Well, if you're talking about what's changed over 30 years, an enormous amount, but probably the biggest thing that's changed is the speed of the underlying computing hardware.
Back in the 1980s when I was at Caltech, we would build very small neural networks, train them to do simple things.
I played with hot field memories to... memorize telephone numbers and stuff like that.
But we were limited because we had computers that were literally 100,000 times slower than what we have today.
And the two things that really enabled the revolution in AI that we have today is more powerful competing hardware and large labeled data sets.
And NVIDIA has been at the forefront of that first one.
We've basically provided the hardware that enables AI.
Right. You've been with the company for about nine years now.
At what point did the shift to GPU computing kind of really, you know, for you kind of really, was there a light bulb moment Was it sort of this inevitable tide that eventually the capacity to build the hardware sort of caught up to the theory?
Well, the shift to GPU computing actually happened before I came to NVIDIA, but it was as a relation to a lot of work I was doing at Stanford at the time.
We had a project at Stanford called Stream Processing, and we were building both stream processors for graphics and stream processors aimed at And a couple of things happened.
One is through a number of people I knew at NVIDIA.
They actually hired me in 2003 to come here as a consultant and work on NV50, the chip that later became G80 when it was announced.
And then they did The really smart thing and hired one of the best graduate students out of the project I was leading at Stanford, a guy named Ian Buck, who's now the general manager of our Tesla business unit.
But for his PhD thesis at Stanford, he had done a language called Brook, which was the language for our streaming supercomputer and had ported it to GPUs.
And when he came here to NVIDIA, he worked with a guy named John Nichols, And improved upon Brooke and came up with the CUDA language that we launched in 2006.
Sure. And that really spurred the revolution in GPU computing because the GPUs had the computational resources and CUDA unlocked it and made it easy for people to harness that to solve real problems.
Let's look for a moment at kind of a day in the life of Bill Dally.
You're the chief scientist here. You said you've been working with the company for a number of years now before kind of taking that role.
What do you do all day? What's a typical day or maybe it's a broader scope, a typical month like for you?
Yeah, I say no day is typical. I do different things.
I try to carve off a certain amount of time to do research myself just to continue to stay sharp and because it's fun.
These days, I've actually been working on applying reinforcement learning to the path planning algorithms for our self-driving cars, basically treating it as a multiplayer cooperative game.
It's similar to AlphaGo, but in a game of Go, you're trying to make the other player lose.
When you're out on the highway, you're not really trying to make the other drivers lose.
You all have a mutual objective of not hitting each other.
I don't know. I was on the 880 today. The traffic in the Bay Area is brutal, but still people are trying to cooperate.
They're perhaps not as nice as they used to be.
But also, you're trying to get where you're going fast, and that may be incompatible with somebody else getting where they're going fast.
But to plan the trajectory of a car over a five or 10 second period, which is really the critical period, for decision-making, you need to predict what the other cars are gonna do.
And what they're gonna do is not independent of what you do.
If you move over to the edge of your lane and turn on your turn signal, The car next to you is either nice and they'll back off and let you in, or they're mean and they'll cut you off.
And so once you classify what kind of driver they are, You can now make predictions and use that to plot an optimal path for your car, which is cognizant of how other people are going to react in terms of your actions. prototyping up a path planner.
And the goal would be eventually if it works out to try to, you know, merge that into our autonomous vehicle offerings.
Now, autonomous vehicle is obviously one of the big areas that NVIDIA is involved in and And AI is kind of a buzzword these days, but that AI is focused on.
What other areas of either pure research or perhaps industry Are you and your team also working on it?
NVIDIA research really covers all technologies that are relevant to NVIDIA, starting with circuits.
We design better flip-flops because that's what stores the bits in our circuits.
GPUs, better signaling to be able to move bits from one place to another faster at lower energy. up through the architecture ways of making our streaming multiprocessors operate faster and more efficiently.
Programming systems, we're trying to simplify the programming of GPUs.
Networking, the NV switch that's in the DGX2 that ties things together is a project that started in video research.
All those parts are what I call the supply side of video research.
They supply the technology that will make GPUs better.
And then on the demand side of NVIDIA research, the programming systems, people really sort of straddle that.
They also develop algorithms and libraries that drive demand. but we have people doing computer vision people doing ai people doing graphics As they develop new algorithms, people use those algorithms and they want to buy GPUs to run them on, so it drives the demand for GPUs.
You said earlier when we were talking about your, I don't want to say beginnings, but your previous history back at Caltech, The hardware now is orders of magnitude faster and more powerful than it was back then. obviously the hardware and the software, the programming languages are symbiotic kind of thing.
At this point, sort of in the industry, Is one still trying to catch up to the other?
Are we kind of waiting? Is the next big breakthrough going to be more predicated on hardware getting faster? or programming languages becoming more sophisticated, easier to use?
Can you kind of see a hurdle we're trying to get over in the future?
Yeah, it's interesting. I don't think the programming languages really changed that much during that period.
The hardware has changed enormously. A lot of that period...
We were in the full bore of Moore's law where processors were getting 10X faster every five years.
That sort of ended about 20 years through that 30-year period, right?
So over those 20 years, things got, you know, 10,000 times faster.
And then they haven't gotten much faster since then.
But the programming languages aren't that different.
So at Caltech, I worked on a parallel machine called the Caltech Cosmic Cube. that we programmed in a language called Cosmic C, that the first approximation is almost identical to MPI. which is how people program parallel computers today.
So the languages haven't changed much. What has changed is the sophistication of a lot of the algorithms.
And I think the algorithms and the hardware Run step in step because if the algorithms get too far ahead, you can't run them.
And then once you get the more capable hardware, people think up algorithms that can use it.
We're talking today with Dr. Bill Dally.
Bill is the chief scientist at NVIDIA. He's been... working on neural networks far longer than most of us have known what a neural network is.
We're going to shift gears and talk about what's coming up in a couple of weeks from when we're recording this. which is the CVPR conference starting June 18th in Salt Lake City.
We mentioned at the top, NVIDIA Research is presenting, I think it's 19 papers and posters.
So first off, congratulations to you and your colleagues on that.
Let's run through some of what you'll be talking about at the conference.
And let's start with super slow-mo. The session is titled super slow-mo high quality estimation of multiple intermediate frames for video interpolation.
Maybe you can decode that, so to speak, and tell the listeners what that's all about.
Yeah, so if you want to see a motion happen in slow motion, normally you would have to use a high-speed camera to record it.
You would want to record it at 1,000 frames per second.
But if you simply took it with a normal camera and you recorded it at 30 frames or 24 frames per second, and you want to generate those intermediate frames, so, you know, the person making the diving catch looks very realistic.
You need to basically predict the motion, where that person would be in every frame, And also, if there's any occlusion, if they're making the diving catch behind a pole or behind another player, you need to decide, okay, this guy's in the foreground.
You're not going to see him as he – parts of him as he goes behind that.
And what this work does, which was done in our perception of learning research group, is it uses a neural network to simultaneously learn the motion and occlusion of those intermediate frames.
So you train it on a certain number of images where you have the ground truth, and then you can give it just the key frames, just the frames on either end.
And it will produce all of the intermediate frames that look as if you had taken them with a high-speed camera.
Now, not to put you on the spot, but how accurate is it or how do you when you talk about the so-called accuracy of something like this, how do you sort of describe your progress?
Yeah, so they have the numbers in the paper, and I haven't looked at them recently, so I would have to take some time off and actually dig the numbers up. get the results.
But to me, the proof is in looking at the video.
When I look at the, you know, one of these slow motion videos, I can't tell that it wasn't taken with a real, real high speed camera.
Could this sort of technique be applied retroactively to video that was produced months or years ago?
Oh, absolutely. As long as you have a video, it will learn the motion. and be able to interpolate those frames.
Well, our producer is telling me not to use the word Zapruder on mic, so I won't.
But I've got some ideas we can talk about later.
Okay, let's go through the list here of the speaking sessions.
Splatnet. which is my favorite title, Sparse Lattice Networks for Point Cloud Processing.
Yeah, so again, I'd have to dig up the papers to give you a lot of work in this, but this is if you have a bunch of point samples You want to be able to take these point samples and feed them into a neural network.
And if you do it kind of in a naive way...
It takes way too many resources because you would have to take all of space where you might or might not have a point. and have an input, your convolutional network there.
And so what this network does is it's more intelligent about how it samples space and only really applies to convolutions where you actually have points.
So it's a more efficient way of processing volumetric data.
So we touched a little bit on autonomous vehicles, one of the areas that NVIDIA is kind of heavily involved in, and your background working with neural networks and computer hardware for years and years.
Looking forward, where, and this is a big question, but where is all this heading or where should it be heading?
We've had guests on the show recently. The most recent episode that was recorded, and these go up sometimes in different order, was talking about...
Google demoed something called Duplex recently where they claim to have synthesized a human voice that basically I don't want to say fooled, but convinced the person on the other end they were speaking to a human was really a machine.
And so in the popular culture, in the mainstream media, you see a lot of sort of, you know, and speculation about, you know, the robots rising and all that sort of thing.
You're in the thick of it and you have been for a long time.
When you think about deep learning, machine learning, AI, the work you're doing, looking down the road, do you have a vision or just collectively, is there a vision?
Yeah, I think that there's a collective NVIDIA vision, or maybe it's my vision.
Perhaps the two are the same. But first of all, I'm not worried about general AI.
That's not what we're building. What we're building today are very specific technologies.
Typically perceptual systems that take an input and produce an output and do it very well, often better than a human could do. whether it's a self-driving car that reads its sensors and produces an output to control the automobile, or a system which listens to some speech and responds to make your hair appointment.
Whatever it's doing, it's a perceptual system.
As I look forward, there's two really key things we need to do to keep this AI revolution going and to have it cover even more of human experience than it is now.
I think we're really just at the tip of the iceberg.
The number of things we're applying AI to, there's so much more that it could do But standing in the way are a couple of things.
One is continuing to scale the hardware performance.
With the end of Moore's law, that's gotten increasingly difficult.
We can't count on the process technology to give us essentially anything.
It has to be better architecture, better algorithms that are going to fuel that.
And so a lot of what we're doing in video research is trying to understand how we can you know, squeeze the last bits of efficiency out of the hardware and come up with better algorithms that use that more efficiently so we can continue to deliver the performance which is needed to scale the value of these systems.
The other is to try to understand how we can do more with less data.
One of the big bottlenecks in AI is the need for labeled data.
It's really sort of the limiting factor in almost every application you look at.
But very often what you'll have is you'll have a lot of samples of very boring data.
Take the autonomous vehicle. You've got a lot of straight driving down Highway 101 without anything very interesting happening.
And then you have very few samples of something interesting happening.
There's somebody dropping a paint can off the back of the truck in front of you, or you know, a ball bouncing into the road or these things.
And so the question is, how can you extract the most value out of even a single sample, whereas the conventional wisdom in AI is that you need many examples of something for the network to learn it.
How can you learn a lot from a single example or even to be able to generalize from an example you've seen to one that you haven't seen?
And learn from that and to pull out a lot of details rather than just one label for a particular image.
Once we've learned a little bit, can we take unlabeled images and learn from them by taking the knowledge we have, putting them on there, and only asking the human labeler to fill in the unknowns.
And so to me, those two big things we need to do to push the revolution forward is more efficiency out of the hardware to keep the revolution going in the absence of of Moore's law and ways of squeezing more efficiency out of the data so we don't have to put particularly so much manual labor into labeling data.
So we're doing a lot of things in NVIDIA research to try to close the gaps on both of those.
So when you started with NVIDIA full-time in 2009, there was a research group in place then.
We were talking before and you said roughly about 10 people.
Now, if I understand correctly, that group's about 200 strong.
So tell us about that time, about your work and sort of building the research organization, what that means to NVIDIA now, because obviously things have changed quite a bit.
NVIDIA, always known as a graphics company, still obviously is and I think always will be But obviously now AI is a big part of what's going on and that comes from the research.
So tell us a little bit about how NVIDIA research has evolved.
That's a good question. It's good to think back to sort of 2009.
So I inherited a group that was doing almost entirely computer graphics and a little bit of GPU computing.
And very shortly into the process, I took the larger part of that group, which was all working on ray tracing.
And they had sort of completed their project and wanted to productize it and move them out of NVIDIA Research to the content and technology group, it was then kind of a clean field of what to do.
And after a lot of conversations with Jensen and with David Kirk, who is my predecessors as chief scientist, it was really clear that we needed to invest in the areas that would make a big difference to NVIDIA. in the future and try to, you know, distinguish what the product groups are doing from what we're doing in research.
In fact, I set the goal for NVIDIA research is to do excellent research and there are a number of ways you can measure that and to make a difference for the company.
And if you look at a lot of industrial research labs, many of them do one or the other, but very few of them do both at the same time.
One common failure mode is they do great research.
They publish lots of papers. but there's a barrier between the research group and the product group and no ideas and research ever influenced product.
And others help the product groups out a lot, but they're actually doing advanced development.
They wind up not actually publishing papers or really advancing the state of the art.
They're more taking... known technology and applying it to the company's problems, which is valuable, but it's not research.
And so I think we've been very successful in pointing out what areas are relevant to the company, building a circuit group because we needed really great circuits, building an architecture group. because you know architecture is core to getting efficiency out of our gpus you know building um you know the programming systems group because we need to program gpus more efficiently i actually created a networking group the group that produced the the predecessor of mv switch because it was clear to me at the time that connecting GPUs was going to be an important thing going forward.
And then in addition to sort of recruiting the people and building these different groups on both the supply side and the demand side, We had to create a culture, which was a culture of you have to do research.
And there's a number of ways we measure that.
We try to push people to publish. Because peer review is a great quality control measure.
I mean, if you put your work up for the scrutiny of somebody who's a real expert in the field ripping it apart... it's a very humbling experience that makes you better right as a result of that we you know people are really doing great research we insist on publishing in the top tier venues, things like CVPR.
And then we also push people to try to say who's We identify two people in the company.
One is, who is the consumer of this? And then often, who is the champion?
Who may be different than the consumer? And to give you an example of that, we developed a memory technology once in our circuit group where the consumer was, in our VLSI organization, the people who are responsible for memories.
But they weren't really excited about this.
They had lots of work to do. They didn't, you know, say, oh, another memory, just what I need.
And it turned out there was a person in the GPU organization who this memory solved a problem.
So they became the champion, and they would then go to the guy in the VLSI group and say, what do you need?
I'll give you whatever you need, but you have to take this memory and productize it.
And so we actually at the beginning of a research project try to find out who these people are and it's often the same person and get them involved because if you wait until the end, it's too late.
Because very often there's a number of subtle constraints.
You could do the project going left or going right, and it doesn't really matter.
It's incidental. But if you go left, these guys can't use it.
And if you go right, they can't. But you don't know that unless you get them involved up front where you're making all of these decisions. early engagement with the product groups, but we make sure we're far enough out that it actually is research.
And the result has been really exciting.
We've developed a number of things over the years that have made a huge impact on NVIDIA's products.
So is it often the researchers themselves who will go to sort of their manager with an idea for a new undertaking and then if it kind of holds water enough.
Then you'll get somebody from the product group involved.
Do you ever get requests coming from the product side to further investigate certain areas?
Yeah, all of the above. I mean, I think the best projects in NVIDIA research are grassroots or bottom-up and individual researchers who really know the technology best get excited about something and start pursuing it.
But they're also trying to think of, as they pursue it, how can I bridge the gap and how can I take this very basic research idea and make it of use to NVIDIA?
As you were speaking, I was thinking to ask, what do you look for in a young researcher?
But maybe a better way to frame it is, I'm sure we've got some folks listening who are You know, researchers right now kind of starting out their careers in school right now.
Any advice? you would give to them or somebody who's looking at sort of bridging that gap between really just having a passion for their work and their field but knowing not only do I have to make a living, but in the current world, there's a huge opportunity to develop a great career working in industry.
Any words of wisdom or advice you would pass on to them?
To me, the one thing that really, I think, helps people succeed is if they're very broad.
A lot of people when they are in school and they're studying, and actually the PhD programs push you in this direction, wound up becoming very narrow.
They know a lot about their one little area of the world.
But I think very often being able to bridge things requires making connections and being able to be broad enough that you understand how this thing you're doing, you know, on deep learning algorithms can impact you know, something over here in VLSI hardware, and now you can connect the two together and do something that you couldn't have done on either side.
Right, right. A few weeks, it was probably a few months ago now at GTC, I had the chance to sit down with Brian Catanzaro and talk a little bit about NVIDIA Research and his work there.
And The topic of AI came up and NVIDIA's evolution from being a graphics company and gaming company to getting into AI.
And this perception that exists some places in the world, I guess, that, oh, they just got lucky.
Obviously, it wasn't just luck. Somebody well before NVIDIA became the AI company Somebody in NVIDIA research obviously had an inkling that, hey, the stuff we're working on, this could be applied to AI also.
Let's look into that. What can you say about that or any stories you can tell looking back or kind of was there a pivotal moment when the collective group and the research org thought, hey, wait a second.
I think there is a little story I can tell you about this.
So the intermediate step from going to graphics to AI was GPU computing.
And I think I told you a little bit about that already with-
Ian Buck coming and working with John Nichols on CUDA and launching the features in G80 and GT200 to support But the next real step along the path I think happened when I actually had a breakfast with Andrew Ng.
So one thing I do is I try to keep my connections from Stanford up because I think it's a great way of getting sort of technical intelligence, what's going on in the world that could be relevant.
And so I try to, you know, talk to people regularly, have breakfast, lunch, coffee or whatever with them.
So I was having breakfast with Andrew Ng.
It was probably late 2010. And at the time, he was actually working at Google Brain.
It was just at the tail end of that. on this big project to basically do unsupervised learning to recognize things and images on the internet.
It's probably best known as finding cats. on the internet because it would find a whole lot of images that looked alike.
And it turns out they were all cats. And to do this, it took 16,000 CPUs in the Google cloud.
Google's probably one of the few companies at that point in time that had the computational resources to do that.
And so he was telling me about this over breakfast.
I said, that's really neat. I didn't realize you could even do that. compared to the stuff I'd done with neural networks back at Caltech that seemed nearly impossible.
And so it struck me that, gee, maybe the computational resources are here now that we can do this sort of thing.
And I was wondering... gee, why don't we use GPUs for this?
So I suggested a joint project to Andrew that he was very excited about because he wanted to use GPUs but kind of lacked the expertise.
And so... I forget exactly the method of doing this recruiting, but I sort of got Brian Catanzaro to work with Andrew on this.
Brian at the time was actually a programming system researcher working – on a project called Copperhead.
It's sort of when Python meets CUDA. He was excited about deep learning, had all the right background to do this.
And so we made his assignment to work with Andrew on this.
I immediately, I shouldn't say immediately, but after he wrote the software, which was part of that project, which actually evolved into QDNN and kind of launched the company's whole endeavor into – into deep warning, Andrew hired him away.
And at that point, I regretted introducing him to Andrew.
But actually, it worked out real well. He was a great evangelist when he was at Baidu.
That's after Andrew went to Baidu. Right. for using GPUs for this.
And then now he's back. And so we got the best of both worlds.
But I was very annoyed at Andrew for a little while for hiring him away from me.
But that was sort of how we got started in deep learning is starting from that breakfast and then kicking off the joint project with Andrew having Brian collaborate with them.
It eventually all comes back to, I guess, people who you know or even better just sort of. how you treat people along the way and, you know, to be able to call them up and have breakfast, you never know what's going to come out of it.
Before we wrap up here, we usually kind of wind up the podcast.
And often we're speaking to people outside on video, so perhaps it's a little different. looking towards the future.
This whole conversation has been about the future in a lot of ways.
Where do you see the field or NVIDIA's work in particular, or even to hone in more your own work in the field?
You mentioned you're still researching, which is great.
Where do you see it headed over the next, let's say, three to five years?
That's a good question. You know, I think that there's always as work moves forward, there's the evolutionary component and there's a revolutionary component.
And in research, we try to focus on the revolutionary part.
And so on the hardware side, we're going to see drastically more efficient ways of doing deep learning, especially on the inference side where the bulk of the computation happens.
And I think this is going to come from co-design of the circuits, the representations and the models, the networks that people are using for these.
And I think that to complement that, we're going to see more powerful networks, more that we're able to train with much less data so that we can apply them to areas without the huge data curation and labeling costs that we have today for these new applications.
And I think that that's going to end up being very powerful and it's going to transform almost every aspect of human life.
I'm literally without words from imagining every aspect of human life and the optimism I see your face.
I think it's translating over the audio, but in case it's not, Dr. Daly looks optimistic and that's all I need, frankly.
Bill Dally, thank you so much. I know you have a very busy schedule, so we really appreciate you taking some time. to sit down and talk with us.
Obviously, best of luck to you and the entire research org at CVPR and going forward.
And thank you. Oh, thank you very much.