Hello, and welcome to the Nvidia AI Podcast.
I'm your host, Noah Kravitz. We're going to get meta for a second here, so hang on.
Back in 2003, Nick Bostrom of the University of Oxford wrote a seminal paper exploring the question of whether or not we're actually living in a computer simulation.
Since then, philosophers and scientists of all sorts have been arguing over the answer.
Whether or not we're living in a simulation ourselves, what's not in question is the utility of computer simulations in our own reality. our ability to create ever more realistic computerized simulations of human beings, digital humans, if you will, serves us in fields ranging from entertainment and gaming to design, architecture, and engineering. simulating not only what humans will look like in a given environment, but also how they'll behave, or how we'll behave, I should say. lets us create more useful, efficient, and safer designs across all sectors of industry and life.
Here to explain a little bit more and pull the curtain back a bit on the cutting-edge work his team is doing on digital humans is Simon Newton.
Simon is director of graphics and AI at NVIDIA, where he leads the digital human efforts, which I'll let him explain in just a second.
Before joining Nvidia, Simon spent more than 21 years in the visual effects industry, working on the art and technology that powered a number of high profile games and movies you're no doubt familiar with. you're with.
Simon, welcome and thanks for taking the time to join the NVIDIA AI podcast.
Thank you, Noah. Thanks for the introduction and very happy to be here.
So we were talking bit before we hit record.
And there are about a million and one things to talk about here.
I was listening. I'll start by saying that anybody listening to this, Go check out the, it was called Digital Humans for Digital Twins.
Did I get that right? The GTC session that just went up?
Yeah, that's correct. Go check that out.
It's fascinating on a number of levels, technical and the sort of meta stuff that I like to think about.
But let me start by at the beginning here, kind of the basics.
Can you explain a little bit about what Digital Human is and what the work that your team has been doing is all about?
Sure. Digital human, in essence, is a digital version of ourselves inside the virtual world.
So it could be something whether a cartoonish avatar of yourself or something more of a realistic version of yourself. in the digital world.
It could be movies, it could be games, it could be in VR environments, could be in You're as a virtual version of yourself as a teacher.
It can be any of these forms and it's growing.
And so what are some of the applications?
I touched a little bit, and I was really just kind of cribbing from the GTC session I listened to.
But what are some of the applications that folks might be familiar with for digital humans?
I mentioned gaming and entertainment, and it made me think of one of the The first episodes of this show that I had the fortune to do when I first started hosting was with an artist talking about using AI to render zombie armies for video games and kind of taking that,
That sort of grunt work off of the visual team and letting the computer kind of churn out these digital zombies.
Is it that kind of stuff? And then... What are some of the more recent things that are being done with digital humans?
So I think digital human for the past two decades or so Everybody knows in terms of the entertainment industry, whether it's games or movies, That's been used a lot.
The goals there are a lot more for visual satisfaction, But I think over time, and especially since I joined NVIDIA, I've realized there are a ton more needs and usage for digital human outside of the entertainment sector.
For example, medical simulations, education, as I mentioned, and a new avenue into the AI-based, AI-driven digital human assistance. where you see a customer service that's represented as a digital human versus just a voice or in the past for graphics.
But that gets more into simulation versus just a visual result.
Those haven't really come about till now is because it's simulating digital human is much harder and the computational needs are much higher.
We're getting close to the point where that's possible.
So when you talk about simulation as opposed to just a visual representation of What does that entail?
Again, from the talk, I know there was discussion about human behavior and sort of simulating the way that people might behave in a given environment.
But can you talk a little bit more about what that means?
So I'll take Digital Twin as an example.
Digital Twin, it's basically virtualizing a real company process.
It could be a factory, it could be some sort of a building that you're trying to build.
But basically the idea is that, and I'll take factory as an example, that is the most clear.
So a factory, like a car manufacturer, has many assembly lines, many workers.
And the idea of Digital Twin is simulating that one-to-one inside the computer first.
That is for safety reasons, for cost reasons, and for many other reasons.
And in simulation, what that means is Now, if I have a robot that can pick up a ton of metal, that is fully simulated that inside the computer, all the mechanics and the work of actually this robot arm is real, meaning Real as in I'm using this solver in the computer to solve what power does it need to pick up this one ton of metal.
And so the same thing applies to digital human.
If I am simulating the workers within the assembly line, All the motion, all the marks, all the timing has to be realistic, not just visually appealing, but not plausible in reality.
How do you capture that data to start with?
How do you determine what is realistic for human behavior?
If I'm on an assembly line, I don't know, 100 feet down from a robot that's moving a ton of metal at a time.
What kinds of things are important to capture and how do you capture and put into the system what those humans are doing?
It is a really good question. In computer graphics, there's multiple discipline.
Animation is one of them. For things like CG rendering, where here's a picture, here's a CG version of it, it's very easy to compare.
For motion, it gets more subjective. We do start by a lot of our simulation by using and training from real-world data so we could motion capture various different peoples and different proportions and things like that.
But the system itself learns those behavior and not only learn the motion of these human behavior, but also the physics, what it takes to move the joints or muscles to mimic those behaviors.
So we're getting closer to simulating those motions with actual physics solvers.
And so it's not just purely say, hey, I'm going to mocap a person and just put this on the skeleton.
So that's where a lot of things, that's where we're heading towards in terms of simulating.
But there's still a lot of work to do. I can only imagine.
Let's take your example of a factory a little bit further.
So when you've created a simulation, a realistic simulation, in terms of the behavior and the physics and the other elements that are crucial to a factory floor.
What would you, or maybe in this case, it's the customer, Automaker X, what would they do with the simulation and how would they use it to inform? you know, real world decisions going forward?
That's a really good question. So a lot of the factories, let's say, The people who are in charge of making sure there's a successful assembly line, the flow that keeps consistent and keeps going, they have to time very carefully what each process takes.
So if you're putting together a coffee machine, you have the jar, the mechanics that's inside the coffee machine, the electronics and all that.
Somebody has to weld it, put it together, put the screws on, and each of those have to be calculated fairly accurately.
Right now, those are all so-called role-played by people.
So they kind of try it out and do that. Okay.
Of course, there's a bit of inefficiency with that and inaccuracy.
With Digital Human, they can try different scenarios out. what makes sense, what's safe.
In a factory like with bigger machinery, the robot arms could be in the way.
So they might not be able to make this 15 second round trip that they need to.
They get to work all that out in the computer version.
Okay. It almost sounds like choreographing a very complex dance between humans and machinery and other things to kind of make sure.
When I go to get my part I need or whatever it is, I don't get hit in the head by this swinging robot arm.
Absolutely. It is like that. So Simon, my kind of nascent understanding is that digital human is already affecting like really kind of a wide swath of industry and society.
Can you give the listeners a little bit of an idea of all the different areas of life that your work is impacting already?
Yes, definitely. One of the things I brought up earlier was in the past two decades, a lot of digital human focus has been more visual.
After joining to NVIDIA and having many different types of industries of customers talking to us about digital human, And I've learned and saw that the usage is actually much, much wider than entertainment and a lot of the exposure I've had in the past.
A lot of it comes from medical people training for doctors and nurses. or even sometimes patients for recovery and consulting.
Interesting. They're all actually looking at digital human.
I'll come back to this because there's a good example for that.
Education, just training, for kids because sometimes you know they might want to be playful with um with a more special type of character than a real person, and it keeps their engagement high. as well as some of these cases we've talked about for digital twins and all that.
Also, there's use cases for a lot more AI driven assistance. as well as streaming characters.
I think there's more and more people who would like to have a virtual presence of themselves in social media, Twitch, games, Discord, and all of that.
Right, right. So we're looking to see how we can make that easier for people to do.
So the one example I wanted to give was we came across a medical education company that they create training videos for just medical professionals.
Okay. One of the project they're getting into is that sometimes certain medical condition patients, they prefer talking to a virtual email.
In some of these cases, they could be disfigured and they're not comfortable.
They're trying to build digital humans for use cases like that.
One of the condition was that it has to be believable.
But they had an overwhelming response that some of these people prefer talking to a digital human for their consulting.
So that would not be something I would have thought of.
Yeah. It makes me wonder what believable means or will evolve into meaning in that kind of a situation where you know, it's somebody or a cohort expressing we're actually more comfortable with something that's digital But kind of, I don't know, it seems sort of by nature of the term digital human and this idea of believable that there has to be some sort of human based element. kind of mixed into it.
And it just makes me wonder, you know, a couple, a year, five years, whatever it is down the road when the technology has evolved more, what that you know, particular most comfortable for this medical cohort, or even to your earlier example for kids. you know, and having that comfort with a digital character, what that'll look like going forward.
It's the beginning of wild times. Yeah, absolutely. and that those things you just mentioned gives a great example of how diverse digital humanists and the needs from stylized to realistic.
You might want to talk to somebody you feel like they're emoting and having believable conversations. to just, hey, I wanna have fun.
Does the work that you're doing either now or going forward... I was wondering before about the technical aspects of applying some of this stuff obviously to 3D rendering, but then into virtual and kind of extended reality spaces.
But what we were just talking about made me think a little bit about social robotics and the idea of physical robotics. machines in the world that are able to emote or otherwise kind of express things that tap into a sort of particular strand of human believability.
Is the work that you're doing, is that applicable also to robots and physical worlds? renderings, if that's not a good term.
You know what I mean? Yeah, it is, okay.
Yes, absolutely. We think that, and this we didn't get into too much, is the conversational AI aspect.
To train an AI so you can interact with is definitely a big part of what we're thinking for robotics. and how you train something virtually that can create and you can teach it a bunch of these behaviors in conversing. having the right understanding of what you just taught me and I can give you the right response. and just even the polish of that personality right it is the same core that drives the human but can be applicable to many things.
We have an actual robotics lab, for example.
Sure. That is for this exact purpose where we can simulate and train the computer, build the real thing that we know it's going to work.
But that's where the simulation has to be very accurate.
That also has to do with safety because Robots need to know how to work around humans.
Otherwise, yeah. Disaster, yeah. Yeah, so.
Great. Yeah, very much so, yeah. So let's talk about this idea of a digital twin.
We talked about the title of the GTC session, Digital Human for Digital Twin.
I'm wondering right off the bat, and I thought this is where you were going with it, but you didn't.
Can I create... a twin of myself with this technology and put it into a situation and let it run wild and see what, you know, how would I behave if I was a...
Assembling coffee machines or what's the, what's the idea behind digital twin?
So a lot of that title. is uh it can it can actually be misleading it's it's a lot of people misunderstood as in this is the twin of myself right but in It actually is a fairly new term that I'm aware of that came up these few years.
And it actually has nothing to do with digital human for that term.
It's really a simulation It's like you're making a twin of some other real environment in the computer.
That's really good. Got it. Okay. And there's a human lives within or simulates within digital twin is usually mostly what happens.
What you mentioned about fully simulating a virtual agent in environment is one of our goals, is to get to that point where It has enough knowledge.
It understands the environment. It understands certain rules in physics that it can be on its own and see what happens.
Now, I don't want to make—well, no, I'm going to go ahead and play the role of making what might be a little bit of an ignorant leap to a question, but—
This sounds like it must be different, but it makes me think of pursuing general AI.
In other words, if We're talking about the goal being to be able to create a digital version of a You know, a person who not just doesn't just look like a person and move like a person, but actually has behaviors that are realistic.
How similar or different is that to kind of pursuing general AI?
It's a little different in that we ourselves as humans judging other humans or digital humans. are the pickiest critique.
I think if we're tasking an AI just to say recognize dogs and cats, It's a very specific domain training and it does these task oriented things well.
But for another somewhat organic agent that lives inside the computer that you wanted to understand and know how to behave on its own.
It is much much more complex. You can almost think of many AI working together.
There's a lot of both thinking decisions, behaviors in humans, that have many, many mechanism working together seamlessly to create that behavior.
So even just the notion of looking real and moving real. is a very hard goal.
100%. And I didn't mean, I hope I didn't downplay the complexity of that.
Just that itself. So that itself having AI to help a lot.
AI I think makes it plausible to even attempt at what we're trying to do.
But it's much more complex, I guess, compared to general AI.
Our guest today is Simon Yoon. Simon is Director of Graphics and AI at NVIDIA, where he's leading something called Digital Human, which we've been talking about.
You have a background in visual effects, and I teased at the top you've worked in motion pictures. and gaming and other aspects of entertainment.
Can you talk a little bit about how AI and all the technologies we put under that umbrella have changed the kind of work you do.
And I know that you've been with NVIDIA for a couple of years now, is that right?
So there may have been sort of a forced kind of jump or change in the kind of work you do, kind of moving into this role and on this side of the industry.
But how has, in the two decades plus you've been doing this kind of work, how has AI played a part and kind of advanced or change or whatever the right word is.
You know, all of this work, but particularly The visual aspects, right?
The making things look more realistic and how they look and how they move and that kind of stuff.
Yeah, definitely. So before I talk about that, I need to explain a little bit, I think, what some of our goals are.
A lot of people are working on some great solutions for digital human.
What makes sense for NVIDIA for us to be thinking about this problem, We think, especially my background in visual effects, creating a realistic digital human is a very time-consuming and laborious process.
It requires a lot of just throwing people.
There's a lot of great technology behind it, but it is such a complex problem that it does end up to require a lot of just artists and TDs and technical people to put a lot of efforts, sometimes sequence to sequence, frame by frame, shot by shot,
And so that itself screams an opportunity to potentially make it easier and make it better.
And that's really, I think, where NVIDIA's vision is.
It's three pillars of experience in AI simulation and real-time graphics really is a good foundation for tackling some of these problems.
And we look at the problem as, it's a very giant problem.
We don't see it as a, It's going to be a collaboration with a lot of different people to get to solve some of these problems.
And especially as the coming of the metaverse as, as we see it.
Yeah. Yeah. How this gets into AI is it has to do a lot with the, how do we make it easier and simpler?
How do we potentially democratize? 3D creation, digital human creation.
How do we potentially do that? And I don't think we can do any of that without the help of AI.
This gets into my experience and how that has changed since more exposure of using AI, understanding of it, and also the difficulty of it.
It can do some things that are very magical, I'll quote Mr. Jensen Huang for a second.
We are at a point where programs can write programs that no humans could write.
And as cliche as that may sound, it is very true.
You cannot write a program that can do image recognition better than an AI system.
You cannot do that manually. Same thing with behaviors or how to automate animation.
How do you or how do you accelerate Muscle simulation that takes days and weeks to real time.
That is all really strong aspects of deep learning and machine learning.
It's something that you can't do. Now, the tough part about leveraging AI is complexity. it is, as I mentioned earlier, it's very good at doing a particular thing you wanted to do.
If you train how some motion work for running, walking, or locomotion, it doesn't mean it'll know how to get into fighting mode or how to tumble on the ground and things like that.
So those are the kind of breakthroughs we are also looking at in terms of adaptive AI and smarter, more intelligent, probably networks of AI systems.
You teased the metaverse. Can you tell us a little bit about the metaverse and Audio Two-Face and...
I'm probably mispronouncing that, audio to face, but I like audio to face.
It sounds like an evil character, a deep fake villain out there somewhere.
Tell us a little bit about that and about... I know that the Digital Human... group has been working kind of under the radar for a while.
But Audio Two-Face is your first kind of product that you're publicly launching.
Tell us a little bit about that. Yeah, that's correct.
So audience to face was something that there was a MV research paper that came out in 2017.
And when I joined Nvidia, I felt that was.
That paper had a lot of potential, but also there were things that can also be better so that we can truly productize it.
And then right about that time, we had a lot of game developers that come talk to us and say, They're pretty much at the verge at the end of what they can do manually in terms of creating these games and large worlds.
This large world in game is very analogous to metaverse.
We'll talk about that. But basically, for example, this one game developer mentioned they did motion capture straight for two and a half years for a single game.
They have more objects in this world than they can imagine. manually tracked.
They can't even create enough just the environments, props, and all this, not to mention humans.
That's where games are heading and their next game is just going to be two or three times bigger.
So they asked, hey, NVIDIA, can you guys come up with the technology potentially to leverage AI to automatically generate some of these, you know, whether it's the models or behaviors and things like that.
So that's, Out of some of the existing needs and our goals for simplifying and accelerating how people create digital human, we came up with how to push forward with audio to face.
It's a technology that basically you can create facial animation based on your voice. and it's specifically for lip sync voice-based animation.
We do a lot of talking. A lot of digital characters do a lot of talking and Right now, the current technology to do that is very time consuming, very domain expert specific.
Not any person without 3D knowledge can do it.
So those are the motivation to create something that And we have an example of this is there's a six-year-old girl from one of our co-worker who's singing happy birthday as a rhino.
Nice. for dad. And that's a perfect example of what the kind of accessibility we want. our technology to have is to have more people joined this 3D world and created it.
I mentioned off air that earlier this morning, I showed my own kids who were, my youngest is eight, The demo, there's a very cool little animation, kind of Bohemian Rhapsody style with three heads with the rapping. uh showing off audio to face uh folks can go check that out on the video site it's great But now I'm imagining, oh, I should download the open beta and let them go nuts with, you know, see what they can come up with a rhino head or I'm sure.
There'll be a Roblox-style character in there before too long if I turn him loose.
So tell us about the metaverse, then. So...
There's been many discussion about the metaverse where you have one or many virtual worlds that can bridge together and where a digital human can exist in this world, having its own society, having its own it's a live social platform in the digital realm.
And a lot of it, already has been pushing towards that direction.
You know, Fortnite has, from Epic, have created some great examples.
Yeah. There are concerts within fortnights and things like that have invited many people and the response was outstanding.
Like a lot of people really enjoyed Disney medium.
And we think that's just going to get bigger and bigger and more.
It would even be less of a specific, uh, game or game world, but more general, like Facebook, if not even bigger.
The other thing that I mentioned in the GTC talk about Digital Twin is that I think there's also a growing trend for enterprise metaverse.
And that actually might hit, in my opinion, even sooner than the commercial or than the social metaverse.
Yeah, the commercial stuff always gets the headlines, but the the enterprises where a lot of this stuff really takes off first.
Yeah. Yeah. Because there's actual business needs for this right now.
People are, I think COVID definitely accelerated this, but I think everybody's heading this way already where It's a global world now, and we have many people collaborating from different cities.
And so, you know, being able to truly have an experience where you feel like you're there and being able to visualize Not only from a visual perspective, but affect the simulation of a factory you're building somewhere, or certain processes or we have actually some clients asking, they're trying to design buildings and architectures together virtually.
And they want to make sure that things like audio to face is something they can use because a lot of design language isn't spoken.
It's by looking at the expressions and gestures and things like that.
That's what they're used to in a meeting.
And so there's more people who are working on this problem than you know, a lot of public have talked about.
So I think this idea of metaverse for enterprise is very real and it's coming We could talk about this all day, but I want to ask you, kind of throw a technical question, kind of a broad technical question at you.
As you think about the work, and let's just cap it to kind of since you've joined NVIDIA and have been working on Digital Human.
Think about the work that your team has been doing.
Was there a technical moment, either... an incredibly difficult challenge that you were able to get past, or maybe something that kind of surprised you.
Something that kind of leaps out as sort of a technical kind of watershed or big moment related to the work you've been doing on Digital Human that you could share with the audience.
Yeah, there's many. I'm sure, yeah. That's what we're trying to figure out.
Democratization hasn't happened because it's a very hard problem.
We can have cool ideas, but to make it something that you know, our parents can use or our kids can use is a different level.
Yes. But I think it's a really great question.
There is one thing I can think of that was a good Realization when we're trying to productize audio to face and it's it's also a testbed for an AI product. to make sure the customers don't have to train in AI to use it.
It was part of the design and to make it easy.
If you ask a person, even though they're manually animating a face and it's a slow process, but they can get to the finish line.
But if they have to change a parameter and wait for a couple hours every time, that's automatically not going to work.
So by thinking how to make some of these AI technology productizable, make it a reality that can really help people.
I think that combining traditional or current computer graphics technology seamlessly in a way with AI
Is something that you know we actually came up with and tried and so far it's been working pretty well, having pretty good positive feedback in our open beta.
And so I think opening the mind, and this is both a conclusion that It's difficult to combine some AI technology seamlessly in 3D graphics, but we are open more to that and so far, if we are careful in how we combine them, it can yield to some really fantastic results.
And that's something that I think people tend to think of AI as its own thing that can solve everything.
Right. I think combining both simulation and also real-time graphics together, and it doesn't have to be one solution.
It's a hybrid of many. and that design of it is very important so i think that was a good realization There are a million, maybe in my head they're related, maybe they're not, I don't know, a million questions I want to ask you about VR and AR and XR and all that, which just means we're going to have to invite you to come back on the show sometime if you're game.
I'd love to. Excellent. For now, are there places people can go?
We mentioned Audio2Face. There's a great little animated demo or video clip demo, and also the open beta.
There's that GTC session that we've been talking about, Where are some places that folks can go online if they want to dig in more? to the work that you and your team are doing and even get their hands on or their eyes on any of the kind of more technical resources as well.
So I think definitely the path to GTC. There's some great talks about Desert Human, especially this past April one.
And a little bit of a plug. Yeah, please.
Stay tuned, you know, for SIGGRAPH. Okay.
We have a lot more plans. Excellent. Good.
We like plugs. That's good. All right. Well, Simon Yoon, thank you.
And again, let's, let's pick this thread back up.
Maybe, you know, down the line when some more stuff is out in the public, because there's, So much to get into here and all kinds of bad metaverse puns that I'm just dying to make, but I'm trying to buy my time.
But thank you for taking the time to come on the show and best of luck to you and your team and everything you're doing.
Seems like one of these things that we're catching the beginning of the bandwagon, if you will, before lots of folks get hip to... to this stuff that's probably impacting their lives already and they just don't know about it.
Yeah, definitely. Thank you. This has been fantastic.
Thank you. Thank you.