Hello and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. We're coming to you from GTC 2024 in San Jose, California, and we're here to talk about nerfs.
No, not foam footballs and dart guns, but neural radiance fields.
What is this kind of nerf? It's a technology that might just be changing the nature of images forever.
Here to explain more is Michael Rubloff.
Michael is the founder and managing editor of RadianceFields.com, a news site covering the progression of Radiance Fields-based technology. including neural radiance fields, aka NERFs, and something called 3D Gaussian splatting that I'll leave to Michael to explain.
Michael, thanks so much for taking time out of GTC to join the AI podcast.
Of course. Thank you so much for having me.
So first things first, goofy football jokes aside, what is a Nerf?
What does that mean? Yeah, so essentially you can think of a nerf as they allow you to take a series of 2D images or video and...
And what you can do from that is you can actually create a hyper-realistic 3D model.
And what that allows for is once you have it created, it's like a photograph, but it's Perfect from any imaginable angle.
Composition is no longer a bottleneck. You can do whatever it is that you would like with that file, and it will look lifelike.
So if I were to take, I don't know how many, two, three, five, pictures of the two of us sitting here right now in this podcast room, if you will.
I could put those together into a nerf, and then I would have a file that...
I can look at it from different perspectives.
I can sort of move through. How does that work from the user experience side?
Yeah, so typically the recommended amount is somewhere between like, I'd say 40 to 100 images.
It's really easy from like a video because then you can, you know, you can just slice up individual frames from that video.
But there are some methods, actually, that are going all the way down to three images and it's actually able to reconstruct, which is just...
Mine boy. I'm too used to, you know, few shot and zero shot learning and things like that.
So I'm like one image. Yes. Yeah. They are getting there.
They are getting there. And there's been a ton of amazing work, one called Reconfusion from Google, which is just shocking.
It can go down as far. as three images and it's still very compelling.
But yeah, so once you actually have gone ahead and taken your images or video you would run it through a radiance field pipeline, whether that's a nerf-based one or a Gaussian spotting-based one. several cloud-based options where it's just drag and drop your images and it does all the work for you.
The resulting image or resulting file, yeah, you have autonomy over it and you can kind of... experience that, whatever you've captured from whatever angle that you would like or whatever your use case might be for it.
How are they created? Like, without getting too technical about it, can you kind of give an overview of kind of what's going on behind the scenes to put these together?
Sure. So once you have your initial images, the first step on both nerfs and Gaussian splatting is running it through something called structure from motion. where essentially you're taking all the images and kind of aligning them in a space with one another.
So it's kind of taking a look, saying like if... image X is over here and image Y is over here.
Here's how they overlap and converge with one another.
So that's the baseline approach for And from there, they each have their own training methods where nerfs have a neural network involved in the training of them, whereas Gaussian splatting uses rasterization.
Okay. And I guess that kind of begs the question of beyond what you just said, or maybe that's it, but what's the difference between a nerve and a Gaussian splotting?
Yeah, so nerfs essentially were created first.
They were found as a joint effort between the University of Berkeley and Google.
Okay. So, nerfs have an implicit representation for them, and they're trained through a neural network.
So there's a lot of work being done to get higher and higher frame rates associated with that.
Whereas Gaussian splatting uses just direct rasterization.
And so you're able to have a much more efficient rendering pipeline where you you can really easily get 100 plus FPS and you can use them with a lot of different methods as well.
So they're very compatible with 3.js. and React 3 Fiber where you can see them being used in website design now and being on platforms like Spline.
Very cool. And so that kind of gets into the next question, which is how are they being used now?
I saw some examples on, I don't know if it was the NVIDIA developer blog, And they used a kind of an obscure song that my wife and I like a lot.
So it was perfect, right? But it was like a – it was a nerf, I believe, of a couple walking down memory lane.
They were walking outside and, you know – the foliage around them.
And you were able to, I was able to sort of zoom around from different points of view, see their front, see the back, look at the trees, that kind of thing.
But beyond sort of a demo scene like that, how are nerfs being used out in the world?
Yeah, so that specific one actually is of my parents that I took.
Very cool. And so I took because that's one of the major use cases for me.
You know, I want to be able to document my life not only in two dimensions, but I want to have a hyper realistic three dimensional. moment in time frozen.
And so for me, you know, that's one of the personal use cases.
But on a more commercial basis where you're seeing a lot of the early adoption is in the media and entertainment world as well as the gaming world too.
So yeah, For instance, Shutterstock has been putting together a library of... radiance fields where essentially what you're able to do so say hypothetically you are wanting to film in Grand Central Terminal But you cannot afford to shut down all the traffic and all the trains and all the foot traffic through that to film.
What you can do is using a radiance field, if you capture it once, you're able to then bring that file into Unreal Engine and into a virtual production environment.
And then you can film infinitely. And there's no more rush outside of the actual rental rate.
And you can go and get the shot that you actually need.
And that's where it's starting to get adopted pretty early on.
And similarly... in the gaming side of things through generative radiance fields because you are able to create these from text and images. and newly video as well.
Now you're able to drop these assets that take, you know, I think the fastest method currently takes about half a second to create a full 3D model.
Wow. And you can put that straight into one of the game engines.
Right. Yeah. And so let me sort of play this back and see if I'm grasping it correctly.
So if I were to go to Grand Central and do my very short road to holy speed and my very quick shoot and come away with enough images to create a nerf.
I would then be able in a virtual production environment to kind of create scenes or put elements into scenes from all of these different points of view, not just from a single perspective?
Is that kind of the big... Yes. Yeah, that's correct.
Where essentially you're able to sync up the nerf to the actual camera and then you can use the full virtual production pipeline to go ahead and create.
Right, right, right. Wow. This is something I probably should have asked you at the top.
So listeners, forgive me for not having a more... scientific inquiry kind of way of organizing my own thoughts.
But a radiance field, what does that term mean?
That's a great question. So essentially, you could think about a radiance field as a, well, if you just break it down into the two simple words, words.
The radiance is just what that individual color would look like based upon your viewing direction.
So say if you're looking at like a, you know, a say, a glass or something, and you can see that there's an actual reflection there.
Depending on how you look at it and what radiance fields offer, is something called view-dependent effects.
So as you move your head around, just as you do in real life, light changes, light shifts and reflects.
And just like that, radiance fields are able to model that effect.
So you could think of radiance fields as the actual shift in colors at a given space.
So knowing that radiance at a specific point You could take a look at that as being radiance, whereas it's being contained inside of a field.
Got it. Okay. So you said something a minute ago about generative AI, using generative AI to create robots. radiance fields.
And you may have used a term that slipped my mind, forgive me.
Is it the same basic principle as, you know, doing a text-to-image using a text-to-image model, you know, ChatGPT or DALI, stable diffusion, whatever it is.
Is it that same principle that I enter a text prompt and then the system can create A radiance field or is it creating a series of images or can you get more, you know, kind of more control than that over it?
Yeah, so exactly. What it does is it will create a series of images of a singular object And actually, in some cases, they're starting to release some papers where they're creating multiple objects.
But each object itself is either a nerf or, say, a Gaussian splatting.
And from that, that's what's actually being used to train the actual resulting 3D model.
But they're able to do that in a fraction of a second.
Right. So earlier this week, you hosted a session at GTC.
I unfortunately wasn't able to make it. I wasn't in town just yet.
We were talking about it before we hit record.
And you spoke to, if I've got this right, some of the artistic implications, possibilities of around using nerfs and then also some more kind of business enterprise oriented applications. how did that go?
What kinds of things did you talk about?
And then I'm kind of curious what the audience reaction was either that night or kind of more generally, you know, what, um, As people learn about nerves and Gaussian splatting, what kind of the reaction is and does it spark imagination and sort of what are some of the implications?
Yeah, I was surprised by how many people actually came out to attend.
Yeah, great. It seemed like there was an extreme amount of interest in terms of just visualizing itself.
So I had to roughly... 20-minute video of just looping different examples of ratings fields that I've created and some of the people in the community have as well.
And so I think there was a lot of interest across a wide variety of industries where I spoke to professors, I spoke to people working on offshore. drilling sites.
I spoke to physicians, people of really diverse backgrounds and use cases, but I think all of them will be affected by radiance fields.
And is the interest in some of those, because I want to ask you about sort of artistic creative implications as well, but we'll put a pin in that for a second.
Is the interest from, say, a physician or the offshore drilling site makes me think of... use cases of robots and drones, you know, to be able to go to places more safely than sending a human to inspect something.
And is it along those lines of being able to create a radiance field and then – from a, you know, quote, safe environment, be able to inspect different aspects of the offshore site from different angles?
Is it that kind of thing or is it something totally different?
Yeah, no, it actually is quite similar, where if you have a predetermined camera path or you give the necessary information to the model, it will be able to create a hyper-realistic view of what it sees and from that, you can then flag for a human if they need to go and make a visit for actual maintenance or repairs.
And so you're able to really give a hyper-realistic look for that specific use case for asynchronous maintenance.
Right, right. Are there implications with VR and extended reality and augmented reality?
Yes, yes. And so that was actually one of the demos that we were showing.
VR applications because they still retain their view-dependent effects when you're in VR.
And so, you know, as you move around a scene, we as humans expect light to behave in a certain way.
And with this, that continues to hold true.
And with Radiance Fields, you can actually walk through the entire scene.
And so it is, I think, the closest thing to actually stepping back into a moment in time that we have.
Yeah, amazing. I'm speaking with Michael Rubloff.
Michael is the founder and managing editor of RadianceFields.com, a website that's covering the progression of Radiance Fields-based technologies.
And we've been talking about them, about neural radiance fields, NERFs, and 3D gaussian splatting, these techniques that allow us to stitch together 2D images and create... hyper-realistic 3D model, 3D environment that we can do all these different things with.
I mentioned, wanted to ask you about some of the artistic implications, and this might be a weird leaping off point, so redirect me if it is, but...
Recently, we recorded a podcast with a woman from a company called Kubrick's. and I'm gonna get this wrong, forgive me, but it's basically kind of like a digital soundstage, sort of an advanced digital green screen type of thing that you can use in filmmaking, video making and, you know, powered by generative AI, kind of similar things.
And I remember asking her, what advice do you have for burgeoning filmmakers who are interested in creating films but wondering how to go about it in this age with all these AI tools now becoming available and advancing so quickly.
And her answer was not what I expected, but it was really interesting.
She said, well, the first thing you should do when you're thinking about using generative AI in filmmaking is really delve into your own subconscious.
And if I understood her correctly, I think she was talking about, you know, the capabilities of what types of images and moving images you can create with generative AI tools at your disposal go well beyond what you could create without them.
And you're not limited to capturing reality, so to speak.
You know, you can create reality, people have been able to do with technology for a while now, but easier, faster, perhaps better.
What do radiance fields do or what do you think that they are doing and could do for creative applications?
Yeah, I see a very large creative opportunity for Radiance Fields going forwards.
I think that... They allow people to take larger risks or be able to actually – I actually wouldn't even call them risks because –
What you can do with them is if you get captured, you can film in post.
You don't actually need to film up front.
You just need to capture it. And I think that that allows a lot more thinking about where you're not being constrained for time. to think like, what are the actual camera movements that we want, or what's the best way to actually tell this story?
And you have the ability to align and say, if you are a director, you can go to your director of photography and say, here's the exact camera movement I'm trying to convey in this, or here's the exact thing I'm trying to show.
And I think that that's going to really supercharge productions in the same way.
I think that it's also going to allow more stories to be told in places that we've never been.
Because, you know, we can be transported to these places and be exposed to an actual lifelike interpretation. of uh wherever you know you want to take an audience and and The same way, I think for students and for independent filmmakers, it represents such a massive opportunity. because you will have the ability to go to locations and tell stories in locations that you may have always dreamed of, you know, shutting down the Las Vegas Strip, for instance.
But now you're actually going to be able to do Right, right.
Where are we at in terms of the technology when it comes to the resolution of images and particularly in the backgrounds? of the images?
Are we just constrained by, you know, the quality of your camera and the available compute?
Or is it more intricate than that? No, I think that where we are right now, it's fascinating.
There's a meta reality labs paper that released late last year called VR Nerf.
And essentially what they did was they created this camera rig, effectively known as the Eiffel Tower. where it has 22, I think, Sony A9 cameras strapped onto it all facing different directions.
And from that, what they would do is they'd go into a room and then they'd push it through the room.
And each camera would take nine bracketed exposure shots. across, you know, different exposures.
And then each one of those images would be compiled into a single HDR image, and then that image would be trained in the nerf.
Okay. And... The resulting image quality of that, you know, it approaches the IMAX level quality.
Wow. to be reconstructed. It's not a bottleneck in terms of the visual fidelity.
It's able to handle that. It's more of a compute issue right now, but obviously as time goes on, we're going to get more and more efficient computers. computers as well and so it's it's more a proof of concept to me than anything is that you know when we get to that level that's the floor Yeah, right, right, right.
Beyond capturing moments in time for your own use, are you doing other things yourself with the technology right now?
Yeah, so I've been doing some consulting work for businesses that want to implement the Radiance Field-based technology into their offerings, and I can't. talk too much about the work, but one of the ones I'm really excited to be working on is Shutterstock. where it's just, you know, how do we create these assets that are available for people to actually use today?
Right, right. Amazing. So kind of out in the mainstream world, so to speak, right, in the pop culture world, I would imagine...
That we've probably seen nerfs in action and just didn't realize it or didn't know how to name it.
Is that off base or are there some examples of things floating out there that listeners might have seen?
Yeah, there have actually been some pretty high profile uses of radiance fields where earlier this year, the Phoenix Suns, actually the entire team was nerfed.
And so there are nerfs of Kevin Durant, Devin Booker. the entire team, and they actually are using it as part of their introductory video for this season.
Oh, like during the games and starting lineups?
Yes. Yeah. Oh, cool. Yes. And it really like showcases a way that, you know, if you as a business can help bring fans closer to the action because, you know, You have these lifelike interpretations that are doing the most insane camera movements that, you know... And so is it like...
Katie's going up for a shot and the camera sort of seems to, and I don't know if the motion stops or not, but the camera sort of stops and then tracks sort of around him from a different angle, like that kind of thing.
Yeah, that's actually extremely close to one of the examples where he's about to, you know, dunk and it kind of flies up around him and circles around, which would be, you know, very difficult to go ahead and create.
But, you know, Riden's fields make that actually very, you know, surprisingly easy.
Yeah, amazing. Yeah. And there's also been a lot of other high profile use cases within the music industry.
So Zayn Malik, for instance, has a music video for Love Like This, which I think got something like 10 million views in the first 24 hours, which is just... insane.
There's a music video for R.L. Grimes' Pour Your Heart Out, which is actually comprised of over 700 individual nerfs.
Oh, wow. It is just an insane endeavor.
And every single shot in that music video is a nerf.
There's also Usher for his most recent single, I think it's called Ruin, has a few different examples of Gaussian spotting in it.
And then Polo G has a song that just released called Saris and Ferraris that also are using Gaussian splatting.
And then there's also Chris Brown has one as well and J. Cole and Drake put out a music video I think within the last like week or two and there is a Nerf hidden in there nice and I don't know if this counts or not, but in Jensen's most recent keynote, there actually is a nerf in the background in one of his... on the slides.
All right. Should we make it a contest for listeners to spot it, or do you want to call it out so people can go look?
If you guys want to pause it and see if you guys can spot it from the two hours.
This is great. You can go queue up your YouTube playlist with all the music videos you just mentioned.
Yes. And then, you know, rewind the podcast, listen back, you're spotting the nerfs.
Yes. And then when you're done with the pod, go back, rewatch Jensen's keynote and see if you can find it.
Yes, but it's in this slide where I think he's taking a look at the different modalities in which NVIDIA operates.
And it's under the 3D tab. And it's like a coastal cliff view.
And it was actually taken by one of my good friends, Jonathan Stevens.
And so it was really cool. That was the first time I think that we've seen nerfs being featured in the keynote.
Right. Excellent. Shout out to Jonathan.
Michael, closing thoughts. Nerfs, radiance fields, 3D Gaussian splatting.
For the listener who, you know, never heard of this stuff before, listen to our conversation, you know, ring some bells, spark some ideas in their head. thinking about going out and exploring some of the music videos we talked about. do you think this is going to go over the next, whatever the time period is, couple of years, 10 years, generations?
Are all of our photos going to become...
3D multi-perspective models going forward.
Are we forever living in a world of 2D images and then we figure out ways to make them more like 3D models.
Where's the future of imaging and sort of post-processing headed?
Yeah, it's a great question. And my opinion is that we now have the ability to no longer be constrained to 2D.
Yeah. And 2D is not actually how we experience our lives.
And to me, I feel like it should not be the final frontier of imaging. and now that we have the technology to do so I think it's really time to begin exploring how we can actually document our lives in a lifelike way to the way that we actually experience life.
Because not only can you create static nerfs or Gaussian splatting files, but you can also create dynamic versions of them too.
And so if you can imagine an analogy of static nerfs to photos, you can also do the same with videos.
And so I think that we're really entering into an age where imaging – is not the same as it has been for the last, since the inception of photography, you know, obviously it progressed significantly, but I,
I think that now the technology is there where we can just take a fundamental leap forwards into an entirely new dimension.
Come to GTC and you leave in a new dimension.
That's how it works. Michael? Michael Rubloff, thank you so much for stopping by the podcast.
Your website, again, is called radiancefields.com.
Are there other places for people who want to follow your work, learn more about the space you direct them to go?
Other websites, social media accounts, anything?
Yeah, so all my social media handles are just at Radiance Fields.
And so that's just generally LinkedIn, Twitter. are the two big ones that I mainly post on.
But yeah, I would just encourage all listeners just to try downloading some of the platforms themselves and some of the really good ones to get started.
You can take a look at Luma.ai. Polycam.
If you're on Windows, you can download PostShot or Nerf Studio.
They're all free right now. It's not as bad as you would imagine to actually capture everything.
It's actually quite straightforward and is pretty forgiving, so...
Yeah, give it a try. If you can take a picture, you can make an ARF.
Exactly. Excellent. Thank you again. Pleasure talking to you.
Thank you so much. Thank you. Thank you.