Hello and welcome to the NVIDIA AI podcast.
I'm your host, Noah Kravitz. Data is the fuel that makes artificial intelligence go.
Training, machine learning, and AI systems requires data, and the quality of datasets has a big impact on the results you get out of your systems.
Compiling quality real-world data for AI and ML use can be difficult and expensive.
That's where synthetic data comes in. Our guest today is Dr. Nathan Kuntz, founder and CEO of Rendered AI, a platform as a service for creating synthetic data to train AI models.
Nathan is a physicist by training, holds a PhD from Duke University, and previously founded Chimeta, a hybrid satellite cellular network company.
So let's get into it. Nathan, welcome and thanks so much for joining the NVIDIA AI podcast.
Yeah, it's super exciting to be here, Noah, and thanks for having me.
Oh, thank you. And extra thanks for dialing in live on location from a conference down in Southern California.
The background noise, listeners, if you hear it, it's just added excitement to the whole here at the thick of it, making the future happen.
So we appreciate you finding the time to join us.
It's my pleasure. I'm here at the Esri User Conference.
It's exciting. There's 14,000 people here trying to figure out how to use how we understand the world to make it a better place.
It's fun to see a place for synthetic data in that process.
Absolutely. So let's get into that then.
Maybe you can start by telling the audience about your company Rendered and as part of that, about synthetic data, what it is and how it's being used out in the field these days.
Yeah, absolutely. So maybe you sort of spoke to the problem we're traveling to solve, but maybe if I could... unzip that just a little bit and get into it.
That'll help contextualize what we do with synthetic data.
So go for it. So you're spot on that the data that we have to train AI really is the dominant factor at this point in the performance.
A lot of the models being used, and even the MLOps tools for doing things like hyperparameter tuning, have started to standardize And so what's left is, hey, do I have access to the best possible data to train algorithms and then also do things like test them and understand their performance?
And there's a number of barriers to getting access to that data.
And I think you described... data collection as one, and that's definitely true, you know, but we also get questions like, well, hey, but you know, social media i got pictures coming in all the time i got massive we're all we're all drowning in data surely Surely there's enough.
And so there's really three major themes that prevent people from getting what they need.
The first is just rare events and edge cases.
It just turns out that a lot of times the things that are most important to find or see or respond to are things that are rare.
You know, you don't see many flipped over cars on the freeway, but you better know what to do if that happens, right?
Yeah. You don't want to go around collecting that data.
It's a bad idea. But it's the same in other kind of maybe seemingly more mundane circumstances.
If you're running a Six Sigma manufacturing plant, the errors in that manufacturing line really crucial to be able to train computer vision systems to detect them, but very, very rare.
The second is that we often think about, especially in computer vision, we think about RGB imagery and sort of cell phone images.
And that has left behind a wide variety of the sensors that we actually use in industrial and high-value applications.
So things like infrared imagery, radar imagery, these things are very difficult to annotate, right?
Even if you have lots of data, you can't just farm it out to people and they could say, oh, you know, that's a cat, that's a dog.
It's much harder to determine that. And then once you start doing that in things like satellite imagery, that becomes even more difficult.
The final, and I think maybe the most interesting and most profound one is that What we do with AI today is all predicated on the use of cameras that were designed for a different purpose.
And so if you flip that whole engineering question on its head and you start saying, hey, What camera do I need to design to solve a certain AI problem or a certain detection problem? you suddenly realize you couldn't possibly have that data set.
The sensor doesn't exist yet. You can't even effectively pose that question effectively.
Where synthetic data comes in is starting to produce those data sets.
In our case, really focused on based on a physics-based simulation, so that we can address all of those problem areas and not just improve performance scores for AI, but really drive this whole process of engineering and business discovery, delivery, quality assurance for AI, which I think is going to be crucial for driving that in future.
So there are a bunch of ways we could go with this, but I think they're all going to get us to the same place at the end.
I'll pick one if this isn't the best place to keep going.
Pick a different one. I trust you. What's kind of, you mentioned a couple of them, but is there a range of industries that Rendered works with And that's kind of getting into a larger question.
I'm interested in what types of applications and industries synthetic data is a good fit for.
And then where synthetic data isn't a good fit, at least not yet.
Yeah, no. Okay. So that's a good question.
So where is synthetic data useful and not?
Where are we seeing the market? And then where does Render play?
There's a kind of few different questions packed in there.
So let's start with the market. There certainly has been a lot of use of synthetic data around autonomous vehicles.
And what we're seeing now is that there's a desire to leverage that kind of capability across a wide variety of other markets.
We see very strong demand in geospatial.
So people launching satellites using satellite imagery Along the entire value stream, we've seen a lot of demand for synthetic data in industrial inspection and manufacturing.
We see it showing up pretty strongly in security images and trying to identify behaviors and insurance companies starting to do similar, wanting to be able to monitor assets in order to provide insurance products.
And then I think one of the most interesting kind of long-term applications is actually in medical.
And that's a super early market for synthetic data, but there are... early signs that actually that ability to produce datasets that represent edge cases, diseases, and then also crucially remove things like bias. in datasets and performance.
There's a lot of interest there. Those are areas where we've seen strong market pool.
From a company standpoint, we've really tried to emphasize our platform as being agnostic.
The product that we sell is not specific to any one market.
It really tries to abstract away the core things that you need no matter what you're doing.
In synthetic data, you need a content management solution almost always, right?
You need some ability to deploy en masse information. in the cloud, you need to be able to task that system efficiently and run experiments if you're a data scientist with lots of different diversity.
You need things like data provenance that you have.
100 data sets or more, you can still figure out what you did in each of them.
And then post-processed. That's our product, and it's not specific to one.
Industry or application area. But then on top of that, we've been building out content in some of these different verticals.
And we have had a focus there on satellite imagery.
Just because that's where we see really strong demand and limited access to other options today.
Let's say that I'm a potential customer or newly signed up customer working in the satellite geospatial imagery field broadly.
How do I use the platform? When I log in, am I prompted to... And I'm oversimplifying here, but I can sort of...
Put in a query saying, you know, I need synthetic data to train my satellite system for XYZ. and it's spun up in kind of a bespoke matter, or how does that process work?
Yeah, the reality is that most customers have sort of their own specific needs.
So we don't try to provide one synthetic data application that meets everybody's needs.
But we have underlying simulation tools and integrations that cover basically all of the different imaging modalities used.
So whether that's kind of RGB and atmospheric modeling or hyperspectral imaging, synthetic aperture radar, or infrared. point solutions for each of those.
And what we do with customers when they sign on is we try to provide them examples of the types of data that they're most interested in, And then usually we work with them to either extend those applications to address their specific need,
Where we can do that or they can do that, it's an open platform.
So they can build on top of things we've built and extend them.
Or we put together a program where we'll deliver the content, some specific content that they might need to solve their problem.
Right. And the big question is, how does the process work?
How is synthetic data created? But maybe the thing I'm trying to get at, I'm interested in sort of the limitations of...
I don't know, not so much how Rendered, I mean, Rendered's platform and what Rendered's platform you know it to be, you know, good at and useful for and then things that just aren't kind of within the purview right now for whatever reason.
But then more broadly, synthetic data as a tool, what are some of the things that it's better at right now and then some of the use cases or whatever the reasons are that it's just not suited for.
Yeah, absolutely. So let's see if I can unpack that a little bit, as you say.
We focus on physics-based synthetic data.
The implication in saying that is that we have some way to introduce new information to the system through our understanding of physical processes.
And so where we shine are those systems where that's clear.
So like we have, We're really good at modeling light interacting.
We're really good at modeling radar interactions.
There are other systems where they're not physical processes.
If you think about Human network interactions, that's a different type of process.
You can model it stochastically. There's other ways to go about it.
But there you have to make different assumptions.
And so it's harder to introduce fundamentally new information For synthetic data.
And so what you see is actually that the industry has bifurcated somewhat where you have physics-based synthetic data being used for things like rare events, edge cases, stuff where you don't have access to any data and you can create it de novo. based on physics knowledge.
And then in other circumstances, you can use generative algorithms to sort of create more data of things that you've already seen.
And that can be helpful when you don't have a physical process to model, but unfortunately it has a limitation of you don't get that kind of rare event in edge case coverage.
Right. How do you know internally if the synthetic data you're creating is good, is good enough to be, you know?
Such an important question. I'm so glad that you asked that question.
Because... There's so much of this, like, ah, it's a really pretty picture.
It looks good. And you're like, right. But that has almost nothing to do with anything.
And so, like, let's start with, okay, what do we even mean by good, right?
Right. So good can mean... You know, it looks nice.
It looks real. Those are actually two different statements.
Good can mean it effectively mimics the sensor. that you're trying to model, but it's a margin of error.
Good could mean that we've effectively modeled our ground truth.
Remember, you have to figure out what you're simulating in the first place.
Right, right, right, right. You know, those are the ground truths.
You're simulating. Right. And so it's a complex question.
We have taken it upon ourselves to try to at least integrate what we think are the best of class tools. for answering some of these questions.
And a lot of times what we find people want to do is be able to answer that question by comparing data sets, by saying, hey, real data set, doesn't have everything I want in it, but I want to know if the synthetic data set is similar.
Yeah. So what we do is we use tools like UMAP analysis, which some of your listeners may be familiar with.
But what that does is it says, hey, here are the relevant features in this image, and I'm going to extract those, and then I'm going to essentially try to reduce the dimensionality of how we can present those so that we can see whether or not, for instance, synthetic data clusters differently than real data,
And expose that. And then when you see that, there's all sorts of things you can do.
One, you might say, oh, shoot. you know, synthetic data is not very good here.
It doesn't, it clusters completely separately.
This is a bad idea. It's not going to work.
Sometimes when that happens, you can overcome that problem through what's called domain adaptation.
We've had a lot of success with, for instance, Gambit's domain adaptation.
The other thing that can happen, this is really interesting, sometimes you'll see that part of the synthetic data doesn't overlap, so you'll have an area of overlap and an area that's not.
And then the user has to decide, well, is the area that's not matching to the real data, is it not doing that because...
It's not good synthetic data? Or is it because there's actually something in there that wasn't represented in the real data?
Right, right, right. And so that's why I think it's actually a really bad idea often to try to just say like, okay, here's your figure of merit.
But we do need to be able to quantify. the efficacy of synthetic data.
And that's a great way to do it. Ultimately, in our experience, We do a combination of that and also training detection algorithms on synthetic data and testing their efficacy.
And those two become this proxy for, is the synthetic data good?
As you were talking, I was thinking about things like physics engines for, you know, 3D modeling and gaming and animations.
And then as you were kind of getting into checking the efficacy of synthetic data and that whole question of what is good, I was thinking about some tools that have come out on the net for public use recently in the past month or two.
Dolly Mini and Mid Journey, you know, tools where you can input a text description of an image. and then the system will generate images.
In my head, and I tend to do this outside of the bounds of what actually makes sense to practitioners, so forgive me, right?
But in my head, I'm thinking, well, wait, are these things synthetic data?
Are these things – they're generating something that if I type in –
Phil Collins playing the drums on a school bus.
It looks sort of like that, but it's obviously not a picture that came from reality.
Is there any parallel towards these sorts of generate image generation systems that are kind of getting played with a little bit in the scientific, if you will, process of creating synthetic data, or are they two very separate domains?
No, so it's a really good question. There were two questions, I think, in there.
One is integrating other physics-based solvers into What does that look like?
And then what do we do with these other generators?
Yeah. You know, part of the problem with this podcast for the guests, especially, and probably the listeners, my apologies, but for me, it's great.
Because as I work out my understanding, I tend to ask seven or eight questions at once.
Well, I'm going to try to, I think that those are connected in a deep way.
And so I'm going to try to answer them together in a way that's as least confusing as possible.
Hopefully it shows some clarity here. Okay, so first of all, there's no single physics solver for something.
There are game engines out there. Omniverse has a tremendous rendering engine capability as well as some ground truth generating capabilities. you know, 3D modeling capabilities out there, lots of different physics solvers for electromagnetic, radar, and otherwise.
What we have tried to do with our platform is build a way that this can be integrated in with the things that sort of promote. a rendering engine to a synthetic data capability.
So the ability to deploy that en masse in the cloud without needing a a cloud infrastructure team, the ability to task that system, whether that be through API or visual interface, data provenance tools, the librarianship, you know, there's just a lot that takes you from, okay, I've got sort of like game engine too.
I can not only produce synthetic data, but use it effectively over time.
So then there's these other approaches for generating, let's say, imagery that are generative.
So not using—I use that term to distinguish from physics-based in the sense that they are using existing images to create more images.
Right. And in some ways, we use a generative tool in our platform today for doing domain adaptation.
We use CycleGAN tool. for doing exactly that.
And so we're using our previous knowledges of images and how sensors perform to modify some of our imagery.
And that's a natural part of even physics-based data.
And I think what you're going to find over time is that boundary between What are we physically simulating?
And what are we generating based on previous examples?
Just gets really turbulent. I think that's some of where we'll see synthetic data continue to grow because you want both sides.
Of course you want to learn from examples that you have when you have them.
And on the other side, you do need to be able to introduce new information, and that's where physics-based simulation comes in.
Right. And so here's some examples, right?
There are GAN-based tools, generative tools that will take semantic input and produce 3D models.
Yeah, right. Based on previous 3D models.
Well, there's no reason that you can't then integrate those 3D models into a physics-based regression.
Right. And then from there, you get an image out, but you might still want to take that image and pass it through a cycle GAN to do a match to a particular sensor output.
And then if you want to get even crazier, right, we're increasingly using AI in the physics-based simulations themselves, right?
So rather than solve something that's... based directly off of, let's say, Maxwell's equations.
We can start to actually use AI to make approximate physical solutions.
And so you have a few different places where you're sort of going in and out and toying with this boundary of...
Where am I using things that I know because I've seen them before?
And where am I using things that I know because I have physics equations that represent the behavior?
Our guest today is Dr. Nathan Kuntz. Nathan is the founder and CEO of Render.ai, a platform as a service for creating synthetic data train AI on it.
I want to switch gears for a second, Nathan, and talk a little bit about your own background and how you came to be a running this company that's going in and out of pushing the boundaries of what's aesthetic and what's generated and all of uh this really fascinating stuff that goes into synthetic data.
But let's go back. You have a background in physics.
You mentioned, Richard, at the top, you have a PhD from Duke in physics.
Have you always been interested in physics, or how did your journey kind of get started?
Yeah, I tell people sometimes to kind of get a PhD in physics, you need... two things.
You need to be deeply curious and you need to have a high pain tolerance.
If you do those things, you'll be fine. So I sometimes joke, right, I studied physics for a long time because I wanted to know how the world works.
And then later I started building businesses because I wanted to know how the world actually worked.
These are the rules that govern and this is how things actually get done.
Yeah. Exactly. Exactly. So. My first company was based on work that I did with my PhD.
It was based in the field of metamaterials.
So if you've heard of like the metamaterial cloak, and that work happening at Duke University.
I did early work there and then was really interested in how we could commercialize those techniques.
I was recruited to a place called Intellectual Ventures.
I got to hang out with people like Nathan Biervold, the former CTO of Microsoft.
Yeah, yeah. and even Bill Gates in building those concepts into Kaimeta, which was, as you said, satellite communications, but really based around new electronically scanning antenna technology.
That was my first baby. I was the CTO there for a couple of years and then CEO for four years. and brought its first products and networks to market.
I left there in 2018, did some venture work trying to actually a lot of a lot of the work I was doing was people that were coming out of grad school starting companies and trying to say, hey.
We don't all have to have the same scars. a few ways where you might be able to avoid it.
Like that's hot. Don't touch that. But this is a good idea.
The pain metaphors are building up here.
I love it. But I had, having come out of the satellite industry, I had all these friends who were working on launching constellations and all had this sort of foundational business problem where they would have a sensor that really ought to be good at detecting X.
But to go from that to building a business, they need to be able to prove to a wide variety of customers Not only that the data was capable of that, but that they could build an analytics product on top of it that would make it easy for people to get to those insights.
And they needed to do that well enough that they could then get the investment community behind them so they could actually launch the satellites.
And so I was seeing that that problem that is the sort of next generation systems engineering problem happen in real time and saying, you know, this is not just a satellite problem.
This is going to be everything in AI facing this problem.
And so I actually wrote a white paper on what we would need from an infrastructure layer to address it.
What are the pieces that are going to be reusable?
How would you actually make that happen?
I shared that with some of my friends. They convinced me to start a company.
And how long ago was that? How long has Rendered been around?
We're coming up on three years now. I guess three years in August.
Yeah. So we spend a lot of time really trying to understand that workflow.
Because I actually don't come from a data science background.
I come from a physics background. And so I really want to understand not just how we make pretty pictures, or physics-based simulations, but how we made that relevant to this new community and set of needs.
And so we did that exploration first and then have spent a lot of time building the platform, using it with some select customers.
But then we didn't launch the platform publicly until February. of this year.
Just a few months ago, yeah. If I asked you to name, you know, name one, Biggest surprise in the three years since Rendered has been a thing, at least in your mind, or your friends convinced you to make it a thing?
Something along the way that's really kind of, as you look back, surprised you about doing this?
Boy, that's a good question. What has been the biggest surprise?
Here's one of the things that's been most interesting.
When you talk about synthetic data, you're usually starting by talking to someone who's interested in sort of improving their AP scores.
So data scientists are saying, hey, I want to improve this figure of merit.
And what's been surprising is that we...
We get a very positive response from them, but you get an even stronger response from the business development teams at those companies.
Because they're saying, hey, I need to know whether or not I can sell this.
Does that make sense? And I don't have... the ability to prove whether or not that's possible today.
And so what has been surprising is I think we started with And so I'd continue to have, where I got developer focus and engineer focus to our product.
But it's been surprising how big the demand has been from that business development sales side of the house, which needs synthetic data more than you might think. yeah and so looking ahead to the next you know next three years of renders life or five years or whatever time frame kind of makes sense Another two-part question for you, but you're used to these, but no.
And pick either one first. Where's Rendered headed?
Are there specific things that you feel comfortable talking about publicly? specific things you're focusing on, some milestones you're trying to get to, that sort of thing.
And then... broader base to your point about how, oh, it's not just, you know, folks who are putting sensors on satellite guidance systems who need synthetic data.
It's actually everybody who needs it. Where do you see, and the field sounds like such a big... ambiguous word to use in this context.
But where do you see the role of synthetic data, the field? headed in that same, you know, three years, five years, whatever that near-term timeframe is.
Okay, so good question. So rendered AI from the beginning set out to be the infrastructure layer for synthetic data.
Right. And that continues to be our North Star, our guiding direction.
The reality is the market today is very early.
And so customers need support that is domain specific, right?
They don't usually. come with their own simulation tools.
They don't come with their own synthetic data engineers that are sort of It's ready to pick up the infrastructure for it.
And so we offer a lot. to help them with that today.
Where I think the market is going to head is it's going to be much more common Any company that's using AI, you're going to step in and it's going to be like, this is our synthetic data engineering team because they'll have to.
In the sense that you don't step into any product company and then expect them not to be using CAD.
You have to. You have to have that ability to do physics.
Modeling, you have to have that ability to integrate this domain-specific knowledge that your company builds up into these engineering tools.
And so I think that we'll see a general trend that way in the coming years as it becomes as synthetic data becomes more commonly used.
And ultimately, I think we're sort of already seeing there's a sort of fuzzy edges between what is synthetic data?
What is a digital twin? And what is just, you know, a Monte Carlo analysis or kind of stochastic simulation.
And I think that what has started to happen is that there's just so many circumstances now where It's not good enough to simulate something once or how it could happen in a particular way.
You have to simulate. wide varieties of things because inevitably you're using them in complex environments often that use AI. that all sort of moves you in that direction from what might be called simulation alone to synthetic data where you're doing large batches and large amounts of diverse compute.
This is one of those conversations that really makes me, it lands on a note of like, all right, we're going to have to check back in.
Yeah. Get an update on where this is all headed because it feels like, you know, the opening lines of an epic story here.
I think we'll, yeah, we'll see a lot from the synthetic data industry.
And I think you'll see people from all sorts of different fields sort of move in the direction of like, you know, we're playing synthetic data.
In the meantime, for folks who got a taste and want to learn more, rendered.ai is the company website.
Are there other places to go, a company blog, your own personal channels, social channels, white papers, wherever it is? where folks who want to find out more about synthetic data and specifically about your company can go online.
Yeah, absolutely. So, you know, the website render.ai is a great place to start.
You can sign up for a trial account there if you want to just play with the tool.
We've got a little literally dropping toys in toy boxes example application.
Nice. want to play with, encourage you to do that.
You can go to support.render.ai if you want to sort of learn more about the nitty gritty and what's in the background.
And then we have actually our developer framework is open sourced and up on GitHub.
And you can find links to that on the support pages as well.
So encourage anybody who's interested to go and check that stuff out.
And if you... If you have any questions, you know, just reach out.
We're constantly monitoring those people who send us inbounds through the website.
And then we're also pretty active on social.
I guess most active on LinkedIn just because professional community types of folks that we work with.
But we're also out there on Twitter and other things as well.
Excellent. Well, Nathan, thank you. This has been fascinating.
And the best possible way leaves me with many more multi-part questions than I came in with, which is great.
It's what we want to do. And thanks for taking the time, finding a space at the conference to talk with us and we'll let you get back to it.
But, um, Enjoy the conference. Best of luck there.
And obviously, best of luck with everything that Rendered is doing.
Thank you very much. And thanks again for having me on the show.
Thank you. Thank you.