Hello and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. My guest today is Naveen Rao, the CEO and co-founder of MosaicML.
MosaicML is part of NVIDIA's Inception program, which helps startups grow by giving them access to cutting edge tech and NVIDIA's expertise.
MosaicML is on a mission to help the AI community improve prediction accuracy, lower costs, and save time.
They're doing it by providing tools that make it easy to train and deploy big AI models using your own data in a secure environment.
Naveen is here to tell us about why he founded Mosaic ML and why the company's work is vital to the future of machine learning.
And I think we've also caught him on the heels of some product announcements.
So we'll have some new stuff to talk about as well.
Let's get right into it. Naveen, welcome, and thank you so much for joining the NVIDIA AI Podcast.
Noah, thanks for having me on. I'm really excited to be here.
Great. So let's get into it. Maybe we can start by you telling the audience a bit about Mosaic ML, why you started the company, what the company does.
And then, as I mentioned, you've got some news, so we can get into that as well.
Yeah, so before the news, I think one of the things we saw very early on, and when I say very early on, this space moves so fast, it was only a little over two years ago. was we believe the capabilities of large language models, large models in general, we're going to explode and we're really going to bring new capabilities, new experiences to humans.
It's a new tool. However, if we want to be responsible about that, we found that it's not a great world where only a few have access to these methods.
We don't. think anything's bad about people investing a lot and building the best possible models.
I think that's actually great for the research community.
It's great for pushing things forward. But the reality is these new technologies have the possibility of concentrating power in ways like we couldn't do before.
Every technology transition does that. And the best way to mitigate harm in those scenarios is really to distribute those capabilities.
It's actually very similar to how software development kind of blew up in the 80s when the PC became cheap and everyone could learn to code.
I learned to code. when I was a kid during that time.
And I think we're seeing something similar now.
We started this company really to bring accessibility to the latest state of the art methods to a greater number of organizations.
What were the big barriers? The big barriers are really difficulty.
When you want to wield 500 GPUs and make them train a big model, it's really hard, right?
There's a lot of software. There are things that go wrong.
And you need packages to make this work well.
And that just didn't exist in the tooling.
The other part of it is cost. Really, it's a very deep question of how do I... make a model train to the same performance with the least number of computing flops and the lowest cost in time.
And really, that's where we spent a lot of our energy at the beginning.
And, you know, we've actually shown we've become the sort of active standard for showing costs.
GPT-3 models for less than 400K, diffusion models for less than 50K.
We publish lots of detailed blogs on this.
And I think... This has been taken up by the community and I'm really happy that we're able to enable a great number of people with these technologies.
So what are your products? What are the services that you offer to the community?
Who's your customer? How would they go about using them?
Yeah, great question. So our customers are ML engineers.
You can think of our tools as supercharging their capabilities.
And we can do it within the confines of their cloud tenancy.
So everything stays totally private. and we never see data.
So we really believe in data privacy. Data privacy is actually, I think, intrinsic to creating the world we want. within AI, IP ownership, data privacy.
So we sell to the ML engineer. We have a few different products that we sell to them.
One is an inference API. it's similar to OpenAI or any of the other inference APIs out there.
However, the models that are served by it are curated set, curated by us of open source models.
Those models, as they improve, we add the newest ones and people get access to it.
So we kind of move at the speed of community there.
That's actually something we just launched this week.
We launched it on Wednesday. All right.
Congratulations. Thank you. And that same service is actually available in the private scenario I described within the cloud tenancy of a customer on their own custom models.
So they can take open models, they can customize them for their applications or even train their own models from scratch and we can serve them.
Then one step back from that is fine tuning and training.
We actually have a managed service for GPUs, again, within any cloud.
We can run on all the major clouds and some of the new smaller clouds that are coming out. manage the whole process soup to nuts, you start with your data, we actually have templates for models that they can start from either pre trained or completely naive. and customers have a choice and we make that process go as smoothly as possible for them.
So for those who are maybe a little less familiar with the details about...
They've been following the AI space and obviously over the past six months in particular, There's been an explosion in the mainstream and the idea of a large language model is a little more familiar to a lot of people.
But when you're talking about a customer, a business customer, an ML engineer, who wants to train their own model versus using, you know, open AI or one of the others that's, that's out there, uh, for the public and for customers to use.
Why would they do that? And what goes into that?
And, you know, you touched upon the costs of training a big model and keeping it running, maintaining it.
Why would a business customer choose to train their own model as opposed to using one of the publicly available ones.
Yeah. I mean, I think the real world is always a mix of different solutions and there are scenarios where, you know, you want to just use something that's out there.
I think if you want control over the model behavior, like really tight control, you want to have your own model.
It's like insourcing software development.
That's one aspect. The other aspect is data ownership. train a generative model on a data set, it can memorize parts of that data set.
That's just a feature, actually, of the models.
It's a good thing. What it means is that wherever those weights of the models go, the data has gone.
So you kind of need to control where the weights go.
And that only happens when you own the model.
And so really... Respecting data privacy, if you believe that that's intrinsic part of this new AI world, then you do need to think about your critical data and where it goes within a model and how that model is served to the end point.
And so we give... our users control over that process.
And also just quick iteration. If you want to incorporate new data, You have some sort of a data pipeline that's built for your application.
Again, you need to control how the model behaves.
So what are some of the things that MosaicML offers in comparison to some of your competitors? some of the cool perks, maybe some of the interesting use cases that Mosaic ML offers.
Some of the perks, I mean, we, like I said, move at the pace of the community as new things come out there.
We attempt to make it as economically feasible as possible and easy to use. we work across different sorts of models, like diffusion models and large language models.
We actually trained from scratch our own diffusion model.
Like I said, it costs about $50,000 in compute costs, which is pretty accessible, I think.
I mean, still it's... No, but it's a five-figure number, not a six, seven, or eight-figure number.
Exactly. That was a big thing for us, trying to get it under 100K.
Yeah. And I think what that enables is people to try to customize these models for their application.
I think that that's really what we offer is sort of that we give you clay.
And you can choose to make a teacup, you can choose to make, you know, whatever little widget you want.
I think that's what's special about us. We're not end application company, we're a platform company.
So we provide tools to enable a whole host of applications within different spaces.
How long has the company been around? When was it founded?
It was founded in January of 2021. So just over two years.
Just over two years. So during that two years, what were some of the... challenges or even just interesting developmental milestones, technical problems, things that you had to figure out solutions for.
Oh boy, so many. How long do we have, right?
Yeah, I know, right? I think a big challenge before was how do we go to market?
How do we sell this stuff? How do we capture value that we're creating? with our pricing and align our pricing with our customers.
Because I think the best business models are the ones where people pay more because it aligns to their value, right?
And so, and it's our value prop to them.
This is not simple, I think. we experimented with a bunch of things that were really complicated as every startup does.
And we've, We've settled on a few simpler ones.
What's really interesting to me is that when you go through this process and you iterate a few times, you actually come to the realization that Just about every AI application is kind of reselling compute in some form.
Right, right. an inference api is reselling a gpu time training is reselling gpu time and it's really how do you add value to that process and price align to it and Really, that was the biggest learning.
I think how we make things work in different clouds and actually the variability between different clouds, they're using, these are all NVIDIA A100 GPUs. basic GPU computing element, but how that's deployed and the software wears around is actually quite different.
So we had to engineer solutions to make that each cloud performance, even though they're still running that core or GPU.
How big is the team now? 55 people, something like that.
And are you in growth mode? Are you kind of in... you know, go to market and handle customers and kind of see how that's happening.
What's the stage of the company? Yeah, we're very much growing our customer base.
I mean, we went to market a few months ago.
So we're at the beginning of the year, end of the last year.
And yeah, it's kind of exploded, which you might have imagined.
Sure. Yeah. Wonderful. I mean, honestly, the biggest...
The biggest barrier to growth is us in terms of closing deals and servicing those deals.
So that's a good problem to have. And of course, we're trying to solve that.
Our stack is highly automated, which has made things really nice.
The dark side of large scale stuff is that GPUs fail sometimes.
Nodes fail, networks fail, and dealing with those failures gracefully and keeping the process going and hiding those details from the user who frankly doesn't care. right right right um that's actually a lot of the work that's that's gone in uh recently to make this whole process stable and actually this morning we released our open source models, our state of the art open source models Better than any other open source model out there.
Very long context. We believe it's a great starting point for customers to build off of.
And that's a culmination of all the work we've done. in terms of engineering our stack to be reliable.
The research work we've done in making long context work and stable training recipes work, all of this stuff.
And I think that's why we're very proud of the stuff we released this morning.
That's great. What models did you release?
We call it the Mosaic Pre-trained Transformers, so MPT family.
And we actually just released a 7 billion parameter base model with a few different variants.
One variant is tuned for chatting, so it's really easy and nice to chat with it.
We did put an interface up on Hugging Face that people can chat with.
It's a little overwhelmed right now. We do have an inference service that will be, we already launched it this week, but we're going to be moving our models to for higher scale very soon.
So people can get a better experience there.
We also have a instruction following, instruction fine-tuned version.
One of the more exciting versions, I think, is what we call the Storyteller.
That one is long context. What context is, is essentially the prompts, right?
Right. We actually fed it the entire contents of the Great Gatsby, 67,000 tokens.
And we asked it to write the epilogue and it did.
I posted it on Twitter. You posted it. Okay.
So I was going to ask what happens, but people, what's your Twitter handle?
Naveen G. Rao. Okay. So if you've been wondering all these years, what happens later?
David G. Rao on Twitter. We can get an answer for you.
You happy with the answer? I'll just ask that.
Yeah, it's actually kind of creative, right?
Sure. We didn't know it spit out, but it actually produced some interesting stuff.
In fact, I think we put some of it in our blog as well. when we release the model so you can check it out there too but uh it's just a fun use case right i mean i think this the way we see it is again we're bringing We're building clay, right?
We're bringing that clay to the market. And what people do with it, I'm excited to see. you know yeah new capabilities it's actually kind of similar to nvidia's roots it's like You bring computing that can do new stuff.
What does it do? Well, people unleash all kinds of creativity like AI itself.
Absolutely, yeah. So that makes me think something you said kind of at the beginning of our chat here, and it's kind of been a theme of the world, a theme of the advancement of technology. computing and ai certainly in recent years and we've talked about it on on this podcast over the years democratization of all of these tools and you know going from gpu's rate for ai people have gpus and their machines at home for gaming for other purposes oh i could use those to start messing around with ai and you had these people doing all these things in the past years, and now, obviously, you know, the explosion of... these LLMs into the wild, people having access, doing all kinds of stuff with them.
You said something at the beginning about it being important that power held in these models not get too concentrated. and that you know we we continue to uphold this democratization of access to ai as being important Can you talk a little bit more about that, kind of what that means to you and then what it means to how Mosaic ML goes about its daily business?
Yeah, so AI, the reason it's so, I don't know, concentrating in terms of its power is that it literally can become a lens on which we see, which we view all data.
You know, people talked a lot about bias in AI and models.
The reality is every model is biased. Every human is biased.
And that's actually not a bad thing. It's how we make sense of the world.
That's how AI makes sense of the world. Those biases, you know, I subscribe to a viewpoint that there's no absolute moral truth.
There are relative to what we feel as society, right?
Right. And what we found as humans is the best way to do that is through some kind of democratic or distributed process, right?
We vote for laws. We vote for people to represent us in our government.
And that represents a majority. Yes, there's always people that are unhappy about it, but I think AI is actually very similar in that you need many people putting their biases into models that provide their lens, their perspective.
And the market will decide what perspectives are useful, what are less useful, whatever, right?
But I think I don't want to create a world where there are no models that don't disagree with me.
I have a view on things. I have my own biases, of course, like we all do, but we need to enable everybody. even the people that disagree with us to create those lenses on data.
I think that's very important in core to what we do.
And second part of your question, how do we How do we go about doing this on a day to day?
I mean, really bringing the cost down, making things in the 100,000 range, that's a lot more accessible to enterprises.
I want to enable many enterprises with their own LLMs that are built for their purpose.
You know, ChatGPT, all the API vendors out there, they've What was a magical thing about it is that it just... It lit up everyone's imagination.
And my high school-aged kids... told me about it.
And I was like, you know, that's what I do.
And they're like, oh, that's pretty cool.
So finally they, right. That's a good moment.
It was a great moment. uh but so it drove this awareness but now i think when you want to build a an ai to help you with healthcare.
It's very different than entertaining someone by a chat.
If you want something that will advise you on your retirement plan, also very different.
If we want models that can model proteomics or genomics completely, not completely different architecture model, but different training.
So we need to move these capabilities into a number of different places.
And that's what we do every day is we actually talk to organizations we try to understand what their problems are and see how our tools can help solve them.
I'm speaking with Naveen Rao. Naveen is the CEO and co-founder of MosaicML.
Mosaic ML is NVIDIA's inception program, which works with startups, giving them access to NVIDIA tech and expertise and advisement.
And as we've been talking about, MosaicML just released, as we're recording this in early May, just released a 7 billion parameter open source model as part of a a family of open source models.
We've been talking about the different use cases for MosaicML's tech.
I want to shift gears for a second, Naveen, and Ask about you.
You mentioned learning to code as a kid in the 80s.
And I know you were at Intel for several years along your path here.
Now, how did you get into coding and tech and working in tech and eventually into machine learning?
Oh, so you're going way back. We don't have to go back, you know, further than you want to.
It's okay. I'll date myself a little bit.
So... My older brother and my dad are both big geeks.
My dad's always been a tinkerer. I think we just sort of got, I guess, blessed or cursed or whatever you want to call it with that same desire.
And, uh, My brother was really into programming in the 70s even.
And we actually had a personal computer when I was three years old in 1978.
The Texas Instruments 99-4A. Yeah, TI-99, yeah.
So I actually learned to code in elementary school.
I learned to code Logo. Right, right. I remember Logo with the turtle, yeah.
Yeah, you could drive the turtle around.
You could make it draw stuff. It was pretty cool.
And then quickly learned basic after that.
And really, it was almost like a game at that point.
It was something fun. I could automate stuff and people thought it was magic back then.
Like, oh, I can print something out over and over again.
Yeah. it's a little program like that i actually started to build games and stuff when i was a kid yeah and uh you know fast forward to college and you know i was a computer science and electrical engineering major loved building stuff with computers back then.
It was the very beginnings of the internet.
And really, one of the things I was fascinated by... kind of started from sci-fi was artificial intelligence.
The concept of a machine that is synthetically intelligent.
And I even did research in the 90s on neuromorphic machines, Carver Mead's work.
And, you know, really fascinated with this premise.
And the thing that really grabbed my attention was the fact that our brains run on 20 watts of energy.
The latest NVIDIA GPU runs on 700 watts for card or for chip.
So I think it's still something that's a big gap, but it was just mind blowing to me that like everything we are is 20 watts of energy.
Right. So I came out to Silicon Valley, was in the startup scene, all that kind of stuff, learned how to build chips, learned how to write software. really well as a professional, did it for 10 years.
And I was like, okay, maybe it's time to go back to that interest.
And I did this kind of, my family thought it was crazy because I had a very nice career And I quit my job and went back to school to get a PhD in computational neuroscience.
Okay. And that was really driven by this desire to...
I want to look back on my life at the end and think I did something.
Yeah, absolutely. I nodded and said, okay.
And then I thought, wait a second, computational neuroscience, let's unpack that a little What does that mean?
It means understanding the computational underpinnings of how the brain processes information.
So it's related to biophysics, it's related to physiology, behavior.
It's kind of the middle of all of these things.
And so it's really like, how does information come to the brain?
How is it represented? that has a process to affect output, which is movement. worked in a motor control lab where we did neural prosthetics literally decoding signals from neurons that were listened to in real time actually my lab was the first one to do uh human neural prosthetics, people that were or quadriplegic actually had an implant.
We could decode their thoughts around movement.
Right, right, yeah. With a robotic arm. Pretty cool.
How long ago was this? I started there in 2007.
Okay. Have you kept up with the field? I mean, to some degree, you know, it's like there's only so many hours in a day.
No, sure. Yeah. I asked because it was a few years ago now, but we did an episode with A startup working on similar things, they sort of described it as a USB port for the body.
The founders had a buddy who... needed a prosthetic leg, I believe it was.
And they kind of saw what he got and thought, oh, we can build something better than this, you know, and sort of went on with it.
So it's a That field and the idea of bridging kind of the body with AI, to use the term very broadly, you know, has always been fascinating to me.
So, oh, absolutely. I mean, really the premise and everyone in our lab kind of felt this way is that we can really crack how information is represented in the brain.
We can actually start to make communication faster. and more rich between humans and what's interesting is that when you actually start to analyze this is total tangent but we start to analyze how information comes out of a person, right?
We speak, we move, we move our faces. It's all motor output.
It's actually a lot of information. It's really hard to do better than it.
Right, right. Just pure bandwidth-wise, right?
Evolution has created something quite amazing with humans.
But after I finished that, I actually was a research scientist at Qualcomm researching back to neuromorphic architectures, how we can use some brain inspiration of computation and actually build better computers.
And that's when deep learning kind of started to take off.
I knew it from the academic field. I saw Jan Wacoon speak years ago at my school and everyone thought he was crazy. back then, but you know, bless him and the others for, for, for keeping at it because really they, they knew that there was more here.
And, uh, GPUs were getting dense enough and parallel computing was getting good enough that that's when the moment happened.
And 2012 was kind of that moment. And we recognized very early that we need a new computing architecture.
GPUs actually weren't it at that time. And my previous company was called Nirvana. which was acquired by Intel.
That's how I ended up at Intel, the first AI chip company. in this new spat spate of the chip companies.
And, uh, we were, we were competing with Nvidia and really, I think, I like to hope that we pushed NVIDIA towards architectures that we see today, which enable all the cool stuff we're seeing like large language models.
So to put you on the spot, since you said evolution, and we've been talking about the idea of the brain inspiring these computational models, Where do you think we're at?
It feels... to me is sort of an observer, but I've been, you know, observing, having the good fortune to talk to people like you for, for, more than half a decade now on this show and working in the industry longer than that.
But it really feels like we're in this just very intense, pivotal moment, or I don't know if pivotal is the right word, but the The pace is accelerating and it's hard to kind of know how much of it is because of the hype around Gen AI specifically and the chat models making this technology accessible to so many more people in a way that just brings it into the consciousness.
Where do you see the technology, the impact of the technology on the kind of greater world What's the moment that we're in?
Is it sort of the more of kind of the tail end of the past several years of the increases in compute and the democratization of tools?
Is it kind of more the beginning of this, you know, entirely new era in...
I mean, information and computing. Where do you kind of, if you sit and take stock of where we're at right now, what do you see?
I think we're at the very beginning. You know, it's like humanity just got wings.
There was something that has changed, right?
And it... That it's funny because like, I know we were living it and we were like, oh my God, you know, things are moving fast.
And like 10 years, it's taken 10 years to get to this point.
And so it feels slow, but it's like, the reality is, 10 years ago to now is actually just the blink of an eye.
And there's been a constant building of technologies upon each other that work.
I think that's the critical piece that people seem to miss when they're like, oh, it's all AI hype and it died, and now it's coming back. it's BS or whatever.
The reality is we hit upon something where we can build large scale optimization systems.
And now that's gotten better. And we've learned how to make them even better, more efficient.
Now, my company has really been focused on efficiency of computing flops. really, how do we get to that 20 watts, right?
That everybody can realize. And I think that's a largely a, it is a physical problem in terms of a physical substrate, silicon or whatever.
So there's a lot to be, refined there, but there's also a lot to be refined on how these learning systems work and how they scale.
Yeah. human brain uses each each computing flop it has extremely efficiently that's an algorithmic problem so i think we got to look across the whole stack still and We're just now figuring it out.
I mean, a large language model takes millions of watts to train.
It's nowhere close to capabilities of a human, even though now it starts getting, I don't know, sometimes indistinguishable.
So I think we're very much at the beginning.
So I wanted to ask you, you mentioned before we hit record here, that you recently published a blog about H100 performance.
Yeah, so the GPU has been something that's been intrinsic to the new capabilities we have.
We were actually very excited to see costs come down and time come down.
Mm-hmm. we were very keen to see the new stuff coming out from nvidia and you know and we benchmarked it and we actually found both of those things to be true about 3x faster And that means everything that takes three weeks will take one week.
Great. Yeah. And the other important thing is you can do it cheaper.
So we found, you know, 30 to 50% cheaper performance per dollar.
So even performance normalized and to cost normalized rather, we saw a pretty big gain here.
That's fantastic. cost savings. So again, back to the whole idea of accessibility, this is it.
Right. is moore's law is part of it new architectural features are part of it software stack is part of it so all of it coming together is really driving the price down to a point where it's gonna be just much more accessible by the community.
That's great. And so for Mosaic ML in particular, you mentioned in a go-to-market phase, customer acquisition, growing what you're doing, And you just dropped a bunch of exciting announcements.
So asking what's next feels a little, I don't mean to push you too far ahead, But as you look ahead to the rest of this year, whatever the timeframe is, what's the roadmap look like for your company?
Yeah, I think enabling people to get to the best possible model and the shortest amount of time and least money is kind of what we what our mission is, right?
And so We look at different parts of that stack.
We are optimizing inference. That's what we put out there and making that cheaper, you know, enabling more applications. we're optimizing the training process.
But what we found is a big part of optimizing training process is actually how you pick what data goes into your models.
This is something that we are looking forward on how to do really well, because ultimately it's going to start with a really like raw data.
Data is observation. How do we morph that into something that's suitable for learning and then feed it to a stack that's extremely efficient and then serve it?
Actually, if you saw the announcement, again, things moving very fast.
Last week from our customer, Replit. Replit is a distributed IDE company.
They built a state-of-the-art code completion model on our platform.
They started with their own data along with some open data. two people on their side using a Databricks pipeline, and then our tools built a state-of-the-art model in less than a week.
That's phenomenal. It's just amazing. Awesome.
Right. That is the story for us. Like we love it because it's like, that's what we want to do.
Right. Everybody could do this. Right. Wild.
Naveen, for folks who would like to learn more, obviously the company has a website.
Is there a research website? landing page or their social media accounts.
We mentioned the epilogue of the Great Gatsby.
But for folks who want to dig in more to the technical side of what you're doing, perhaps look into working with Mosaic ML, where should they go online to find out more?
Yeah, mosaicml.com is our main website. Our blogs actually have a lot of detail in what we do.
We are transparent in how we work. And all of the details are there.
So anything to do with our model release, our inference, releases, any other technology, our streaming data loader service, all of these things are all detailed in our blogs.
You can also take a look at some of the solution pages, but for a technical audience, the blogs are really where you want to go.
Perfect. Well, Naveen, it's been a pleasure.
Thank you for coming on and talking about what you're doing and kind of sharing Looking ahead to the big picture, even on the very close heels of, you know, some big announcements for your company as well.
Appreciate you. be in game to talk about the future in that way.
And we wish you all the best of luck with Mosaic ML and everything else you're doing.
Wonderful. Thanks for having me on. Really appreciate it.
Thank you.