Agents can do three things.
They can access your files, they can access the internet, and then now they can write custom code and execute it.
You should really only let an agent do two of those three things.
If you can access your files and you can write custom code, you don't want internet access, because that's one is a vulnerability, right?
If you have access to internet and your file system, you should know the full scope of what that agent's capable of doing.
Otherwise, malware can get injected or something that can happen.
And so that's a lot of what we've been thinking about is like you know, how do we both enable this because it's clearly the future, but then also, you know, what are these enforcement points that we can start to like protect.
All right, welcome to the Latent Space podcast in the Chroma Studio.
Welcome to all the guests here.
We're back with our guest host, Vibhu.
Welcome.
Good to have you back.
And our friends, Netter and Kyle from NVIDIA.
Welcome.
Yeah, thanks for having us.
Yeah, thank you.
Actually, I don't even know your titles.
Architect something of Dynamo.
Yeah, I'm one of the engineering leaders and architects of Dynamo.
And you're director of something developers.
Yeah.
You're the developers, developers, developers guy at Nvidia.
Open source, agent marketing, brev, and like dev tools and stuff.
Yeah.
And we're kind of recording this ahead of NVIDIA GTC, which is coming to town again or taking over town, which we'll all be at.
And we'll talk a little bit about your sessions and stuff.
Yeah, we're super excited for it.
One of my favorite memories for Natter, you always do marketing stunts.
And while you were Rev, you had this surfboard that you went down to GTC with.
Yeah.
And like, Nvidia apparently liked it so much that they bought you.
What was that like?
Yeah, our logo was a shocker.
We were always just kind of trying to keep true to who we were.
I think some of the startups, you're trying to pretend that you're a bigger, more mature company than you are.
And it was actually Evan Conrad from SF Compute who was just like you guys are like- Previous guest yeah.
Oh, really?
Amazing, yeah.
He was just like, guys, you're two dudes in a room.
Why are you pretending that you're not?
And so then we were like, okay, let's make the logo Ashoka.
We brought surfboards to our booth to GTC and the energy was great.
Some palm trees too.
They actually poked out over the walls so you could see the bread booth and no one else, just from very far away.
Oh, so you remember it back then?
Yeah, I remember it.
Pre-acquisition, I was like, oh, those guys are cool.
Dude, that makes sense.
So we signed up really last minute.
And so we had the last booth.
It was all the way in the corner.
And so I was worried that no one was going to come.
So that's why we had the palm trees.
We really came in with the surfboards.
We even had one of our investors bring her dog.
And then she was just walking the dog around to try to bring energy towards our booth.
Yeah, Steph.
Yeah, yeah, she's the best. you know, as a conference organizer, I love that, right?
Like.
It's like everyone who sponsors a conference comes does their booth.
They're like we are changing the future of AI or something.
Some generic bullshit.
And like, no, like actually try to stand out, make it fun, right?
And people still remember it after three years.
You know what's so funny?
I'll give you this clip if you wanna add it in.
But my wife was at the time fiance, she was in medical school and she came to help us because it was like a big moment for us, and so we bought this Cricut.
It's like a vinyl printer, because how else are we gonna label the surfboard?
So we got a surfboard, luckily was able to purchase that on the company card.
We got a cricket and it was just like fine tuning for enterprises or something like that that we put on the on the surfboard.
And it's 1 a.m. the day before we go to GTC.
She's helping me put these like vinyl stickers on.
And she goes, you son of a she's like, if you pull this off, you son of a bitch.
And so pretty much after the acquisition, I stitched that within the acquisition.
I sent it to our family group chat.
Yeah, yeah, yeah.
No.
Well, she made a good choice there.
Was that basically the origin story for Launchables?
Maybe we should explain what Breve is.
Yeah.
I mean, Breve is just a developer tool that makes it really easy to get a GPU.
So we connect a bunch of different GPU sources.
So the basics of it is, how quickly can we SSH you into a GPU?
Whenever we would talk to users, they wanted a GPU, they wanted an A100.
If you go to any Cloud provisioning page, usually it's like three pages of forms or in the form somewhere there's a drop-down and in the drop-down there's some weird code that you know to translate to an A100.
And I remember just thinking, like every time someone says they want an A100, like the piece of text that they're telling me that they want is like stuffed away in the corner.
And so we were like, what if the biggest piece of text was what the user's asking for?
And so when you go to Brev, it's just big GPU chips with the type that you want.
With beautiful animations that you worked on pre like, pre you can like.
Now you can just prompt it.
But back in the day yeah, i was actually really proud of that because uh, it was an, i made it in figma yeah, and then i found i was like really struggling to figure out how to turn it from like figma to react.
So what it actually is is just an svg And I have all the styles.
And so when you change the chip, whether it's active or not, it changes the SVG code.
And that somehow looks like it's animating, but we just had the transition slow.
But it's just a JavaScript function to change the underlying SVG.
And that was how I ended up figuring out how to move it from Figma.
But yeah, that's artisan.
Speaking of marketing stunts though, he actually used those SVGs, or kind of used those SVGs to make these cards.
Oh, yeah.
Like a GPU gift card.
Yes.
He handed out everywhere.
That was actually my first impression of that.
Yeah.
Yeah, I think I still have one of them.
They look great.
Yeah, i have a ton of them still actually in our garage, but just they don't have labels.
We should honestly like bring them back.
But um, i found this old printing press here actually just around the corner on vaness, and um, it's a third generation san francisco shop and so i come in an excited startup founder trying to like, and they just have this crazy old machinery and i'm in awe because the like, the whole building, is so physical.
Like you're seeing these machines, they have like pedals to like move these saws and whatever.
I don't know what this machinery is, but I saw all three generations.
Like there's like the grandpa, the father and the son.
And the son was like around my age.
It's the holy trinity.
So I just took the same SVG and we just like printed it and it's foil printing.
So they make a mold that's like an inverse of, like the A100, and then they put the foil on it and then they press it into the paper.
And I remember once we got them, he was like, hey, don't forget about us.
I guess like early Apple and Cisco's first business cards were all made there.
And so he was like yeah, we get like the startup businesses, but then as they mature they kind of go somewhere else.
And so I actually, I think we were talking with marketing about like using them.
We should go back and make some cards.
Yeah, 100%.
Yeah, yeah.
I remember as a very, very small breadth investor, I was like why are we spending time doing these stunts for GPUs?
I think as a typical cloud hardware person, you go into AWS, you pick T5 XXL whatever, and it's just from a list and you look at the specs.
Why animate this GPU?
And I do think it just shows the level of care that goes throughout Rev and also Dynamo.
And NVIDIA.
I think that's the thing that struck me most when we first came in was the amount of passion that everyone has.
You know you talk to.
You talk to Kyle, you talk to like every.
Every VP that I've met in video goes so close to the metal.
Like I remember it was almost a year ago and like my VP asked me, he's like, Hey, what's cursor?
And like, are you using it?
And if so, why?
And I'm just like surprised at this.
And he downloaded cursor and he was asking me to help him, like use it, and i thought that was or like just show him what.
You know why we were using it.
And so um, the amount of care that i think everyone has and the uh appreciate passion and appreciation for the moment right, this is a very unique time uh, so it's really cool to see everyone uh really like uh appreciate that Yeah.
One thing I wanted to do, before we move over to research topics and the stuff that Kyle's working on, is just tell the story of the acquisition.
Not many people have been through an acquisition with NVIDIA.
What's it like?
Yeah, just anything you'd like to say.
I mean, it's a crazy experience.
I think you know we were the thing that was the most exciting for us was our goal was just to make it easier for developers.
We wanted to find access to GPUs, make it easier to do that.
And then all actually your question about launchable.
So launchables was just make one click like one click deploys for any software on top of the GPU.
And so What we really liked about NVIDIA was that it felt like we just got a lot more resources to do all of that.
I think NVIDIA's goal is to make things as easy for developers as possible, so there was a really nice synergy there.
I think that, when it comes to an acquisition, I think the amount that the soul of the products align I think is going to be is going to speak to the success of the acquisition.
And so in many ways feels like we're home.
This is a really great outcome for us.
Like we, you know, I love brev.nvidia.com.
Like you should, you should use it.
It's a front page for GPUs.
If you want GPUs, you go there.
And it's like internally is growing very quickly.
I don't remember.
You said some stats there, right?
Yeah, I wish I had the exact numbers, but internally, externally, it's been growing really quickly.
We've been working with a bunch of partners, with a bunch of different customers and ISVs.
If you have a solution that you want someone that runs on a GPU and you want people to use it quickly, we can bundle it up in a launchable and make it a one-click run. if you're doing things and you want just like a sandbox or something to run on right like open claw huge moment super exciting our uh and we'll talk into it more but you know internally people want to run this and you know we have to be really careful from the security implications do we let this run on the corporate network Security's guidance was, hey, run this on breath.
It's in, you know, it's a VM.
It's sitting in the cloud.
It's off the corporate network.
It's isolated.
And so that's been our stance, internally and externally, about how to even run something like OpenClaw, while we figure out how to run these things securely.
But yeah, yeah.
I think you almost were the right team at the right time, when NVIDIA is starting to invest a lot more in developer experience, or whatever you call it UX, or I don't know what you call it software.
Obviously, NVIDIA is always invested in software, but this is a different audience.
It's a wider developer base.
Yeah, you know, it's funny.
It's like, it's not... So what is it called internally?
What is this that people should be aware that is going on there?
Like developer experience?
Yeah, is this called developer experience?
Or is there like a broader strategy here?
NVIDIA always wants to make a good developer experience.
The thing is, you know, a lot of the technology is just really complicated.
Like it's not, it's, you know, I think...
The thing that's been really growing, or the AI is growing, is having a huge moment, not because, like let's say, data scientists in 2018 were quiet then and are much louder now.
The pie is right.
There's a whole bunch of new audiences.
My mom's wondering what she's doing.
My sister's learned, like taught herself how to code, like the.
You know, I actually think just generally, AI is a big equalizer and you're seeing a more like technologically literate society.
I guess, like everyone's, everyone's learning how to code.
There isn't really an excuse for that.
And so building a good UX means that you really understand who your end user is.
And when your end user becomes such a wide variety of people, then you have to almost like reinvent the practice, right?
And actually build more developer UX.
Because there are tiers of developer base that were added.
The hackers that are building on top of OpenClaw, for example, have never used GPU.
They don't know what CUDA is.
They just want to run something.
You need new UX that is not just hey, how do you program something in CUDA and run it?
And then when deep learning was getting big, we built Torch.
But recently the amount of layers that are added to that developer stack has just exploded because AI has become ubiquitous.
Everyone's using it in different ways.
It's moving fast in every direction, vertical, horizontal.
You guys, you even take it down to hardware.
Like the DGX Spark.
You know, it's basically the same system as just throwing it up on big GPU clusters.
Yeah, yeah, yeah.
It's a Blackwell.
Yeah, we saw the preview at the last year's UTC and that was one of the better performing videos of our NVIDIA coverage so far.
Awesome.
This will beat it.
That was actually fun.
Yeah, even when DGX Spark was first coming out, getting to be involved in that from the beginning of the developer experience and it just comes back.
You were involved?
Yeah.
Yeah.
Yeah.
Yeah.
I mean, it was just like I got an email.
We just got thrown into the loop and suddenly yeah, it was actually really funny because I'm still pretty fresh from the acquisition and I'm getting an email from a bunch of the engineering VPs about, like the new hardware, GPU chip or not chip, but just GPU system that we're putting out.
And I'm like, okay, cool.
Natter's now involved with this for the UX.
I'm like, what am I going to do here?
So I remember the first meeting.
I was just like kind of quiet as I was hearing engineering VPs talk about what this box could be, what it could do, how we should use it.
And I remember one of the first ideas that people were ideaing was like oh, the first thing that it was like I think a quote was like the first thing someone's gonna wanna do with this is get two of them and run a Kubernetes cluster on top of them.
And I was like, oh, I think I know why I'm here.
I was like, the first thing we're doing is easy SSH into the machine.
And you're just kind of like scoping it down of like.
Once you can do that, like the person who wants to run a Kubernetes cluster onto Sparks has a higher propensity for pain, then then you know someone who buys it and wants to run open claw right now.
Right, if you can make sure that that's as effortless as possible, then the rest becomes easy.
So there's a tool called nvidia sync.
It just makes the ssh connection really simple.
So you know, if you think about it, like if you have a mac or a PC or whatever, if you have a laptop and you buy this GPU and you want to use it, you should be able to use it like it's a GPU in the cloud, right?
But there's all this friction of like, how do you actually get into that?
That's part of Brev's value proposition is just.
You know, there's a CLI that wraps SSH and makes it simple.
And so our goal is just get you into that machine really easily.
And one thing we just launched at CES, it's in It's still in like early access.
We're ironing out some kinks, but it should be ready by GTC.
You can register your Spark on Rev.
And so now if you.
Like remote managed.
Yeah.
Local hardware.
Single pane of glass.
Yeah.
Because Rev can already manage other clouds anyway.
And you use the Spark on Brev as well, right?
Yeah, exactly.
So you set it up at home, you can run the command on it and then it gets.
It's essentially.
It'll appear in your Brev account.
And then you can take your laptop to a Starbucks or to a cafe and you will continue to use your.
You can continue to use your Spark, just like any other cloud node.
On Brev.
And it's just like a pre-provisioned.
Yeah, exactly.
Tiny little data center.
One more thing before we move on to Kyle.
Just have so many Jensen stories and I just love mining Jensen stories.
My favorite so far is SOL.
What is SOL?
I think of all the lessons I've learned, that one's definitely my favorite.
I can always stick with you.
Yeah.
In your startup, everything's existential, right?
We've run out of money.
We were on the risk of losing payroll.
We've had to contract our team because we ran out of money.
And so because of that, you're really always forcing yourself to like understand the root cause of everything.
If you get a date, if you get a timeline, you know exactly why that date or timeline is there.
You're pushing every boundary and like you're not just say, you're not just accepting like a no, just because.
And so as you start to introduce more layers, as you start to become a much larger organization, SOL is essentially like what is the physics right?
The speed of light moves at a certain speed.
So if light's moving slower, then you know something's in the way.
So before trying to like layer reality back in of like, why can't this be delivered at some date?
Let's just understand the physics.
What is the theoretical limit to like how fast this can go?
And then start to tell me why.
Because otherwise, people will start telling you why something can't be done.
But actually, I think any great leader's goal is just to create urgency.
There's an infinite- Yeah, create impelling events, right?
SOL as a term in NVIDIA is used to sort of like instigate a compelling event.
You say, this is done.
How do we get there?
What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here?
And it helps you just break through a bunch of noise.
Yeah.
One thing I'm unclear about is, can only Jensen use the SOL card?
Oh, no, no, no.
Get the bullshit out.
Because obviously it's Jensen.
But can someone else be like, no?
Frontline engineers use it.
Okay.
Yeah.
Every, I think it's not so much about like, get the bullshit out.
It's like, it's like, give me the root understanding, right?
Like if you tell me something takes three weeks, first principles.
Yeah.
The first principles.
It's like, what's, what's the, what, like, why is it three weeks?
What is the actual uh, like you know Yeah, what's the actual limit of why this is gonna take three weeks?
If, let's say, you wanted to buy a new computer and someone told you it's gonna be here in five days, what's the SOL?
Well, the SOL is I could walk into a Best Buy and pick it up for you.
Right.
So then anything that's like beyond that is, and is that practical?
Is that how we're going to, you know, let's say give everyone in the company a laptop?
Like, obviously not.
So then like, that's the SOL.
And then it's like, okay, well, if we have to get more than 10, suddenly there might be some, right.
And so now we can kind of piece the reality back.
So, so this is the Paul Graham do things that don't scale.
Yeah.
And this is also the, what people would now call be high agency.
It's actually really interesting because there's a second hardware angle to SOL that doesn't come up for all the orgs.
So SOL is used culturally for everything.
I'm also mining for that.
I think that can be annoying sometimes.
When someone keeps going, SOL, SOL, SOL.
And you're like, guys, we have to be stable.
We have to fucking plan.
It's an interesting balance.
Yeah, i encountered that with like, actually just with alec right, because we have, we have a new conference uh uh, so we need to launch.
We have, we have goals of what we want to launch by uh, by the conference, and like yeah, at the end of the day, is this gta?
Well, this is like.
So we I mean, we did it for CES, we did it for GTCDC before that, we're doing it for GTC San Jose.
So I mean, like every you know we have, we have a new moment and we want to launch something, and we want to do so at SOL.
And that does mean that some, there's some level of prioritization that needs to happen.
It is difficult, right?
I think you have to be careful with what you're pushing.
Stability is important, and that should be factored into SOL.
SOL isn't just like build everything and let it break.
That's part of the conversation, as you're laying layering in all the details.
One of them might be Hey, we could build this, but then it's not going to be stable, for X Y, Z reasons.
And so that was like one of our conversations for CES was, you know Hey, like we, we can get this into early access uh, registering your spark with brev.
But there are a lot of things that we need to do in order to feel really comfortable from a security perspective, right?
There's a lot of networking involved before we deliver that to users.
So it's like, okay, let's get this to a point where we can at least let people experiment with it.
We had it in a booth, we had it in Jensen's keynote, and then let's go iron out all the networking kinks.
And that's not easy.
And so that can come later.
And so that was the way that we layered that back in.
But it's not really about like saying like you don't have to do the maintenance or operational work.
It's more about saying, you know, it's kind of like highlights how progress is incremental, right?
Like what is the minimum thing that we can get to?
And then there's SOL for like every component after that.
But there's the SOL to get you to the starting line, right?
And that's usually how it's asked.
On the other side, you know, like, SOL came out of, like, hardware at NVIDIA, right?
So SOL is like literally, if we ran the accelerator or the GPU with like at basically full speed, with like no other constraints, like how fast would we be able to make a program go?
Yeah, yeah.
Right?
Uh, so in in training that, like you know, then you work back to like some percentage of like MFU, for example.
Yeah.
That's a, that's a great example.
So like there's an, there's an SOL MFU, and then there's, like you know, what's practically achievable.
Cool, should we move on to sort of Kyle's side?
Kyle, you're sort of coming more from the data science world.
And I mean I always whenever I meet someone who's done working tabular stuff graph, neural networks, time series, these are basically.
When I go to NeurIPS, I go to ICML, I walk the back halls, there's always like a small group of graph people, small group of tabular people.
And like, there's no one there.
And like, it's very, like, you know what I mean?
Like, it's important, interesting work if you care about solving the problems that they solve.
But everyone else is just LLMs all the time.
Yeah.
I mean, it's like, it's like the black hole, right?
Has the event horizon reached this yet in nerves?
But like, you know, those are, those are transformers too.
And, and those are also like interesting things.
Anyway, I just want to spend a little bit of time on on those that background before we go into Dynamo proper.
Yeah, sure.
I took a different path to Nvidia than Adder.
I joined six years ago, seven, if you count when I was an intern.
So I joined NVIDIA right out of college and the first thing I jumped into was not what I'd done during internship, which was some stuff for autonomous vehicles like heavyweight object detection.
I jumped into know something.
I'm like recommenders, this is popular um, and oh yeah, you did rexus yeah yeah, i mean, that was the tabular data at the time.
Right, you have tables of, like you know, audience qualities and item qualities and you're trying to figure out like, which member of the audience matches which item or, more practically, which item matches which member of the audience.
And, And at the time, really it was like we were trying to enable recommenders, which had historically been a little bit of a CPU-based workflow, into something that ran really well in GPUs.
And it's since been done.
There are a bunch of libraries for Exis. that run on GPUs.
The common models, like Deep Learning Recommendation Model which came out of Meta, and the Wide and Deep Model which was released by Google, were very accelerated by GPUs using the fast HPM on the chips, especially to do vector lookups.
But it was very interesting at the time and super, super relevant, because we were starting to get this explosion of feeds and things that required recommenders to just actively be on all the time.
Sort of transitioned that a little bit towards graph neural networks when I discovered them because I was like okay, you can actually use graph neural networks to represent, like relationships between people items, concepts.
And that interested me.
So I jumped into that at NVIDIA and got really involved for like two-ish years.
Yeah.
And something I learned from Brian Cannizzaro is that you can just kind of choose your own path in nvidia oh my god yeah, which is not a normal big corp thing.
Yeah, you have a lane, you stay in your lane.
I think probably the reason why I enjoy being in a big company from a startup guy.
Yeah.
It feels like a big game of pickup basketball.
Like you know, if you put one, if you want to play basketball, you just go up to the court and you're like Hey look, we're going to play this game and we need three.
And you just like find your three.
Um, that's honestly for every new initiative.
Uh, that's what it feels like.
Yeah.
It also, like, shows, right?
Like, NVIDIA is just releasing state-of-the-art stuff in every domain.
Like, okay, you expect foundation models with Nemotron.
Voice just randomly, like popped your parakeet, comes out.
Another one.
Uh, the voice team has always been producing.
Yeah, there's always just every other domain of paper that comes out, data set that comes out.
It's like, i mean, it also stems back to what nvidia has to do right, you have to make chips years before they're actually produced, right?
So you need to know, you need to really focus.
The design process starts like exactly three to five years before the chip gets to the market.
Yeah, I'm curious more about what that's like, right?
So like you have specialist teams.
Is it just like you know people find an interest, you go in, you go deep on whatever and that kind of feeds back into you know?
Okay, we expect predictions.
Like the internals at NVIDIA must be crazy, right?
You know, you must...
Not even without selling to people.
You have your own predictions of where things are going and they're very based, very grounded right.
Yeah, it's really interesting.
So there's two things I think that Amedia does which are quite interesting.
One is we really index into passion.
There's a big sort of organizational top sound push to ensure that people are working on the things that they're passionate about.
So if someone proposes something that's interesting, many times they can just email someone like way up the chain that they would find this relevant and say like hey, can I go work with this?
That's actually like I worked at a big company for a couple of years before starting on my startup journey.
And like, it felt very weird if you were to like email out of chain, if that makes sense.
The emails at Nvidia are like mosh pits shoot and it's just like 60 people just whatever, and like there's this thing at messy like reply all.
Oh, it's insane, it's insane agents help, you know, mix the context, but but that's actually like um, I've actually so this is a weird thing where I used to be like why would we send emails?
We have Slack.
I am the entire, I'm the exact opposite.
I feel so bad for anyone who's like messaging me on Slack because I'm so unresponsive.
I'm email maxing now.
Email is perfect.
Oh man, we can't work together.
Email is great because important threads get bumped back up, right?
And so Slack doesn't do that.
So I just have like this casino going off on the right or on the left.
And like I don't know which thread was from where or what, but like the threads get and then also just like the subject.
So you can have like working threads.
I think what's difficult is like when you're small.
If it's not 40000 people, I think Slack will work fine.
But there's, I don't know what the inflection point is.
There is going to be a point where that becomes really messy and you'll actually prefer having email because you can have working threads.
You can CC more than nine people in a thread.
You can fork stuff.
You can fork stuff, which is super nice.
And just like, yeah.
And so, but that is part of where you can, propose a plan, you can also just like start.
Honestly, momentum is the only authority, right?
So like if you can just start to make a little bit of progress and show someone something, and then they can try it.
That's, I think, what's been, you know, I think, the most effective way to push anything forward.
And that's both at Nvidia and I think just generally.
Yeah, there's the other concept that like is explored a lot in video, which is this idea of a zero billion dollar business.
Like like, market creation is a big thing.
At nvidia like, you want to go and start a zero billion dollar business, jensen says we're completely happy investing in zero billion dollar markets.
We don't care if this creates revenue.
It's important for us to know about this market.
We think it will be important in the future.
It can be zero billion dollars for a while.
I'm probably mangling his words here But, like you know like, I'll give an example.
NVIDIA has been working on autonomous driving for a long time.
Like an NVIDIA car?
No.
They've used the Mercedes, right?
They're on the HQ.
And I think it finally just got licensed out.
Now they're starting to be used quite a bit.
But for 10 years, you've been seeing Mercedes with NVIDIA logos.
If you're in the South Bay near Santa Clara, it's actually pretty common.
You know, zero billion dollar markets are a thing, like you know, Jensen.
I mean, OK, look, cars are not a zero billion dollar market, but yeah, that's a bad example.
I think he's messaging zero today, but.
Or even like internally, right?
Like it's like an org doesn't have to ruthlessly find revenue very quickly to justify their existence right.
Like a lot of the important research, a lot of the important technology being developed.
That's kind of where
Research is very ideologically free at NVIDIA.
Like they can pursue things that Will you research officially?
I was never in research officially.
I was always in engineering.
I'm in an org called Deep Warning Algorithms, which is basically just how do we make things that are relevant to deep warning go fast.
That sounds freaking cool.
And I think a lot of that is underappreciated, right?
Like, time series.
This week, Google put out a new time series paper.
REXIS.
Semantic ID started applying transformers, LLMs to REXIS.
And when you think the scale of companies deploying these, right?
Amazon recommendations, Google web search.
It's huge scale, and you want fast.
Actually, there's a fun moment that brought me full circle.
Amazon ads recently gave a talk where they talked about using Dynamo for generative recommendation, which was super weirdly cathartic for me.
Oh my God.
I've supplanted what I was working on.
You're using LLMs now to do what I was doing five years ago.
Yeah, yeah, yeah.
Amazing.
Let's go right into Dynamo.
Maybe introduce it to the top down.
Yeah sure, i think at this point a lot of people are familiar with the term of inference.
Like, funnily enough, like i, i went from, you know, inference being like a really niche topic to being something that's like discussed on like normal people's twitter feed.
It's on billboards here.
Yeah very, very strange driving driving, seeing just an inference ad on 101.
But inference at scale is becoming a lot more important.
We have these moments like OpenClaw, where you have these agents that take lots and lots of tokens but produce incredible results.
There are many different aspects of test time scaling so that you can use more inference to generate a better result than if you were to use a short amount of inference.
There's reasoning, there's re-querying, there's adding agency to the model, allowing it to call tools and use skills.
Dyno sort of came about at NVIDIA, because myself and a couple others we're sort of talking about these concepts that, like you know, you have inference engines like VLM SGLang TensorFlow TLM, and they have one single copy.
They think about things as one single copy, one replica, one version of the model.
But when you're actually serving things at scale, you can't just scale up that replica, because you end up with performance problems.
There's a scaling limit to scaling up replicas.
So you actually have to scale out to use maybe some Kubernetes terminology.
We realized that there was a lot of potential optimization that we could do in scaling out and building systems for data center scale inference.
So Dynamo is this data center scale inference engine that sits on top of the frameworks like VLM, SQLang and TensorFlow TLM and just makes things go faster.
Because you can leverage the economy of scale, the fact that you have KV cache, which we can define a little bit later.
In all of these machines that is unique and you want to figure out the ways to maximize your cache hits.
Or you want to employ new techniques in inference, like disaggregation, which Dynamo introduced to the world in March.
Not introduced academic topic beforehand, but we're, you know, one of the first frameworks to start, you know supporting it, and we want to like sort of combine all these techniques into sort of a modular framework that allows you to accelerate your inference at scale.
By the way, kyle and i became friends on my first date in video and i always love because, like he always teaches me New things.
By the way, this is why I wanted to put two of you together.
I was like, yeah, this is going to be good.
It's very different.
We've talked to each other a bunch.
Actually, you asked, why can't we scale up?
Yeah.
You said model replicas.
Yeah, so scale up means assigning more- Heavier.
Yeah, heavier, like making things heavier, adding more GPUs, adding more CPUs.
Scale out is just like having a barrier saying I'm going to duplicate my representation of the model or representation of this microservice or something and I'm going to replicate it many times to handle the load.
And the reason that you can't scale up past some points is there are hardware bounds and algorithmic bounds on that type of scaling.
So I'll give you a good example that's very trivial.
Let's say you're on an H100.
The maximum NVLink domain for H100 for most DJX H100s is eight GPUs.
So if you scaled up past that, you're going to have to figure out ways to handle the fact that now for the GPUs to communicate, you have to do it over InfiniBand, which is still very fast, but it is not as fast as MVLink.
Is it like one order of magnitude, like hundreds?
It's about an order of magnitude.
Not terrible.
Yeah.
I need to remember the data sheet here about 500 gigabytes a second unidirectional for NVLink and about 50 gigabytes a second unidirectional for InfiniBand.
It depends on the generation.
I just want to set this up for people who are not familiar with these kinds of like layers and the transfer speeds.
Also maybe even just going like a few steps back before that, like most people are very familiar with you.
See, you know you can use on your laptop whatever these STLM, VLM.
You know you can just run inference, you can run it on that laptop.
You can run on laptop.
Then you get to okay uh, models got pretty big right glm5, they doubled the size.
So uh, what do you do when you have to go from okay, i can get 128 gigs of memory, i can run it on a spark.
Then you have to go multi-gpu okay, multi-gpu.
There's some support there.
Now, if i'm a company and i don't have like I'm not hiring the best researchers for this right,
But I need to go multi-node, right?
I have a lot of servers.
Well, okay, now there's efficiency problems, right?
You can have multiple H100 nodes, but how do you do that efficiently?
Yeah, how do you represent them?
How do you choose how to represent the model?
That's a hard question everyone asks.
How do you size?
Like, oh, I want to run GLM5, which just came out.
New model.
There have been four of them in the past week, by the way.
A bunch of new models.
You know why, right?
Deep secrets.
No comment.
Yeah, but GLM5, right?
We have this new model.
It's of a large size.
And you have to figure out how to both scale up and scale out, right?
Because you have to find the right representation that you care about.
I mean, everyone does this differently.
Let's be very clear.
Everyone figures this out in their own path.
I feel like a lot of AI or ML even is like this.
I think people think you know there was some tweet a few months ago that was like why hasn't fine tuning as a service taken off?
That might be me.
It might have been you.
But people want it to be such an easy recipe to follow.
But even if you look at an MLE model- It's specific to you.
Yeah, yeah.
And the model and the situation.
And there's so much tinkering.
When you see a model that has however many experts in the MLE model, it's like why that many experts?
I don't know.
I don't know, they tried a bunch of things and that one seemed to do better.
And I think, when it comes to how you're serving inference, you have a bunch of decisions to make and you can always argue that you can take something and make it more optimal, but I think it's this internal calibration and appetite for continued calibration.
And that doesn't mean people aren't taking a shot at this.
Tinker from Thinking Machines, RL as a service.
It also gets even harder when you try to do big model training.
We're not the best at training MOEs when they're pre-trained.
We saw this with Llama 3.
They're trained in such a sparse way that Meta knows there's going to be a bunch of inference done on these.
They'll open source it.
But it's very trained for what meta infrastructure wants.
Right, they want to, they want to inference it a lot.
Now the the question to basically think about is okay, say you want to serve a chat application, a coding copilot right, you're doing a layer of rl, you're serving a model for x amount of people.
It's a chat model, a coding model.
So dynamo, you know back to that.
Yeah sorry, so we we sort of like jumped off of you know, on that topic, everyone has like their own journey and i like to think of it as defined by like what is the model you need?
What is the accuracy you need?
Actually, i talked to nana about this earlier.
There's, there's three axes you care about.
What is the quality they're able to produce?
So like are you accurate enough or can you complete the task with enough you know performance, high enough performance.
Yeah, there's cost.
Can you serve the model or serve your workflow?
Because it's not just the model anymore.
It's the workflow.
It's the multi turn with an agent cheaply enough.
And then can you serve it fast enough?
And we're seeing all three of these play out.
We saw new models from OpenAI that are faster.
You have these new fast versions of models.
You can change the amount of thinking to change the amount of quality, produce more tokens but at a higher cost and a higher latency.
Really, when you start this journey of trying to figure out how you want to host a model, you think about three things.
What is the model I need to serve?
How many times do I need to call it?
What is the input sequence link?
What does the workflow look like on top of it?
What is the SLA?
What is the latency SLA that I need to achieve?
Because there's usually some.
This is usually like a constant.
You know the SLA that you need to hit.
And then like you try and find the lowest cost version that hits all of these constraints.
Usually you know you start with those things and you say you kind of do like a bit of experimentation across some common configurations.
You change the tensor parallel size, which is a form of parallelism.
I'd say it goes even deeper.
First kind of thing will model.
Yeah, it's like it's like a multi-step design process.
Because, as you said, you can choose a smaller model and then do more test time scaling and it'll equate the quality of a larger model, because you're doing the test time scaling or you're adding a harness or something.
So yes, it goes way deeper than that.
But from the performance perspective, once you get to the model you need to host, you look at that and you say hey, I have this model.
I need to serve it at this speed.
What is the right configuration for that?
You guys see the recent.
There's a paper I just saw, like a few days ago, that if you run the same prompt twice you're getting like double digits.
Just try it again.
Yeah, exactly.
But the key thing there is you give the context of the failed try, right?
So it takes a shot.
And this has been like, you know, basic guidance for quite a while.
Just try again.
Because you know, if you try it, just try again.
Did you try again?
All advice in life.
It's a paper from google, if i'm not mistaken right.
I think it's like a seven page little short paper.
Yeah yeah, the title is very cute and it's just like yeah, just try again, give it.
It has context.
You just like say like hey, like you know like take, take a little bit more, take a little bit more information.
Trying to fail, and That basic concept has gone pretty deep.
There's like self-distillation RL, where you do self-distillation, you do RL and you have past failure and you know, that gives some signal.
So people take, try it again, not strong enough.
For listeners who listen to here, Veebo actually and I, we run a second YouTube channel for our paper club.
Oh, that's awesome.
Veebo just covered this.
Self-dissolution and all that.
That's why he's so up to speed on it.
Off to check it out.
It's just a good practice.
Like everyone needs like a paper club where, like you, just read papers together and the social pressure just kind of forces you.
There's like a big inference reading group.
I feel so bad every time I, he put it on like on our, he shared it.
One of your guys is big in that.
I forget.
Ishan.
Ishan.
Ishan's on my team.
Actually, funny.
There's an employee transfer between us.
Ishan worked for Natter at Brev, and now he's on my team.
He was our head of AI and then Yeah, once we got in- Because I'm always looking for like okay, can I start another podcast that only does that thing.
And Ishan was like, I was trying to like nudge Ishan into like, is there something here?
I mean, I don't think there's new infant techniques every day.
So it's like- You would actually be surprised.
The amount of blog posts you see-
There was a period where it was like Medusa, Hydra, Eagle.
Now we have new forms of specular decoding.
What are you excited about?
It's exciting when you guys put out something like Nemotron, because I remember the paper on this, Nemotron 3 the amount of post-training, the amount of tokens that the GPU rich can just train on.
And it was a hybrid state space model.
Yeah, it's co-designed for the hardware.
Yeah, co-designed for the hardware.
And one of the things was always the state space models don't scale as well when you do a conversion or whatever the performance.
And you guys are like, no, just keep training.
And Nemotron chose a lot of that.
Also something cool about Nibitron, it was released in layers, if you will, very similar to Dynamo.
It was released as aggregate.
The pre-training, post-training data sets are released.
The recipes on how to do it are released.
The model itself is released.
Benefit from us churning on the GPUs.
But there are companies like ServiceNow took the data set and they trained their own model.
And we were super excited and celebrated that work.
And Zoom, the frontier model.
Zoom is AGI.
I think, you know, also just to add, like, a lot of models don't put out base models.
And if there's that, why is fine tuning not taken off?
You know, you can do your own post-training, but you guys put out base model?
I think you put out everything.
I believe so.
I don't know about base.
Base can be cancelable.
Base can be cancelable?
Yeah.
Safety training?
Did we get a full picture of Dynamo?
I don't know if we- What I'd love is, you mentioned the three axes.
Break it down of what's pre-filled decode and what are the optimizations that we can get with Dynamo.
Yeah, that's a great point.
To summarize on that three-axis problem, there are three things that determine whether or not something can be done with inference.
Cost, quality, latency.
Dynamo is supposed to be there to provide you the runtime that allows you to pull levers to mix it up and move around the Pareto frontier or the Pareto surface that determines.
Is this actually possible with inference and AI today?
It gives you the knobs.
Yeah, exactly.
It gives you the knobs.
And one thing that we use a lot in contemporary inference and is starting to pick up in general knowledge is this concept of disaggregation.
So historically, models would be hosted with a single inference engine.
And that inference engine would sort of ping pong between two phases.
There's pre-fill where you're reading the sequence, generating KV cache, which is basically just a set of vectors that represent the sequence, and then using that KV cache to generate new tokens, which is called decode.
And some brilliant researchers across multiple different papers essentially made the realization that if you separate these two phases, you actually gain some benefits.
Those benefits are basically, A, you don't have to worry about step synchronous scheduling.
So the way that an inference engine works is you do one step And then you finish it, and then you start scheduling the next step.
It's not fully asynchronous.
And the problem with that is you would have essentially pre-fill and decode are actually very different, in terms of both their resource requirements and sometimes their runtime.
So you would have pre-fill.
That would block decode steps because you'd still be pre-filling and you couldn't schedule because the step has to end.
So you remove that scheduling issue.
And then you also allow yourself to split the work into two different types of pools.
So pre-fill typically, and this changes as model architecture changes, pre-fill is right now compute bound most of the time.
If the sequence is sufficiently long, it's compute bound.
On the decode side, because you're doing a full pass over all the weights and the entire sequence every time you do a decode step and you don't have the quadratic computation of KV KV cache.
It's usually memory bound, because you're retrieving a linear amount of memory and you're doing a linear amount of compute, as opposed to pre-fill, where you retrieve a linear amount of memory and then use a quadratic.
You know, it's funny.
Someone, uh, XO labs did a really cool demo where for the DJX spark, which has a lot more compute, you can do the pre the compute hungry pre-fill on a DJX spark and then do the, uh, decode on a, on a Mac.
That's faster.
Yeah.
You can do machine stratification.
With our future generations of hardware.
We actually announced with Rubin this new accelerator that is pre-fill specific.
It's called Rubin CPX.
So i have a question when you do this scale out, is scaling out easier with dynamo because when you need a new node you can dedicate it to either the pre-fill or uh decode.
Yeah, so dynamo actually has like a kubernetes uh component in it called grove that allows you to do this like crazy scaling specialization.
It has like this hot it's a representation that I don't want to go too deep into Kubernetes here.
But there was a previous way that you would launch multi-node work.
It's called Leader Worker Set.
It's in the Kubernetes standard.
And Leader Worker Set is great.
It served a lot of people super well for a long period of time.
But one of the things that it struggles with is representing a set of cases where you have a multi-node replica that has a pair prefill and decode or it's not paired, but it has a second stage that has a ratio that changes over time.
Prefill and decode are two different things.
As your workload changes, the amount of prefill you'll need to do may change.
The amount of decode that you'll need to do might change.
Let's say you start getting insanely long queries.
That probably means that your prefill scales harder because you're hitting this quadratic scaling growth.
And then for listeners like, pre-fill will be long input, decode will be long output, for example, right.
Yeah.
So like decode, decode scale.
I mean decode is funny because the amount of tokens that you produce scales with the output length, but the amount of work that you do per step scales with the amount of tokens in the context.
Yes.
So both scales with the input and the output.
That's true.
But on the pre-filled decode side.
If suddenly the amount of work you're doing on the decode side stays about the same or scales a little bit, and then the pre-filled side jumps up a lot, you actually don't want that ratio to be the same.
You want it to change over time.
So Dynamo has a set of components that, A, tell you how to scale.
It tells you how many pre-filled workers and decoded workers it thinks you should have and also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduling on your actual hardware, on your computer infrastructure.
Not going to lie, I feel a little embarrassed for being proud of my SVG function earlier.
It was really cute.
It's all engineering.
It's all engineering.
Sort of technical.
One thing I'm kind of just curious about, you see at a systems level everything going on here.
And we're scaling it up in multi-distributed systems.
I think one thing that's like kind of of the moment right now is people are asking is there any SOL sort of upper bounds in terms of like let's call it, just call it context length for one for a better word.
But you can break it down however you like.
I just think like Well yeah, I mean like clearly you can engage in hybrid architectures and throw in some state space models in there all you want, but it still looks very attention heavy.
Yes uh yeah, long context is attention heavy.
I mean, we have these, these hybrid models um, and most, most models like cap out at a million context and that's it like for the last two years, has been it.
Yeah, the model hardware context, co-design thing that we're seeing these days is actually super interesting.
It's like my passion, like my secret side passion.
We see models like Kimi or GPT-OSS.
I'm going to use these because I know specific things about these models.
So Kimi 2 comes out.
And it's an interesting model.
It's like a deep-seek style architecture.
It is MLA.
It's basically deep-seek scaled a little bit differently and obviously trained differently as well.
But they talked about why they made the design choices.
For context, Kimi has more experts but fewer attention heads and, I believe, a slightly smaller attention like dimension, but I need to remember, I need to check that.
It doesn't matter.
But they discussed this actually at length in a blog post on Zhihu, which is like, or Zhipu, which is like, Reddit.
Yeah.
Chinese Reddit.
Yeah.
So it's actually an incredible blog post.
Like, all the MLSIS people that I've seen on GPU are, like, very brilliant.
But they talk about like the creators of KimiK2.
Actually like talked about it on there in a blog post.
And they say, we actually did an experiment, right?
Attention scales with the number of heads.
Obviously, like if you have 64 heads versus 32 heads, you do half the work of attention.
Attention is you still scale quadratically, but you do have to work.
And they made a very specific sort of barter in their system, in their architecture.
They basically said, hey, what if we gave it more experts?
So we're going to use more memory capacity, but we keep the amount of activated experts the same.
We increase the experts' ,, so we have fewer experts.
The ratio of experts activated to number of experts is smaller and we decrease the number of attention heads.
And kind of for context, what we had been seeing was you make models sparser instead.
So no one was really touching heads.
You're just having- Well, they implicitly made it sparser.
Yeah, for Kimmy, they did.
They also made it sparser.
But basically what we were seeing was people were at the level of, okay, there's a sparsity ratio.
You want more total parameters, less active, and that's sparsity.
But what you see from papers, like you know the labs, like Moonshot DeepSeek, they go to the level of okay, outside of just number of experts.
You can also change how many attention heads and less attention layers.
Uh, more attention layers.
Yes, yes, yes.
So and that's all basically coming back to just tie together is like hardware model co-design, which is
Harder model, model context for co-design.
Yeah.
If you were training a model that was really short context or is good at super short context tasks, you may design it in a way such that you don't care about attention scaling because it hasn't hit that the turning point where the quadratic curve takes over.
Why do you consider attention or context as a separate part of the co-design?
I would imagine hardware, or just how I would have thought of it, as hardware model co-design would be hardware model context co-design.
Because the harness and the context that is produced by the harness is a part of the model once it's trained in.
Like, even though towards the end you'll do long context, you're not changing architecture through training.
I mean, you can try.
You're saying everyone's training the harness into the model?
I would say to some degree... Or there's co-design for the harness.
I know there's a small amount, but I feel like not everyone has gone full send on this.
I think it's important to internalize the harness that you think the model will be running into the model.
Interesting.
And bash is like the universal harness.
I'll give... an example here, just like an easy proof.
If you can train against a harness and you're using that harness for everything, wouldn't you just train with the harness to ensure that you get the best possible quality out of?
Well, I can provide a counterargument, which is you want to provide a generally useful model for other people to plug into their harnesses.
Yeah, but harnesses can be open source, right?
Yes, that's effectively what's happening with Codex.
But you may want a different search tool, and then you may have to name it differently.
I don't know how much people have pushed on this, but can you train a model?
Would it be have people compared training a model for the uh for the harness, versus like post-training for for?
I think it's the same thing.
It's the same thing.
Okay.
It's just extra post-training.
I see.
And so I mean Cognition, does this Cursor, does this?
Where you just have to like if your tool is slightly different, either force your tool to be like the tool that they train for, or like undo their training for their tool and then retrain.
It's really annoying.
I would hope that eventually we hit a certain level of generality with respect to .
This is not AGI.
This is really stupid.
Learn my tool, bitch.
I don't know if I can say that.
I think what my point kind of is is that there's like I look at slopes of the scaling laws and like this slope is not working man.
We're at a million token context.
OK, maybe next year, two million.
We're not going to 100 trillion.
You know, like this, this is.
There's so many interesting things.
This doesn't work.
This doesn't work.
What's kind of funny is, whenever there, i feel like we always want to see a trend that we can predict, but every time something's come it's been like a leapfrog.
So i imagine i i don't know how we go from one to two, but i imagine what.
What's likely to happen is we break through that from some new.
Yeah, there's actually there's an interesting formalization of this.
There's an essay it's a pretty interesting essay by leopold aschenbrenner called situational awareness.
Okay yes, he introduces a concept, awareness called an unhoveler right, so you know leopold, in this essay details hey, i want to get you know like i want to get to this point in intelligence and i think that it is four orders of magnitude worth of like compute and data and training away.
And he says oh yeah, I think data centers can scale up by about this much.
I think that you can scale up the data and some other things by this much.
But one of the things that makes the rest of that order of magnitude growth possible is these un-Hobblers, these scientific discoveries that are discovered during model architecture search or training that really really, really impact how you are able to scale.
A good example of this might be that we see a lot of models that are and this is probably a very tiny unhobbler, but is important for the performance perspective.
We see a lot of models that are trained with multi-token prediction natively during pre-training.
And per DeepSeek.
In their paper they say hey, this actually helped us ensure more stable convergence.
But there are unhobblers that are like that, and then there are rather large unhobblers.
Architecturally, a lot of our models, we have different types of attention.
And one of the problems with attention is you have a lot of KV.
But people have found different forms of attention, like group query attention and MLA in DeepSeq, multi-head-laden attention that decrease the burden that KV has on the model which allows you to grow longer in context.
MARK MANDELMAN- Yeah, and that was very drastic for DeepSeq.
Yeah, for context, like the total.
I think the total context length of DeepSeq is 128000 tokens or it might be 256000 with rope extension.
That entire context, I think it's 128,000, fits into eight gigabytes.
And previously context like, I think the, the llama, four or five B context of a similar size was like 40 or 80 gigabytes in the same precision.
Yeah.
So like those are hobblers like really decrease the stuff of that size.
And i wouldn't be surprised if we do see the ability to like break through to like 10 million, 20 million, 100 million context through the.
An unhoveler showing up i see, and it's just science.
So more deep learning algorithms.
He's playing pick up and he has room for two.
I could actually give you an example of a theory, not a theory here, but something theoretical.
A hobbler that you're excited about?
A hobbler that I haven't seen.
It could be a tar pit and it could just not work.
But I would be really excited to see a model that does pre-fill and decode differently.
So a model that does pre-fill locally, document-wise pre-fill like a dozen in chunks, and then you do decode globally across the entire sequence.
Logically, to me it doesn't seem like you would necessarily need to have KV be associative between documents that have no mutual association, but that places a lot of burden on decode and pure attention within the decode phase to make those connections, since the KV is static at that point.
And you see other techniques that are interesting like this too.
But if you're able to do that, you know, if pre-fill becomes local and decode is still global, you solve that pre-fill quadratic scaling problem because you have a bunch of like small chunks that you pre-fill independently.
Okay.
All right.
Well, let's wait and see.
But I think it'll be pretty exciting.
Fingers crossed.
Yeah, fingers crossed.
Yeah.
I'm excited for like pre-fail decode on separate hardware.
So like Grok acquisition, right?
Can we decode on the Grok?
Can we get super fast?
I don't think I'm allowed to comment on this.
Mark is going to shoot arrows at us.
I'm super excited to see the team come in and, like you know, I've gotten the pleasure of working with some of the GROK people coming in.
So, um, you know, I know Sonny, we've had him, uh, at the same conference that you're at.
Um, and, uh, I think you're, you guys are going to be doing some sessions at GTC.
I don't know if you want, this is a good place to plug them.
Yeah, yeah.
So I can't speak to any LPU-related sessions at GDC.
I have no idea about that.
Oh, no, no.
You.
On the Grok side, yeah.
I use the associative NVIDIA U.
On the NVIDIA Dynamo side, there are a large number of sessions.
For those that aren't aware, you can actually search all of these sessions for GTC online.
Just go to the GTC website.
I don't know what the URL is, but Go there and you can just look up Dynamo and you'll get all the sessions.
There are about 20.
There are a couple that are hosted by the Dynamo team.
There are a couple that are hosted by people that use Dynamo that want to show off the results they've been able to get.
But there are two that I'm really excited about.
One is just the general Dynamo tutorial.
And this is the, you know, I'm going out with Harry, who's our lead product manager for Dynamo.
And we're sort of talking about like how to use Dynamo to get better performance and also like where we see Dynamo going in the future.
And then there's another session that I'm doing with one of our agents teams at Nvidia to talk about sort of the future of agents in production inference.
Um, so we're talking about there's like this new horizon with respect to agents because we have these harnesses that actually impart structure among upon calls, like if you compare it, like you know, the past and the present with respect to how LM calls work.
In the early days when there were chatbots, every call was very different.
There was basically no structure.
You could assume that if it was conversational there might be some implicit structure, because you have a multi-turn conversation.
But agents, you have this harness that abides by rules.
So it imparts direct structure onto the context.
And you see this.
There was an interesting Twitter post about how Cloud Code structures its context so that you get as many cache hits as possible.
And I think it was by one of the PMs for Cloud Code.
And he wrote about it. that type of structure that the harness can impart actually goes hand in hand with the inference co-design.
So I'm doing a talk.
I don't know the session name or the session number, but I'm doing a talk.
You can look me up by name on the GTC website on how we accelerate agents and where we see specific optimizations for agents going in Dynamo and in inference in general.
Yeah, I think there's only 1 p.m. for Cloud Code, and it's kind of woo.
The rest, there's DevRel, there's Boris.
Maybe it was DevRel.
Exactly.
I mean, let's go into agents.
I think this was the last part of the discussion we planned.
How have we not talked about agents?
Well, we scheduled it.
I was like, okay, let's have cohesive sections.
I mean, there's the big news, right?
NVIDIA is a huge deployment of Codecs.
Yeah, NVIDIA uses everything.
It uses Cursor and it uses Codecs.
That's a pretty big deployment, right?
That's tens of thousands of people I'm curious what it's like.
Yeah, I mean it goes back to the mosh pit of emails we kind of mentioned earlier, or just how fluid the org feels.
So when there's new technology, people will just email it out and everyone will try it.
And if it's making people's lives easier, it'll spread like wildfire.
A lot of times Jensen will get it and he'll be like, Let's make this work across the company.
Let's make this work right now.
Honestly, if I was a startup, I feel like a cool hack.
If you have something that's going to save an invidians time, they'll spread it to a couple in the same thing, right?
It'll just spread like wildfire.
Careful before your email blows up from startups, by the way.
Well, you gotta know the person, right?
But no, I... Yeah, so, I mean, I love using Codex.
It's been a ton of fun.
I've been using it personally.
I've been using it at work.
It's been...
Yeah, I don't know.
It's been great to see the rollout.
Something really funny, on the day that we got Codex and Cloud Code Access, I found this person.
His name's Carlos at the company.
He wrote an Outlook CLI.
Oh, yeah.
And just the CLI for email.
And this was- I've been using that.
Yeah, maybe like four or five weeks ago.
And so once I got Codex access, I installed the CLI.
It had a skill.
And I just asked it to go through all of my emails, which it's very messy.
So if I don't respond to your email, I'm really sorry.
But I asked it to give me a summary highlight any escalations that I should look at, put any thread that it thinks I should respond to in a folder and then archive everything.
And it did.
So if I missed your email, it's because it didn't give me the power.
So I should put a prompt injection in my emails too, yeah.
What you should do is just FaceTime.
Yeah, my SLA is highest on FaceTime.
But it was magic.
And so I sent it in a big email thread to like 500 people.
A bunch of folks tried it out.
I started like FaceTiming whoever I could at the company to get them set up with this.
Yeah.
That specific example, you guys deal with like some pretty sensitive emails.
Yeah.
Is there a security review with this?
Because, like, one guy made it for himself, but, like, it's not meant for all NVIDIA users.
The security team at NVIDIA is incredible.
Like, shout out to them.
They're trying to.
We have an amazing security team because they're progressive and they know that this is really important technology and we have to bring it in.
If you think about, like, if you work at a big company, your laptop's usually very locked down.
You can only access certain things.
NVIDIA engineers have, like, those restrictions aren't there.
So you're expected to understand the risks when you try things out.
And so very quickly, you know, made sure to chime in security on what we were doing.
There's actually a lot that we've been thinking about, especially with OpenClaw, right?
Like there's, you know, agents can do three things.
Yeah.
I mean, agents can do three things.
They can access your files, they can access the internet, and then now they can write custom code and execute it.
And you really only let an agent do two of those three things.
If you can access your files and you can write custom code, you don't want internet access, because that's one to see for vulnerability, right?
If you have access to internet and your file system, you should know the full scope of what that agent's capable of doing.
Otherwise, you know, malware can get injected or something that can happen.
And so that's a lot of what we've been thinking about is like you know, how do we both enable this because it's clearly the future, but then also, you know, what are these enforcement points that we can start to like protect.
And there's certainly directive of like hey, we have a company account or a company agreement with OpenAI.
We use OpenAI models here, or choose whatever.
No, no.
So I would never put any company data in a model.
That's not either that we don't.
Even It has the most security.
Yeah.
Yeah.
Like how... You know, obviously you could run your own models.
You have Nimotran and... We have an internal cluster.
So, you know, of course...
Yeah.
Yeah.
I think we're Dynamo's first customer.
Actually, there's a funny story about how I got the experience that informed what we needed for Dynamo.
At one point.
There's a website called buildNVIDIAcom and also, for us inferenceNVIDIAcom, that allows people to try models.
It gives an API service.
You can call the model with a REST API, and you get a response.
I ran the model side for that.
And it was at one point the largest inference deployment and still may actually be the largest inference deployment.
I've since like handed that off to some people and they're doing wonderful.
This is an extremely under known or less known resource.
Build.mv.com.
You can get any of these open source models and it's rate limited, but it's free.
So it's perfect for hackers to say.
And and and and the SLA on getting models day zero models up is like a day.
Yeah.
Like they're incredibly good at like figuring out the right way to host the model to get it up there.
As soon as it comes up, you ran this.
Yeah.
I ran, I ran it a long time ago.
It was originally called Nvidia AI playground.
Then it was called AI foundation and then it was called build out a video call, and I ran the model side of it.
So there were, there was a large multi-organizational team.
I ran which models should we post?
How should we host them?
And what's the proportion of them?
And then, of course, there was an SRE team that made sure that things ran well and scaled the models as well.
But I ran model.
How do we get the model to silicon?
And then you know which also worked with our product team, determined like which models were important a very long time ago.
Yeah, there's also like a middle ground in between there.
This is like for the hacker try anything.
There's the brev console, then there's dynamo.
There was also nim's right.
Yeah, i remember it had its little moment like a year or two ago.
NIM is how enterprises can take any of this technology and run it with support and all of that.
And so that includes Dynamo, that includes, you know, I don't know all of our other optimizations that are packaged up for enterprise.
Yeah.
Yeah.
Anyway.
So you got a bunch of experience like running the sort of internal inference gateway playground.
Yeah.
And Bill also built, helped build NVIDIA's first internal, like, VS code thing.
You call that MB code?
FRANCESC CAMPOY- It's an extension.
Yeah, it was a VS code.
FRANCESC CAMPOY- The fork VS code.
MARK MANDELMANN- We joke, absolutely not.
FRANCESC CAMPOY- Just a while back, we should have a fork VS code hackathon where you .
MARK MANDELMANN- It's always the best fork VS code.
FRANCESC CAMPOY- We are doing a good thing.
MARK MANDELMANN- How do you make a billion dollars?
Someone from VS Code was there.
And he was somewhat down to get involved.
And I was like, oh, you should do that.
That's what I said.
Then the cool thing became for Chrome Hackathon.
And now IDs are not cool.
MARK MANDELMANN- I was talking to Joseph from RoboFlow, the partner in crime.
We were talking about how with the new Alpamayo model.
So NVIDIA just released an open source, the Mercedes cars that you saw drive.
Sounds crazy.
Yeah, released.
Will you open source a autonomous driving model?
So I already, yeah.
So we were thinking like, could we hackathon a driverless car?
Like I have my old car, let's just try it, we'll take it, take it down like click trailer with a treasure island in the middle of the day.
That's why i just see, like everyone yeah, like how many?
How many cameras do we need?
Right, like one two three four, i don't know, maybe five states.
I don't like yeah, but um, i think we're gonna try.
You just do with us.
We could even have a race.
It's like the first person to automate their driving.
We do have an autonomy track at World's Fair.
Waymo was there.
NVIDIA did send people those for Groot.
Because you didn't have the driving thing yet.
Yeah, that's cool.
I think Karma also has a version of this.
Karma, yeah, they have open source driving.
They've done a fun hackathon on it.
He and I have talked about it, because what I really want is a Tesla with Tesla-level self-driving, but as a smart car like a two-seater.
That's basically a wheelchair with a roof.
I don't think they make them in Dubai.
The demand has been there.
Is this really five years?
Yeah.
Really?
Yeah.
They were this manufacturer.
I thought it was one of those things where we'll see someone buy the brand and it'll be revived.
I would buy it.
Go.
Someone hears this go buy your car.
Yeah, that's crazy, because that they're like i think my camera says a mercedes uh, and they're saying you're used to make them.
Yeah, i don't know, i feel like they own the brand and you out.
Your dream might come true enough.
Okay, every time i try to park in san francisco, i yeah, i have to buy a smart car, because like 20 of the parking lots in san francisco only fit smart cars Yeah.
Really?
This comes from someone that like basically does not drive.
That's sort of the, the Vespa was a life hack.
Yeah, exactly.
You know what happened to the Vespa?
I used to have this yellow Vespa.
I left it outside the hacker house when we moved out.
It was always there, and then a month ago, it's not there anymore.
I've been meaning to do it.
I don't know.
It's actually been here.
It's like a TV.
You forgot about it.
Yeah.
No, this is probably a hazard.
And speaking of hackathons, I also wanted to give a big shout out to the world's shortest hackathon.
Let's go.
You did it twice?
We did it a handful of times.
Yeah, there's going to be one at GTC.
Oh, we're doing that?
Oh, what?
Pretty much, we have a bunch of challenges that we haven't released.
And you get to bring your agent to come and attempt to go through those challenges.
It's like the zero-minute hackathon idea, where you just bring your agent.
I had a long time ago.
You just bring your agent, and then you press the Go button.
You're not allowed to code.
It's just the agent doing .
It's a good hidden evil, right?
Yeah.
You make a JRO, and you make .
I will love to see from Cognition or someone else be like, come, bring your agent.
Like, drop it in.
MARK MANDELMANN, Because you don't know what you like to prove.
Will it be, you know, operate a browser, order a pizza?
Will it just seem like a snake game, you know?
FRANCESC CAMPOY, And you don't know what the task is.
MARK MANDELMANN, You don't know what the task is.
Or just, like, you don't even know what the judging categories are.
And then you give it the judging categories, like, try and win as much as possible.
FRANCESC CAMPOY, It's great, though.
It turns into, like, yeah, so let's build something on Dino Party.
MARK MANDELMANN, It's a great person to speak to.
Funny story, actually.
We have a couple of people at NVIDIA.
We've been working with security to bring agents really close to compute.
So we now have stuff where you can tell Dynamo go write some experience with Dynamo on xCluster and just try it right now.
Queue up.
Once you get queued, you know send this request load.
And we've actually been able to like, just like you know, like one shot problems, like we used to have this problem where, like you know, with dynamo, you have to like find the right configurations and we, you know, sort of do it automatically for some parts of it, but you have to like a good initial configuration that you want to use and we've just had like an agent, just completely one shot that it goes, it gets the compute, it like runs.
A couple of experiments is like this is the best.
These are part of the Pareto frontier.
Go run this.
And then we just like give that to people and it's like faster than anything that they have.
Agent UX and agent marketing are super important.
There's something we've been thinking a lot about.
Alec is like redoing the entire Rev CLI so that you can fetch all the different compute types that are available.
I don't know, it's going to be really soon, but then you can just browse what GPUs are available and then provision one SSH to it right there and you can pipe all the commands.
But I think it goes back to like the Alec CLI.
Like if you, you know, coding agents are.
It's kind of funny.
I feel like coding agents have been so much more effective than general purpose agents and i think a large part of that is it just has access to the terminal, like you said, and that means it has access to everything that you've installed into your terminal.
It can run, so you know it would write code and it can compile the code and if there are errors it can fix it.
It can run your suite of tests, because that's all just in your terminal.
And so that you know, for the idea of what caught me really excited about the Outlook CLI, we're now just turning through building CLIs for the entire, like for the entire business suite.
Slug holding.
Slug.
Also a workday CLI.
I've also done that for myself.
Really?
Yeah, yeah.
Um, we're gonna, we're gonna open source all of this.
Um and like yeah, all the the, I mean they're just they're they're, they're yeah.
CLIs for the business applications.
We would love for someone to run with this and like, build like I don't know, like open CLI foundation or something.
Yeah.
We, NVIDIA would love to support, uh, anyone that's doing this.
Like every dev tool should really have good CLI support at this point.
Like at one point it was, you want your docs to be. accessible by an LLM, right?
You want LLM, good dog.
No, everything needs some CLI tool.
Yeah, it's kind of funny, right?
Computing began with a terminal, with a shell, but we said that it's not empathetic to humans.
So we built these nice user interfaces.
And then now we have LLMs navigating our user interfaces and ironically, we're not empathetic to the machine anymore.
Just give the LLM access to the shell.
One thing that slightly makes it uncomfortable is, why do we have to build CLIs?
Why can't we just expose APIs?
I have an interesting answer to this.
So there are a couple of reasons.
Portability is one issue.
Sometimes APIs are not discoverable or reachable by some types of things.
There's some element of locality, right.
Like the CLI is like literally, you interfacing with your local system, which is a little bit different.
You could still do it by API.
But like there's this highlighting of like, what is the difference between like a CLI and an MCP right?
Like they kind of occupy the same purposes and you call them.
It does something on the system and that's done.
I think that in pre-training, there's just an enormous amount of command line data.
Yeah.
Yeah, like even let's ignore let's, let's ignore RL, like you're doing no harness.
You're doing no harness by string.
Just the amount of like CLI versus API documentation for just like navigating this world of the CLI in your file system through that is just enormous.
Yeah yeah right, i think there's a couple of things too, like if let's say we want to.
So one um, i think your intuition is right.
The cli is just wrapping the api right, so functionally right yeah, and i think it's nice because one, you're being very uh, specific and pedantic, even um of what, and and that's really good because you're describing the problem space uh, so you know what the um I don't know.
I don't want to call it like the space for vulnerability.
You know what network calls you're making.
It's not arbitrary.
And that's not decided on the fly.
That's like pre-decided, which is important from a security perspective.
But then if you were to write a bunch of API requests, you would probably do that.
I don't know.
Would the model like use Python to do so?
I kind of like that everything, like a CLI is just bash because it's ubiquitous.
Like it's just there and you don't have to make sure that there's certain environment variables that are set up.
Like, if your Python version is different than my Python version, we're using the same model to go do the same thing.
Is it going to write like different code?
It probably would.
And so it's kind of nice to go work, right?
We're human as well.
I think just like making those decisions happen ahead of time versus, yeah.
One last thing on this sort of agent, I guess, maybe co-location or whatever you call it.
One pattern I'm tracking for this year.
I always try to think about what's the theme of this year going to be.
Last year, definitely coding agents.
This year is definitely coding agents breaking out of containment into brothering.
So you're going to rent a human?
Yeah, I'm on there.
Are you really?
Yeah, I'm like $5,000.
I'll do anything.
Really?
I think so.
I need my balance from Costco.
But I think the best part is only the agent can book me, you know?
It's very usually, it's just like another labor marketplace.
Mechanical Turk was this.
So I have a weird story with why I did it.
So back to your example of just giving agent access to compute, right?
Yeah.
You guys are GPU rich at NVIDIA.
Yeah.
I hooked up... He's not shy about it.
I have a 24-7 agent running.
I hooked up to RunPod.
It doesn't shut down instances.
And I'm like, I've tried prompting you.
I've given you instructions.
Shut down when you're done.
It's like, I need to keep it warm.
I'll need it soon.
And it's horrible on time estimates too.
Because they realize it's like, yeah, I'll need it in 45 minutes.
45 minutes, I'll shut it down.
45 minutes of human time is actually three minutes of agent time.
So it's like, I'm booting it up.
I'm waiting.
I'll just leave it on all night.
And Modo's good at shutting down after some inactivity.
I had it on my local server, like a little dual GPU thing.
It just stays on.
I have a little space heater at home now.
But careful.
So basically, they don't care about the concept of money.
Just burn it.
I need it.
It's useful.
DGX Spark would be really nice.
I think I'm looking at it as it's super useful for agents because yeah, you buy it once, you plug it in and it can rip.
I'm going to make an NVIDIA ad here. okay the blackwell like rtx 6000 cards pro pro are only like i think it's eight thousand dollars slightly cheaper yeah well it's much it's much cheaper than the data center cards yeah and it's got 96 gigabytes of vram so if you and your your your crew want to go like run a local agent for you you know you you in the home I feel like it's got a significant amount of VRAM.
I've thought about purchasing this and running it in my basement, except my neighbors would hate me.
It's just a single, like, two, three slot GPU.
Yeah, it's a VCIE.
It's a GPU.
You can go buy that.
I mean, the big difference against the RTX gaming GPUs is, I mean, obviously, it's like BlackBull.
It's a pro GPU.
And it has a lot of VRAM, which means you can run pretty large models on it.
You can stack four of them for the max queue in a system because that's a beast.
It's beefy.
You can run, what is that, 96?
Zeke or anything, 96.
You can run a lot of Zeke.
But also, they are slow.
I mean, performance of speed will be somewhat slower relative to API.
Oh yeah, that's true.
So again, big learning.
Economy of scale allows you to do things that allow you to get both speed and throughput.
Like you can run, I'll give you an example.
There's an optimization called wide EP.
I'm not going to go into it fully, but like it featured heavily in inference max for deep seek.
And there's a great set of stories from NVIDIA and from Semi Analysis about why YEP is important.
But for MOE models, it's basically essential.
And you run it, a level of parallelism, the level of scale-up parallelism used for it is 32.
So it goes beyond that eight barrier and it really really, really is important to have that MVL72 GB200 MVLink to serve at scale.
And it's like, I don't remember the cost improvement.
I think against Hopper, right?
Against Hopper.
With this MVL72 system you're getting like 35 times cheaper per token for like a lot of the curve.
Yeah.
Which is crazy.
Yeah.
And normalized per GPU, obviously because part of the GPU is cost or the code that GPU is part of the cost.
One thing I'm exploring is this year is also the year of the sub-agent, where you have the main agent, but then that also kicks off tools, which are in themselves agents that have limited agents.
And such a model context, whatever, right?
Different prompts.
So, for example, one thing that Condition does is before you kick off a search.
They do like a fast context model where you kick off April, you just search across the code base.
How's all that?
That is better than indexing a lot of the times, not all the times.
You should still index for some picks.
But the idea that agents should be able to command sub-agents and probably run them maybe close to inference is why I don't know if that's architecturally possible.
Yeah, we're thinking about that for Dynamo.
That's our big theme for the year.
Because, like, if you can design that into your stuff, then a lot more people will use it.
Right now it's like just kind of theoretical because you do pay a lot of like back and forth coordination costs.
I think you'll net speed up though, right?
Like, even at a basic level, speculative decoding, you're running a small model.
You're running two instances, but it's net.
That is one example, yes.
Yeah.
But this is like a little bit like different with like agents.
Agents, yeah.
This is not spectacular.
I think there's like a summarization of that trend that I like to do or I like to say to my team.
It's like, this is the year.
So there are two things.
This is the year system as model where, instead of having a single model be a thing, you have a system of models and components that are working together to emulate the black box model.
So when you make an API call to something that's like a multi-agent, in the background it still looks like an API call to a model.
You're still getting back to it.
Right.
Under the hood.
Yeah, under the hood, it's like a billion different models.
And that's a lot of complexity, right?
With Dynamo and with other libraries and media, we're looking at how to manage that complexity.
Yeah, it's funny.
We actually, for CES, we just released a model router.
So for DGX Spark, where you can have a local model that's running on the Spark and then also a foundation model, And then the model router decides when to send queries to which one.
So it's no longer this either-or.
It's use the best of everything that's available to you.
You have a good post-training model that's running when you need it.
MARK MANDELMANN- Is there also the breadth functionality of being able to manage the Spark?
Oh, that'd be cool.
Oh, yeah.
I actually have a question I'd like to extend and flip over.
How much longer do you guys think agents are going to be running?
Because that's one thing I've been throwing around.
What happens when- I mean, always on.
It even affects back to the pre-fail decode, right?
Yeah.
Codex is I'd say compared to cloud code, it's much longer at tasks.
Like that thing will like to run six, seven, eight hours.
I'll run it overnight and I'll go back and I have like a little crappy, uh, logging software I use.
And there's just times where it wants to like I'm going to go deep on research and it'll eat up 80000 tokens.
Go on another, go on another.
Just, just eat through tokens.
And you know, that's part of it.
Like at the end, it does, it does hit a long task.
And I think you only see that, that experience, right?
Yeah.
There's insatiable demand for tokens, and every improvement that comes kind of just makes our demand even higher.
It's kind of funny, right?
Like, if you have like a teammate and you ask them to do a task and they're like should I save some effort and not think too hard about this task?
I'm like, fuck no.
I mean, my favorite was like, you can have four shots, right?
Like, the original codex before the app, why do one call?
Like, give it four attempts.
Just use all the tokens like that, right?
Try more.
It's like the meta index, right, is the thing that tracks how long models are able to run.
I expect that we'll just see like log linear, if not log super linear growth.
We will see, before the end of the year, an agent that is capable of running for longer than 24 hours with like self-consistency the entire time.
I would also poke at different domains having different desires, right?
Like at its consumer level, I'm getting slightly frustrated at 20 minutes per basic query.
Sure, you can optimize, you know, six, eight hour.
I don't see myself shooting off many one week agents, right?
Someone doing like okay GPU, kernel research, or medical or biological, like you know, in those domains sure shoot off a lot that take them out of.
So like I think it will be somewhat domain-specific, because you also really need to train that in right.
You know what's funny?
What if I was doing your taxes?
Right?
Like, that's taxing.
Yeah, it's kind of, yeah, okay.
Yeah, exactly.
Get it right.
I wonder if, like maybe you're supposed to say sort of like speculative decoding is like your agent figuring out what you might be, prompting it the next day at night, and like pre-fetching.
Yeah, you can already do that.
Yeah, really?
Branch prediction.
Oh, no, that's too low level, but yes.
Sorry.
Yeah, yeah, yeah.
One question I got to get is so we actually did record a part with the meter folks who sat right here.
Their chart is the human equivalent work hours of work, rather than how long the agents themselves are being autonomous.
And there's a huge difference, right?
Like, human work, five hours.
Agent work, 30 minutes.
Like, it's actually 30 minutes, not...
Yeah, .
So that chart that you see is them estimating what the human equivalent replacement is.
I think actually Anthropic released a more recent chart that showed cloud code autonomy from their production traffic numbers.
And that was 20 to 45 minutes.
That's roughly where we are.
So yeah, that's the sort of realistic thing.
I mean, I do think like there's experimental setups where you can just like sort of Ralph Williams, like just prompt it to keep going when it stops.
And obviously, that can go arbitrarily long.
I feel like from my experience around yeah, I guess 20 to 40 minutes seems right for when I'm using, like Codex or Cloud Code.
But then, like I always try to just like, if I want to spin up like a new, there's a net new project.
I'll often start with Replit.
And, like, it'll, yeah, inferred, I believe.
Yeah, like spin up, like they're new, like from the V3 agent, like it'll spin up a web browser and like click around and discover new bugs and just keep churning.
So I think, like, my longest was, like, over an hour that I had been churning.
I think before we see super long running, I think there's going to be a bit of an efficiency hit.
So
Sure, you can take an hour and go down paths, but you also want to be more efficient.
You want to be smarter in your reasoning, right?
So I think that'll actually go down before we go back up.
You don't want to scale non-optimized systems just for the heck of it.
As much as I love saying use all the tokens, you know, they are expensive.
Like going from dense to reasoning models, that's an added cost, right?
You're paying for a lot of tokens and it doesn't make sense to just scale stuff that's not optimized.
So there's always that little balance.
Yeah.
But, you know, I think you'll see both sides of it.
Yeah, so 2023 was super exciting.
I think if you were in SF, you were like okay, I know this is going to be a huge world-changing moment, but it seemed like, you know, no one had known yet.
And maybe even before, was it 2022, maybe?
Yeah, I would say like Rune had this tweet where, like everyone was in SF from like 2021 to 2023, like understood what it was like to be like RD.
Totally.
Yeah, 2021.
That's when I made my first OpenAI account.
Yeah, it was crazy.
And I remember it was so funny because at the time, SF had not been doing well.
So pretty much what it felt like was the concentration of founders in the city had risen because where my neighbors were used to doing a bunch of stuff, those people had all left.
So the only people that were still in the city were people that really wanted to build.
It was cheap tech.
Yeah, it was also way cheaper.
I feel really bad anyone who is trying to get rent now.
But Celo, they had a huge office.
So blockchain, it took over the old Casper building.
Yeah, they had the showroom and they had the, I think it was like the back warehouse.
It was a huge office.
It's right across from Opening Eyes in New Orleans.
Yeah, it was in the original arena.
I named the arena because of it.
FRANCESC CAMPOY- Yeah, yeah.
And so it was really exciting because like RoboFlow, I think, I forgot.
MARK MANDELMANN- Mitlify.
FRANCESC CAMPOY- Yeah, Mitlify.
Brad was there.
You guys were there.
I remember that was actually, it was there that you bought the ai.engineer domain.
MARK MANDELMANN- Yeah.
I didn't know what I was going to do in AI.
I wanted to do something.
And it was kind of this, it was a really fun moment where we were kind of all in this solo space.
And I don't know, it was a really cool community, especially being so early.
Yeah.
And then you got me early cruise access.
Oh, yeah.
So there was a golden period of time that both cruise and Waymo's were just free.
Yeah, always.
I mean, they're so back.
Celo's opened again.
So nature zoops.
Zoops is doing it.
Zoops and RoboTaxi.
Yeah, so.
Totally.
But yeah, and so it's actually really cool that you guys have this studio so close to Celo.
This rock climbing gym right around the corner.
Oh, yeah.
So yeah, it's an awesome block.
Well yeah, just a little bit of San Francisco.
But I do think like one thing I try to do with the podcast is like bring like what it's like to be in San Francisco to the rest of the world.
And also just like maybe give El Tepo Taqueria.
Yeah, my favorite tacos in the city.
Yeah, steak and shrimp.
I know, it's very good.
And I guess what it's like to be in San Francisco, I think, is just everyone seems to be super supportive.
Sometimes I feel like the city believes in you more than you do.
And even, I don't know if you remember, but I remember posting my first blog post.
And I had met you on Twitter.
And you gave me like an hour of your time super randomly and you kind of coached me through writing content for developers.
And I was trying really hard not to come off salesy or plug myself.
And so I kind of stripped all personality out of the blog post.
And you brought that out.
You're like, people don't, it's okay to talk about what you're doing.
Like you don't have to be weird about it.
And I remember just that.
I think that really helped me kind of figure out what our voice is and not shy away from it.
And so always really grateful for you.
Now you inject your voice into like everything.
It's actually a huge advantage to be very genuine about what you care about.
Yeah.
Imagine some random person DMs you and is like, can you give me feedback on this blog post?
And it's pretty boring.
And you're like, fine.
He looks interesting.
I'll just do a Zoom call.
And then you meet this guy.
He's so energetic.
Just be right there.
I think people are trained to write a certain way in school and they never see there's a broader world out there.
Writing is thinking, and everyone thinks differently.
So you might as well just write your way.
Cool.
Well, thank you for indulging with us.
Really broad, breaking discussion.
But I love like you guys are sort of like the young faces of NVIDIA, with so much energy but also a lot of technical depth.
And I think people learn a lot from this session.
So thank you.
This was awesome.
Thank you guys.
Thank you for everything that you've done.
Yeah, the podcast, all the above.
And see you at GTC.
Yeah, we're looking forward to it.
Yeah.
Cool.
That's awesome.
Thank you guys.
Thank you.