Hey, everyone.
Welcome to the Latent Space Podcast.
This is Alessio, founder of Kernel Labs, and I'm joined by Swix, editor of Latent Space.
Hey, hey, hey.
And today we're in the studio with Jay Cooper of Railway.
Conductor of Railway.
Conductor of Railway, yeah.
Choo-choo.
Do you actually have that anywhere on your business card?
Well, I don't have a business card.
We're not that big, yet at some point I will.
I got handed a nice business card from the super micro folks and I was like damn, that's actually like pretty official.
They're coming back.
Business cards, yeah.
Yeah, they're cool, they're hip, they're jiggy, but yeah, the whole conductor thing, like we call some of our volunteer moderators conductors, you know yeah, so It's a good one.
It's a good one.
Like we're trying to figure out what we want to call each other internally.
And there's like varying levels of thought.
Some people are like, oh, it's super cringe.
Like just don't like you don't need a name for like, you know, people internally.
And some people are like, oh, yeah, we want to call each other like this thing or whatever.
I was like, we still don't have a really good one, you know.
Yeah.
We've got like new rail recruits.
We've got like trainiacs.
We've got like, nothing's like really.
I like trainiacs.
Yeah.
Railwayians.
Okay.
So, well, for those who don't know what is railway, let's give people a crisp definition up front.
Yeah.
Railway is the easiest way to ship anything.
You just go to the canvas or you talk with Claude and you say deploy Postgres instance, deploy my GitHub repository, run this code, et cetera.
Right.
And you'll just be up and away to the races.
Right.
Yeah.
You've got nice animation on the landing page.
Oh, well, thank you.
None of my work, by the way.
They don't let me touch any of the design stuff anymore.
But yeah, we want to make it really easy for not just to like deploy things but for you to almost like evolve applications over time.
Like we believe that most of the tooling right now is kind of like stacked up, like you're stacking entropy on top of entropy on top of entropy, right.
So you have like Docker and Cube and then like Ansible scripts and all of these other things, right?
And if we can kind of like version all of your software for you and keep track of all the changes, then we can make it actually trivial for you to clone environments you know, fork into a parallel universe, get copies of like production data, get copies of like any of your services, make those changes, validate those changes, collapse it in, without kind of having to just like reproduce everything across a staging environment or all of those other things.
Right?
Yeah.
Amazing.
One thing I was looking at your background, right?
Like Bloomberg, Uber.
There's nothing immediately that stands out to me as like okay, this guy's going to found like the next great platform as a service.
What prepared you for real way?
It's almost like a curiosity to just like ever go deeper.
Right.
And so, like you know, started out on like front end stuff, you know like working on the like Wolfram, like Web Mathematica.
Yes.
Like porting it over there.
And then you know briefly, moving to Bloomberg and then moving towards Uber and like distributed systems and kind of like taking all the jump bikes kind of systems and moving them over to a distributed system built on top of Cadence.
It's like the pre-temporal.
Yeah, the pre-temporal, temporal.
Which, by the way, I'm happy to talk about pros and cons.
Yeah, I think like it's like – But let's do the roadway story.
And so like it's just been a continual step of like I –
I want this experience, whether it is like walking up to like a bike and just unlocking it and like having it be like frictionless, to like work, or whatever.
And then like necessitating the like depth required to go in and make that happen.
Right.
Like a lot of the work that I do and a lot of the team does is like it's all in service of that experience.
Right.
And like we fundamentally don't care, like how deep we have to go whatever, like we will swim to the bottom of the swimming pool to go and get the experience.
Right.
And I think that's what a lot of of, you know, kind of the trajectory was right and so it's not like i have a physics phd or whatever i did like an ecs degree you know it's just it's always been about just trying to figure out that next step of like how do we get there right um and that's like what's led to you know starting railway for that experience and then like moving all the way to bare metal data centers right like you know i was adding patches to the kernel this week right just to like get the experience there because i'm like see it and, like, how much better it can be, right?
You added patches to the Linux kernel this week?
Yeah, well, not upstream.
That's a flex.
Railpack?
No, this is different.
This is the OS on top of Railpack.
Yeah, no, this is, like, this is the actual kernel.
But it's always literally just what do we have to do to get that experience and just like figure it out, right?
Like, anything is figureoutable, right?
Like, you'll just figure it out, you know?
Would you send the patch upstream?
Or is it just because, like, it doesn't fit?
Maybe.
It's like, we have to work out the experience for us internally.
It has to do a lot with... the storage layer that we're building for some of the agentic stuff.
So maybe it'll be useful to people upstream, but it's deeply useful for us internally.
I mean, you mentioned open source before, so I'm just kind of curious about...
How you think about starting from open source and then coding agents.
Let you do a lot more from forks of it.
I think the it's funny because, like I think, GitHub's original sin is that it's like almost a series of broken pointers.
It's like you have essentially this thing and then you clone it and then, OK, great.
Like I just lost that whole upstream.
Right.
How do we make it trivial for people to modify really, really small pieces of it?
Right.
And you did like.
You think of Git almost in this, like discrete sense of like I've either made a change and I've merged upstream or I haven't right.
What would it look like if it was like percentage based or a little bit more non-deterministic, or anything else like that?
More of like a stream of changes that you kind of like traversed as a user, more as kind of like a percentage of this is rolled out in general and it's been rolled all the way up right.
You know we have the open source, like kickback program and allowing you to deploy those templates, because we almost want to make it trivial for people to like go and version these shards over time.
It solves like a really, really large problem in terms of authentication authorization, security.
Like you know, NPM has that thing where you can almost define hey, don't take any new packages, or whatever.
Like the ideal end state is actually like you should roll out progressively to the users who have the minimum impact zone for any of these things and just continually roll up.
Right.
Like JP Morgan or something else like that should probably be the last one on the patch line for that.
Right.
For all of our sakes.
Right.
Like because we have all of our, you know.
Money or livelihood, all of those other things.
It's okay if, like Johnny Vibecoder, gets like a broken patch or something else like that, because ultimately there's so much entropy in the system that you do have to.
You do have to roll, like rubber has to be rolled at some point.
Like you have to test at varying levels, right?
So yeah, a little diversion from wherever we started, but you know.
So I just wanted to pull up this glorious chart you say, which is basically your usage or number.
Daily signups, I think.
Daily signups?
Yeah, yeah, yeah.
So you started six years ago and like a slow grind.
Slow grind, yeah.
And now obviously you're on a rocket ship.
You say don't doubt your fight and don't quit.
But like maybe, if you want to pick out like certain points that were like sort of key inflections of the company, that might be fun.
Oh, yeah.
Well, I mean at the start.
It's basically like how do you get your first hundred users?
Like hell or high water.
Right.
And so, like starting in, you know we had a website and we had a support link, and the support link was the Discord channel.
And you just showed up there and I had notifications on.
I had two monitors.
I had the monitor I was working on and then I had the other monitor.
And if anybody came in, I was like.
Oh, hey, how's it going?
Like, you know, and it was like super rare or whatever.
So trying to get those initial like first hundred users to like actually kind of come back to it.
And that's, I think, where you can kind of like see the really like in between January 2021 and 2022, like probably the middle, like they're kind of right.
And that's like the start.
And then you ultimately end up building a consultancy factory of like users wanted all of these things in general.
And so you kind of have to go back to the board a little bit and be like well, what is the actual product offering that I want to build on top of these?
And I think like incidentally, it's funny, like I think VCs really want like charts that like always look like this or whatever.
Right.
But I think in reality, you actually don't want. charts that look like that.
Most companies, I think, or at least for us, there's been periods of like, expansion of like.
Okay, we're going to go and add these features to like.
Go in and test these use cases.
And then there's been periods of like compaction where we're saying like okay, how do we have?
If the experience we have is really, really good, how do we make it significantly better?
Maybe we're even stripping out features that don't fit our ICP anymore.
How do we go in and do that?
And I think throughout this whole chart, you can see a lot of those things.
The boom in the 2022 to 2023 is we had a free tier and everybody under the sun was using it and all those other things.
A lot of Reddit bots and stuff.
Yeah, right?
And like I think there's a thing that's really really tough to like teach people or tell people about is like when you build an open product on the internet where anybody can sign up.
The internet is a horrible place that has like so many things like.
You kind of go through these periods of, like, well, how do I reach as many people as possible?
And then like, how do I fit in exactly the use case for the people who are really really going to matter and are going to be really really excited about specifically this thing?
Right,
And we go back and forth internally.
And then there's like what is that?
A two-year period of like making the actual business work in general, right.
So like free-tier era losing what I think half a million dollars a month on.
Like you know, we're making On like a 20 million bank account.
Yeah.
Yeah.
And like a 20 million bank account with like I don't know, like maybe 50000 a month in revenue or something else.
Like that is horrible.
I don't know.
But anyways, you have to kind of go through and be like cool, like we have an experience that people love in general, but like the business has to work.
Right.
And I think there's like, I guess, two schools of thought.
You can continually run the horrible business all the way up in general and have bad margins, or you can actually go back and kind of make it work right.
And for us, you know...
We've always really wanted to have like a super lean team, right?
So we're 35 people right now.
You know, it's very, very small.
We have like, what, 3 million users?
Supporting 3 million already?
Yeah, yeah.
Holy shit.
Because we're adding like 100,000 users a week right now, right?
So it's like, it's growing really fast, right?
But we've always wanted to have a really, really lean team.
Like we didn't want to, just like add headcount for the sake of headcount, just like throw bodies at these problems.
We want to build like systems, right?
And it's really really hard to build systems when you're kind of in that expansion phase because you're just adding stuff to the system in general because people are asking for it or things are breaking in general right.
We basically were like, all right, Like, you know, we're going to we're going to cut it for now.
Like, we're just we can't support this.
Like these free users that like we want, like we want to reach as many people as possible, because we believe that you know software.
Is this really really important thing where, if you can kind of like create something, it's become really difficult to create things in a physical world?
So it's really important to make it really easy for people to build things in a virtual world so that people have access to creation.
Right.
And so we want to reach as many people as possible.
But there's kind of like legs on that journey.
So we basically had to kind of close off the free free kind of users for a little while, rebuild the business, make sure that worked in general.
Right.
And then I think you can kind of like see The building of that in general.
Right.
And then I think you see kind of some divots in those charts.
Right.
Like if you actually follow between, I think, 2025 and 2026, it's either summer or winter.
That's basically it.
Right.
And either people go on holidays with their family or they go on a holiday.
Oh, it affects that much.
Yeah.
Yeah.
Yeah.
Well, because it's like.
It's kind of B2C.
It's kind of B2B in general.
Right.
And so you have a lot of these users where like, they're shipping constantly and then you know they'll kind of like stop or whatever.
Right.
And so maybe for summer or like, maybe like our activation curve is like now, we see a lot of people like activating in the weekday.
Right.
Because we have a lot more like business users in general.
So that gets a lot less active. sheer, so to speak, right?
And it kind of like smooths out over time, you know?
Yeah.
Is there any point at which you started prioritizing AI developments or agent development?
I think, like, so we've prioritized almost, like, Agentic as, like, a top-of-funnel thing.
And probably over the last like six months, we've probably deeply prioritized like Agentic as a mechanism to go and build and deploy things, just because we believe fundamentally like the the curve is so sheer and that is the way that people are going to go and build and deploy software.
It almost fundamentally doesn't matter if this iscom or not, because we're all on the Internet now anyways.
If agents are going to go and deploy a bunch of things and we hit an inference wall at some point then, like at some point, we will go in and fix those problems.
But like that will be kind of the dominant species over the next like 10 years is we've moved from assembly to C to C, to JavaScript, to now like words, right.
And you're going to need to be able to close that loop, right?
But that's where it goes, you know?
When you say this is .com, do you mean, like, buying the domain or...
No, no, no, no, no.
I mean like actually, just like you know, they had a bunch of run up in the dot com era for companies because they were like the Internet is really, really important.
And then you hit kind of like bottlenecks fundamental laws of physics, math didn't work, all of those other things.
And everybody kind of like, you know. went back down to the earth.
Right.
But at the end of the day it didn't matter, because the Internet is like so, so impactful for our lives that if you operate on a long enough time horizon that you should be like, you should just build these things anyways, because you can see where that's going.
Right.
And that's where I fundamentally believe a lot of the agent stuff is right.
And we can talk about a little bit of it later.
But you're going to get to a point where you're running thousands of these agents like in parallel, right?
Like, one, what's the inference cost for that?
What's the compute cost?
How are you going to make that efficient?
All of those other things.
But, like, two, how do you coordinate all this stuff?
Like, we have issues coordinating humans in general, right?
We don't have good tooling for that.
And now we're starting to figure out, it's like, oh, like, how do you get... agents to coordinate?
How do you go and get them to be able to like safely version changes, or like for them to know when to like, put their hand up to get somebody to intervene right?
Otherwise, it just becomes like an interrupt factory that's like crazy, you know?
Well, so maybe we'll go right on the technical side of things.
Yeah, yeah.
What are the core, like infrastructure or architectural beliefs of real way that allow you to do what you do?
Yeah, I think the primitives matter a lot for us, like a lot, a lot.
We need to be able to do network compute and storage and orchestration all kind of around it.
You kind of need control over a lot of those things.
We've talked a lot about how we don't really use Kube like Kubernetes, because we want the higher order of control to be able to go in and place workloads in very, very specific places, right.
The reason for that is, like you know, it's kind of the thing we talked about previously.
But, like you have to be very, very efficient with these agents like memory reuse, all of those other things, or you're going to massively, massively blow up your cost structure.
Right.
Also incidentally, being able to rack and stack your own servers and and build your own metal.
It unlocks a level of like performance one, but like to cost where you can say oh, those experiences that you want to offer, where you're running a thousand agents in parallel, are not like massively cost prohibitive.
Right.
Because if you look at just like token use right now or compute use or anything else like that, those things are blowing up massively.
Right.
Over time, those things are going to have to get a lot and lot more efficient.
You can get a lot of almost like back of the napkin balance sheet margin whatever you want to call it to to kind of make those experiences like solid by building your own metal right.
And so kind of to the earlier point of like we've always tried to go a little bit deeper every time to make that experience.
It's all in the service of offering that differentiated experience to as many people as like humanly possible you know.
Yeah, you have a data center in Singapore.
Yeah, so we have two in every other region now.
Singapore, we're adding a second one in Q3, so, yep.
So, like, what's it like?
I mean, I've never built a data center.
Yeah, well, we'll have to, like, go to one or whatever.
Is it just go to, like, Equinox and say, hey, I want some stocks?
Yeah, so, yeah, I mean, I can run... Equinix.
Equinix, yeah.
Yeah, Equinox.
I mean, you can put a data center in the Steam room and get nice and hot or whatever.
But yeah, you basically just go and you say, hey, listen, I want power and I want a cage.
And they're like, great, here, this is what it's going to be.
And then...
You rent the cage for a period of time and then you have to fill the cage with racks servers and then hook up internet to it, right?
That's realistically all there is.
And then you handle everything else, right?
Yeah, you just handle everything else, right?
And what's the math versus, obviously, the clouds?
Yeah.
Our payback period when we go to Metal.
If we rent it in the cloud, our payback period is about three months.
It's crazy.
It's nuts.
Yeah.
And that's like four years worth of like depreciated hardware.
Right.
And so I think it's like you're going to see a lot of this, almost like compute crunch, so to speak, because a lot of the hyperscalers are buying up a lot of stuff.
Like we're working directly with and resellers, and directly with people who are building these machines, like Supermicro Dell, all of those other things to go in and get these things working.
But upstream, there's a bunch of supply stuff.
It was funny because when we raised our last round, in between...
Basically deploying the capital for the servers and actually I think even now the amount of money that we've raised is less than the amount of money that we have in the bank plus what the value of the servers are, because the servers have actually appreciated in value, because RAM has gone up in general right.
So it's kind of nuts just in terms of like how valuable hardware is.
And all of this stuff is, right?
If you look at especially a lot of the like hyperscalers, like what they deployed, like 80 billion of like capital expenditures, like this year.
And, like, into next, it's going to be, like, more in general, right?
There's a massive, massive scale, like, infrastructure build-outs.
And you can look at that and be like wow, that's crazy that they're spending like way more than the Manhattan Project.
But, like again, if you go back to every person that's going to run, you know dozens hundreds whatever, of agents in parallel.
You should spend more than that.
You have no conceptual idea of how much compute is required to go in and make that experience happen.
Even if you're deeply efficient, even if you're sharing resources, even if you're doing all of these things correctly, and that doesn't even count inference.
How do you plan on the build out?
Like.
I mean, the growth chart is so vertical that you know like are you usually 100 utilization rate as soon as you're live with these tracks.
Like, how far ahead are you?
Like, we still maintain, like, cloud presence for, like, bursting, essentially.
And so what we can do is, you know we work with AWS and GCP and a few of those.
You know other clouds.
Like we can just rent and then the moment we kind of get space or power or whatever, you almost just like, compact those off the cloud right.
Like we, Because we started on the clouds and then we built a system to allow us to migrate to our own metal.
And so there's nothing that says you can't just continually do that again, which is exactly what we do right now, right.
And so we never want to be in a spot where essentially we are, you know, compute constrained, right?
Right.
And at the start of the year like we actually got to a point where we were compute constrained because the one upstream provider that we were actually working with wasn't able to give us quota at the rate that we needed to.
And the hardware was like slower.
Right.
And so we had to do a bunch of different stuff.
I spent a weekend rebuilding our entire like network, like overlay, essentially so that we could straddle five different clouds.
Right.
Yeah.
Oracle, AWS, ourselves, GCP and like one other one.
Right.
And we can do more than that now.
Right.
But you know, we got into a spot where, like we were just trying to like pack instances tight because we couldn't get the amount of compute that we needed.
Right.
And it was really unfortunate because, as a result, like some of we had a few like reliability kind of things which are now kind of past us.
But it was all a result of this kind of like.
There was a tweet that I made where I got in trouble because I was trying to point it out but I accidentally caught the Superbase folks in the crossfire.
But the tweet was about.
It's really, really difficult and it's going to become more and more difficult to acquire compute at the rate that these models need to acquire compute right.
And we got bit by it, which is, you know, fair and reasonable in the karma scheme of me, you know, trying to point it out.
So, yeah.
How do you think about pricing, knowing that you might not have Eurometal available at all time?
Like are you pricing, assuming that you'll need to like, pay yourself extra margins if you had to end up going in the cloud?
Because we've built out our our metal data centers, like our margins on metal are like quite high for the like 70.
And so we can actually deeply subsidize the cloud business if we want to scale at a reasonable rate.
And so we have a few different like It's actually very fun from an operations perspective because you have a few different levers on how you can go and scale it.
You have the metal, which actually makes your margins.
You have the cloud burst, et cetera.
You have debt you can use to buy servers in general.
So it's a very interesting operational problem to basically say... okay, we have this much cash.
Oh, and then you've obviously venture capital that you can raise on top of it, right?
And so you have this much cash.
How much money should we raise?
How quickly can we go and deploy it, et cetera, if we can scale revenues basically as quickly as we can scale compute, provided we continue to make it trivially easy for people to go and build and deploy.
The faster you can close this loop...
And and the more operational excellent you are with the capital, like just the faster your business.
It's just a basically like straight linear kind of like yeah, deployment rate on some of that stuff.
Yeah.
I think in first startups, raising debt is a tool that people don't utilize enough or know enough about.
What can you tell us about that?
Is it secured against your CPUs or what?
Yeah, it's just secured against our hardware.
What rates do you get?
We just pay like prime at whatever it is, plus we can refinance any of the debt as it goes down.
The terms are pretty good from that perspective.
I think like The unfortunate thing is Twitter has no nuance or whatever, so they're like venture debt bad or whatever.
It's like, well, no, as with all things... It's not venture debt.
Yeah, yeah, or whatever.
It's data-centered debt.
Yeah, it's data-centered debt, right?
But yeah, I think there's specific tools in specific areas where you can be very, very deliberate about not just using one specific tool as a hammer, like venture capital is a hammer for everything.
You just have to kind of like go out and explore it and figure out how it kind of like Yeah, VC is the most expensive financing you can get.
Yeah, yeah, yeah.
I think.
Incidentally, I think also people think about VC, completely wrong from a raising capital perspective.
Okay, tell us how it's...
Yeah well, I think most people are like okay well, how do I raise as much money as possible from whoever is probably the best I can get at that point in time?
And I think that's kind of close to right, but I think what you should be doing, or at least what we've tried to go in and do is like... try and figure out what almost unfair advantage you can buy with that equity because it's the cheapest equity or it's the most expensive kind of equity you're going to give away at that point in time, assuming your company is going to get better and better and better.
And how do you use that to like go in and work with somebody who is stellar and who's going to go in and compliment you right?
Like, you know, Yeah, like Series A. So, lucky.
Yeah, right?
Like, you know, great.
I've never started a company.
Raised from lucky.
He's got good advice.
I can text him all the time.
He's really fast, et cetera.
Like, awesome, right?
Then you kind of like, move on and you kind of like, you know, worked with you know John and Jordan in Jordan at Unusual right.
And they were like, yeah, you roughly know what you're doing in building a product.
Like, we're just going to mostly, like, leave you alone and be totally available for advice.
Amazing.
Awesome.
Get to Series A. Business is a total, you know, operational tire fire, right?
Because we just don't know how to scale a business, right?
Go and work with Erica and, you know, Jordan's over at Redpoint, so... bonus.
Like, you know, we get to work with them continually.
Right.
And then now moving into, like you know, raised from TQ and FPV, like we're moving into the enterprises now.
Right.
And like feeding into air.
Right.
So every step of the way we've kind of moved towards like who can we partner at this specific time?
Who's going to help us unlock that next section of the journey?
Because Guess what?
I just I don't know enterprise sales.
I can roughly like, eyeball it and be like yeah, as an engineer, I think these are the kind of features that we're going to roughly go in and need.
And we have some wonderful people who are going to help us internally.
But you really want to work with those people, like at the boardroom dynamic level are going to be like oh yeah, we're all aligned.
And that's obviously what we want to go in and do.
And we can spend our time basically saying how do we, how do we win this, versus like bickering about strategy.
Right.
No, I just had to pull up some beautiful data center charts.
Yeah.
I feel like you've done others.
I just couldn't find them.
Well, these are good.
I mean, like, they all kind of look the same.
Like, the server's in a rack, right?
Look at our box.
Yeah, exactly.
This is our box.
Such a gorgeous box.
Yeah, it's like, do you want to see more racks?
It's like, oh, yeah.
It's like, you know.
I want the Jay Cooper signature edition.
Yeah, yeah.
We actually have plans internally.
Yeah.
So it'll be fun.
We've got a few different promos that we're going to do and like stunts for the year.
So those will be fun.
Yeah.
You had a tweet about data centers in space just before we wrap this section.
Yes.
Why no data centers in space, man?
Why you hate so much?
Okay, so it's not no data centers in space, because actually I think like my hot take is like I think this is solvable.
I've just never seen anybody solve it, right?
Because you need to, like... No, no, no.
You said, how are you going to dissipate that much heat in a vacuum?
You're making a physics claim.
Yeah, yeah, yeah.
Well, because I haven't seen anybody like prove how you're going to go and dissipate that much heat in a vacuum, right?
Like it doesn't mean that it's not possible.
It just means that like nobody's kind of put it up there.
Pardon?
Astrophage.
I don't know what that is.
The Martian thing.
Okay, you're very lucky.
Okay.
Yeah, that's fair.
But yeah, I don't know.
I mean, it could work in general.
Right.
But I think a lot of people and I think incidentally, this is probably what you have to sort of do is like they're putting almost the cart before the horse, is like oh yeah, we're going to put data centers in space.
It's like, OK, but how?
It's like, well, we have some period of time to basically figure it out.
Right.
It's like, it's like, you know, in the Martian, where they're like oh, how are we going to like intercept?
Yeah.
Oh, okay, right.
It's like, how are we going to do that?
It's like, well, we'll figure it out.
We have however long to go in and figure that out, you know?
Yeah, yeah, yeah.
Making a bet on like human invention is weird, because you just have to blind trust that it can be solved.
100%, right?
I feel like physics and there's like some first principles, bounds that you can put on like maybe not.
Yeah, I know, right?
Maybe you're asking to travel time here or break some fundamental thermodynamic law.
Yeah.
And I don't know how VCs do this incidentally too, because it's like how do you know what's like basically not possible, and like is a grift versus like is possible, but like sounds completely insane right, and you're like, oh cool, like you know, we're gonna put data centers in space.
It's like okay, coin flip as to whether that's like one or the other, you just don't know, I guess.
And I guess you'll know in like 10 years.
Cool.
That's one cycle.
Okay.
Moving back to agents.
I think the branching that you do, the fast spin up and orchestration, it's kind of like the pre-work. that happen to be exactly what agents want?
Yeah.
What do agents want differently than humans?
What do agents want differently than humans?
I think they want the ability to version things.
So it's not actually that different.
There's just... almost like slight deviations in terms of how it kind of materializes, right?
So agents want a way to be able to go in and test changes incrementally, right?
Like we have feature flags as like engineers or whatever, right?
Like, is there any reason why they can't just use feature flags, right?
Yeah.
I don't think so.
I think there's ways that you can just go in and do that.
They want version control.
Is there ways we can use Git or not Git?
I think that one is realistically completely up in the air.
I do think that's something Ultimately outside Git will emerge in terms of how we're going to go into version, a lot of these things over time.
They need observability.
You need to be able to go in and essentially query what happened at what point in time, which steps failed traces logs metrics, all of those other things.
They need network compute and storage.
They need the ability to write files, save files, iterate on files snapshots, file system, all of those other things right.
And so I think a lot of the stuff that we roughly needed is very, very kind of in line with a lot of the stuff that agents also need right.
And so the branching and forking stuff, it's not different.
We're just... moving 1,000 times quicker than we used to.
Some of these things look like you really need something massively, massively different, but it's just.
You need something massively better than what currently existed.
You need orchestration. you need something massively better than Q, right?
You need like networking, you need something probably better than Envoy, right?
Like, and it just goes all the way down the stack, essentially in terms of well, if the workload profile doesn't change so much as it gets like massively, massively compressed because you need to do thousands of these things, what assumptions change, right?
Like, that CD is going to melt, right?
Like, you know, you need to replace it with something, right?
And then I think you can go all the way down the stack and basically say okay well, that part has to change, and that part has to change, and that part has to change.
And the interesting thing about the kind of like super exponential curve is that you have to build your systems in such a way where You can rip out those parts at any point in time because a new bottleneck might emerge, because you know you start getting really, really good at like parallel agents.
Right.
And then that's that's kind of where the new bottleneck is.
Right.
And that breaks a different part of your system.
Right.
So I think it's very much like similar kind of stuff, that kind of like humans have needed.
You just need at a 1000x scale, right.
So, like, how do you do code review in the age of the agents, right?
I guess this is more of a question.
You throw more agents at it.
We don't.
Yeah, right?
But then, like, who reviews things for, like, CVEs and, like, all of those other things?
More agents.
Right.
And then that's how we hit the inference wall at some point.
Right.
And you can continually throw agents and agents and agents at that problem.
Right.
But like you know, I think there's, I think there's a limit to like the amount of agents you can kind of throw at a problem.
You already had a CLI before it was cool, I guess.
CLIs have always been cool, by the way.
How has the shape of what you're exposing changed, if at all?
Yeah, so I think the CLI changes because the way that we think about this is like how do you give Claude or Codex or chat or like whatever, like any of these models, almost like, And like a CLI is a single command when you think about it, right?
It's like, okay, well, you're going to do a deploy or whatever, right?
You're going to get logs, you know, whatever, right?
Like things that were prohibitively annoying to humans are not actually prohibitively annoying to agents.
They're really, really nice, right?
And so, if I wanted to hand you a CLI and I said, hey, guess what?
The CLI has...
40 arguments and 600 flags.
You'd be like, wow, that's crazy.
Like, I'm never going to use all those things in general, right?
But you hand it to an agent and you say, hey, there's 40 arguments and 600 flags.
He'd be like... oh, yeah, this is excellent.
I have so many handles that I can go in and kind of work on with this, right?
And so I think incidentally, if you're going to go in and try and expose things for agents over that mechanism, you want to just basically have as many handles as possible where they can get information, query additional dynamic information and then see how it can close that loop as quickly as possible.
Most of the problems right now are actually just how do you close loop as quickly as possible?
Where does the agent get stuck and how can you go and remove that?
That's why incidentally, telemetry is very, very important, because if you can tell where the agent gets stuck From the CLI and you say hey listen, like 12 of people are actually getting deviated from the happy path because of this thing.
And now I go and add this arg and that drives it down to 2%.
You've massively increased the like rate of the loop closing for a lot of people in general.
Right.
So that's kind of the way that we think about not just the CLI, but every point in the dashboard.
Right.
Like it is a user journey from.
I hear about Rayleigh.
I go and get something deployed.
I get my first green build, whatever, aha moment.
I see an endpoint.
I see some logs.
I see whatever.
And then I go in and iterate, right?
And then I go in and iterate.
Loop is indefinite and infinite until the end of time, right?
It's basically like user wants to deploy a new thing.
User wants to deploy new Postgres instances.
User wants to change their code.
User wants to iterate all over time, right?
And so if you just focus on a lot of those iteration loops and figuring out what's blocking that loop from closing as quickly as possible, like
One of the things we talk about internally is you never ever, ever want to be waiting on compute anymore.
You always want to be waiting on intelligence, right?
And if you're waiting on compute, there's a bottleneck that needs to be destroyed there because at some point that bottleneck will be so so, so large that some other workflow will kind of emerge to go in and change a lot of that stuff.
And I think incidentally, like you know, we've built a really, really awesome product where you can push code and then you build the code and all those other things right.
Like push pull, whatever kind of like loop.
I just fundamentally believe it's going to go away right.
Like it's.
We're going to get to a point where You make a small change in production that changes version across your entire kind of infrastructure.
You're working alongside.
You know copy and write versions of your database, all of your infrastructure.
And then you merge it in and instantaneously it's live.
Because that's the holy grail of loops.
But that push-pull rebuild thing is a point of friction that we're removing entirely from our loops.
Yeah, it's incredibly fast.
So if anyone hasn't tried it, that fast feedback is great.
You know, my hot take is that.
You know, Railway was kind of famous for its canvas which sort of visualizes your infrastructure unless you manipulate it visually.
But that was for humans.
And actually now for the next phase in growth, like really CLI is more important than Canvas, which is what you were famous for.
Yeah.
So I think the Canvas is funny because like, It's actually just a mechanism to show you changes over time.
But I think you're totally right in the sense that, like we have previously used it a lot as an input and its goal moving forward is actually a lot more like an output.
What I mean by that is.
You would go to the canvas and you'd make some changes and all these other things, whatever, right?
And you'd see them and, you know, your agents or your infrastructure would evolve over time, right?
Now you just have a bunch of agents that, like they have access to CLI and they can go in and make those changes in general, right.
And so the canvas actually, instead of becoming this like input thing where you're like, oh cool, like how do I go in and make this happen?
It's actually just more of an output thing.
It basically says what information...
Yeah.
What information does the human need at this point in time to make suitable decisions about?
Like control requests of.
Do I approve this?
Do I not approve this right?
Like that's realistically all the canvas becomes at that point in general, right?
And also a way – and I think this is important and I think this is – lost on a lot of people who are like building some of these, like Canvas experiences.
It has to be almost like an anchor for your context.
It has to be like a port in the storm.
It has to be like.
You have to think basically about it as like layers and like a file system, almost to like.
Get to the next spot right.
And so you have all your infrastructure and like this is why the canvas starts is like it's just a project, right.
And then you have a drill down chart, right?
Like it's like I'm breaking down into these services or this like section that just is like a function or code or anything else like that, because you want to actually be able to represent, you know, The entire thing, not just in your head, but in this, in this canvas, so that other people can also get that representation, so that they can think on the same wavelength as you, so that they can move as as quickly.
Right.
I think a lot of orgs, especially as they scale, they get in trouble because all that context lives in somebody's head basically.
And it's like, oh, how does this microservice work?
It's like.
I have no idea, go ask this specific person.
Then you have entire categories and classes of products that are built around.
How do you do context discovery at all of these things?
I think a lot of that stuff gets melted in terms of if you can have a really solid hierarchy and you can infinitely nest services, infinitely less nest code, infinitely less context, infinitely nest all these things all the way down.
That's what allows you to build these structures up over time.
And I think it's also what's going to allow us to like build I've written a bit about this like these, like hyper structures, like things that are way, way bigger.
And like, you know, you look at the Golden Gate Bridge and you're like, how, how do we build that?
Like, you know, there's that whole meme of like, oh, how do we build this?
Like, we lost the technology.
We don't know how.
We don't know how anymore.
Right.
It's like well yeah, I mean to some extent yes, because a lot of the coordination that we do, that built those things like has evolved right.
And like has changed and there's new things that we've lost, almost like some of the art of like building that structure as we've just like jammed everything into Slack right.
And we're just like, everything happens through Slack and it's just- in Discord.
Yeah, well, it's the same point.
It doesn't really matter.
It's just like message passing and interrupts message passing and interrupts message passing and interrupts right.
So you're arguing that there should be something better, more structured?
Than Slack?
Yeah.
Yeah.
Okay, for sure.
I think Slack, and incidentally, I think Discord's awful too.
This is the equivalent of my mom test, right?
Like, what have you done that has your solution to this?
So internally we built that a tool called Central Station that allows us to go in and aggregate all the context from all of our users.
So every piece of feedback, every piece of customer support, every single thing like that, gets aggregated into what we call clusters.
If you have an incident brewing or anything else like that, now we can go and determine how many users are affected, all of those other things, et cetera.
And then we can actually break off a discussion based on that.
And I think a lot of that is actually a lot more and more helpful and more correct in terms of, instead of like having just these like long running channels where you're just like, which channel should I put this thing in?
Right.
Like if you can dynamic aggregate that information and dynamically route it to the right person based on the context.
Right.
We know we know internally like these four people are pretty close on networking.
Right.
And so if we see like OK, we've got a networking thing, you can roughly like, drill it down to like those four people.
Right.
And if you're saying like, oh OK cool, it's actually with this part, you can just go like look at the commits.
Right.
And this is like no longer a manual process internally.
Like this is the whole point of why we built if you go to like station or help.railway.com.
There's a whole reason we built this thing.
Right.
It's because we wanted to figure out how we're going to go in and scale with, like a massive massive, massive amount of leverage to go and aggregate all this feedback.
You know.
This is built in-house?
Yep.
Okay.
So, and then I remember helping out on this one with Angelo in 2023.
Yeah.
You scale a lot with a very small team.
Yeah.
Yeah.
So we're like 10 times bigger now.
Oh my God.
You have your full developer count here?
Okay, all right.
I can just like cron this and then just have your life.
Well, you don't even have to cron it.
We suppose this is like a pub subable thing.
So go to railway.com slash stats.
Oh, there you go.
Yeah, that's your board.
And so it's like all real-time metrics for all of this stuff.
There's a way to get this as like a JSON too somewhere if you care or anything else like that.
Go look it up.
Yeah.
But yeah, we're big on like trying to build everything in public, talk about a lot of stuff we're working on.
You know, like we've had some issues or whatever in the past and we're like hey cool, like here's how we're fixing these things.
Like we've.
You know, we've got both compliments as well as some flack for incident reports and, like always trying to like make them better over time, just to like talk with people.
Right.
Yeah.
Yeah.
Any, obviously you had a big one recently.
I liked that it was only scoped to 3000.
You use, presumably use, Central Station like any talk, talking through, like what happens, and I guess, how do you, how do you address it?
You know, internally as a team.
Yeah.
So internally, I think this one like really, really sucked.
You know it was.
It was like to do with an upstream provider that didn't.
They didn't do the behavior that they said they were documenting, which is unfortunate given they like, wrote the RFC on how the behavior should work.
But we rolled those things out and then Central Station kind of caught that initially, where we had a couple users being like oh like, caches aren't invalidating for some of this stuff, right.
And so... turn it off immediately, et cetera.
Right.
But when you go and kind of roll out to those like that like a large user base of like three million people,
Right.
You know, like you have a lot of different disparate behaviors that that can kind of come up.
Right.
And so Try as we will.
We tested those things in staging.
We have tests for them.
All of this other stuff.
Unfortunately we hit an edge case there and we've incidentally gone and hardened a lot of those systems and now we can make a lot of that stuff better.
But yeah, it was a tough one, unfortunately.
Yeah, I always wonder how the private disclosures are supposed to work.
If people find an issue, are they supposed to contact you first?
When you run a platform, these things are going to happen.
And what channels should people pursue to quietly resolve it before it becomes a much bigger incident?
Yeah.
So I think there's like there's responsible disclosure.
We kind of err on the side of like we'd rather over disclose and know that you know that something is wrong, versus almost like having your provider gaslight you.
And so, yeah, you know, we've we've we've kind of we've erred on the side of like sharing those things kind of more publicly, even if they go and impact a small subset of those users.
Right.
And that's kind of just a decision that we've made internally.
It's under like we have four values.
One of them is honor.
And so like what's the honorable thing to go in and do?
It's like.
Well, you notify, you notify people you know to the widest degree in which they may have been affected or there was an issue, or whatever.
And then we kind of confront that head on and be like, why did that happen?
What can we do better in the future?
All of those things kind of like that, you know?
So, yeah, not the whole user base.
And that's because of incremental rollouts and- Yeah, progressive rollouts and stuff like that, right.
Interesting.
Yeah, yeah.
I feel like that should just be the norm at all large platforms, right?
Yeah.
Oh, it totally should.
And a variety of companies, it totally is, right?
There's a whole quote of Meta runs 10,000 versions of different versions of Meta in general.
And to our earlier point about agents right, they need the same thing.
They need to build a shadow traffic, they need to build all these other.
I think we've built so much ceremony around like production is sacred, all of these other things that like.
We need to get to a point where it's just trivially easy to test different behaviors right in a safe environment, because then you can make those mistakes in an environment that's like safe in general, right?
So You mentioned somebody brought it up
Do you see a world in which these things get automatically caught, not necessarily by your agent, but like your customer agent?
You know what I mean?
That, like the cash invalidation thing, seems like a pretty easy thing to check if you know to look for it.
It's hard because then you almost need well for us to like determine it like we need, almost we'd have to hook in with, like your observability infrastructure in general right.
This is like why we almost have the template loop on the platform is to be able to kind of roll those things out progressively where you say Hey listen, you know I can roll this out to like Johnny Vibe Coder initially right.
Or I can push a shard and you can almost like consume that at your own leisure and be like oh okay, I'm going to update to this specific version, right.
Or have this kind of like roll out over a period of weeks where you're pushing a new version and then it goes to You know, 01 percent of people, 1 percent of people early, like whatever, and then rolls out all the way there.
Right.
That's the kind of like non-deterministic version control that we've kind of like talked about earlier.
So, yeah, 100 percent.
Right.
And I do believe that like, that's where most things should go, go towards, because I think ultimately,
Most companies end up building that stage rollout system in-house, right?
And it's just the same thing, built again and again, and again at every single one of these different companies.
So there's a massive opportunity to consolidate a lot of developer stack.
You should have a free tier, like the model providers give you free tokens if you let them use the data.
Like, we'll give you free compute if you're like the number one shark that goes out and you, let us plug into your observability.
Yeah.
Like, incidentally, we do that, right?
And that's why the you know we talked about yeah, we talked about you know the impact of that on like 3000 people or whatever.
We start with the kind of lower impact token.
Like the, you know, larger companies that are on the platform, right?
Like they're the last ultimately, that should receive those kind of rollouts so that they have a version of the platform that's like deeply, deeply stable right.
I have three services, so I'm sure I get the first one.
You can nuke my thing at any time, man.
I guess my other question is, there's all these SRE agent companies.
There's the observability people also want to have agents that fix your upstream problems.
You have your own agent in the canvas now that you can chat with?
Yeah.
How do you kind of see that play out?
It's almost like the stacking entropy thing in general, right?
Like I think, if you don't have the primitives to make iterating in production safe, it becomes very, very difficult, right.
And so if you're an observability provider and you're like, oh... here's this fix to this error.
Assume 80 percent of those, they're probably actually good.
They're going to make sense, etc.
But then the last 20 percent of that long tail of complex issues in general, Ultimately rolling those changes out.
If you just kind of let somebody say like oh cool, this looks good and just like stamps it, there's an opportunity for you to have an issue or an incident or anything else like that.
And I think that's why it's really really important to have those kind of like forked environments in general and people have staging, et cetera.
It always ends up deviating from prod right.
And so you need the primitives and the workflows and the experience built in our mind as a first-party thing on the platform so that you can fork any point at any service at any point in time, so that you can almost like
You know, I think I consider the canvas almost as like a little like sheet of transparency paper.
And the agent is kind of like this little guy that you push up and it's like should be able to like pop up in the canvas, and it should be like, oh cool.
Like, well, I need to copy that service.
I need to copy that service so I can test these two things.
Right.
That's my hypothesis as like an agent or whatever.
Okay, cool.
I can go in and do that.
Looks good for all this stuff.
Ideally, I get a read-only copy of production.
Anything that's PII et cetera is kind of like marked as like a transform when we automatically clone that database or go for a copy on, write version of it or read from it.
And it just makes those changes.
It says, does does this actually work?
Right.
Like as close to production as possible.
Right.
Because ultimately, that's how close you have to be, or you just have a massive amount of drift where oh, I've changed this thing.
And then it just kind of gets out of sort.
Right.
The system gets a lot more unstable.
And I think that's like what you see with a lot of these kind of almost massive systems that these companies built on top of, like Docker for local and then like Cube for production and like this specific thing for whatever right.
It's like all of that complexity ends up getting to a point where...
It slows down the developers yes, but it just gets to a point where it's so unstable at scale that it becomes hard for people to go and iterate and make those changes right.
And so we want to compress a lot of that stuff way down and just say like as close to prod as you could possibly be, that's where we want to be right.
Yeah.
I was texting Erica for questions and she says actually you were originally not a believer in AISRE.
Oh, yeah, yeah.
I mean, I've kind of... Have you come around on it?
Yeah.
Well, I flipped.
I'm actually still not a believer on the AISRE, because I believe that you need the primitives to make those things safe.
And if you just unleash an AISRE on your production infrastructure and you don't have like safe primitives for like copying volumes, making sure that this is fine, it's going to nuke your production database.
Like it's not a matter of if, it's a matter of when it's going to nuke that database, right?
I'm a big believer in making those kind of like loops safe in general.
I think I was a pretty deep, like almost...
I don't want to say AI skeptic until like 2023.
And then 2024.
I've kind of like it was like okay, like maybe I can make this thing roughly do it, et cetera.
2025, I was like, okay, now I can like hold this, et cetera.
And then, like over the whole Christmas break I think you just saw like I guess winter break, but you just massive, like everybody, came back and like, oh my God, it's almost impossible to.
Here's you on the cloud docs.
Yeah.
Cloud bot.
But it's gotten to a point where it's almost like it's harder to hold it wrong than it is to hold it right.
You know, and it's like you know there's that scene in like Avengers or whatever, where Vision's like it's terribly well balanced, you know, like when he picks up Thor's hammer or whatever like.
You're like damn like this thing, just kind of like self-balances and like works quite well from that perspective.
So, yeah, I'm a deep believer at this point in terms of that will be the dominant species, right?
Again, you know, assembly, C, C++, JavaScript, words, right?
Yeah, it feels like a big jump.
Yeah, it feels like a big jump, and it is too, right?
And I think like there's, It's not like you abandon like CPU-based discrete logic in general and just move straight to fuzzy logic.
You need both, right?
So your skills should call code or applications or like whatever, some sort of like static structure and you can use the skills to kind of distill what the almost like procedure should be or like how the code should act.
Right.
I'm kind of coming to this thesis, which is you need three points essentially, which is you need a clear spec of what, what defines the system.
You need the code and then you need the tests.
Right.
And I think when you say this thesis out loud, it's like well, if you've been in engineering for any amount of time, you're like well no, like yeah, of course, like that's a RFC, like a request for comment.
That's tests and that's your code, right?
But they all matter a lot and having them all be actually together so that they can reinforce each other and say well, the spec and the tests match, but the code doesn't.
Let me reconcile that.
Oh, okay.
Now, the tests and the spec match, let me go and reconcile this other thing, right?
And you can kind of move through that period of basically saying well, this is fuzzy, and these two are either discrete in the case of tests, or slightly fuzzy, slightly discrete in the case of code right.
And that's kind of your iteration loop.
I think that's also incidentally, where you're seeing a lot of people be like software factories and I want to write this doc and like how to go and reconcile all this other stuff, which I think is a bit of architectural astronomy, if you like.
Don't actually go in and implement it.
But I do think generally that's kind of that loop is kind of where most things are going to ultimately end up.
Yeah, for listeners.
We've been talking about this on the pod for three years.
The holy trinity of specs and tests.
Oh, okay.
Itamar Freeman from Kodo is the reference for people who want to look it up.
Nice.
One thing I do want to mention, just on the open cloud thing, is also the idea that you can self-modify.
Which is kind of interesting.
I don't know how exactly Railway would support it, but I do have my OpenClaw and I just tell it that it has the Railway CLI.
It can do whatever.
And in theory you can just whatever capabilities and new infra you need.
You can just call the Railway CLI, provision it and add it to itself.
And so the agent can modify its own infra, which I think is great.
Yeah, it's nuts.
We have a loop that I've kind of set up, which is you put the railway CLI on top of something that runs on top of railway right.
And so you're essentially authenticated as whatever the current box is in general, and you can make any sort of changes to it.
And then you just call railway deploy, and it deploys itself.
It's just like, oh, cool, I need to go and spin up this instance in this environment.
I already exist in this environment.
Excellent.
I've got access to a Postgres instance now.
This is where we want to go with a lot of the agentic, almost self-replicating infrastructure.
Is that's your loop?
Like, you iterate in production.
That's your loop, right?
You're going to just continue to make some sort of change, and either it will work and you're going to want to go in and merge it and say cool, that's great.
Like put it into our upstream or it will not work and you can just kind of throw it away et cetera, right?
How do you go and make those throwaway copies like as trivial as possible to spin up, run super cheap, et cetera?
I think the era of like I have an AWS instance and I'm going to, you know, get four vCPU and 16 gigs of RAM.
It's going to get like completely destroyed.
Right.
Because it's like if you do that for for agents or anything else like that, you now need a thousand of those machines.
Right.
Like it's so prohibitively cost, expensive versus.
Like you know, we've spent a ton of time trying to figure out how do we go in and make these things deploys, whatever you want to call them.
You know, CloudSphere's got the, like, isolates.
Everybody's, like, calls it Spambox.
Like, whatever.
Like that atomic unit of deploy only pay for what you use.
Spin up instantaneously close loop as quickly as possible.
Because if the system can self-replicate the system and it can do so safely and say this is my environment, I'm making these changes, et cetera, it can come back with hey, does this look good?
Like.
This is a new state of infrastructure.
Given this prompt, I think I've solved this problem right.
And then you can go back to the agent and say actually like, looks a little bit different, goes and does the loop again and you're like cool excellent apply yeah um, i think i think that's retroactively obvious kind of like the the most uh, uh useful kind.
I don't know any other comments on like, just like agent deployment on uh, railway.
No, I mean, it's getting better every day.
And I'm on X or Twitter or whatever you want to call it.
And you can always yell at me about the experience not working as well as it should, because there's plenty of things that should work way, way better.
I was going to say.
I think at this stage in the juncture, when people want the massively or embarrassingly parallel compute they usually talk serverless.
And I feel like there's a new serverless that has emerged compared to the last generation, the previous five years of serverless.
You're kind of in that new bucket.
I don't know if you have comparisons or philosophical differences that you want to call out.
No, I think it's like, as you kind of mentioned, it's somewhere in between, right?
It's like the ability to run stateful long-running, like you want to call them workflows, you want to call them executions, you want to call them whatever.
Which like Vercel has fluid execution.
And then Cloudflare has some container thing.
Yeah.
Google has always had the app runner.
App runner and... The new one.
Yeah.
I forget a bunch of them.
Yeah.
Yeah.
I think like that's kind of where everything roughly And this is why we've been working on it for the last like six years.
It's like we just believe, like you do need access to a computer, like a box, that speaks Linux, right?
So that you can deploy the things that you want to go in and deploy on it, right?
Like other things are going to.
I mean, they're going to change the almost like surface area of what you can kind of go in and build.
And for us, we're always like No, like.
Users need a computer and they need.
They need to be able to deploy anything that they truly want.
Right.
And that's why we focused on a long time for a long time on those primitives.
Right.
Of like network compute and storage.
Right.
Because if we can give you those things and we can expose them to you and allow you to run these things indefinitely,
Right.
That's, of course, like. where we believe that it's going to go in general, right?
And so I think you're seeing right now where again the whole like Twitter has no nuance, where it's like servers, right.
It's servers.
It's like, no, it's like, it's always somewhere in the middle, you know?
Like, it's always some sort of convergence of, well...
I want to run it for a long time, but also I don't want to like provision this resource statically or pay for just things that I'm not using or anything else like that.
And that's always been our thesis from like day one.
It's like pay only for what you use, run it indefinitely.
It is just like full Linux basically.
Yeah.
Yeah.
I think that's why I like the Vercel naming of fluid.
It's like, well, it's fluid.
It's flexible.
Another milestone.
And then I wanted to ask one more technical question, which is the Heroku official deprecation.
Basically, you are one of the presumptive new Herokus.
New Heroku has been a category for like... as long as I've been in developer tooling.
Yeah, right.
It's finally happening.
Yeah.
What was that like when, you know, is there any behind the scenes of like, well, this is the moment?
Yeah, I mean, you just have so many people just like you're, just like, like you were running stuff on here, like you as this company, like it's crazy that like you whatever like name that you would know is running this thing, and then you're coming to us be like yeah, we kind of like want to like move a lot of the stuff off, or whatever.
Like Okay, cool.
But yeah, it's kind of just nuts.
Any behind the scenes on why did Salesforce let Heroku kind of just stagnate?
Well, I mean, I can only guess, I guess, right?
I mean, I think it's just hard when it's not your business.
Like the business of Salesforce is to build a really, really good CRM, you know, right?
And like that's their focus, right?
They should be really, really focused on building a really, really great CRM.
And then you acquire this business as a compute business.
That's kind of an offshoot of your business in general, right.
And I think like, you know, a lot of the early meta folks have talked a lot about like focus. right?
And like I think Boz has a whole like write-up that he's done, basically where he talks about in the early days of meta we had no money and like we were forced to get focused right.
And then we basically turned on the money.
This is all, like you know me um verbatim yeah, rephrasing or whatever.
We turned on the money tree and then we had no reason to like, not like have focus, because we just had infinite money where we could go and split all of our focus right, but that ends up diluting your product.
It ends up like making these things where you kind of have these offshoots, where you're just like is that the focus of the business, right?
And it ultimately ends up not being, if it's not, the core of your business, right.
And so to me it's.
It's like kind of no wonder that, like it languished in general, right.
Because it just wasn't the core focus of the business.
And I think that a lot of companies get in trouble with this when they kind of like split out their focus in general, because it means that you're almost like fighting a like multi-fronted war, trying to like compete with all these things and not just like compete with them externally but compete with them internally for like alignment and like where are we going?
What are we doing?
What is our purpose? here, right?
Like if you're you know, if you're really really you know like Salesforce built and you're like hey listen, I love Salesforce and I really want to like work on all those things.
Like you know and you're mission driven, which is like the aspiration for a company in general of like, why do people work on things right?
It's like they want to work on something interesting, right?
Like, Heroku is off to the side.
It's like it's not the core of the business.
Right.
And so to get those resourcing, you know like budget or focus or alignment or whatever, internally it's just pushed away.
Right.
So it was it was literally just a matter of time for it to happen in our mind.
Right.
Yeah.
And then kudos for them to like actually call it out instead of just letting it be unknown or hanging in the air.
Yeah well, their whole release was a little bit odd because they, like you know, they kind of called it out, they did the Our Incredible Journey.
Yeah, they didn't say they were like shutting it down, but they're like yeah Yeah, yeah.
So, yeah.
And then you know behind the scenes I think they issued some stuff to people being like hey yeah, you should like close these accounts down.
Like, we are going to go in and defragate this and, like, remove it every time.
So yeah, I mean, it's just like, and it's crazy because, like some of my first deployment experiences were like on Heroku.
I learned to code on Heroku, man.
I had a freaking alias in my bash for like Heroku deployment.
Yeah, right?
Like you start with, like dragging stuff into an FTP server and then, like you move on to, like trying to get a deploy working.
You're like, how do I go in and make this happen?
And it's like...
Heroku, right?
Did you know about Heroku packs?
Yeah, exactly.
And you learn about all this and it was the on-ramp for us, right?
You know, but the wheel turns regardless, right?
Like there's new stuff that's emerging and like we're very, very happy to like, almost like continue to like, you know, carry the torch on for a lot of that stuff.
But we don't want to be the new Heroku.
We want to be the way in which people are building and deploying software and ultimately, the way that people monetize software over time right.
Yeah.
I mean, it's a big crown to be a new Heroku.
Like, there's like 50 companies that fought for this.
Oh, yeah.
Everybody's kind of like, you know, holding some portion of this being like.
But yeah, I think you know for us we're just happy to go in and support people, companies etc.
The platform works a bit differently.
So it's like, you know, it's obviously kind of almost the similar kind of like game loop.
Yeah, exactly.
Right.
But we've been quite dogmatic in terms of where we believe these things are going to go in terms of primitives.
You know, the agents kind of fan off all of those other things, right?
And so some things will fit and then some things will.
You know, you have to change a few of their workflows, et cetera.
Like we don't have, and what's that feature that people really love?
Pipelines?
Heroku?
Yeah, right.
We have some approximation of it with the environment system in general, right?
But yeah, so it's been super exciting.
We've got a ton of people that we're able to go and support.
And it's growing a lot.
Yeah.
Any other technical?
I have one more about Temporal.
Okay, so Temporal, I have sold my shares.
You are a power user.
You're one of our earliest customers.
I met you through Temporal or something.
You're a big temporal business.
Like your business is built on temporal.
You have complaints.
I think this is the neutral, most neutral, most informed conversation that anyone will ever hear about temporal without someone working at the company.
Yeah, that's fair.
It's the two of us.
Yeah, yeah, yeah.
No, I think that's fair.
Yeah, so I have used Temporal for almost like 10 years now, right?
Because like Cadence, Uber, all of those other things like that.
Just give people a scale of what Cadence is at Uber.
People don't know.
Yeah, so Cadence was the precursor to Temporal and it powers all of the trip actions, the rides, the like when you rent a jump bike or a scooter or anything else like that, or a car.
It's like you're running these workflows for a period of time and you're basically saying This ride will run for an indefinite period, until it like, finishes.
Right.
And you can go and attach information, whether it's like, oh, you paused it in this zone.
And so, you know, you need to add this dollar charge to like the bill or anything else like that.
And then when you end the trip, like your workflow is done and that whole experience behind the scenes.
I don't know about today in general, but it was like powered by by cadence at that point in time.
And so it's a really, really like And I used to say like.
It's like imagine if you could program the entire user journey top down as one function.
Yeah.
Yeah.
And it's such a powerful idea and it's so, so important.
It's also incidentally, so important for the next phase of the agentic journey where like, you want an agent to do a specific task and then you want, it to like, be complete or incomplete on that task and then move on to the next thing.
Right,
Like you need a way to be able to go in and manage these workflows.
You need a way to be able to go and manage these workflows dynamically.
And I think for me, Temporal was always like really, really, really great.
In theory.
And it was really, really great when you got it working the way that you wanted to in production.
It's just it required you to like model that entire journey in your head.
And if you didn't have the entire journey in your head, you could put yourself in a spot where you would cause like issues where, like replaying the state of the entire workflow, like causes, like a non-determinism.
Yeah, because it works on like deterministic workflow history.
Yeah, exactly, right?
And so it's very, very easy.
It's like the way that I kind of like would describe it is like, well, it's a jet engine, right?
Like, if you know how to like go in and operate, if you know how to go in and run it, all of those other things right.
But you can't hand it to...
People who are trying to build things that end up being complicated but don't have that whole kind of like state in their head, right?
So if you have a large, like we run our whole deployment pipeline on top of it, right?
And so that's like a reasonably complicated workflow, right?
Like there's pre-commit hooks, there's signaling, there's queuing, there's all of this other stuff in general.
We ran into the same thing at Uber where, as you try to express this large workflow, as you mentioned, going all the way down got more and more complicated and it got more and more states in the state machine that you had to like, map the state machine back to like the workflow.
It's a lot of ifs, right?
Yeah, exactly.
If this, if that.
Yeah, and so at Uber we built a system for, you know, doing the state machine and like testing the state machine and all that other stuff, and we've started to like go and build some of those things here because, like it's grown, you know quite heavily right.
But it's, like...
It's such a like – I don't want to say love-hate relationship, because that's like too broad in general,
Like when it works really, really well, it works like super, super well, right?
But then you run into a spot where you just like somebody who hasn't interacted with the system or doesn't have the full context of the system, goes and puts something in the system that – invalidate some of the state or causes a non-determinism issue, or spins off a ton of activities, or anything else like that.
And then you have to kind of keep track of like almost underlying SRE knobs, of like oh we have, you know, the amount of activity slots in this thing.
Right.
It's like well, these should just scale with, like memory vCPU, all of those other things in general.
Right.
So it ends up becoming a bit of a bear to kind of scale out in general.
Yeah.
So you need like a very capable sysadmin running things behind the scenes for you.
Yeah.
Yeah.
If you were to move off, what would you do?
I think we would build our own workflow engine.
We have a few internally that we've kind of like worked on.
So, yeah, because it's like, yeah.
This is one of those things where, like you know, this is one of those classes of things where, like you, typically wouldn't vibe code it.
But I'm wondering if you can.
Well, I don't think you should vibe code it still.
Like you still want to run like depth and tests and stuff like that, like to make sure that, like you.
Yeah.
I mean, you know, like, it's not like Turbo had to invent that from scratch either, right?
No.
So like, there's libraries for those things that you can run.
And like, on top of that, it's just a state machine and, you know, that you have to really map out.
But ultimately you define those abstractions that you want and you run into a state machine and that's it.
Yeah, it's very, very doable.
So yeah, I think the workflow stuff is very, very interesting.
Like there's a few really cool companies that I think like Restate's doing some neat stuff here.
So you're very tied into JavaScript.
You're like a JavaScript maxi.
Internally, we have JavaScript, we have TypeScript, we have Rust, and we have Go.
Those are three languages, right?
We don't add any more stuff.
Actually, that's not true.
We have a little bit of C because we write BPF code and it's hooks and stuff like that.
But those are the kind of languages.
Is this for the side container things, side car stuff?
No.
Well, so this is for the networking stack as well as the volumes and stuff like that.
So, yeah, but it's like.
Yeah, we used the TypeScript stuff a lot because it's like what powers the dashboard.
But we're going to move a lot of the kind of workflow stuff off of the kind of dashboard stack into actually the infrastructure stack.
We're just recently.
Yeah, don't power things on front end, guys.
Even though it's free compute.
Yep.
Yeah, yeah.
Cool.
Any other technical infrastructure cool stuff?
Railpacks?
I don't know if that's still Yeah.
I mean, we built an engine for determining dependencies based on your source code, which is super cool.
It's called Railpack.
We built the first version called Nixpacks, which is on top of Nix.
And then, yeah, we moved.
People have been trying to get me to adopt Nix and NixOS for like four years.
Yeah.
Is it ever going to be a thing?
I don't know.
We were super excited about it in general, but it has a bunch of different pain points in general.
Because, if you just think of it, it's a stack of version source code or it's a stack of version binary at specific slices in time, right.
And so if you want version X and version Y, you end up bloating a lot of your kind of like package, like space, right?
Which blows up the size of your images and makes it really really difficult for like really real world workloads.
I think if you...
You content, address it, you cache it.
There's a lot of optimizations that in theory you should be able to do.
In theory, yes.
And what happens ultimately is you have a large enough user base and you have a disparate enough set of machines that you kind of run into the problem that there's a paper that Meta released XFAAS, their internal kind of serverless system.
It ends up being very, very difficult to go in and do that at scale, unless you break out specific runtimes basically, which we did not want to go in and do right.
Because we wanted to truly allow you to deploy anything, right?
Which was our initial kind of thing with Nix.
But we've moved towards some interesting stuff that I think we'll be able to talk about a little bit later.
That we've built for doing context-addressable file systems to be able to lazy load anything from any point and then just page that into memory.
Amazing.
Okay.
Yeah, it's going to be fun.
The whole future is very, very bright.
It's crazy.
It's going to be nuts.
Uh, okay.
Founder journey stuff.
Yeah.
And your, uh, cloud usage, you tweeted, you're going to spend 300 K this month.
Yeah.
I think we got it all.
I think we got 200 agents across the company.
Yeah, you only have 35 people.
So I'm sure they're not all spending 10K a month.
What's kind of the distribution?
I think I'm at about 25 in general.
And then we have some power users kind of all the way down.
We came back from the winter break and I was basically like, if you're writing code by hand, you are doing this wrong, right.
Like they're.
The tools are good enough at this point.
Like that, you move extremely, extremely quickly.
And like, yes, there are issues and pain points and all these other things.
But you should be reviewing the code that you are writing instead of trying to go in and write it by hand.
All of those architectural patterns, all of those other things, you're not going to throw them in the garbage or whatever.
Actually, they matter more now than any other time, but you just shouldn't spend your time generating code that you would write.
Like if you know how to go in and write it, just like ask the agent to go in and write it and then reconcile it until it looks like you would have written it yourself, right?
And I think incidentally, like people misconstrue my propensity to like push people towards agents for like hey, we're growing really, really fast and we've had some kind of like bumps in real life.
They're not necessarily related in terms of that.
But I think people should really really understand like the tools are good enough for you to be able to move extremely, extremely quickly to build things way, way larger than you could have possibly built before.
Right.
And so to our point about way earlier, but like How do you cool data centers in space?
It's like, well, I don't know, actually.
Right.
But you're at a point now with, with software, you can actually be like well, how would I build block storage from scratch?
How would I go in and do these things?
I have ideas because I've got history.
I've read all these papers in general.
Right.
Let me go in and work them out in general.
And let me build like massive test benches with, like things, thousands of tests right because they're free to, they're free to author right now.
Right to go in and make sure that, like this system can now can can be built right.
And i think that if you're not using the the kind of ai systems to almost like speed run your roadmap, to like go in and figure out where you need to go in and be to reconcile your existing system onto the future, then you're kind of missing a large point of of what is currently happening right now right.
Because you can just template out anything and validate it on the side for free, right?
What's the path to spend 3 million a month?
Like, is it bound by, like, ideas and things that the customers can absorb?
I think for most companies, it's actually bound by deployment at this point in time.
And I think that's why we've seen a lot of like a massive boon in terms of like users trying like, not just users like companies.
Like you know, Fortune 50 is, like you know below, et cetera like going and being like.
How do we get our developers to like go in and move quicker?
Right.
I think you're probably going to hit your CFO before you hit any of these limits in general, because they're going to look at this and be like there's an eye-watering amount of like money being spent on these tokens.
Like, I think, I don't know which, I think it was the Uber season.
It was like, blew our token budget for the entire year or whatever, right?
And so, inference costs have to come down, but...
They're also, you know, they were inference constrained at this point in time.
Right.
And so you're going to almost get this like price discovery of like what makes sense for an org to go and adopt.
And I think what you're going to end up with is actually you're going to almost like end up with the like F1 driver concept, which is, if you have a, if you have somebody who's like really, really adept at these things, it makes sense to go and put them into like a three million dollar car or whatever.
Right.
But if you're not, then it probably doesn't actually make sense for you to go in and do that.
And we're gonna take a few of these people and say, You can drive the F1 car.
We need to go in this general direction, figure out if this works and almost go ahead and prototype it right.
And so we've done a few of those things.
We've vastly accelerated our roadmap in terms of oh, we thought we were going to be able to go in and ship this thing in the next few years, but actually we can probably ship it in the next few months.
Now, right.
Because we're saying oh, validated it out, it works.
You don't have to even like build it incrementally.
We can have skip steps to like go and just move towards where our vision is for a lot of the stuff, and I think that that's kind of where we end up with a lot of it, you know.
Yeah, I think a lot of people are realizing the roadmap doesn't always have a business impact.
And so it's like, oh, it's too expensive to run these tokens.
But like if your roadmap was actually built to make more money, by the time you built the whole thing, you would have some sort of token pricing for it.
The same way you do with sales.
Yeah.
Like you, will spend a billion dollars in sales.
If you knew, you would get two billion dollars of revenue.
Exactly, right?
And I think the really naive way to go in and measure this is almost like your percentage of tokens that end up in production.
Right.
And so if you can measure that you are getting this level of impact because those tokens are ending up in production, that's awesome right.
But I think the kind of burden of proof is now going to kind of like arise.
And you see it internally too on our stuff.
Like we have a growing number of pull requests that like, or like haven't yet been merged, right?
And you're just like, okay, how do you get this into production, right?
And so it's really about like how quickly you can go and kind of build and deploy that software right.
Which is exciting because we build and deploy software, you know, right?
So, yeah.
Yeah.
The SDLC is changing and it's something that both of us are super interested in exploring as well.
One of my thesis, or it's not my thesis, it's the pull request is dying.
It's going to be the prompt request.
And then beyond that, code review is also kind of dying because you really need to if you have all the other systems in place.
What else is changing about the SDLC?
What else is different?
Well, I think the... AISRE.
AISRE the tools to make.
So the AISRE is like one of those things where it's like you know, it's a pie-in-the-sky aspirational.
What does it take to get an AISRE?
By the way, you should expose your tooling to your customers at some point.
Yeah.
Which tooling?
Central Command Center.
Oh, Central Station?
Yeah, yeah.
So we have it for template maintainers, right?
So template maintainers can deploy and maintain templates and they get feedback on a lot of that stuff right?
And so we're 100% going to go in and expose those things incrementally.
Yeah, but clustering around incidents.
Everyone has a version of that, but I don't think anyone solved it.
Yeah.
Right.
And I don't say we've solved it internally, but it's gotten.
It's gotten so good that, like now, we can see those incidents forming like pretty quickly.
Yeah.
Yeah.
Yeah.
Yeah.
So at some point, those will be things that either somebody else goes and builds or we go in and build.
But we've always built stuff that was purpose-built for us and, if it made sense and there was a way to go in and make it useful for users or monetize it or make sure that that loop becomes like a profit center instead of a cost center like, we want to go in and do that at some point right, um so, but yeah, poor press is definitely dying.
Do you do first party uh, feature flagging and uh, incremental rollout type stuff as well?
So we have a feature flagging engine that we built internally.
That at some point we will.
We will, because i don't see it as a user.
Yeah yeah yeah um, so like that, that would be.
That's good right, how come you didn't give us what you have?
Well, Because we have to beta test it.
We actually care a lot, a lot, a lot about the quality of the things.
There's plenty of stuff that we've used internally and then we've got it to a point where it doesn't make its way entirely through the journey because it fails.
It's like this holds for one service. but it doesn't hold for multiple services, right?
So we'd have to go and build these things for multiple services to go and make this work, right?
And we know for a fact that if we release this thing, we'd have to go and rebuild this thing again and again, and again.
And some things are worth doing to go in and do that.
But a lot of them are basically like that also, that kind of just informs our roadmap of okay well, like for us to go and make that actually a bit easier, we can do a few of these things first and then we get to that experience, right.
We don't wanna dilute the experience by basically saying like, oh yeah, this works, but only for this service, right?
Unless it's like a very, very core initiative, which is, like you know, over the next like few months, we're going to roll out a few things that are like OK, it works for a single service, and then it works for multiple services.
That works multiple service across the environment.
Right.
But you have to be very, very deliberate about those things.
Otherwise, you end up with a bunch of broken, disparate experiences, which ultimately end up creating a ton of support load because people are like how do I use this feature?
How do I go in and do this other stuff?
Right.
So it's kind of the thing earlier about, like you expand your company in general to to get those like features and then you almost compact it, smooth out those things.
So the experience is like really, really stellar.
Like we were talking in the hallway earlier, where you're like, oh my God, it's gotten so much better.
And I'm like, oh man, like just internally, we're like, damn, this part really sucks.
And we got to make this significantly, significantly better.
No, I can attest, you know, over the last three years that I've watched you build in real way.
But yeah no, I would call to listeners if you're not aware.
Like the importance of feature flagging, it's a very big part of Uber culture.
Yep.
So much so that they have too many feature flags, and then they have another thing to remove feature flags.
Yep.
100%.
What was it called?
There's a paper about this.
There's a flagger and there's been another one.
There's a thing that like Facebook has gatekeeper.
Yeah.
So they're really important.
And agents are going to need this.
That's like the fundamental thing behind, you know, like just incremental rollouts.
Yeah.
Opening up acquired stat sig.
And basically GPT-5 is just routing and flagging through different models.
And it's super important, right?
Because if you assume...
The software development lifecycle is 100 going to go in and change.
But it's going to change because we're trying to do things a thousand times faster and a thousand times more concurrent than we were currently doing, right.
This is routing.
Yeah, right?
And so what ends up becoming important at scale?
You know, before I even you know started Railway, I actually built a feature flagging product.
I tried to go in and sell it to people, right?
Okay.
Because I was like, oh, it's an easier version of LaunchDarkly or whatever, right?
And I ran into this situation, which is like anybody who's small enough to adopt your technology, doesn't care about feature flags, right?
And then anybody who's large enough to try and actually need feature flags needs so much scale that you have to build out all the existing infrastructure, So end up scrapping that.
But it's what is old is new again, because now companies are trying to move really, really quickly.
But you can't just YOLO this like vibe coded thing straight into production.
You need to basically say, hey, here's my blast radius.
Here's my impact.
Here's my like whatever.
I want to shadow it for these users.
Right.
Right.
Like, you're going to need those tools that ultimately those larger companies ended up having to go in and build to maintain their structures.
Everything's just going to get compressed by like a thousand X so that everybody can go and do that and everybody can build those structures really, really quickly.
Right.
And that's like exactly where we're at right now is like you're compressing the software development lifecycle and then we're going to expand it and add way more new things to it.
Yeah.
Yeah.
And then the other term that comes to mind with when this kind of discussion happens for me, for newer developers who haven't heard this term cattle, not pets.
Yeah.
Right.
Because like your prod, people treat it like a pet.
Like it has a name.
I have to keep it alive.
But when it's cattle, you can just mass farm and you can like roll out and you can, like you know, portion out parts of them and kill them or whatever.
Yeah exactly, I actually think that maybe that's the hot take, but I think that that's actually going to change and I think you can move towards having pets so long as you have a and this is going to be a jump so long as you have a cloning machine for your pets.
If you can snapshot every single thing at every frame, then it actually doesn't matter if that thing gets obliterated, because you have some sort of snapshot of it, right?
All of the things that we have built right now are to essentially block out any sort of changes or alterations or whatever from that, like hermetically sealed DevOps, like line or whatever.
It's like okay well, you have to write a Docker file because I only need these specific instance, like only this specific cut of the file system, et cetera right.
What if you just had the whole file system?
What if you just snapshot it?
What if you lazily load the entirety of the file system, right?
Then you can get around this problem entirely.
You don't need the ceremony of those you know, having a Docker file or like having Ansible script or like having all of these other things.
You can just iterate on that loop and then like snapshot it.
It's like, is this the right loop?
Is this the right thing at this point in time?
Okay, cool.
Like now I'm going to go and merge it in production.
Like go merge the file system.
Yeah, it's gonna be really fun.
Yeah, this is like a whole other kind of worms.
But like I think the number of things that are stateful in a VM, I think if you just kind of catalog them and just like develop dedicated solutions for solving each of them, you can actually kind of cut this problem down a lot.
And it's surprising that people weren't really trying until now.
Yeah.
Well, so it's, it's surprising.
I mean, it's always been surprising to me because these are the things that we've worked on.
Cause they're just like, I'm like, it's so obvious.
You need them.
Everyone in theory needs them.
And then like the big clouds don't do them.
So you're like, it's impossible.
Yeah.
Yeah, exactly, right?
You're like oh well they, you know, Meta has all the people who write, like you know eBPF code and they're like doing something with them.
You know, but like you need that kind of stuff to solve these problems, right.
And like it's like whatever is required, however deep we have to go in and like get to, like solve those problems right.
Like, all the way down to, like, the kernel's TCP IP stack, right?
Like, we're going to go and figure that out.
Is there something that we need to go in and modify, to like go in and make that work for the mental model that we have for the universe moving forward?
Like, yeah, 100%, we're going to go in and do it.
We'll just keep going our way down.
It's super fun.
It's so much fun.
I have to literally peel myself away from the fun, interesting problems that we have to make sure that we can scale the company in a way that works.
And there's so many different fun, interesting problems, whether it is how do you get the information from the customer to... support to the person who built the thing internally right or it's like how do you do safe iteration or how do you get context like from the dashboard to users or like how do you drill down all the way to the infrastructure layer how do you manage orchestration it's like a real time operating system versus a feedback control system right like it's just so fun you know yeah I mean, speaking of maybe you talk about the founder side, you're famously like, you know, the YC, the SF consensus is you go to YC, you get a co-founder, you get to do all these things.
You have done none of that.
Yeah, I've like done a lot of different things in general.
In the elevator you were like actually co-founder.
It kind of makes sense if one person is the tech person, the other is the business person.
But you have to contain all those multitudes yourself.
Yeah.
How do you do it?
Oh, okay.
I was going to ask, is there a question in there?
Yeah.
The question is, what the hell?
How do you do it?
The question is, how are you alive right now?
Yeah.
Well, I mean, yeah.
I mean, just try to get eight hours of sleep.
You know, like...
Is there like a balance that you ideally like 50-50, 30-30-30?
Like what's the mental model that you use?
There's no balance.
There's like you just have to think about all these things and be obsessed with all of these things.
Like whether it is being obsessed with, like how do people think about your product from a go-to-market perspective or being obsessed from a perspective of like well, if I can make this change at the kernel level, then I can make it so that the user's SSH connection never drops.
Because that's what I want.
I want a universe in which I can go and snapshot all these things, and it looks exactly like you would just kind of iterate on a VM, right?
And I think you just have to be obsessed with all those things like at every layer of the stack.
And I think that's what makes it easier for me.
I think some people, like they're obsessed with different portions of the kind of like journey, like the company, like whatever right.
And I think that that's when you can get really really good, almost like cohesion, by like segmenting out these things right.
And so, you know, in The Elevator I was talking about, like you know, you have a technical kind of like person, et cetera, and then you have the customer kind of like person in general right,
And I think like if you can segment those lines out really, really well and you can be very, very clear about what your areas of ownership are for yourself or your, you know company or like just where you're going to operate you're going to have a good time right.
If you can't be clear about those things, right?
And this is why I was saying like two is the worst number of co-founders is because you have no tiebreaker right.
You basically are like, well, I disagree on this thing and I disagree on this thing, right?
It's like, well, how do you resolve that, right?
Well, usually someone's CEO, right?
Right, exactly, right?
Then you're like, okay, you have to tiebreaker.
Yeah, totally.
I mean, listen, it's hard.
It's hard every single way you cut it, right?
It's hard.
It's hard if you get help.
It's hard if you do it yourself.
It's just hard to like run things, roughly speaking, right?
But it's so rewarding.
It's so fun, you know?
What have you found useful, like a coach, any advice that has been really helpful?
I like to write a lot.
I got in trouble.
I get in trouble a lot for my Twitter.
There's a pattern.
Who do you get in trouble with?
The people on Twitter.
Oh, okay.
I was talking about it and I was like, hey, if you...
If you're working weekends, you're kind of messing up your planning, roughly, right?
And I've gone kind of back and forth on that, right?
Because I think actually right now we're kind of at an extenuating time in general, where it actually makes sense to like work.
Right.
Because the goals are pretty clear in my mind.
Right.
And so if you have the vision and you know where you're going, you should work a little bit harder to distill that vision and go and do those things.
But if you don't have the like we're like, I think we should be going this journal.
I'm not 100 percent certain.
I want to get a little bit of clarity.
I think what you need to do is you need to like disconnect and you take your weekends like very, very seriously.
You need to write about things.
Where are you?
What do you wanna do?
Where do you wanna go?
What problems are you trying to go and solve?
And like, think about a lot of these things, right?
So, you know, like, Writing is important.
Sitting down.
I don't like the word meditation or whatever, but whatever gets you into the state of your mental clarity.
That's the thing that's really, really important when you're trying to go on these journeys of saying
Well, we're here and we really need to be here in general, or like we're here and I think we need to be roughly in this kind of like space for this to like work.
Right.
So those are the things.
And then you know disconnect, hang out with the people that you love, and then like work super, super hard when when, you're like
You know, like I try and work like sunup to sundown, Monday to Friday, all out in general.
And then I try and disconnect on Saturday and then I come back to work on Sunday afternoon.
Right.
And then I do my writing plan for the week, all those other things.
And it works really, really well for me.
But another hot take is like most advice is to be digested and to be thrown out the window.
And if it's helpful, it'll come back, right?
If it's helpful, you'll have kind of like learned it over time through experience or anything else like that.
But yeah, you mentioned like the kind of standard, you know, YC advice, all of those other things.
We have a lot like.
We've made failure as a society very, very expensive and it makes it difficult for people to kind of trot off the paths right.
Yeah, makes sense.
Any other softbooks you want to get on?
Like anything that you have not tweeted and got in trouble with, that you want to preview to the world.
No, I think the agent stuff is just like...
It's crazy.
It's going to be the dominant way in which people are doing pretty much everything, right?
Provided we can, of course, get the amount of inference required for that to go and happen.
But over the next like 10 years, right?
You see a fundamental shift in terms of how people are thinking about even just authoring the logic that's in their head, right?
Yeah.
My, you know.
Maybe one way of phrasing this is if all birds can become a GPU provider, so can railway.
Yeah.
I think there's a lot of art in us actually not becoming a GPU provider.
I think you're defined almost more by the things that you don't do than the things that you do, because it's really really easy for you to just say yes to a bunch of different things, right.
Yeah.
And I think like it's going to be very, very interesting to watch.
You know, I think Anthropic is like an amazing company and like super, super stellar.
And they're moving into a variety of different zones, right?
They're moving into, like, the Figma kind of, like, stuff that they're after.
Yes, as a recording.
They've got Claude.
My career was on Figma's board and then they removed him, like Monday, and then they launched this today.
Yeah, yeah.
Yeah, so I mean, things move very, very fast right now.
But yeah, it's just going to be the way in which people are operating.
Okay, so your answer is focus, no GPUs for now.
Yeah, focus.
But never say never.
Yeah, right.
Like I can tell you for a fact that we will not be doing GPUs now, but we 100 will be doing GPUs at some point in the future.
And that's not like me leaking our roadmap because we don't have plans to go and do GPUs.
It's just a function of at some point you need flops, right?
Like, at some point you want like, if you're fully vertically integrated and you want to make it really, really trivial for people to go and iterate and build and deploy things, you need access to this core piece of fundamental logic, right?
So, yeah.
Yeah.
And then, at some point, presumably your own data center.
Traffic is a minority of your workload right now.
But is there a majority or you just kind of completely turn off?
Oh, at some point we got to 100% data center.
Like our own data centers.
Yeah, yeah, yeah.
And right now it's the vast majority of the stuff that exists on our bare metal data centers.
Okay.
So you're already there, the vast majority.
Yeah, yeah, yeah, right.
I didn't know the extent of the transition.
Yeah, totally.
It was completed at some point, and then we grew so fast that we had to basically go and scale back on that.
Take us back.
Yeah, basically.
Sorry, Google Cloud.
Yeah, it was funny.
It was funny.
We got to on the Datadog dashboard.
It got to 100 and then it divvied it back down into the 90s or whatever, because we were like
Yeah, you're adding capacity.
Yeah.
Yeah, it's interesting.
You're literally building a new cloud that's independent and people assume that that could never happen post the AWS.
Yeah, and it's hard, right?
Like you know, we were going to, you know, figure out a bunch of different things to make sure that the platform is deeply, deeply reliable.
But you have to break ground on a lot of new things when you basically decide you're going to build a cloud from scratch but not copy the hyperscalers.
We've been very, very deliberate to invent our own infrastructure from scratch, based on reading a ton of papers in general,
But like, almost like promising to ourselves that we wouldn't copy somebody else's homework, right?
Because we were saying, hey, listen, if we copy somebody else, we lose.
Like you're just gonna become them over time, right?
And so you have to have a core thesis about, like why does this business need to go and exist at this point in time?
And for us it's always been about the activation energy to get something, to go and deploy it in production at any of the hyperscalers, as right now is far too high right.
And we believe that it should be instantaneous.
We believe that there should be no friction in between what your thought is and reality.
That kind of comes out that you can share with your friends, right?
And so that's what we're kind of like building toward again at every layer of the stack, like if we got to go down to energy, we'll go down to energy at some point.
It matters a lot for us from the experience of giving people access to this tooling, because it's gated behind
It's not even just gated for regular kind of like these citizen developers that are now vibe coding.
It's like you have multiple layers.
You have the citizen developer, you have the front-end developer, you have the back-end developer, you have a DevOps person.
You have all of these layers right.
And they all need to go in and disappear so people can just like ship like that.
Amazing.
All right.
That's the future of cloud.
Thank you.
Thank you for having me.
It's been wonderful.