We thought that we were looking at the form factor of AI, which is you're talking back and forth to something.
The real ultimate end state of AI, and thus AI agents, is these are autonomous things that run in the background on your behalf and executing real work for you.
The more work that it's doing without you having to intervene, the more agentic it's becoming.
Somehow it produces output that it feeds back into itself.
It's literally just the ampersand in Linux, which is it's a background hash.
And it's like the worst assistant in the world.
And agentification is just hiring a lot of these really bad interns.
What exactly is an agent?
And how will agents change the way we work?
To unpack this question, we brought together three people with deep but very different advantage points.
Aaron Levy, co-founder and CEO of Box.
Steven Sanofsky, former Microsoft exec and A16Z board partner, and Martin Casado, general partner here at A16Z.
From the old school idea of agents as just background tasks to today's vision of fully autonomous systems.
We'll explore what this means for coding enterprise workflows and how whole industries might reorganize around agents.
Let's get into it.
I thought I'd start this wide-ranging podcast by asking the very simple, very provocative question what is an agent?
Oh, boy.
To who?
Steven.
Oh, to me?
Oh, I'll go for it.
So I actually have a very old-person view of what an agent is, which is it's literally just the ampersand in Linux, which is it's a background sound.
Because, like, you type something into O3, and then it's like, hey, I'm trying this out.
Oh, wait, I need a password.
Can't do that.
And it's like the worst assistant in the world.
And really it's just because they need to entertain you while it's taking a long time to answer your prompt.
And so that's my old-person view of what an agent And agentification is just hiring a lot of these really bad interns.
They're getting better.
They are, but they still don't remember if I have a password to nature.
Is it possible you guys just had bad interns in, like, the 80s and 90s?
We had terrible interns.
I have, like, a very high esteem for interns, so...
But now a real answer.
No, no, I mean, I think collectively we're seeing what these are becoming.
So, if you think about two years ago, the post-ChatGPT moment, we thought that we were looking at the form factor of AI, which is you're talking back and forth to something.
And I think to Steven's point, the real ultimate end state of AI, and thus AI agents, is these are autonomous things that run in the background on your behalf and executing real work for you.
And you're ideally in an ideal world interacting with them actually relatively little relative to the amount of value that they're creating.
And so there's some kind of metric where the more work that it's doing without you having to intervene, the more agentic it's becoming.
And I think that's sort of the paradigm that we're seeing.
The only addition I'd have, in addition to long running, which I agree, is that somehow it produces output that it feeds back into itself as input which you can actually do long running inference.
Like, you can make a video that's really long running, but it's just basically a single shot video and you just throw more compute at it.
I think there's like technical limitations if you start feeding the input back in, because we're not quite sure how to contain that too.
And so...
I think you can measure things based on how long they run, and you can also measure it by how many times it's actually taken its own guidance, which would be kind of more of an agency.
Yeah, because I do think it's important that in this transition Look what Aaron described is where we're going to be.
It's just that what are the interesting steps that happen along the way?
Because we are going to need, for the time being it, to stop and say am I heading in the right direction or not?
There is this thing where you just don't want to waste your time on the clock while it's churning away way off in the wrong direction.
Yeah.
So the question is, to what extent do they have their own agency, which to me means they've spit something out and they've kind of consumed it back up again, and it's still a sensible thing.
Which, by the way, as you start thinking of these things in distribution, it's actually a very difficult thing to do, because it doesn't know if it's going to be spitting something out that's still in distribution when it brings it back in.
They don't have that self-reflection.
So I think there's actually a very kind of technical question here of to what extent we can make these things have independent agency.
But we can make them long run pretty easily.
Yeah, yeah, we're good at the long run.
What you get back is, yeah.
Yeah, I mean, I think the interesting thing is how the ecosystem is sort of solving problems or mitigating then the issues.
Like you're seeing sort of this logical division of the agents.
So they might be long running, but they're not actually trying to do everything.
And so the more that you subdivide the tasks out, then actually the more that they can go pretty far on a single task without getting kind of totally lost on what they're working on.
Well, Unix is going to prove to be right, which is like you're going to want to break things up into much smaller granularity and tools.
And I think to other points that you've made on X, like you're going to want to divide things up so that it's like an expert in this thing.
Yeah.
And then it might be a different let's just say body of code where you go and ask you know, are you good at this thing?
Let me get your answer on this part of the problem.
Yeah, it's kind of interesting.
I don't know how much you've plotted this, but like the conversation on AGI has sort of evolved very clearly in the past like six months.
And I think that the consensus was maybe not even consensus.
What some of the view was, let's say, two years ago, is this sort of monolithic system that's just super intelligent and it solves all things.
And now, if you kind of fast forward to today and let's say whatever we agree kind of state of the art is, It's sort of looking like that's probably not going to work for a variety of reasons, at least in today's architecture.
So then, what do you have is maybe a system of many agents, and those agents have to become very, very deep experts in a particular set of tasks.
And then somehow you're orchestrating those agents together.
And then now you have two different types of problems.
One has to go deep.
The other has to be really good at orchestration.
And that maybe is how you end up solving some of these issues over the long run.
I just think it's very difficult to think cleanly about this.
I've still yet to see a system where they perform very well and you don't draw a circle that doesn't have a human being in it somewhere.
Oh, yeah.
So in a sense, the G often seems to be coming from The general seems to be, So I just Listen.
These things are tremendously good at increasing productivity of humans.
At some point maybe they'll increase productivity without humans, but until then it's just very hard for me to actually talk cleanly.
Well, and...
It's so important for people to get past sort of the anthropomorphization of AI, because that's what's holding everybody back.
Like AGI is about robot fantasy land, and that leads to all the nonsense about destroying jobs and blah blah, blah.
And none of that is helpful because you have to.
Then you dig yourself out of that hole to just explain wow, it's really really good at writing a case study.
Right.
Right.
Which it writes a better case study than all the people that work for it.
But it doesn't know who to write it about.
It doesn't know what necessarily you want to emphasize.
It doesn't know what the budget is, what's needed, how many words.
But it also turns out like AGI just does an awful lot of work.
So, for example, someone asked me recently.
They say well, are you worried that if we have AGI, then you'll no longer be investing in software companies?
I'm like, well, I mean, you're AGI.
Right.
I'm still investing in software companies, right?
And so like just because your AGI says nothing about economic equilibrium or economic feasibility et cetera, so like just the term AGI does basically infinite work for every kind of fear we have and maybe every hope that we have.
And the more we tie it down to like not only it solves a class of problems, but the economics pencil out yes or no, we can actually have a more sensible discussion, which I actually I think is finally entering the discourse.
I think we're actually talking a lot more sensibly now than we were a year ago.
And so when people say things or the AI 2027 paper, when they talk about sort of automated research or recursive self-improvement, does that feel like fiction or fantasy?
Or does it feel like, or is it thinking, that even with those things we're sort of nowhere near peak software and there would just be unlimited sort of demand?
I think you've got to go first for each question.
I need you to anchor us in reality and then we can deviate.
Well, I think that first, I'm just not a fan right now of buying into anything by year.
Because whatever year you want to buy into in 2027, we're just going to be having a fight over what we meant by the metrics.
And it just turns into like OKRs for an industry, which is just like a ridiculous place to be.
That's really funny.
I think that everything takes 10 years, but you can't predict anything in 10 years.
So how do you even reconcile that?
And I think that you just have to recognize that we're on an exponential curve.
So no one's predictive powers work.
And it's just going to keep happening.
It's not going to plateau.
It's not going to... all of a sudden we're done.
That's what makes this a different platform shift.
If you look at the progress, and that's the same.
That went through with storage, that went through with bandwidth, that went through with productivity on computing, on connectivity around the world.
Because it's exponential, you can't predict it and it's just folly to sit around and try to predict.
Now you could do science fiction and you could say in the future, when we all have our personal AI with all this other stuff, and then that's great.
But then you say it's going to happen in 2029.
You're an idiot.
Yes.
That sounds totally correct because basically three years ago you would not have been able to conceive of Cloud Code or cursor or name your background agent writing code.
So it's like, what is the point of having some date at which you're naming something?
And so we've actually seen probably vastly more progress in the past just two years of actual applied AI than we would have thought.
And yet, does it matter that one or two of the predictions didn't play out?
No.
So I think it's probably more interesting to think about like where is the technology?
From more of a classic Moore's Law standpoint.
Like, how much compute do we have?
How much data are we working through?
How powerful are these models?
Let me ask you like as semi-old I mean like nobody after AI collapsed and machine translation and machine vision failed, you couldn't find anybody who thought that those would become solved problems.
Yeah, totally.
Or after neural nets imploded and like literally you were teaching.
Or expert systems.
Or expert systems.
But you were teaching and if you tried to teach neural nets like, the students would rebel because you were wasting everybody's time.
In 1999, like Hinton couldn't get funded.
Trying to do neural nets.
Grad school was this three-volume history of artificial intelligence thing.
Neural nets was like eight pages.
You know, ironically, I remember when ML was the cool thing and neural nets was the old thing.
And now like, you know, ML is like the old thing and neural nets are the cool thing.
Right, or NLP.
And so the fact that, so we will return to all of these problems that couldn't be solved.
Like even like this, everyone's favorite one.
Oh, it doesn't understand math.
Like, okay, that is a solvable problem because math is solvable.
Like there's just no one put the math layer in to understand what a number was and, to you know, hard code it and just build in an expert system for math, which is actually a well-understood thing, because we've had maxima since like 1975, you know.
I think it's important to, like, maybe for us to describe how hard it is to predict anything, right?
So let's take recursive self-improvement.
This is one of my favorite ones.
And then, of course, you look at that and you're like... It works.
Right?
So I guess you know, like from an intuitive lay perspective, every time you have a box with an arrow back in it, you're like okay, we're done right.
But, like if you know anything about nonlinear control theory, answering that question is one of the most difficult questions that we know in all of technical sciences, right?
Like, does it converge?
Does it diverge?
Like, does it asymptote, right?
So...
For example, you could recursively self-improve if you're doing basic search but you asymptote right.
And so, like saying, recursive self-improvement from like a deeply tactical perspective says almost nothing.
It says.
But unfortunately, because we tend to anthropomorphize AI, we say recursive self-improvement and all of a sudden we're like.
And then it like, overcomes energy boundaries and human intelligence.
Well, that's how it goes from being a toddler to being, like, an eight-year-old.
It's just because it...
It recursively solves the truth, right?
I mean, the reality is like nonlinear control systems, which are feedback loops that are adaptive.
We don't even have the math for a relatively simple system to understand what happens.
You have to actually know the distributions that come out and go into them.
And so these things are going to improve.
They're going to continue to improve.
Maybe they'll improve themselves.
But just because they do improve themselves doesn't mean they can continue to do it.
And this is kind of part of this entire journey as we're learning about these systems.
Again.
The good news is, I think we're talking a lot more sensibly now than we were a year ago, and hopefully that will continue.
Hopefully the discourse can recursively self-improve so we're just more sensible.
Well, the good news is that's involving humans, so we don't actually have to worry.
But I think that...
I mean, you must be seeing this even with customers.
I mean take the conversation about hallucinations and things like that.
How dramatically that's altered in just the past two years, say?
Yeah, on two dimensions, actually.
So on one dimension, the problem of hallucinations has improved.
So, as the models get better, as our understanding of how do you, you know whether it's RAG or whatever you know, even the problem of actually the efficacy of the context window has improved.
So you have the technical improvements, you know, kind of across the stack.
And equally, you have a kind of a cultural understanding to some degree within the enterprise as to like okay actually no, these are non-deterministic systems.
They're probabilistic.
So you're starting to see almost a culture shift, which is okay.
You can actually implement AI in essentially more and more critical use cases, because the employees that are using those systems understand that they do actually have to do the work to verify it.
And then the only question is is what is that ratio of time it took to verify versus if I had done it myself, and how much efficiency gain for whatever that workflow is?
But we're going from probably like two and a half years ago, where there was this instant excitement as to oh my God, this is going to be the greatest thing of all time, to a reality check within three to six months, because everybody was like hallucination is going to be the massive kind of problem, to now, a couple years later, after that, which is like okay, we're seeing the hallucination rates shrink, we're seeing the quality of the outputs increase, and we understand that you do have to go and review the work that these AI agents are doing.
And that takes on a different form depending on the use case.
So in the form of coding, that means you just have to go review the code.
Which you had to do anyway.
People seem to be forgetting.
You had to do anyway.
But there was probably at least a little bit of theory as to what part you should go review with extra level of detail, because you kind of knew the person you were working with.
It also implicitly limits the value of AI, which people are uncomfortable with.
Right.
It just basically says it helps people that will know more than the AI does.
And as soon as it knows more than, you know, like it starts to actually kind of bisect the utility.
Yeah.
Yeah.
Basically, it's super interesting, which is the experts are now becoming.
The productivity of an expert is outpacing everything else.
Which is?
Which was this?
You know?
I think we could have probably predicted it based on historical events.
And I think you've got some good theories about how you know the skill, the type of skills that are, that you know the kind of the right user for these models, for the kind of use case.
So we're seeing that where the expert engineers are like I don't mind that it's a slot machine, where I'm pulling it and I see what comes out, because I know I can still get 10x productivity.
It gives me good ideas.
Yeah, and I get it good enough that it's worth that productivity gain.
Whereas if you were not an expert engineer and you did this slot machine, you probably would try and go and deploy all the ones that were also wrong.
And you actually don't know which lever to pull.
A big thing is literally knowing what to ask for and what language to use will get you a better response.
I think that this is just an incredibly important point that you're making, and it really gets to the heart of what it means to use a tool.
Like, you know, you put me in front of, like, a 12-inch chop saw and say, like, go fix the fence.
Really, really bad idea.
I mean, I could go buy one.
I could cruise the home.
And I'm like, dang, man, I don't have a DeWalt.
And I could buy it. but it's really not a particularly good idea.
And I think that how these platform shifts happen and why there's so much excitement over coding is that well, the best way for a platform shift to take hold is it's the experts that are the closest you have to.
But what's neat is that was like the OG place for computer clubs.
Oh, nice.
Like in the early 1990s and the late 80s, Like if you ever wanted to meet the computer and you would go and like this is like Halt and Catch Fire.
And it's like a bunch of people with soldering irons and shit.
And like they're, that's who, and you know, when it didn't work, when something was broken, that wasn't like oh man, these things are terrible.
I'm wasting all my time.
That was like the whole meeting.
It was like, who could get, like, one of these new discrete graphics cards to actually work?
And debug the driver.
Can anyone print?
Is there anyone in this room who can print in this new thing called PostScript?
And I think that's what's really happening right now.
And so, first, it's obvious it should happen with development and coding first.
Yeah, because they're the most forgiving and the most understanding of, like what's a bug, what's a thing that can never get fixed.
And the thing to watch for is no one is saying that coding can't get fixed.
Like whatever it's generating, that's bad for like a 2x coder rather than a 10x coder.
No one is saying well, that'll never be fixed.
And then the next thing that's going to happen is going to be what I think is just going to be like the creation of words, like the marketing document, the positioning document, all of this long form stuff where, if you're really good at that job, you know the right questions to ask.
You know what looks good.
And then you can get really domain specific, like on the next level is like oh, I need to understand like a competitor which then is using real information from the internet in real time, not just statistical.
And then you're like, well, they already know what the competitor does.
Right.
And then my favorite scenario is the one that just constantly, just has these aha moments is attack this thing.
I'm not interested in you adding em dashes and making it a little bit better.
I just want to know what did I miss?
You said one, I think, on this last one about like, here's my earning statement.
Yeah, for people that's the thing you read, that you read after to the analyst that now like attack it like an analyst and there's like 6000 hours per company of analyst questions.
It knows what they're gonna.
They only ask like three questions anyway, expense line, you know and I feel like Do not watch this if you're an analyst.
This is not any advice about being an analyst.
But this is what's really going to happen with writing, and then it's going to happen with PowerPoint and slides, and then it's going to happen with video.
It's really important to call out, which is you're getting the consensus, mean response and so, in the limit, it's offloading a lot of kind of busy work.
If you're a professional, like you're a professional, you actually know all of these things.
You just don't have the time to go through all of it yeah, and you may not remember it so.
So in a way it's it's it's, it's productivity helpful, but it's not, you know, solving some problems of where you know you are a particular expert in, and this is maybe why, for those that are non-expert, it's a little bit more threatening because it can do that job.
Yeah well, maybe to bridge a few and probably throw in a different tangent, like so Stephen, you're asking, like so where is the enterprise now?
So that was the coding piece.
I think where you're seeing this is, you know, kind of clear understanding, which is okay.
What I'm going to get out will be correlated to what I put in.
So how precise I put the prompt.
What like?
I think prompting doesn't go away anytime soon, simply because the leverage you get on the set of instructions you're going to give the AI at the start is still going to be massive.
The prompting went away, what would you end up with?
Well, I mean two years ago, I think people were like you'll just tell the AGI what you want it to produce.
There's just one prompt, like you unbox it and you say, go do something.
Be a software engineer.
No, literally, that was like an open debate.
And it was like no, you're probably missing the fact that what is in my head is going to be unbelievably germane to the thing that I'm trying to produce.
And like, I have to somehow give you that context.
Like there's no world where you have that context without me telling it to you.
And now you're seeing it.
Like you're seeing these incredibly unhinged prompts, which are like pages long.
And the output you're getting from that is actually like way better than if you didn't give it that context.
So I think there's a clear understanding of that side on the enterprise use cases and then a clear understanding that you've got to go and review it.
And then on this point about like well, you know What is?
We forget that formal languages came out of natural languages.
For a reason.
We didn't start with formal languages like oh, it's much easier to speak in English than to speak in English.
It's the opposite.
We have this natural language.
We're like, it's very tough to convey the information that I want to.
You and I are both experts.
We understand the solution space, so let's communicate more efficiently.
So to think that this somehow wouldn't happen... And that's what jargon is.
Of course.
Jargon is just a formalized way that people who have domain expertise talk to each other.
That's exactly right.
So when does the style of work change?
Changed because of the tool versus the tool, sort of adapted to the style of work.
And so what I'm starting we're like only in day one of this, but what I'm starting to see kind of some patterns emerge, which is we thought agents would go and learn how we work and then automate that.
And so basically agents conform to how we work.
The question is, when is the moment when we conform to how agents are best used?
Yeah.
You're seeing this in a couple of areas.
You're seeing this in engineering to start with, which is like people are saying okay, I'm going to have agents and then sub-agents for parts of the code base, and then I'm going to give them read me files that the agents read and then I'm going to actually optimize my code base for the agent, as opposed to the other way around in other forms of knowledge work.
Within how we use Box with our AI product.
You're starting to see people basically tell the agent it's complete job and the agent is almost dictating the workflow in the future, as opposed to it's just mapping to the existing workflow.
So I don't know like, what the history is on this of like, when does the work pattern itself shift because of what the technology is capable of?
But I think probably where this goes has to be some version of that, which is it's not going to just be the agents just plop into how we currently do our work and then just automate everything.
I do think you start to change what the work is itself and then agents actually go in and accelerate that.
Well, as important as that is, it's actually more important.
Because what happens is to reuse the word in a different way this anthropomorphization of work.
What happens is that the first tools actually anthropomorphize the work.
If you go back, this is every single evolution of computing.
How long did it take for Steve Jobs to get rid of the number buttons on a smartphone?
Like, they still had number buttons.
Or like you, look at cars and until Elon got rid of all the controls, everybody kept all of the controls.
I don't want to get in that fight.
But, like what happened with every technology shift, is you know if you were to look at what accounting software looked like in the 60s before IBM said stop why?
We all use double entry, but we need to have people skilled in how computers can do the accounting, not how people can, because we're never going to figure out how to close the books if we have to automate this whole room of people with green eye shades that have a manual process based on how far apart the desks were.
Right, yeah.
And everything that happened with the rise of PCs and personal productivity started off and I always use this example because I've watched it happen like five times now which is the first PCs that did word processing.
The biggest request was how do I fill in like expense reports?
And so this whole world grew up of tractor-fed paper that was pre-printed with the expense report.
And so then software, we wrote all of this code, like, are you using an Avery 2942 expense report?
Or is it a New England Business Systems A397?
And, like you know, and then you had like these adjustments in the print dialogue, like 208 inches, and you moved little things, and then you would print out like ate dinner 22, and that was all you printed.
And then someone said, you know, we could use the computer to actually print the whole thing.
And then like fast forward, and finally Concur said you know why not just take a picture of the receipt?
And then we could do all of it.
And so then the whole thing gets inverted.
And every single business process ended up being like that.
And then there are things that really, really do change the tools.
Like when email came along You know it used to be to prepare an agenda for a meeting somebody would open up Word and type in all the things and then print it out, and everybody would show up at the meeting with this very well-formatted and now and then, like email came out and that whole use case for Word just evaporated.
And then an email agenda became, no formatting nothing, just like.
Here are the eight things we're going to talk about and you show up and everybody's like did you get the agenda?
Yeah.
You know what's interesting about the AI?
One is it's kind of it's like we're seeing the same thing, but vis-a-vis AI.
So nobody really predicted the generative stuff.
And we've had AI for a very long time.
So we had chatbots we've had, you know and so you had these kind of like AI-shaped holes in the enterprise for a long time.
And a lot of the mistakes that we see today is people are taking the generative stuff and trying to kind of cram it into the old models where there's really a new behavior that's emerging.
That's very much more like it used to be.
You'd centrally sell you know AI to some platform team.
And then they would kind of try to get the NLP thing to work or the voice to work for like, talking to people on the phone for support.
And it was this kind of very central.
A lot of the adoption that we see is like much more individual, for example.
And so I just think that there is a bit of a mix match, as we're seeing now that it's getting ironed out too.
Well, and so I think the question is yeah.
Are we in the phase where we're trying to graft the agents and work in basically what we've been doing for 30, 40 years of software?
And is this going to be actually like the first real step function shift we've seen in what the workflow itself should look like?
Oh, we are.
I mean, like if you, you know, remember people like I tried to jam the Internet into Office.
Right.
And it was fun to watch.
But you were not watching.
But everybody around was trying to jam the internet into their product, because that's the only way you could envision it.
And it didn't really, like, you were like, well, where else would the internet go?
Like, there's no word processor on the internet.
Like, there's no spreadsheet on the internet.
And then other people would be like well, let me just try to implement Excel using these seven HTML tags, with no script.
That turned out to not be a really good idea either.
The best was like, let's do PowerPoint.
Well, how do you do it?
You give them five edit controls, tell them their bullet points and then we'll generate a GIF on the back end and send it back to you as the slide.
Okay, that was not... And so there was that whole...
I think actually maybe the main point is just the durability of Office.
It transcends all disruptions.
I like to think it pretty much rises above everything.
But the thing is, is that that's where we are now.
Is everybody... And, you know, like... But do you think... No.
I mean, just to dig a little bit.
So do you think this is... similar to the internet, and that is a consumption layer change.
So I always viewed the internet as very much a consumption layer change.
Like I go to a, you know, instead of going to my computer, I go to the internet.
But otherwise things kind of are the same, where AI has got this weird quirk which, for the first time I can recall, programs are abdicating logic to a third party.
Like, we've always abdicated resources, right?
So we'd be like, okay, I'll use your disks or whatever, but, like, I'm writing the logic.
But this time it feels like we're changing the consumption layer.
So, like you know, when my son, you know, talks to an AI character, you know he's not going to wellsfargocom, he's going to an AI character.
And so, like, that's changing kind of how we're interacting with the computer.
But also these programs are no longer kind of written by a human in the same way.
So I feel like the change is maybe a bit more sophisticated.
Oh, I think, but this is why it's a platform shift and not just an application shift.
Like, where each platform shift changes the abstraction layer with which you interact with computing.
But what that also does is it changes the way what you write the programs to.
Do you remember ever abdicating logic?
Oh, here's a great, here's an example of how disruptive this can be.
The first word processors in the DOS era, the character mode era, they all implemented their own print drivers and clipboard.
So if you were Lotus and you wanted to put a chart into a memo, you couldn't because you didn't have a word processor.
You didn't sell a word processor.
So you actually made a separate program to make something that the leading word processor could consume.
And if you were perfect, your ads said, we support 1,700 printers.
And you won reviews because you had 7,800 and Microsoft had 1,200.
That's a great one actually.
Windows comes along.
If you were trying to enter the word processing business.
Step one I need to hire a team of 17 people to build device drivers for Epson and Okidata and Canon printers, because you can't get them anywhere.
Microsoft came along and, for Windows, built print drivers and a clipboard.
And what happened was a bunch of developers were like, wow, this is cool.
Because now I'm just by myself.
When we did C for Windows We were like the demo.
In fact, at that Cumberley Community Center I would go and I would show brand new Windows programmers in 1990 like hey, you don't have to write print drivers and use the clipboard.
And literally standing ovation of all 10 people at the, But they were like more than happy to let data interchange between products.
Because they were like, that's nothing but opportunity for me.
Can we Like they probably from an emotional standpoint, felt exactly the same way as like a vibe coder does today.
Which is like...
You've just given me the platform, and it was just a print driver.
The writing code for Windows Book was, like, this big.
But the writing a device driver for an Epson printer was this big.
Writing it for a Canon printer was this big.
I'm just actually trying to think of like, but the paradigm shift is the same, which is there's been many times where we've reduced the amount of work a developer takes.
But I just don't remember ever where the programmer advocates logic.
Like, so, for example... SDM didn't?
Not logic.
Like, I would always say what is correct and what's not correct.
I think you undersold it, though.
No, this is the thing, by the way, everybody, that Martin invented and worked on.
But it's a big deal.
Maybe we should post more of your pitch at the time if you're not pitching this.
Well, no, no.
Let me, like, logic specifically, which is, I am writing an app.
My app is, whatever, some vertical SaaS app for a certain customer base.
The answer to the app gives is based on logic that I've written historically, right?
Like, if I run it on the cloud, the cloud is not producing an answer, it's providing resources.
If I'm using your device driver, it's, you know, providing access to a device, resources.
But if I'm like...
Hey, large model, tell me the answer here.
You're actually abdicating application logic.
Maybe you're right.
I think you're almost playing incumbent in the sense of trying to decide this is abdicating the logic and this isn't, when in fact, it really was a huge competitive advantage for WordPerfect.
And they didn't want to give it up.
And they fought against it.
And the number of people who didn't want to do like great clip.
And the next example, of course, is the browser, where people literally gave up.
Like, in Windows or in Mac, you could rasterize anything you wanted.
You wanted a button that you pushed and it spun and it animated like a rainbow.
You could do that in your product.
But then the web came along, and you're like, wow, I have to use a gray button that says submit.
And that was like... We do use a bunch of third-party things.
But it took a long time for those to show up.
And so, early in the internet, the magazines in particular, and the printed media were the ones who absolutely wouldn't go to the internet because they would not give up their ability to format.
And this is another part about the tooling and where what's going to happen with AI is that, like a huge amount of the productivity, software space today is like the preparation of output.
Like Office is basically a format debugger.
Like, all it is is, like, 7,000 commands for how to do kerning and bold and italic.
And, like, it turns out AI not only doesn't care, you could ask it to make whatever you want.
Like, you could just say, I'd like this to be a double index pie chart thing.
That's not a thing, I just.
But you can do that and it will just figure out something that looks like that and you'll go ooh cool.
And this was where to this disempowering and experts and who's not an expert.
When productivity software arose...
The big thing about it was that there were people who figured out how to like make like killer charts.
Like, Bennett Yevans, like, killer chart guy.
And there were people who were like, every meeting started with, how did you make that chart?
Like, I could be on an airplane and somebody would be, like, making a shitty chart.
Oh, so interesting.
So, like...
In this case, the abdication is like actually, what's the way to visually represent the data, which is absolutely an abdication?
Right and it turns out well because like 90 of the people never really got to be expert at doing that task, even though 90 of the tool is about like to.
Even so, what happens is each generation, But the programmer didn't abdicate the logic in this case.
This is the user, we.
But you know what's the user, what's the programmer in that, and and in fact, what the programmer was doing was like like we would invent the thing called wizards or whatever you know, and that would make a whole bunch of choices for you style sheets or whatever and so, in a sense, we were making a bunch of choices for the user which, to the experts, looked like disempowering the experts who were tweaking, and so i.
This is all like this.
There's some Steve Jobs quote that he loves about Schopenhauer how, if you've seen the Conjurer, The grand continuum of If you've seen the Conjurer, it's not a trick anymore.
And I really feel like this is like the third or fourth time that this has happened just in my lifetime of watching this.
So something that's really caught my attention because it's the most senior people I know is that a lot of very senior developers are spinning up a lot of background agents, like code agents, and they're interfacing at like, the GitHub PR level.
Right.
And so it's not obvious to me why you'd do a bunch as opposed to one, and it's not obvious to me why you wouldn't interact directly.
So it feels like something's going on here, but I'm not quite sure what, and I would love your thoughts.
Well, my read on it and then I guess I would kind of sort of throw out what then happens next as a result of this, because to me it's actually a little bit of an epiphany on what the future work design could look like in this world, because engineers back to the prior conversation are just the first to experience this.
But I think what my read from talking to kind of similar folks that are all in on this, is this mix of basically effectively the context rot problem, which is the more that we put in the context window, the more it gets confused, the lossier the answers get.
And so you have to have some kind of way to partition what an agent should work on.
And we see this in building agents internally, which is, you know, the panacea that I think we maybe would have hoped for is like well, you just put a million tokens into the context window and then obviously Oh, so you're saying this is almost like a counter trend to the AGI.
It's almost like the opposite.
It's the opposite, but it only works because the models are so good.
Yeah, but you're giving more things more specific tasks rather than one thing less specific tasks.
Right, but I think this is why it's happening.
So basically, the craziest version of this is I was talking to somebody who who is in startup land and they have to your point, they have all these sub-agents, but what's amazing is it maps one-to-one to each microservice in their code base.
If you just said here's my entire code base, you know, go run wild, you know, to one agent, it will just, you know, produce worse and worse code over time because it's going to have context rot.
It's not going to know exactly what you're trying to do in that one area of the microservice.
But the subagent model seems to be working for that paradigm.
I love this counter pattern because everybody is like they're going to like you know, models will get you know smarter and you'll give them higher level tasks and they'll do things longer.
Yeah.
This is a counter one.
I want to tweet that, but you have more Twitter followers.
We can collectively do it.
But then, so then the question is, okay, so let's just assume this works in engineering.
You have this interesting dynamic, which is well then.
That means that, like some of the coding practices will be pretty different in the future.
We've talked about this idea of, you know, the individual engineer becomes the manager of agents.
So that was already kind of, I think, a well-understood path.
This is like a supercharger of that concept.
Yep, yep.
And then the question is, like, how does that translate to almost every form of work?
Obviously, one, just the sheer leverage now you get is going to be insane.
But I do think the way that you might even organize the work and what the workflows within an organization are inevitably going to change as a result of that.
Oh, but I mean, I think this just gets to the.
You know essentially, that the flow in the workflow has been serialized or linearized, based sometimes on knowledge but other times on tooling.
And so what happens when the tooling changes is you just get this realignment of what's truly serial and what's not.
Like if you're...
If you're planning an event for a company which is still going to keep happening, you know like oh, I have to book the venue, I have to invite all these people, we have to create all these materials.
Well, they're actually not particularly gated on each other.
But if you have an events person, they're gated.
And so now an events person can start spinning up all of these different elements, and then they're going to come back.
Like, I've gotten as far as I can on collateral until I get a logo for this event.
Like, I've gotten as far as I can on invites until I get the date and the time and the venue.
And I think there's no reason why you can't spin off all those in parallel because, of course, how does that happen today?
Well, if you're a company and you use Box and you've done this is your 58th event you know you have a folder called event.
And people take the folder and go event 59.
And they make a copy of it and all the stuff in it.
And well, if you think about that workflow, that's exactly what a series of different background tasks or agents could go do.
And so I think the reason that you could be doing all that encoding is well, there was a natural tendency.
There's a natural way to break that up, because there's a bunch of programming that's not Right.
But there's the other side.
But there's also a bit of an indictment on the ability of you to give a high-level, You know.
It kind of suggests that the human being needs to be, you know, giving them more granular orders.
Otherwise, you know, to start a company, you'd issue one prompt, you'd go to the beach for six months, you'd go back and you'd have a full company.
Which is this almost re-anthropomorphizing effect, which is, like, it turns out we did...
We did kind of figure out division of labor.
We figured it out in the context of a lot of physical, you know kind of analog limits that we clearly had that agents won't have.
But we now, you know, there's no kind of, you know, total free lunch.
So you have this context rot issue, which is that you do actually have to subdivide the tasks at some point.
So then the question is like, what are the rights?
I mean, it may not be a context fraud.
Like, the Occam's razor here is you need to give them specific instructions for specific tests.
And if you give them higher-level instructions independent of context, they just don't know what you want.
Right, and this gets to the formal language part.
Like, at some point, if you tried to use, like the Uber to get the whole thing done, you have to tell it the whole thing.
Yeah, exactly.
And that just seems like a lot of work.
Whereas if you have to tell it less because the part of the model you're using knows more.
It's basically a different way of thinking about templates or a different way of thinking about starting artifacts or scoping the context in a generic world.
But then there's this, I mean...
It might though, be the right architecture in general, if you assume that you know we're never going to get to a point where the model is just 100 perfect, right.
And so it might also be the right kind of architecture design, because at some point you're going to have you don't want an agent or a set of agents to go so far down a path when there was a step that it needed to check in with you on, because there's just a compounding effect of that.
So you do need to kind of subdivide the work also, because if you do have gating, you know moments that are going to have a bunch of dependencies.
The agent does need to know, like at what point should I roll that back up to the user?
Yeah, against the common narrative.
Now that I think about it, it seems that the trend is prompts are getting more complex.
Yes, not less.
And we're seeing more agents, not less, doing more narrow tasks, which is almost this kind of counter-AGI narrative.
It's almost like these are much more specialized and much more deep looking, with much more specific instructions.
And there's sort of a history of this, wow, maybe we can actually solve it if we're specialized.
Like if you take expert systems.
At first they thought expert systems would just be experts and they would just know.
And then, like by the time you got to the actual published research, like at Stanford, it was like this is an expert system in deciding on what type of infectious disease you have.
As long as you have one of these cells, No, literally.
There was a paper that was like there's just one digestive disorder that actually is a metabolic disorder.
I do want to though, because you wouldn't want like.
There is one big difference, which is somehow the model itself is packing in the inherent intelligence or capability to solve all of these problems.
Like, we are benefiting from the fact that at least you can build these all on Cloud 4 and GPT-5.
And that all on a computer, too.
Let me try to demonstrate this one with an old person example on this, one which was like early in the PC era there were words processors and spreadsheets and graphics and databases.
And a lot of people were like, why are there these four programs?
There should only be one program.
My answer to that, which often involved screaming, was, have you been to an office supply store?
Because if you go to an office supply store, there's like paper with numbers, and then there's blank rectangles of paper and then there's transparency paper.
And like, this has been around a really long time.
There's some reason that these are different.
How many minutes did it take you for you to know Google Wave wasn't going to work?
Zero.
Okay, okay, okay.
It was instantaneous.
It was instantaneous.
No, I mean, but this was the thing.
There was a product, an ancient Mac product, that was lauded by the industry, called Claris Works, which was like oh it does, you could have a spreadsheet inside a word processor.
And my first reaction is, have you seen a person use a spreadsheet?
Because their monitor can't be big enough.
So they just want as many cells as you could possibly have.
And you're sitting there saying it has to fit on an 8.5 by 11 sheet of paper on a Mac.
And I think that one of the things that happens is that these lenses that humans bring to specialization like really, really matter.
And if you think about the medical profession and you think about going from a GP to the radiologist, to a specialist, to a nurse practitioner, through the whole series, they're each going to look at and use AI in a different way.
So then the only thing would be okay.
So that was that level of specialization and division of labor emerged over a 100-year period with, you know, alongside tools, but also with, driven by a lot of the physical constraints and realities of how organizations emerge.
So the only question would be in a post-agent world, in 10 years from now, do those divisions of labor look exactly the same?
Or do those shift also because the agents collapse?
You know some of the functions and is there some blurring?
And then is there just a new set of roles?
Like clearly there's a role in a bunch of organizations emerging which is like no, I'm just like.
My role is like I'm the AI productivity person.
And, like I just like, have a way of, you know, creating all new forms of productivity in the organization with AI.
So like, clearly we'll have a bunch of new roles, But is our current division of labor going to also collapse in some interesting ways because of AI?
Well, I think that like, if you actually stick with the medical example, we're just going to wake up and there's going to be way more people with way more specialties.
Right.
And AI will have created more jobs.
And in the interim— Do you think AI causes more specialization over time?
Absolutely.
Yeah.
Because every human is going to be way better.
Right.
And more knowledge will amount.
And I think this is a thing that has really happened with computing that people forget.
Like there used to just be like this morass of marketing.
Right.
And R&D.
Right.
And all of a sudden, there used to just be coding.
And then there was coding and testing and design and product management and program management and usability and research and all of these specialties.
And all of those had their own tools.
Go to a construction site.
I remember growing up, Our neighbors built a house.
We lived in an apartment, and they built a house, and there was Clem, the carpenter.
And you built a house with a guy named Clem who used all the tools and everything.
And now, like you build a house and it's like this 20-person list of subcontractors, all who have whole companies that do nothing but like put in pavers.
You know, and that's what it's going to be.
I mean, there's been a long disaggregation in the history of IT, right?
Like everything in the same sheet metal.
Then you know, disaggregate the OS and the hardware.
Then you disaggregate the apps.
Right.
And then it was kind of interesting like in the last 15 years we saw the app and like independent functions got disaggregated right.
It's like almost everything became like an API would become a company, right? out.
Twilio's like Auth became a company, like PubSub became a company, et cetera.
And so it may very well be the case that every agent becomes like a whole new vertical and a whole new specialization.
And then you can actually build a company or Like.
It may be the case that today, just like with APIs, one company will have a whole bunch of agents.
It may be the case in the future that a third party will provide that agent as an independent company.
Well, it's so... The opportunity, to your point, is really there for that.
Yeah.
Because, like...
It used to be, like, the impedance to creating a company and distributing was infinite.
It used to be ridiculous to think that a single API like Auth could become a company, but then you know, of course.
Or it used to be ridiculous to think you could build a whole company out of signing documents.
Right.
And like not just a whole company.
But then all of a sudden you realize wow, the addressable market for that is huge and it's way bigger than signing because of all the stuff that got done.
That was just Right surrounding.
Baked into a company causing headcount and waste and fraud and abuse.
Well, I think you can kind of underwrite thousands of these companies emerging.
So Jared Freeman had a tweet about basically like go deep on a workflow, you know, basically do the job of some part of the economy payroll specialist and then build an agent for that.
Yeah.
And it's not obvious that there's not literally like a thousand of those.
So by every vertical and every line of department.
I just love this because this is like literally the anti-AGHB.
It's basically following, like the long arc of computer science where, as the market grows, the level, the granularity can create a company.
Well, it's also economic growth.
Like take that payroll example.
Like today, just like...
Salesforce, which is always my favorite example, like the idea of having a productive Salesforce used to just be a consultancy.
Right.
And the only way you could ever fix it was hiring a consultancy to show up and analyze what everybody does and then do a report that says this is how you need to reorg.
And it usually meant go the opposite of whatever you have.
And then they would leave.
And then, you know, people tried, but there was no cloud.
So to build like CRM, you had to do all that consulting work and then roll it out.
And then it was static and you couldn't maintain it.
And then all of a sudden there's like oh, here's Mark Benioff and here's a whole way to do all this.
And not only that, the people actually like it.
And they think they're better at selling because they're using their phone and they're putting in a few notes about this client which helps everybody.
And I think that's what's really going to happen with all this.
And so suddenly something that looks really, really small becomes like a whole thing, because there's no problem with distribution.
There's no problem with customization.
You know, we'll actually have ways to solve security and privacy, just like we solved reliability and things like that.
And I think it's just.
I mean, look at, you know the stuff that you're a world expert in in the stack of internet technology, of networking technologies.
I mean, you would have asked me 15 years ago, would CDN be companies?
I never would have.
I'm like, that doesn't make any sense.
Like, how could you have a company that's a cache?
Yeah, yeah, yeah.
I think that people are probably way too afraid of the model providers kind of eating them.
And I think it was basically a phenomenon in the first wave, which was if you were just doing like basic, like if you had figured out that you could do something on GPT, you know two and three, where it was a text interface that produced more text.
Like, yes, ChatsBG ate you.
Like, that clearly happened.
But basically, since then, most enterprises want kind of applied use cases for AI and AI agents.
And so it's not obvious that the current crop of companies if you're doing AI for healthcare, if you're doing AI for life sciences, if you're doing AI for financial services, if you're doing AI for coding at the right parts of the stack.
AI for coding may be the one asterisk area which will be hyper-competitive, simply because the model companies don't want to use somebody else's product to build their own models.
And so that kind of almost forces them to get really good at AI for coding.
But with that as the one kind of exception, I think.
Basically we're just in a five-year period right now where you're going to have to build agents for every vertical, every domain, and there's a playbook that's starting to emerge of what that needs to look like.
I mean.
So I think there was kind of a technical head fake that happened early on, which was pre-training.
So the pre-training really was a 10 out of 10 technical innovation.
I can't tell you, like two years ago, if somebody was like I, had a friend that was building like their own aging model, post-training aging model.
Like, we're going to make it so good at aging technology.
Like you know, this is a text-to-image model and they wanted to make it so like old people looked really good at it.
And then, of course, the next version of like mini-training, whatever comes out, and it does a better job of it.
And the thing with pre-training was you're just kind of consuming all of the world's existing data.
You're draining all of that energy, and it perfectly generalized, right?
But it feels like technically that's past and now we're more in post-training and RL, which is a lot more domain-specific.
And so...
Well, in the moment that you have access to some set of data that is only just for that enterprise.
And so who gets permission to access that data?
Who gets permission to do the workflow on it?
It's going to be applied companies.
Yeah.
So, yeah, it is.
If we had an infinite number of tokens, then the models would just continue to generalize.
But it's pretty clear that that's not happening.
And so now we're going into which we all understand very well, which is now companies have to choose which domains they go into.
And they've got to solve the long tail of problems there to get access to the data, et cetera.
Right.
And I also think that the shadow having been the shadow, the shadow cast by large companies over, we're going to put you out of business and stomp you.
It's ridiculous.
And it has never in any technology way lived up to the fear that people have.
Look, if you built a new word processor in 1995, you were an idiot.
Mm-hmm.
Like, that was not the thing to go build.
You know, but you know, there was a time, just 10 years earlier, where like companies built standalone spellcheckers.
Like, it was just a thing.
You went to the store and you bought a spellchecker.
And, like, it had more words than the other spellchecker.
And so the thing that's not being said now, which we should do a whole one on, is what is the actual platform?
Yeah, this is a great topic.
It's all well and good to say that the large models will go subsume every application.
The thing is, the minute they start doing that, no one will be in their platform.
Yeah, this is a great topic.
Because like, no developer is going to sit around and say, if you're going to subsume me, then and this is there's a phrase that it's Sherlocking on the Mac and the Apple world to this thing.
And so it does, it has a real chilling effect.
And that's one of the things all the model people are going to learn very, very quickly.
There's a chilling effect, but there's also just I think there really is just a problem of like it's hard to go deep in 50 categories.
Like you just can't.
I think everybody's scared because pre-training was actually the one thing that was good at that.
And then now they have to actually, yeah, I agree.
You do have to like.
At some point it becomes purely just an execution issue, which is like.
I don't know how anybody would set up a company to be able to beat 50 startups across 50 different domains.
No, it's ridiculous.
And, in fact, like it's only good, because what happens is that the big company raises the awareness of a whole category.
And then you just swoop in and you go, to them, I'm just a feature.
Right, yeah.
But to you, this is my whole life.
Right.
And you're going to wait.
I always come back.
There's a whole company that just signs things.
Right.
Like, I cannot believe there's a whole company that just signs things.
I have so much to say about this topic.
I mean even minimally if you graph like the cost to produce so the willingness to pay for an inference versus the cost to serve it.
Something like for most companies for most spaces, 20 of the inferences are 80 of the cost.
Like, actually the problem of the application is just to choose those ones on which tend to be more domain-specific.
Yeah, yeah.
This is the problem of inviting the three of us on here.
We're just getting us to shut up in the trick.
Guys, thank you so much for coming on.
This is fantastic.
Thanks for listening to the A16Z podcast.
If you enjoyed the episode, let us know by leaving a review at ratethispodcast.com slash A16Z.
We've got more great conversations coming your way.
See you next time.
As a reminder, the content here is for informational purposes only, should not be taken as legal business tax or investment advice or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast.
For more details, including a link to our investments, please see a16zcom forward slash disclosures.