I have a very specific concern about AI.
Like.
Generally, I'm very pro technology and I really believe in the sort of like.
The upsides usually outweigh the downsides.
Everything technology can be misused.
You should usually wait, and you should wait until you eventually, as we understand it better, you want to put in regulations.
But regulating early is usually a mistake.
When you do do regulation, you want to be making regulations that are about reducing risk for innovation and actually authorizing more innovation.
Innovation is usually good for us.
I sort of have a very high level syllogism about AI that I've come to believe.
That, I think, is correct.
It's like if we can build, if you consider intelligence to be the capacity to solve problems from a given set of resources to a given goal, we are building things that are more and more intelligent.
Like we've built an intelligence.
It's kind of amazing, actually.
We built something like definably it may not be the smartest intelligence, but it is unintelligence.
It can solve problems.
So we built something that can solve problems like arbitrary problems from arbitrary resources and make arbitrary plans.
At some point as it gets better.
The kinds of problems that we'll be able to solve will include problems programming, chip design, material science, power production all of the things you would need to design an artificial intelligence.
At that point, you will be able to point the thing we've built back at itself.
And this will happen before you get that point with humans in the loop.
It already is happening with humans in the loop.
But that loop will get tighter and tighter and tighter and faster and faster and faster until it can fully self-improve itself, at which point it will get very fast, very quickly.
And a thing that is very, very, very smart, And you generate something that is very intelligent.
And by intelligent, again, I mean this capacity to solve problems.
There's lots of ways.
It's an English word, which means it has no one definition.
But in that sense, we mean that kind of intelligence.
And that kind of intelligence is just an intrinsically very dangerous thing, because intelligence is power.
Human beings are the dominant form of life on this planet, pretty much entirely because we're more smarter than the other creatures.
When was the last time intelligence came into, you know, entity at this level?
Right.
And within humans, I think people get confused between these two different things.
Within the band of human intelligence, intelligence is not the most important thing.
Humans have a lot of other attributes.
But intelligence is a very important thing for people.
Gorillas are more intimidating as a species than we are.
But gorillas did not build this studio.
And humans are like humans are dangerous as hunters, for example, because we're smart.
We do things like you know crowd the woolly mammoths off cliffs and set traps and like uh uh, We have whites in our irises.
The exact balance of white to color is what communicates what we're looking at.
Other primates have black eyes because they don't want the other creatures to see them, because they're trying not to give away information.
We have the white around the iris because... it lets you tell what other people are looking at.
You basically have this eerie ability to do it.
You know exactly what other people are looking at.
Imagine hunting.
These creatures are hunting you and you see one of them and one of them can like indicate to the other one across the way that they just saw a deer by like looking at the deer, and there's no sound exchange at all.
It's just purely this like and like.
You have to be pretty smart to do like big, very strong theory of mind to do that.
If you build something that is a lot smarter than us, not like somewhat smarter again?
Within humans?
The smartest people don't rule the earth obviously, But like as much smarter than we are as we are, than like dogs, right.
Like a big jump.
That thing is intrinsically pretty dangerous because if it gets set on a goal that isn't that like the first instrumental steps, instrumental convergence the first instrumental step towards achieving that goal is we'll go first step one.
If this is easy for you because you're really just that smart well, step one is to take over the planet, right?
It's like, then I just have control over everything.
And then step two is solve my goal.
Can you define instrumental convergence?
For those that didn't sit through three and a half hours of me talking to Eliezer Yudkowsky,
Yeah, so instrumental convergence is this idea that like...
Often, when you're trying to achieve a goal, step one is to achieve a instrumental goal along the way.
So, like, if you want to, like, for example, like... Paperclips is the one that... Yeah, yeah.
No, I'm thinking more like an instrumental convergence.
I'm trying to give an example from where it happens to people in their people's lives.
Oh, in chess.
Your actual goal is to checkmate them.
But, like along the way to checkmaking them, most of the time one of the things you want to do is take their pieces.
Now, you could not take their pieces.
There's probably, I bet, a really good chess player, I guess a kind of mediocre.
One could checkmate them without taking any pieces.
Just, like, trapping their... But, like...
That's like even more impressive.
But, like generally speaking, if you're just trying to win a game of chess, taking their queen, taking their pawns, taking their pieces is a good idea.
It like makes it easier to checkmate them.
And so I can predict something about almost anyone, any good chess player.
They'll take a bunch of their pieces.
They'll take the other person's queen eventually, probably.
In the same way, you can predict that corporations that want to expand into a new market.
There's an instrumental goal.
Step one, they'll probably hire people in that market.
It's just predictable.
And in general, if you want to accomplish most goals, like big goals.
Step one is like accumulate a lot of money and power.
Like, if you have a goal of changing the world, accumulating money and power, or accumulating followers who like, listen to you and will do what you say, these things are like, obviously good first steps.
Even if you didn't even know what the next goal was, they would be good first steps.
And if you know what it is, they're definitely good first step.
Step one, if you can achieve it, along the way to achieving any big goal.
Yeah, paperclips is the traditional one is first, if you can pull this off which, like humans, can't do this.
We don't think of goals like this because we're not capable enough.
But if you could pull it off, Step one would be like well, first I'm just going to like literally just make sure I have total control over everything at all times.
And then step two, I'll like do whatever.
Step two, do the thing.
Easy.
I already have control over everything.
No one can stop me.
I have access to all the resources.
Simple.
And I think people just don't.
It's hard.
It's hard to have.
People don't imagine sufficiently capable as sufficiently capable.
And some people have this idea like, oh, what if we just don't give it goals?
Well, first of all, we are giving it goals.
People are already building agents.
But let's just say we didn't.
And so you ask this oracle, what's the best way for me to accomplish goal X?
And it knows if it takes you literally and it actually answers your question correctly.
The answer to your question will be a thing that causes you to bootstrap an AI that then takes over the world and accomplishes the goal.
That's the most reliable way to accomplish that goal.
I just laid out a chain of argument with a lot of if this then this, if this then this, if this then this.
I know Eliza thinks that we're all doomed for sure.
I buy his doom argument.
I buy the chain and the logic.
I just think that, first of all, I'm less optimistic that the current set of technology is going to get to self-bootstrapping superintelligence.
I'm less optimistic than he is that or pessimistic, whatever.
I'm less sure than he is that when it hits that self-bootstrapping step, that that process will be fast and that we will, that there aren't important new discoveries that will take a long time on top of that that we haven't found.
I'm less sure that there's an idea of alignment.
You could make the AI such that it wants the same things we want and then if you ask it to do the thing, it won't go and do horrible things because it's not dumb and it's aligned.
And if it wants the same things, it knows what you mean.
It's smart.
And it has, it has aligned goals.
Hooray.
Like that'll work great.
I'm uh, Eliza thinks that we're like just alignments, this incredibly hard problem that like, is almost unsolvable when we're doomed.
I'm like, not so sure.
I think it's a more solvable problem than he thinks it is for a variety of reasons.
You know, just, it would take too long to like go into it.
But like my, my belief is that it's easier.
Um, and so, uh, as a result, like my P doom, my probability of doom is like my bid ask spread.
And that's pretty high because I have a lot of uncertainty.
But I would say it's like between like five and 50.
So there's a wide spread.
Which I think Paul Cristiano who handled, you know, a lot of the stuff within open AI.
I think said 25 to 50.
It seems like if you, if you talk to most AI researchers, there's some preponderance of people that give some percentage.
That should cause you to shit your pants.
But it's human-level extinction, I think.
Yeah, yeah.
No, no, it's not just human-level extinction.
Extincting humans is bad enough.
It's like potential destruction of all value in the light code.
Like, not just for us, but for any other species caught in the wake of the explosion.
Like, it's like a universe-destroying bomb.
Like, it's really... If, if, if, if, if, it's really bad.
It's bad in a way that's like... makes global warming not a problem.
It's bad in a way that makes normal kinds of bad not Normally.
I'm like yeah, we'll just roll the dice.
It's fine.
We'll figure it out later.
No, no.
This is not a figure it out later thing.
This is like a big fucking problem.
Southern Manhattan, Miami might go underwater.
Like, OK, but this is we're talking about it.
So why do you think?
I mean, it's like someone figured out how to, invented a way to make like 10x more powerful fusion bond bombs out of like sand and bleach that, like anyone could do at home.
Yeah, it's terrifying.
And and.
I've had enough time with it now that I can laugh about it.
When I first realized it was fucking heart stopping.
When was that?
Uh, probably like it's, it was, it was a dinner I went to before when opening.
I was just basically right after attention is all you need had been written like, and they sort of realized the scaling laws were there and i went to a dinner and someone was there and they were talking about it and they were like i think we were were uh, we actually might be on the the path to build with general ai.
Attention all you need was 2018.
Yeah, i think early 2017 2017 um, and google paper that started all this all off, and uh, And I heard about the problem.
I thought about it.
I'd been like, the AI Doom thing seems plausible.
Whatever.
Like, I don't think an AI is coming anytime soon.
So I'm just not going to think about it that hard yet.
And then I was like, oh, maybe, and then I started thinking about it harder.
And then I was like, oh shit.
Hi, I'm Logan Bartlett, the host of this podcast.
I just wanted to take a quick second to tell you that we have a bunch of killer guests coming on over the course of the next few weeks.
And so if you're enjoying these conversations with both entrepreneurs and investors, please do subscribe to our channel so you don't miss out.
I don't think an AI is coming anytime soon, so I'm just not going to think about it that hard yet.
And then I was like, oh, maybe.
And then I started thinking about it harder.
And then I was like, oh, shit.
Oh, uh-oh.
And so I guess the I believe the proper response is like unfortunately, this isn't the kind of thing where we can stop forever.
And it unfortunately is also the kind of thing where like more time is good.
Like I'm actually.
I'm okay with stretching out the time a little bit, but like ultimately to solve the problem.
I think this is one of my biggest points of divergence with Yudkowsky.
He is a mathematician, philosopher, decision theorist by training.
I am an engineer.
Everything I've ever learned about engineering is the only way you will ever get something that works.
If you need to work on the first try, it's to build lots of prototypes and models at a smaller scale and practice and, practice and practice and start building the thing.
But like smaller, and if there is a world where we survive, it is and everything goes wrong where we build an AI that's smarter than humans and we survive it, it's going to be because We built smaller AIs in that and we actually had lots of as many people smart people as we can working on that and taking the problem seriously.
And so I'm I'm generally I'm in favor of trying to create a fire alarm, where we were like like maybe not AI is bigger than X at some point.
Like trying to like create a, and I actually think there's good reason to believe.
Like nobody wants to end the world.
And this argument is not that hard to understand.
And so I actually think there's a good option for international cooperation and like treaties about some sort of you know the AI test ban treaty about not bigger than X.
At some point.
I don't think we're actually at the point where it just needs to be not bigger than our.
The current AIs are just not that smart yet.
But I think we should be moving towards creating that kind of a, some kind of soft.
I don't know.
We have to figure out what that looks like, because it's trickier to set that than it.
Setting that rule is way harder than it looks.
Writing good policy is hard.
We should be thinking about it now.
I just think we're not ready for it yet.
But in the meantime, on these smaller models, it is good that lots of people are fucking around with them.
It's good that we have more and more people trying to figure out how they work and trying to figure out how you can make them do things, how to figure out how to make them do bad things.
The best way to figure out how to stop a big AI from doing bad things without the best way.
There's a bunch of, it is true.
There's a bunch of failure modes for super intelligent AIs that don't exist in less super intelligent AIs.
And we better not bet on Oh, don't worry.
It works.
It works in the dumb ones.
It'll work.
No, no, that's not how it works.
You can't do that.
But like we will figure out we are figuring out more and more about the principles of how it works.
And if we survive, it's going to be because that process produces a good generalized understanding, a good generalized model of how these kinds of predictive models which I think includes humans, as we are some kind of very complex predictive model with other stuff too, but that's a big part of what a human is.
We'll understand how those work at a deeper level.
We'll have some science of it, actually.
And that's what we need.
We need a science of AIs.
Right now we have an engineering of AIs and no science of AIs.
And we need to use the engineering to bootstrap ourselves into a science of AIs before we build a super intelligent AI, so that it doesn't kill us all.
Why do you think people are struggling with the discourse around this?
Like very smart people seem to object.
It's very, very obvious.
Mood affiliation rules everything around me.
People don't make decisions on a reasoning basis.
I mean, myself included most of the time.
I happen to like, I happen to... find this problem interesting and compelling on its own.
So I spent a lot of time digging into the arguments themselves because I was drawn to it.
But most of the time, I make decisions the same way as everyone else does.
The person pitching me on the Flat Earth thing who's saying a bunch of things.
I don't really listen to them.
They say some things and it triggers like, oh, you're part of that tribe.
Those people generally I don't think think very clearly.
I'm just going to discount everything you're saying and ignore you.
Yeah.
Robin Hanson talks about 9-11 people and it's like you don't argue with them.
You just sort of move on.
Yeah, exactly.
And the AI people sound like religious nuts who are telling you about the end of the doomsday into the world.
And it sadly pattern matches really nicely, right?
Like the AI is like the, you know, is like the antichrist it's coming and you know if, what?
If we're good, the good ai will come and save us from that.
It's like it sounds like like uh, christian rapturists yes um, last of us.
What was it?
Yeah, it uh unfortunately um, reasoning from fictional evidence is it doesn't work, and that mood affiliation reminds me of this is not an argument.
And the earth is not round, because the flat earthers sound crazy.
The earth is round, because you can demonstrably see that the earth is round and go measure that yourself.
And it's true.
And if you make decisions based on anyone who's telling you that doing X will unleash a force which is going to kill us all.
They sound like a bunch of crazy religious people.
Because it's never happened before.
It's never happened before.
Guaranteed, guaranteed the first time that's true, we're all dead.
Because your algorithm always predicts the same thing.
The thing you're going through in your head always predicts the same outcome for anyone who is predicting doom from creating a powerful force beyond human ability.
Now it is true, you should be skeptical in general when someone proposes that, because there are a lot.
There are infinity examples from the past 6000 years of history of people predicting that falsely about things.
And people made imaginary cures for medicine, medicinal cures for like tinctures that were supposed to cure you for a very, very long time.
And then we made one that worked.
And you just can't reason that way.
Sometimes it's new.
Sometimes it's not like before.
Usually it's like before.
And sometimes it's not.
And I am personally convinced this time it is not like before.
And I encourage everyone who's in that mode like the main thing, that what I want you to pay attention to is listen to me, like people like me, listen to people like even Yudkowsky, who I disagree with on the amount of doom, but we're like pro cryonics, pro technology, like technology is going to fix all our problems crazies.
If I have a defect, it's that I am too pro-technology.
I want too little regulation.
Are you doing cryonics, by the way?
I have not signed up yet.
I really should.
It's one of those things on my to-do list.
I'm failing the rationality test.
Yes.
But you're a techno optimist.
I'm a techno optimist.
And most of you should notice that the people who are affected by this particular one are not like the people who generally predict.
We're not Paul Ehrlich.
Paul Ehrlich is full of shit on like, oh the doom is coming, the population bomb.
He is wrong, and weirdly so, and refuses to learn from his lesson that he's wrong again.
He's wrong over and over again.
This is not like that.
Like i am not panicked about global warming, i think i think warming is a real, a real thing.
I don't want to say like i don't want anyone to walk away thinking like i think global warming isn't real, but like i believe i i believe in technology.
The engineers will figure it out, don't worry, it's going to be fine.
I'm like quite sure.
Um, it might have some, Depending on whether we have the political will to get around to it quickly or not.
It might have more or less dire consequences.
But like, we will figure it out eventually.
It will be fine.
We will figure it.
We will get this.
I do.
And it's OK.
So here I am being the guy who's like the techno optimist.
And I am like, no, no, no.
I think this thing, though, actually. maybe maybe a problem.
And I think if you're, if you are rejecting it because we sound like a bunch of crazies.
Just notice that like at least some number of people who are worried about this are not are on.
I'm on your team.
I'm on the Techno Optimist team.
Really think and it's not.
It's not obvious why it's true.
It takes a good deal of engagement with the material to see why it's true.
Because at first, at first, it seems like it shouldn't be that big of a deal, it shouldn't be that big of a problem, but then the more you dig in, the more you're like, oh well actually, but wait a second and it is, and so like I just encourage people to engage with the technical merits of the argument and I always welcome people to come to, to you know, ask questions but also like, if you want to debate no no, I have an idea.
We can align it this way.
Do this and it will create you know, it will create the AI that like, cares about things we care about.
We want your ideas like that's great.
Let's have an argument about that.
Self-improvement won't work.
OK.
If no because, like people have said, self-improvement will work in the past and it never did, then I don't want to hear it, because whatever.
It's not a real argument.
But if you have an actual like argument about how the current learning techniques won't work, for that for some reason, engineering reason
Absolutely.
Let's talk about it.
I'm glad we spent all this time building your credibility as a normal person just go off the rails.
But But seriously, I mean, it's it is something that people are like, genuinely concerned about.
And there's no incentive you have.
Mark Andreessen published something recently about like AI and he had a bunch of different points against it, but the AI killing everyone point was the first one and I think it went back to a lot of discourse around people that profit from this is their business and this is how they make money and therefore that's what, why they're intended to communicate in this way.
You are not i.
I, as far as i know, your business is not in dooms doomsday scenario planning around ai.
I have no financial stake in in either doomsday scenario planning around ai or the other weird like double think 40 chess thing people impute to uh, it is like oh, we're actually trying to build up OpenAI and Google as being super powerful.
And it's all an ego thing about making, Talking about how amazingly great this stuff is so that we can, like raise more money for opening up.
For regulatory capture or whatever.
Yeah, yeah, yeah.
And, like, I don't own any equity in any of these things.
You're not trying to help your 2006 batchmate or 2005 batchmate, Sam Altman.
No.
Sam Altman's really good at raising money.
Yes.
Like, he does not need my help.
Yes.
I promise.
Yeah.
What would you recommend to people?
I mean, if they are curious about all of this stuff, and like I, when people ask me about it because I've had Eliezer on, they're like so now what?
And I'm like, ah, you know, call your congressman.
I don't know.
So there's sort of two paths that you can where you can contribute if you care about this problem.
One is if you're technical and you're technically minded and you think that you want to work on it, like go learn how the eyes work and work on interpretability and work on courage ability.
And can you define those terms?
So interpretability is like understanding what's going on inside the AI.
Which we have very little understanding of right now.
Yeah, and corrigibility is how do you make a decision-making thing that is willing to be corrected by others?
And that's where Eliza has a lot to say about this, because it's kind of a decision theory thing.
How do you get it to sort of like believe that, even though everything it knows seems like this is the best way of doing it?
The other people are telling me that's not right.
And so I'm open to being corrected by their point of view.
How do you give something humility almost in a way, right?
And we have no idea to do that either.
We've made more progress on interpretability than corrigibility, but we've made a little bit of progress on both, and we're making more on interpretability.
And my real pitch for this is it's actually really interesting.
There's just stuff lying around that no one's checked, no one's tried.
It's like when the microscope was invented and suddenly like you could just make a scientific career by like pointing microscopes at things.
And like there were like lots of things to point the microscope at that no one had ever looked at before.
We have a mind you can go look at and examine and experiment with.
We've never had one of those.
And so there's a bunch of like interesting kind of like scientific work.
So I would encourage you to go do scientific work on A.I.'
's go.
Take the AI's other people have built, and then the engineering, and try to see what science you can do to them, particularly around interpreting what's going on inside them or how you get them to accept corrective.
If you're not technically minded and you don't feel like learning how to like actually build a transformer.
There's less obvious how you contribute directly.
But I think mostly like helping correct the debate and correct the tone of the discussion, because this, like the super intelligent AI must kill, it might kill everyone thing gets wrapped together with a bunch of other concerns about AI that are real.
They're things, but they're like normal concerns.
Like, will it just cause discrimination?
Will it- Job loss.
Will it cause job loss?
Like I know that stuff's all like a thing, but like honestly on those things I sort of feel about it the way I feel about most regulation.
We're like, you know, like, It's a little early, probably.
There probably will be a good regulation to write.
We don't know what it looks like.
The area is evolving so fast, very hard to write good regulation.
Don't let that get confused.
Job loss doesn't kill everyone.
Yeah, so there's AI ethics, and then there's AI not-kill-everyoneism.
The AI not kill everyone thing is like, we need the AI to not kill everyone.
Don't let people mush it together.
When you hear people conflating the two, correct the record, because they're just not the same idea.
And they're not the same kind of problem, and they won't be solved with the same kind of tools.
And the danger is that those two get sort of wrapped together into some thing that gets rejected by people who just like well, I don't, who reject the safety AI safety AI, ethics stuff that they think is like well, this doesn't really make sense, which I kind of agree with.
A lot of the people who say like, most of that stuff kind of doesn't make sense to me.
I could see a theory, maybe some of it.
We will throw out the good part with the bad and a lot of that other stuff is true of almost any technology, right?
Yeah, that stuff is just general, that's general technological like, and that's my thing of why I think Silicon Valley in particular has struggled to talk about this.