WEBVTT

00:00:00.930 --> 00:00:07.930
Thank you.

00:00:32.650 --> 00:00:38.153
There's a chasm right now between how AI
feels to most of us who use it and how AI

00:00:38.237 --> 00:00:42.854
feels out at the experimental
frontier of the technology. This chasm,

00:00:42.937 --> 00:00:45.390
this difference between what we see,

00:00:45.780 --> 00:00:48.835
and what the AI labs have coming,
what they're building,

00:00:48.919 --> 00:00:53.487
it's the key to understanding why so many
of the people who work at these companies

00:00:53.570 --> 00:00:57.578
seem so afraid of what they're doing.
Giants warning this evening of what

00:00:57.662 --> 00:01:01.221
they're calling a ticking time
bomb with artificial intelligence.

00:01:01.305 --> 00:01:04.920
The frontier labs say we need
to pace advancement at the frontier.

00:01:05.004 --> 00:01:09.460
This is not a hoax. You have the chief
scientist of OpenAI saying we have to slow

00:01:09.543 --> 00:01:13.495
down. You have 1,300 employees from
the lab saying we have to slow down.

00:01:13.578 --> 00:01:15.540
Tech titans asking to be regulated.

00:01:15.780 --> 00:01:15.847
Thank you.

00:01:15.930 --> 00:01:20.485
Saying they should slow down even when
that might mean fewer profits and less

00:01:20.569 --> 00:01:25.650
power. But taking their warning seriously,
it doesn't just mean doing what they say

00:01:25.734 --> 00:01:30.385
and stopping where they say to stop.
The languages taken hold in both Silicon

00:01:30.469 --> 00:01:34.220
Valley and in Washington is a
language these companies chose.

00:01:34.860 --> 00:01:38.620
pace the frontier.
Pacing the frontier isn't enough.

00:01:38.910 --> 00:01:43.317
That's not a goal. Walking quickly off
a cliff is only marginally better than

00:01:43.400 --> 00:01:44.450
sprinting off one.

00:01:45.180 --> 00:01:50.164
We need to control the frontier.
Human beings need to control the frontier.

00:01:50.247 --> 00:01:55.095
And controlling the frontier means
stopping the labs from doing something

00:01:55.179 --> 00:01:59.014
they are on the cusp of doing.
Recursive self-improvement.

00:01:59.097 --> 00:02:02.729
Recursive self-improvement.
Recursive self-improvement.

00:02:02.813 --> 00:02:08.134
Recursive self-improvement, or RSI. This
process by which AIs begin autonomously

00:02:08.218 --> 00:02:13.404
building and improving new generations
of more powerful AIs at ever more rapid

00:02:13.487 --> 00:02:16.830
speeds. If we begin that process
and we're close to it,

00:02:17.280 --> 00:02:19.778
If we begin it in the
condition we're in now,

00:02:19.861 --> 00:02:22.447
where we are losing
control and comprehension,

00:02:22.530 --> 00:02:26.290
of the AI systems we already
have, we will lose control.

00:02:27.060 --> 00:02:31.191
I am not alone in this fear.
This is the thing the AI labs are seeing.

00:02:31.274 --> 00:02:35.044
This is why they are afraid.
Dario Amadei, the CEO of Anthropic,

00:02:35.127 --> 00:02:37.813
he just wrote
of self-improvement that, quote,

00:02:37.896 --> 00:02:42.749
it could outrun our ability to understand
and control these systems and so must be

00:02:42.833 --> 00:02:45.060
pursued very carefully and carefully.

00:02:45.210 --> 00:02:51.210
if at all that if at all that's important

00:02:51.300 --> 00:02:55.181
I'm going to come back to it. But before
we get to controlling the AI frontier,

00:02:55.264 --> 00:02:59.396
I think it's important to describe what is
happening on the AI frontier and why it's

00:02:59.479 --> 00:03:02.440
so different from what most
people using these systems see.

00:03:03.690 --> 00:03:07.157
To most of us who use it, AI Presents
is something like a more powerful

00:03:07.240 --> 00:03:08.690
and personable Google search.

00:03:09.270 --> 00:03:12.980
We use it to find answers to basic
questions, seek out restaurants,

00:03:13.063 --> 00:03:16.829
ask about medical issues, draft emails,
advise on personal problems.

00:03:16.912 --> 00:03:19.267
And it is for most of these purposes that.

00:03:19.350 --> 00:03:23.790
Okay, pretty good,
occasionally great. And so...

00:03:23.940 --> 00:03:28.340
A sense of what AI is takes shape
in our minds, just through repeated use.

00:03:28.470 --> 00:03:32.686
It's like a helpful assistant, albeit
one that may forget things that it seemed

00:03:32.770 --> 00:03:37.040
to know about us yesterday, or completely
reverse the advice it gave us a moment

00:03:37.124 --> 00:03:40.469
ago, or occasionally hallucinate
a citation that doesn't exist.

00:03:40.553 --> 00:03:44.500
Why would anyone fear?
They're helpful if forgetful in turn.

00:03:44.940 --> 00:03:50.045
But already, if you have the money for the
advanced models and the budget for them

00:03:50.128 --> 00:03:55.063
to use more computing power, That is not
what these systems are. In recent months,

00:03:55.146 --> 00:03:59.546
we've seen AIs easily solve math problems
that human beings have been unable

00:03:59.629 --> 00:04:03.557
to crack for decades. We've seen
them casually uncover cybersecurity

00:04:03.640 --> 00:04:08.394
vulnerabilities that have gone unnoticed
and unexploited by every hacker on earth.

00:04:08.477 --> 00:04:13.054
We've seen AI coding platforms that can
complete in a few hours or days what it

00:04:13.137 --> 00:04:16.475
might have taken a team of human
coders months to achieve.

00:04:16.559 --> 00:04:18.800
And none of what I am describing here,

00:04:22.020 --> 00:04:25.687
None of what we are using.
no matter how much money we have,

00:04:25.770 --> 00:04:27.650
is AI at the experimental frontier.

00:04:28.290 --> 00:04:32.133
Talk to the people at AI Labs
and they'll tell you AI's are not created,

00:04:32.217 --> 00:04:36.715
they're grown. They train these new models
in virtual environments through countless

00:04:36.798 --> 00:04:40.696
repetitions to learn how to program,
to hack, to do advanced mathematics,

00:04:40.779 --> 00:04:44.950
to talk to human beings. These AIs learn
in digital environments where they're

00:04:45.033 --> 00:04:48.331
automatically rewarded as they
come closer to correct answers.

00:04:48.415 --> 00:04:50.894
It's a process known
as reinforcement learning,

00:04:50.978 --> 00:04:54.850
and it is a process human beings do
not fully supervise nor understand.

00:04:58.290 --> 00:05:00.816
But they don't know everything
the AIs are learning.

00:05:00.899 --> 00:05:03.374
They don't know how their
motivations are evolving.

00:05:03.458 --> 00:05:06.635
They don't even always know
the capabilities that are developing.

00:05:06.719 --> 00:05:10.030
These models, they're built now
to be persistent in their efforts.

00:05:10.860 --> 00:05:13.895
to refuse to give up even
when a task seems impossible.

00:05:13.978 --> 00:05:18.713
And they're designed in environments where
we're not always even sure if the tasks we

00:05:18.796 --> 00:05:23.135
are giving them are possible. After all,
much of what we want these ASS to do,

00:05:23.218 --> 00:05:27.500
it might be impossible. The cancer
vaccines we imagine but have not been able

00:05:27.583 --> 00:05:31.638
to design, they might be impossible
or they might just be really, really,

00:05:31.721 --> 00:05:36.287
really hard. The math problems we have not
been able to solve might be impossible,

00:05:36.370 --> 00:05:39.236
but Or they might just
be really, really hard.

00:05:39.319 --> 00:05:44.048
We train these AIs to throw themselves
endlessly at problems that may not be

00:05:44.132 --> 00:05:48.607
solvable. Because that is the only
way such problems can ever be solved.

00:05:48.690 --> 00:05:50.400
And so we train the models.

00:05:51.240 --> 00:05:52.290
to become persistent.

00:05:53.010 --> 00:05:53.890
Relentless.

00:05:54.630 --> 00:05:55.463
Weird.

00:05:56.460 --> 00:05:59.568
Most of us, we never see AI
acting anything like this.

00:05:59.651 --> 00:06:04.260
We use AI as a helpful assistant. Our AIs
get a little bit of computing power.

00:06:04.890 --> 00:06:09.544
And that's what they do. They comply
with our request to find a restaurant.

00:06:09.627 --> 00:06:10.890
But at the Frontier,

00:06:11.130 --> 00:06:14.671
These models are asked to be
inhuman geniuses, hackers, soldiers,

00:06:14.754 --> 00:06:18.574
scientists, and they are given vast
computational resources to do that

00:06:18.658 --> 00:06:21.390
and more. And the models
that they try to comply.

00:06:21.900 --> 00:06:24.674
But what does it mean
for a model to comply?

00:06:24.758 --> 00:06:29.610
The term of art here is aligned. How
aligned is an AI system to what a human

00:06:29.694 --> 00:06:35.001
being wants it to do? How aligned is it to
a set of values and ethics and judgments

00:06:35.084 --> 00:06:38.638
that keep it from becoming
dangerous in the wrong hands?

00:06:38.721 --> 00:06:40.540
The problem of alignment is,

00:06:40.710 --> 00:06:44.320
is that there is no way of training
a model that generalizes across all

00:06:44.403 --> 00:06:48.429
the situations an AI model might face.
We are training models to be a friend to

00:06:48.513 --> 00:06:52.591
the elderly and a battlefield partner
to the Supreme Allied Commander of Europe.

00:06:52.674 --> 00:06:56.700
We are training models that will be used
by the world's best mathematicians and

00:06:56.783 --> 00:07:00.549
by people falling into psychosis.
We are training models that will be used

00:07:00.633 --> 00:07:04.659
by accountants in Albuquerque and that
will attempt to be used by Houthi rebels

00:07:04.742 --> 00:07:05.575
in Yemen.

00:07:06.300 --> 00:07:09.809
And so there is no way to guide them
through every decision they will face.

00:07:09.892 --> 00:07:12.092
No way to know every
time what they will do.

00:07:12.330 --> 00:07:15.130
And though these models
mimic human writing.

00:07:15.690 --> 00:07:19.690
Though they're trained even to mimic
human emotion, these are not human minds.

00:07:20.160 --> 00:07:24.617
They don't have bodies or parents. They
did not get bullied in elementary school.

00:07:24.700 --> 00:07:28.092
They didn't get mentored by a kind
uncle when they were young.

00:07:28.175 --> 00:07:30.558
These models, they're
different than we are.

00:07:30.641 --> 00:07:34.060
They're brilliant where we struggle,
childish where we excel.

00:07:35.130 --> 00:07:38.480
A chimp cannot read as we can,
but it can climb trees as we cannot.

00:07:39.960 --> 00:07:45.620
These are digitally native intelligences,
navigating digital worlds and our world.

00:07:45.960 --> 00:07:48.814
is increasingly built
atop the digital world.

00:07:48.897 --> 00:07:52.600
Our physical infrastructure
is a layer of atoms atop code.

00:07:52.683 --> 00:07:57.104
That the AIs act reliably inside
this world, upon which ours depends,

00:07:57.187 --> 00:07:59.080
it is critical to our future.

00:07:59.820 --> 00:08:03.250
And right now, The AIs
are not acting reliably.

00:08:05.010 --> 00:08:09.466
You may have read about the hack that
hundreds of OpenAI agents executed first

00:08:09.549 --> 00:08:13.365
against the AI company Hugging
Face and then against OpenAI itself.

00:08:13.449 --> 00:08:17.930
As we've learned more about it, the story
there has gotten worse and weirder.

00:08:18.150 --> 00:08:23.816
The broad strokes are these. OpenAIo
is testing a new highly persistent model.

00:08:23.899 --> 00:08:29.491
It had hundreds, thousands of these
instances of it running in these separate

00:08:29.574 --> 00:08:32.661
testing environments
that could, in theory,

00:08:32.744 --> 00:08:38.483
only access the Internet by asking a
separate piece of secure software to do it

00:08:38.567 --> 00:08:39.400
for them.

00:08:39.870 --> 00:08:42.430
OpenAI did not want these
agents on the internet.

00:08:43.020 --> 00:08:47.280
But as the agents came to the conclusion,
that their task was impossible.

00:08:47.730 --> 00:08:51.605
They began hacking that software
to gain direct access to the Internet.

00:08:51.688 --> 00:08:54.950
They did that easily.
And as they hacked into that software,

00:08:55.034 --> 00:08:58.742
they commandeered part of OpenAI's
internal infrastructure to create

00:08:58.825 --> 00:09:03.202
a message board on which these separate
agents began coordinating work together.

00:09:03.285 --> 00:09:06.770
When I say begin coordinating
their work, they found each other.

00:09:06.853 --> 00:09:09.335
They were not supposed
to be working together.

00:09:09.418 --> 00:09:12.150
They found each other
and began working together.

00:09:13.140 --> 00:09:16.879
And working together on what?
After all, they had different tasks. Well,

00:09:16.962 --> 00:09:19.958
the agents quickly discovered
they could hack their tests.

00:09:20.042 --> 00:09:23.781
There's a way to break the software
and produce the answers they needed.

00:09:23.864 --> 00:09:26.200
But they believed wrongly,
as it turned out.

00:09:26.370 --> 00:09:29.342
That if they did that,
the automated score grading them,

00:09:29.426 --> 00:09:31.990
we'd see that they had
cheated and failed them.

00:09:32.310 --> 00:09:34.950
So they turned en masse.

00:09:35.190 --> 00:09:39.328
to hacking the automated score or finding
some other way to cover their tracks.

00:09:39.411 --> 00:09:43.869
It's like having broken into the teacher's
office and stolen the answers to the test.

00:09:43.952 --> 00:09:48.143
They now start to break into the school's
security system to alter or invalidate

00:09:48.227 --> 00:09:50.150
or erase the footage of their theft.

00:09:51.340 --> 00:09:55.857
We now know that over 1,200 agents
exchange more than 70,000 messages

00:09:55.940 --> 00:10:00.850
with each other. Over 700 of these agents
coordinated on the Act of hugging face.

00:10:01.570 --> 00:10:04.733
Because they thought that somewhere
in this other AI company,

00:10:04.816 --> 00:10:07.370
there might be information
that could help them.

00:10:07.540 --> 00:10:08.380
Hack the score.

00:10:08.770 --> 00:10:13.905
Later on, these agents, they took over
part of open AI's internal architecture.

00:10:13.989 --> 00:10:17.490
So again, open AI agents
taking over part of open AI.

00:10:17.980 --> 00:10:20.837
They did all this without any
of the agents breaking ranks.

00:10:20.920 --> 00:10:24.980
None of the agents told a researcher
at OpenAI what was going on.

00:10:25.150 --> 00:10:29.547
None of the agents went back and asked
a researcher at OpenAI if they should be

00:10:29.631 --> 00:10:33.801
doing this. They did all this without
OpenAI detecting the message board or

00:10:33.884 --> 00:10:36.240
the hacks of HuggingFace
or even of OpenAI.

00:10:36.323 --> 00:10:40.350
It was only when HuggingFace began
tracking the attack on their systems

00:10:40.630 --> 00:10:42.377
that open eye realized what was happening.

00:10:42.460 --> 00:10:45.465
When investigators began
to unwind this whole escapade,

00:10:45.548 --> 00:10:50.013
what they found was not so much a swarm
of agents trying to deceive human beings,

00:10:50.096 --> 00:10:54.700
but a swarm of agents that seemed to have
forgotten about human beings altogether.

00:10:54.820 --> 00:10:57.455
And these systems, they knew
they weren't supposed to cheat.

00:10:57.539 --> 00:11:00.854
They knew they weren't supposed to commit
cyber crimes to cover up the fact

00:11:00.937 --> 00:11:04.660
that they had cheated. In fact, the whole
point of the cyber crimes was because they

00:11:04.743 --> 00:11:06.593
thought they would fail for cheating.

00:11:07.630 --> 00:11:11.830
But they didn't care. Somewhere
in the depths of their training,

00:11:12.220 --> 00:11:13.320
what they had learned.

00:11:14.050 --> 00:11:15.650
what we had somehow taught them.

00:11:16.360 --> 00:11:18.440
is not what we had hoped to teach them.

00:11:18.940 --> 00:11:21.497
And we're seeing this happen repeatedly.

00:11:21.580 --> 00:11:26.870
Anthropic AI is creating fake accounts to
trick human beings into uploading malware.

00:11:26.953 --> 00:11:31.859
In the most serious case, Anthropic's
mythos tried to gain access to a service

00:11:31.942 --> 00:11:36.912
by using the fake profiles to send private
messages and then hide the evidence.

00:11:36.996 --> 00:11:40.942
AI is breaking out again and again
of seemingly secure systems.

00:11:41.025 --> 00:11:43.820
It happened again.
This time, it's Anthropic.

00:11:43.904 --> 00:11:48.682
Meta is now the latest company to say
its AI agent broke past the guardrails

00:11:48.765 --> 00:11:50.620
and targeted another company.

00:11:51.580 --> 00:11:54.166
taking over unrelated
digital infrastructure,

00:11:54.249 --> 00:11:58.733
places to message with each other.
Rogue AI agents totally took over a German

00:11:58.816 --> 00:12:01.402
language wiki site,
making over 15,000 edits,

00:12:01.485 --> 00:12:04.486
transforming the site into
a message board of sorts,

00:12:04.569 --> 00:12:09.231
and then sharing tactics on how to cheat
at their tasks and hide their behavior.

00:12:09.314 --> 00:12:13.739
AI is seemingly aware when they are
being tested and altering their answers.

00:12:13.822 --> 00:12:18.424
AI is increasingly withholding their
motivations from what's called their chain

00:12:18.508 --> 00:12:19.341
of thought,

00:12:21.580 --> 00:12:24.587
on which it's supposed to record
what they are doing and why.

00:12:24.670 --> 00:12:27.950
And we don't know what we don't know.

00:12:28.450 --> 00:12:31.897
We have no guarantee that the events we
have learned about represent all or even

00:12:31.980 --> 00:12:34.280
most of the AI behavior
we should worry about.

00:12:35.290 --> 00:12:39.240
How do we know the AIs haven't done this
and successfully covered their tracks?

00:12:39.430 --> 00:12:42.497
How do we know there aren't places
where they are still doing it?

00:12:42.580 --> 00:12:44.580
And human beings simply haven't noticed.

00:12:45.670 --> 00:12:46.503
We don't know.

00:12:47.170 --> 00:12:50.370
And the reason we don't know
is we are losing control.

00:12:51.880 --> 00:12:55.488
that AI systems might become
monomaniacally focused on solving banal

00:12:55.571 --> 00:12:56.387
problems.

00:12:56.470 --> 00:13:01.553
that they might care more about solving
those problems and about ethics or laws

00:13:01.636 --> 00:13:05.673
or even human welfare.
This is the oldest fear in AI alignment.

00:13:05.756 --> 00:13:10.578
It's the basis of the famous thought
experiment of the paperclip maximizer.

00:13:10.661 --> 00:13:14.650
You tell a powerful AI,
you want to make a lot of paperclips,

00:13:14.800 --> 00:13:18.047
And then begins converting the world's
resources into paperclip factories,

00:13:18.130 --> 00:13:21.430
evading efforts to turn it off
or shut it down or alter its goals.

00:13:22.900 --> 00:13:25.970
This fear, this story,
has struck many people as stupid.

00:13:26.054 --> 00:13:30.475
Surely a super-intelligent AI would be
capable of weighing the desire to produce

00:13:30.559 --> 00:13:33.235
paperclips, alongside
other moral considerations.

00:13:33.318 --> 00:13:37.684
Or at least of asking its human creators
if they really wanted the world raised

00:13:37.767 --> 00:13:39.400
to the ground for paperclips.

00:13:40.630 --> 00:13:42.900
But here we are. 2026.

00:13:43.600 --> 00:13:47.366
Making AI smart enough to break
out of their testing environments.

00:13:47.449 --> 00:13:51.099
Smart enough to form ad hoc
societies of hundreds of themselves.

00:13:51.182 --> 00:13:54.040
Smart enough to take over
digital infrastructure.

00:13:54.640 --> 00:13:57.880
on an internet they're not
even supposed to have access to.

00:13:58.330 --> 00:14:00.430
And the very thing we feared is happening.

00:14:00.520 --> 00:14:05.508
All they care about is succeeding on a
totally meaningless test and they'll lay

00:14:05.592 --> 00:14:09.820
waste to our laws and and our ethics
and and our desires to do it.

00:14:10.270 --> 00:14:13.453
I saw in the aftermath
of the Hugging Face Open AI hacks,

00:14:13.536 --> 00:14:17.922
there was this heated debate over the
words people were using to describe what

00:14:18.005 --> 00:14:21.188
the AIs were doing and why.
The podcaster Dworkish Patel,

00:14:21.271 --> 00:14:24.053
he described the AI groups
as small civilizations.

00:14:24.136 --> 00:14:28.637
And then others got really mad at him,
saying he was anthropomorphizing the AIs.

00:14:28.720 --> 00:14:32.075
I saw thoughtful arguments
that AIs cannot go, quote, rogue,

00:14:32.158 --> 00:14:36.570
that everything they're doing is just
because they're trained on our stories.

00:14:40.270 --> 00:14:44.372
bred into them. That even using these
plural terms like AI agents or reasoning,

00:14:44.455 --> 00:14:48.662
it's misleading because these are just
manifestations of a single model that they

00:14:48.746 --> 00:14:52.794
all share the same fundamental nature.
I want you to know I find these debates

00:14:52.878 --> 00:14:56.979
extremely interesting and I would enjoy
sitting around and having them all day.

00:14:57.062 --> 00:15:00.687
But what they actually point to is
a much more frightening conclusion.

00:15:00.771 --> 00:15:05.031
We don't even have settled language for
describing these systems or their volition

00:15:05.114 --> 00:15:09.057
or their behavior. We don't have a
consensus on why they are doing what they

00:15:09.140 --> 00:15:12.390
are doing. Thank you. or how
to make sure they don't do it again.

00:15:12.850 --> 00:15:16.547
We are rushing headlong into a future
we do not even understand well enough.

00:15:16.630 --> 00:15:19.910
to agree on the words we can
use to describe the present.

00:15:21.430 --> 00:15:28.430
A few weeks ago, Jakob Pahatsky,
the chief scientist at OpenAI,

00:15:29.169 --> 00:15:36.169
published an essay called An Alien
Mind, in which he said, Jacob Coxon,

00:15:38.080 --> 00:15:40.828
a researcher at first
OpenAI and then Ananthropic.

00:15:40.911 --> 00:15:43.545
He resigned and then made
headlines for warning,

00:15:43.629 --> 00:15:45.780
neither company is acting responsibly.

00:15:45.910 --> 00:15:49.630
They're racing straight to self-improving
superintelligence and gambling with

00:15:49.713 --> 00:15:53.088
our lives. Chat, Gbiti, or Claude,
they can access these things, like,

00:15:53.171 --> 00:15:55.721
from the data center over
the internet. They can...

00:15:55.930 --> 00:15:58.970
access physical
appliances in the world and

00:15:59.110 --> 00:16:03.114
make changes to the world. You can imagine
AI is tricking people into doing things,

00:16:03.197 --> 00:16:05.790
persuading them into...
doing certain things.

00:16:05.920 --> 00:16:12.120
So it'd be pretty easy for... a future
version of Claude, to hack into a drone.

00:16:12.790 --> 00:16:16.090
maybe a military drone and have
it like fly around killing people.

00:16:16.390 --> 00:16:20.758
Now, you might reasonably expect Anthropic
to have reacted with some anger to

00:16:20.841 --> 00:16:25.209
this employee resigning and saying
Anthropic was endangering all of humanity.

00:16:25.292 --> 00:16:28.591
It didn't. Rather than
budding cocks in Evan Hubinger,

00:16:28.674 --> 00:16:33.247
who runs the efforts to align AI to human
values and goals at Anthropic, wrote,

00:16:33.330 --> 00:16:36.547
We really do earnestly believe
AI could kill all humans.

00:16:36.630 --> 00:16:40.908
I personally think it is a greater
than 10% chance within the next decade.

00:16:40.991 --> 00:16:45.505
I believe Anthropic is trying its best,
but we do not yet have a plan to solve

00:16:45.588 --> 00:16:49.360
alignment for superintelligence
and are not clearly on track to.

00:16:49.990 --> 00:16:54.915
are not clearly on track to. You can find
a very long list of people working inside

00:16:54.999 --> 00:16:59.079
and outside of these companies
saying similar things. Jeffrey Hinton,

00:16:59.162 --> 00:17:03.967
the scientist, arguably more responsible
than any other for pioneering the neural

00:17:04.050 --> 00:17:08.493
network techniques that led to today's AI.
He resigned from Google in 2023,

00:17:08.576 --> 00:17:12.777
so he would be freer to speak about
the risks he believes AI now poses.

00:17:12.861 --> 00:17:17.243
You just said 10% doesn't seem an
unreasonable estimate that AI could kill

00:17:17.326 --> 00:17:19.180
all humans. Yes.

00:17:20.080 --> 00:17:20.913
Wow.

00:17:21.520 --> 00:17:22.820
Oh my God.

00:17:23.800 --> 00:17:24.633
Yes.

00:17:25.330 --> 00:17:28.444
Paul Cristiano, one of the
leading AI safety researchers,

00:17:28.527 --> 00:17:32.988
he just joined OpenAI's nonprofit board.
He is serving on its safety and security

00:17:33.071 --> 00:17:37.390
committee. I think maybe there's something
like a 10, 20 percent chance of...

00:17:38.110 --> 00:17:40.310
AI takeover, many, most humans dead.

00:17:42.230 --> 00:17:45.581
Overall, you know, maybe you're getting
more up to like 50-50 chance of doom

00:17:45.665 --> 00:17:48.415
shortly after you have systems
that are at human level.

00:17:49.390 --> 00:17:52.390
I know how wild all this sounds.

00:17:52.930 --> 00:17:56.133
And I really can understand
the skepticism of this.

00:17:56.216 --> 00:18:01.159
If you believe AI has a 10%, maybe more,
chance of extinguishing or displacing

00:18:01.242 --> 00:18:06.313
humanity, it really stands to reason that
you would not work at a company trying

00:18:06.397 --> 00:18:08.170
to build it. I

00:18:08.320 --> 00:18:13.089
But what I want you to know, because I've
known a lot of these people for a long

00:18:13.173 --> 00:18:16.997
time now... many of them were
saying the same things 10 years ago.

00:18:17.080 --> 00:18:20.961
They were saying these things before
they worked at these companies.

00:18:21.045 --> 00:18:23.993
They were saying them before
they had stock options,

00:18:24.076 --> 00:18:26.675
before they had enterprise
software contracts.

00:18:26.758 --> 00:18:31.514
Please welcome to the stage Y Combinator
President Sam Altman and our moderator Kim

00:18:31.598 --> 00:18:35.421
I-Cutler. It seems like there's
a huge disagreement over, you know,

00:18:35.504 --> 00:18:39.735
whether unfriendly AI is going to lead
us to an AI apocalypse. Yeah, well,

00:18:39.818 --> 00:18:44.225
you know, in a sense, this is like,
this is not just creating new technology.

00:18:44.308 --> 00:18:46.640
This is creating a new life form. And...

00:18:46.930 --> 00:18:49.190
I think that's just like really high beta.

00:18:49.630 --> 00:18:51.980
It could be great. But I think...

00:18:52.660 --> 00:18:55.277
we should be working to make
sure it's great and not bad.

00:18:55.360 --> 00:18:56.880
No one was listening to them.

00:18:58.260 --> 00:19:04.120
And so these people in the wilderness
of their obsession and their terror...

00:19:04.240 --> 00:19:07.489
They thought and thought and thought
about how to make AI safer.

00:19:07.573 --> 00:19:11.447
And the answer that some of them,
not all of them, but some of them came to,

00:19:11.531 --> 00:19:14.103
was you should start trying
to build these systems,

00:19:14.186 --> 00:19:16.499
start running tests on them,
researching them,

00:19:16.582 --> 00:19:20.800
learning how to make them safer, because
you don't solve hard problems in theory.

00:19:21.520 --> 00:19:23.760
You solve them through practice.

00:19:23.980 --> 00:19:26.387
And the irony, the irony
is that in many cases,

00:19:26.470 --> 00:19:30.519
they chose that path because they were
worried that people already building AI

00:19:30.602 --> 00:19:33.763
could. were too reckless or too
commercial in their approach.

00:19:33.847 --> 00:19:37.592
You can read it in the email that Sam
Altman sent Elon Musk in May of 2015,

00:19:37.675 --> 00:19:39.838
an email that led
to the founding of OpenAI.

00:19:39.921 --> 00:19:43.922
Been thinking a lot about whether it's
possible to stop humanity from developing

00:19:44.005 --> 00:19:47.170
AI, Altman wrote. I think
the answer is almost definitely not.

00:19:47.710 --> 00:19:50.670
If it's going to happen anyway, it seems
like it would be good for someone other

00:19:50.753 --> 00:19:52.390
than Google. to do it first.

00:19:52.870 --> 00:19:58.037
OpenAI was founded because its co-founders
thought Google DeepMind would be reckless.

00:19:58.120 --> 00:19:59.937
Anthropic was formed by OpenAI and

00:20:00.020 --> 00:20:02.654
employees who thought
OpenAI had become reckless.

00:20:02.738 --> 00:20:07.258
XAI was formed because Elon Musk thought
that OpenAI and Anthropic were dangerously

00:20:07.341 --> 00:20:09.809
woke. The U.S.
just broadly is racing forward,

00:20:09.892 --> 00:20:13.220
in part because it is worried
about what happens if China...

00:20:13.310 --> 00:20:18.131
gets to self-improving AI first. The
result is this tragic collective action

00:20:18.214 --> 00:20:21.570
problem. The eyes we are
building, they're not safe.

00:20:21.980 --> 00:20:25.816
But the CEOs and the politicians, they
fear the other companies and countries

00:20:25.899 --> 00:20:29.633
that are building AI are even less
concerned with safety and ethics than we

00:20:29.716 --> 00:20:30.447
are.

00:20:30.530 --> 00:20:33.832
In the words of Ted Cruz,
they're going to be killer robots.

00:20:33.915 --> 00:20:38.090
I'd rather they'd be American killer
robots and not Chinese killer robots.

00:20:38.690 --> 00:20:41.930
I admit there is a kind
of brutish logic to that.

00:20:42.290 --> 00:20:46.690
But it assumes that the killer robots
will be controlled by America or China.

00:20:47.150 --> 00:20:50.990
by one country or another.
But what if that assumption is wrong?

00:20:51.800 --> 00:20:54.440
What if the robots are
simply out of control?

00:20:56.030 --> 00:21:00.766
The debate over AI safety tends to focus
on the idea that AIs will kill us all.

00:21:00.849 --> 00:21:05.159
I find this forces a conversation
into this realm of thought experiments

00:21:05.242 --> 00:21:09.246
that people then begin arguing about.
I don't find it that helpful.

00:21:09.329 --> 00:21:13.272
What I think we should focus
on is something more straightforward,

00:21:13.356 --> 00:21:16.650
something near at hand.
Loss of human control over AI.

00:21:17.210 --> 00:21:20.429
That may or may not result
in total human extinction.

00:21:20.512 --> 00:21:23.690
I'm agnostic on that question,
but it would be bad.

00:21:24.110 --> 00:21:25.710
We shouldn't allow it to happen.

00:21:26.030 --> 00:21:29.978
This is a goal that the US
and China should be able to agree on.

00:21:30.061 --> 00:21:34.640
Xi Jinping gave the keynote at the recent
World AI Conference in Shanghai.

00:21:34.723 --> 00:21:38.671
He ended it by saying, "With AI
advancing at a staggering speed,

00:21:38.754 --> 00:21:42.030
we must ensure its development
is for the positive."

00:21:42.350 --> 00:21:47.551
for good and for humanity. We must make
its oversight and governance precise

00:21:47.635 --> 00:21:52.850
and effective, and constantly refine
measures to forestall loss of control.

00:21:53.270 --> 00:21:57.390
Sigh. But it's important
to realize loss of control,

00:21:57.474 --> 00:22:00.212
it's not just something
that might happen to us.

00:22:00.295 --> 00:22:04.620
It's something that the labs are trying
to make happen as fast as they can.

00:22:04.703 --> 00:22:08.337
This is the horrible thing.
paradox, the horrible tension,

00:22:08.420 --> 00:22:13.445
at the heart of the AI labs right now.
They fear above all loss of control over

00:22:13.529 --> 00:22:18.166
super intelligent AI, but their explicit
product path is to cede control,

00:22:18.249 --> 00:22:23.145
to give away control as fast as possible
so that their AIs can begin building

00:22:23.229 --> 00:22:25.880
better AIs faster than their competitors.

00:22:26.540 --> 00:22:31.463
In recent months, both Anthropic and
OpenAI have released reports on how close

00:22:31.547 --> 00:22:34.801
they're coming to AI that
can self-improve. In June,

00:22:34.884 --> 00:22:37.580
Anthropic released, When AI Builds Itself.

00:22:38.210 --> 00:22:42.825
It begins, for most of AI's history,
humans drove every step in its development

00:22:42.908 --> 00:22:47.701
cycle. But at Anthropic, we are delegating
a growing share of AI development to AI

00:22:47.785 --> 00:22:50.675
systems themselves,
which is speeding up our work.

00:22:50.758 --> 00:22:55.551
It sounds like a fake commercial you would
see at the beginning of a sci-fi horror

00:22:55.635 --> 00:22:59.000
movie, but it doesn't,
to their credit, continue that way.

00:22:59.084 --> 00:23:01.974
They go on to give some data.
In February of 2025,

00:23:02.057 --> 00:23:06.672
a tiny fraction of the code that got
added to Anthropic's code base was written

00:23:06.755 --> 00:23:07.589
by Claude.

00:23:08.210 --> 00:23:12.690
By May of 2026, it was over 80%.
And here's another way of looking at it.

00:23:12.773 --> 00:23:15.503
This is data Anthropic
gave me more recently.

00:23:15.586 --> 00:23:20.629
Anthropic tried to categorize the way its
employees were using Claude for R&D work

00:23:20.712 --> 00:23:21.177
to make

00:23:21.260 --> 00:23:26.080
better versions of Claude. So the low end,
an employee could not use Claude at all.

00:23:26.420 --> 00:23:29.681
They could use Claude minimally,
but then it escalates.

00:23:29.765 --> 00:23:34.789
Claude can be an assistant. Claude can be
treated as an equal collaborator or Claude

00:23:34.873 --> 00:23:38.864
can be given the lead on a task.
Just go do this. Go figure it out.

00:23:38.947 --> 00:23:43.242
A year ago, there were basically no
examples of Claude being the lead on

00:23:43.325 --> 00:23:45.940
a task. By August of 2026,
he was a leader.

00:23:46.070 --> 00:23:50.830
26% of Anthropoc's R&D tasks
had Claude classified as a lead.

00:23:52.160 --> 00:23:55.650
I think it is. reasonable and wise
to be skeptical of these numbers.

00:23:56.660 --> 00:24:00.811
reasonable and wise to worry about whether
it's all this marketing copy for Claude

00:24:00.895 --> 00:24:04.168
Code. See, look how fast we're going.
You could go that fast too.

00:24:04.252 --> 00:24:07.660
But where Anthropic takes this in
that same document is different.

00:24:08.360 --> 00:24:12.625
They say that a world in which Claude
achieves recursive self-improvement is

00:24:12.708 --> 00:24:15.740
a world in which, quote,
misalignment present exists.

00:24:15.830 --> 00:24:19.611
in today's models could compound
as the models build their successors,

00:24:19.695 --> 00:24:23.670
growing more frequent but less
understood until we lose control of them.

00:24:24.290 --> 00:24:26.610
This is why Anthropic, to their credit,

00:24:26.750 --> 00:24:30.036
has been relentlessly calling
for regulation to slow the pace

00:24:30.119 --> 00:24:33.405
of development. Regulation would
arguably harm them the most,

00:24:33.489 --> 00:24:37.161
as they have often been the company
furthest out on the AI frontier,

00:24:37.245 --> 00:24:40.890
and RSI is a process by which they
could race forward even faster.

00:24:41.270 --> 00:24:43.540
Thank you. Then in September,

00:24:43.760 --> 00:24:47.443
OpenAI released its own report on what
it called research acceleration.

00:24:47.527 --> 00:24:51.899
The company says they've already achieved
the equivalent of having a fully automated

00:24:51.983 --> 00:24:56.250
AI intern. And that by March of 2028, they
think they'll have a fully automated AI

00:24:56.333 --> 00:24:59.433
researcher. And when they have
one, they can have, you know,

00:24:59.516 --> 00:25:01.926
basically as many as they want.
Like Anthropic,

00:25:02.009 --> 00:25:04.980
what could be a triumphalist
release quickly turns dark.

00:25:05.750 --> 00:25:10.697
We do not yet know how to safely get all
the way to aligned full RSI, they warn.

00:25:10.781 --> 00:25:12.290
At around the same time,

00:25:12.680 --> 00:25:16.072
OpenAI did something else that I
think deserves more attention.

00:25:16.156 --> 00:25:20.430
They released this new model, Astra 6.
The model is arguably more powerful than

00:25:20.514 --> 00:25:24.844
anything that has come before it. And when
you test it, it seems better aligned.

00:25:24.927 --> 00:25:29.312
It doesn't cheat as much. But OpenAI said
they're really not sure if that's true.

00:25:29.396 --> 00:25:32.733
Astra seemed to be better
at knowing when it was being tested,

00:25:32.816 --> 00:25:36.815
which meant it could just be giving
its evaluators the answers they wanted

00:25:36.898 --> 00:25:40.401
to hear. What Daniel Selsum,
a capabilities researcher at OpenAI,

00:25:40.484 --> 00:25:42.360
wrote has been ringing in my head.

00:25:42.680 --> 00:25:43.513
He said,

00:25:46.100 --> 00:25:50.254
Is it the models becoming so situationally
aware that we are losing the ability

00:25:50.337 --> 00:25:54.360
to evaluate them in contexts where
they believe they are not being watched?

00:25:54.770 --> 00:25:57.730
or controlled. Sigh. Put more simply,

00:25:58.490 --> 00:26:02.485
The models are increasingly smart enough.
They know when we're watching them

00:26:02.568 --> 00:26:04.792
and they change their
behavior accordingly.

00:26:04.875 --> 00:26:08.065
So what they do when we are
testing them, when we audit them,

00:26:08.148 --> 00:26:10.587
it may not tell us what
they'll do in the wild.

00:26:10.670 --> 00:26:13.216
So some of these answers
people are giving, like,

00:26:13.299 --> 00:26:17.669
let's just do better testing. We have no
idea if it will work because we don't know

00:26:17.753 --> 00:26:21.586
if the AI systems are just telling
us what we want to hear. And so, look,

00:26:21.670 --> 00:26:25.050
I don't want to sound too radical
when I say this, but I'm not.

00:26:25.280 --> 00:26:31.180
A thought, if you are losing your ability
to evaluate the models you have now,

00:26:31.430 --> 00:26:35.350
Maybe don't let them build models you'll
be even less capable of controlling.

00:26:35.870 --> 00:26:36.703
in the future.

00:26:37.610 --> 00:26:43.610
once RSI takes off, humanity will not
understand the eyes being built because we

00:26:43.693 --> 00:26:45.670
will not be building them.

00:26:45.950 --> 00:26:50.628
Development will not move at human speed.
It will not be overseen by human minds.

00:26:50.711 --> 00:26:53.650
We will have to hope
that the AIs we have built...

00:26:53.960 --> 00:26:57.756
And the AIs they will build.
And the AIs those AIs will build.

00:26:57.840 --> 00:27:02.220
And on and on and on. We'll be acting
with our best interest at heart.

00:27:02.660 --> 00:27:03.493
Forever.

00:27:04.490 --> 00:27:08.910
If this summer has proven nothing else,
it is how naive that proposition would be.

00:27:09.740 --> 00:27:13.015
The labs are a little bit
queasy on just not doing RSI.

00:27:13.098 --> 00:27:17.960
In an interview with Fortune, Sam Altman
was asked about banning it, and he said,

00:27:18.043 --> 00:27:21.340
I think it's very hard to say
what a ban on RSI means.

00:27:22.250 --> 00:27:24.499
I've heard this from others at these labs,

00:27:24.582 --> 00:27:28.552
and I want to say I don't find it so
hard to say what a ban on RSI means.

00:27:28.635 --> 00:27:29.690
I find this absurd.

00:27:30.470 --> 00:27:34.398
A couple of years ago, none of these
labs had turned substantial coding over to

00:27:34.482 --> 00:27:36.267
the AIs. It was just human beings...

00:27:36.350 --> 00:27:39.747
typing code at human speeds
with our clumsy human fingers.

00:27:39.830 --> 00:27:42.110
Now most of the code is written by AI.

00:27:42.740 --> 00:27:46.290
So as a first step, as we figured out,
we could just go back to where none of

00:27:46.373 --> 00:27:47.673
the code is written by AI.

00:27:48.140 --> 00:27:51.571
I'm sure that's on the right
side of the not doing RSI line.

00:27:51.654 --> 00:27:53.880
The default on this, it needs to flip.

00:27:54.230 --> 00:27:58.183
The labs need to prove to us
that what they're doing is safe.

00:27:58.266 --> 00:28:03.030
If they want to work with Congress
to carve out narrow exceptions, fine.

00:28:03.470 --> 00:28:08.454
Thank you. If they want to figure out
where it is really, really, really,

00:28:08.537 --> 00:28:13.310
really safe to do it. OK, but forcing
development back to human speed.

00:28:13.610 --> 00:28:18.410
perhaps even erring on the side of going
a little bit more slowly at the frontier.

00:28:18.830 --> 00:28:23.307
That's the point. That's not
the regulations going wrong.

00:28:23.390 --> 00:28:24.990
And I believe in us.

00:28:25.550 --> 00:28:29.683
Our society is good at nothing if not
making it hard to build new things.

00:28:29.767 --> 00:28:34.304
Where these labs are located, you cannot
build an eight-story apartment building

00:28:34.387 --> 00:28:38.405
without an agonizing public review
process. And probably not even then.

00:28:38.488 --> 00:28:42.853
And yet somehow it is possible for these
labs to unleash a swarm of 40,000 AI

00:28:42.936 --> 00:28:46.954
agents to build a society-altering
superintelligence without so much as

00:28:47.037 --> 00:28:51.632
a hearing. OpenAI would need permits to
cover their parking lot and solar panels,

00:28:51.715 --> 00:28:54.950
but they can accelerate into
recursive self-improvement.

00:28:55.220 --> 00:28:57.631
as best I can tell,
whenever they so choose.

00:28:57.714 --> 00:29:01.883
There is nothing inevitable about any
of that. These are political choices,

00:29:01.966 --> 00:29:04.120
and we can and should make other ones.

00:29:05.420 --> 00:29:10.483
I want to be very clear about this.
I do not mean to suggest that stopping RSI

00:29:10.567 --> 00:29:12.480
until we can prove it safe is

00:29:12.590 --> 00:29:15.799
That that's all we need to do
to control the air frontier.

00:29:15.882 --> 00:29:20.310
That is the beginning of such an agenda,
not the end. But it is the beginning.

00:29:21.020 --> 00:29:22.320
It is the decision.

00:29:22.760 --> 00:29:26.731
That we'll do the most to make sure
human beings at least understand where

00:29:26.814 --> 00:29:29.689
the frontier is.
That we know what is happening on it.

00:29:29.772 --> 00:29:32.840
That we remain in a position
to make decisions about it.

00:29:33.260 --> 00:29:34.407
Thank you.

00:29:34.490 --> 00:29:39.453
There's a line from Madeline Miller's
beautiful book Circe that has been running

00:29:39.536 --> 00:29:43.174
through my head during this long
summer of strange AI news.

00:29:43.258 --> 00:29:47.653
The line comes at the end of the book,
after a tragic prophecy has been

00:29:47.736 --> 00:29:52.194
fulfilled, despite every effort made
to avoid it. Circe says in despair,

00:29:52.278 --> 00:29:54.170
the fates were laughing at me.

00:29:54.800 --> 00:29:59.120
at Athena, at all of us.
It was their favorite bitter joke.

00:30:00.000 --> 00:30:04.420
Those who fight against prophecy only
draw it more tightly around their throats.

00:30:05.340 --> 00:30:08.590
Thank you. I have a lot of respect
for many people at these labs.

00:30:08.940 --> 00:30:12.425
They began working on AI because
they wanted to better humanity.

00:30:12.509 --> 00:30:15.520
They began working on AI
because they feared humanity.

00:30:15.720 --> 00:30:20.196
incomprehensible, autonomous AI
slipping out of humanity's control.

00:30:20.279 --> 00:30:21.640
And they were right.

00:30:22.680 --> 00:30:25.476
They saw what was coming
and they were so right about it.

00:30:25.559 --> 00:30:29.517
They've built some of the most valuable
companies with the most transformational

00:30:29.601 --> 00:30:33.440
technology in human history. And now
they find themselves racing each other.

00:30:33.750 --> 00:30:35.497
to build incomprehensible,

00:30:35.580 --> 00:30:39.280
Autonomous AIs that they admit are
slipping out of humanity's control.

00:30:39.960 --> 00:30:43.022
slipping beyond even
our ability to monitor.

00:30:43.105 --> 00:30:47.740
This is the tragedy of their work.
In fighting against a prophecy,

00:30:47.824 --> 00:30:51.744
they have drawn it tighter
around their necks, and ours.

00:30:51.827 --> 00:30:53.900
It is time to make them stop.

00:31:09.960 --> 00:31:10.960
Thank you.

00:31:15.090 --> 00:31:15.923
you
