1
00:00:00,400 --> 00:00:05,677
Naveen Rao, co-founder and CEO
of Unconventional AI, which is an AI chip

2
00:00:05,760 --> 00:00:09,357
startup. >> Best known for building and
[music] selling two deep tech companies.

3
00:00:09,440 --> 00:00:12,317
>> Naveen is kind
of definitionally outlier founder.

4
00:00:12,400 --> 00:00:15,917
>> When I came there, we had about a $20
million business and it was, you know, 7

5
00:00:16,000 --> 00:00:17,197
or 800 million [music] when I left.

6
00:00:17,280 --> 00:00:19,277
>> think you really understand
something until you can build it.

7
00:00:19,360 --> 00:00:21,357
>> Just because something is tried
does not mean it's wrong. [music]

8
00:00:21,440 --> 00:00:25,677
>> I'm the opposite of an AI doomer.
I think AI is the next evolution of

9
00:00:25,760 --> 00:00:30,197
humanity. >> We need innovation on the
hardware substrate to actually build true

10
00:00:30,280 --> 00:00:36,400
intelligence.
>> Please welcome Naveen Rao.

11
00:00:38,650 --> 00:00:40,670
>> [music]

12
00:00:44,080 --> 00:00:46,757
>> Everyone, great to be here.

13
00:00:46,840 --> 00:00:50,077
Um, you know, switching gears a little
bit to AI now, which you may have heard

14
00:00:50,160 --> 00:00:54,077
a little bit about. Um, I it's super
exciting to be at this conference

15
00:00:54,160 --> 00:00:58,757
specifically because, as was said in the
intro, I'm the opposite of a doomer. I

16
00:00:58,840 --> 00:01:02,357
think AI is one of the most
transformational technologies that

17
00:01:02,440 --> 00:01:06,157
humanity's ever created and will
enable us to get to that next level of

18
00:01:06,240 --> 00:01:10,317
evolution, which I'm here for.
And uh, this is sort of the anti-doomer

19
00:01:10,400 --> 00:01:12,617
conference. So, what let's go.

20
00:01:12,700 --> 00:01:15,877
>> [laughter] >> Um,
so before we get going, I'll tell

21
00:01:15,960 --> 00:01:19,797
you a little bit about myself. Um,
you know, I it's kind of weird. I'm really

22
00:01:19,880 --> 00:01:23,557
right where I wanted to be my whole life.
Uh, this was me at about five or

23
00:01:23,640 --> 00:01:25,117
six years old, something like that. We

24
00:01:25,200 --> 00:01:30,637
had a computer very early on, so I'll
date myself, but uh, this was in 1978.

25
00:01:30,720 --> 00:01:33,997
We got a computer. This is probably
in the early '80s. Uh, learned to program

26
00:01:34,080 --> 00:01:37,037
when I was a little kid. I just thought
it was like a puzzle. Um, you know,

27
00:01:37,120 --> 00:01:41,397
learned to I became an electrical
engineer, um, really because I enjoyed

28
00:01:41,480 --> 00:01:42,837
sci-fi and always wanted to think about

29
00:01:42,920 --> 00:01:45,277
how I could make a intelligent machine.

30
00:01:45,360 --> 00:01:49,277
And uh, you know, then after a career in

31
00:01:49,360 --> 00:01:53,077
uh, building computers, I went back
to school and got a PhD in neuroscience.

32
00:01:53,160 --> 00:01:56,717
And the idea was like, let's go back
back thing. How do we make computers

33
00:01:56,800 --> 00:02:00,277
intelligent? and fortunately the whole
world kind of moved in this direction.

34
00:02:00,360 --> 00:02:02,677
So, you know, as a technologist it's

35
00:02:02,760 --> 00:02:05,397
sort of the dream right now.

36
00:02:05,480 --> 00:02:08,277
Um a little bit about me tech from a

37
00:02:08,360 --> 00:02:10,477
company entrepreneurship standpoint so

38
00:02:10,560 --> 00:02:12,557
like I actually founded the first AI

39
00:02:12,640 --> 00:02:15,597
chip company called Nirvana systems. So,

40
00:02:15,680 --> 00:02:17,917
um this was in 2014. If anyone remembers

41
00:02:18,000 --> 00:02:22,837
back then there was no AI or at least
not in the common vernacular and uh

42
00:02:22,920 --> 00:02:25,837
it was really hard to actually convince
people this is important much less to

43
00:02:25,920 --> 00:02:29,557
build hardware around it. Now, you heard
from Jensen up here like the largest

44
00:02:29,640 --> 00:02:33,837
company in the world a hardware company
because of AI. So, we were early on. Um

45
00:02:33,920 --> 00:02:36,277
I think I sold the company
way too early uh

46
00:02:36,360 --> 00:02:39,197
to Intel but I I ran I started and ran

47
00:02:39,280 --> 00:02:43,717
the AI group at Intel. Um after
I was done with that in 2020, I

48
00:02:43,800 --> 00:02:46,837
actually started thinking about the next
problem which was how do we build bigger

49
00:02:46,920 --> 00:02:51,357
models like the large language models we
talk about today and it was about how do

50
00:02:51,440 --> 00:02:54,437
I build the infrastructure to build
those models. And so, we started

51
00:02:54,520 --> 00:02:58,997
platform rising GPUs and enabling it
to scale and making that easy to use for

52
00:02:59,080 --> 00:03:03,397
other people. And after chat GPT
happened in 20 2022, we were kind of the

53
00:03:03,480 --> 00:03:06,997
best game in town for people to start
building their own models. So, took off

54
00:03:07,080 --> 00:03:08,797
really fast. We decided to actually join

55
00:03:08,880 --> 00:03:10,957
forces with data bricks. Um

56
00:03:11,040 --> 00:03:14,357
that was in 2023 and and uh actually

57
00:03:14,440 --> 00:03:18,877
that's a quarter of the total revenue
of of data bricks today. So, um

58
00:03:18,960 --> 00:03:23,037
a lot of fun doing that whole thing with
uh Ali and team at data bricks. And now

59
00:03:23,120 --> 00:03:26,277
I want to tell you about unconventional
AI which is rethinking the foundations

60
00:03:26,360 --> 00:03:30,477
of how a computer works. So,
there we're going back to first

61
00:03:30,560 --> 00:03:33,557
principles here to really trying
to build a new machine. Computers have

62
00:03:33,640 --> 00:03:37,277
worked a certain way for a long time.
We want to rethink that for the singular

63
00:03:37,360 --> 00:03:40,797
purpose of making something
very power efficient.

64
00:03:40,880 --> 00:03:44,717
And the goal has been within it
was initially within 5 years to get to a

65
00:03:44,800 --> 00:03:48,277
thousand X power efficiency.
I've actually revised this to three and a

66
00:03:48,360 --> 00:03:51,437
half years because things have gone
faster than we anticipated. We've

67
00:03:51,520 --> 00:03:55,597
actually solved very deep scientific
problems quicker because of AI,

68
00:03:55,680 --> 00:03:59,517
interestingly enough. So,
just a little bit how we organize.

69
00:03:59,600 --> 00:04:03,157
Like, we're truly a top-to-bottom company.
We start with theorists. These

70
00:04:03,240 --> 00:04:07,397
are people with like math PhDs and um
you know, theoretical neuroscience,

71
00:04:07,480 --> 00:04:11,357
that kind of thing. They come up with,
you know, concepts that we think would

72
00:04:11,440 --> 00:04:13,957
uh effectively give us more power
efficiency from the perspective of

73
00:04:14,040 --> 00:04:16,317
moving less information around. And we

74
00:04:16,400 --> 00:04:19,397
then translate that into models and that

75
00:04:19,480 --> 00:04:23,917
do real things, trained on real data,
and evaluated against real criteria. So,

76
00:04:24,000 --> 00:04:26,637
it's kind of the rubber hitting
the road of these concepts.

77
00:04:26,720 --> 00:04:29,877
Then eventually, we have to actually
build something physical. So, these are

78
00:04:29,960 --> 00:04:34,837
people who architect a physical circuit,
um actually design those circuits, uh

79
00:04:34,920 --> 00:04:37,797
model them, and see if it actually works.
So, we try to connect this whole

80
00:04:37,880 --> 00:04:40,717
stack together.
Um then eventually, we have to build a

81
00:04:40,800 --> 00:04:44,837
system and and a board and all
that kind of stuff and build a product.

82
00:04:44,920 --> 00:04:49,317
So, is energy really a problem? Uh I'm not
sure how much everyone in this audience

83
00:04:49,400 --> 00:04:51,997
has thought about this, but interestingly
enough, I'll give you some

84
00:04:52,080 --> 00:04:55,277
some data points here. This is one
company. This is just Google. I'm using

85
00:04:55,360 --> 00:04:59,237
Google because Google's actually talked
about this publicly. Um per month, they

86
00:04:59,320 --> 00:05:04,317
cross 3.2 quadrillion tokens.
It's a crazy number. I never even

87
00:05:04,400 --> 00:05:08,157
I never think in quadrillions, but that's
the world we're in today. And if

88
00:05:08,240 --> 00:05:12,437
I just take 10 joules per token of energy,
this is a actually on the lower

89
00:05:12,520 --> 00:05:16,117
end of uh of the energy spectrum for
models, but let's just take that number

90
00:05:16,200 --> 00:05:19,517
and multiply it out. This is 12 gigawatts.

91
00:05:19,600 --> 00:05:24,557
The US puts about 40 gigawatts of energy
into data centers today. And we're about

92
00:05:24,640 --> 00:05:26,317
half of the data center capacity of the

93
00:05:26,400 --> 00:05:29,637
world. And so, you know, we're under 100

94
00:05:29,720 --> 00:05:31,837
gigawatts of of data center energy in

95
00:05:31,920 --> 00:05:34,077
the world today. 12 gigawatts is going

96
00:05:34,160 --> 00:05:37,197
into one company just for AI services.

97
00:05:37,280 --> 00:05:39,437
So, you can imagine if models get

98
00:05:39,520 --> 00:05:44,557
bigger, that that energy goes up. And if
demand grows, which it is, that energy

99
00:05:44,640 --> 00:05:47,917
goes up. So, we're going to run out
of energy pretty fast, in like 3 years or

100
00:05:48,000 --> 00:05:51,117
so is my estimate. So,
really just putting it graphically,

101
00:05:51,200 --> 00:05:54,597
this is what we have. You know,
we have this huge market that's growing

102
00:05:54,680 --> 00:05:58,237
exponentially. Call it a trillion-dollar
market in 2030, maybe it's bigger than

103
00:05:58,320 --> 00:06:03,237
that. And then we have this kind of
linearized energy at the bottom. And so,

104
00:06:03,320 --> 00:06:05,157
you've heard a lot about this today, uh

105
00:06:05,240 --> 00:06:11,640
but this gap is the problem. And we
want to solve that gap with technology.

106
00:06:13,040 --> 00:06:17,037
So, um I don't know if people are aware
of this, but the way we think about data

107
00:06:17,120 --> 00:06:20,237
centers has shifted over the last
several years. It used to be about floor

108
00:06:20,320 --> 00:06:23,717
space. Can I get the floor space? Can
I get the rack space? Then it was about

109
00:06:23,800 --> 00:06:27,357
networking equipment, then it became
about GPUs. Today, it's about energy.

110
00:06:27,440 --> 00:06:30,397
First, you think about energy. I get
the energy contract, and then I have to

111
00:06:30,480 --> 00:06:34,797
figure out how to fill it
and and basically create uh um

112
00:06:34,880 --> 00:06:36,957
uh infrastructure out of out of GPUs and

113
00:06:37,040 --> 00:06:40,277
things like this. About 50% of the cost

114
00:06:40,360 --> 00:06:44,597
of serving a token, so every time you
try something on ChatGPT, 50% of that

115
00:06:44,680 --> 00:06:47,677
cost is energy. The rest
of it is the CAPEX of the

116
00:06:47,760 --> 00:06:52,197
hardware and the floor space and all
that kind of stuff. So, today we sort of

117
00:06:52,280 --> 00:06:54,877
think about it, oops,
we kind of think about it as

118
00:06:54,960 --> 00:06:58,797
I get I get a a power contract, and I
need to monetize every watt. And simply

119
00:06:58,880 --> 00:07:00,877
put, our business case is pretty easy.

120
00:07:00,960 --> 00:07:04,757
We're going to monetize that 1,000x
better than existing hardware.

121
00:07:04,840 --> 00:07:07,797
So, then the question becomes, okay,
great, that all makes sense, but can we

122
00:07:07,880 --> 00:07:11,237
actually do it? How do we build
a better, more efficient computer?

123
00:07:11,320 --> 00:07:14,957
Well, biology actually provides some
proof for us here. So, the human brain,

124
00:07:15,040 --> 00:07:18,717
you may have heard this, runs on about
20 watts of energy. And what's even more

125
00:07:18,800 --> 00:07:22,037
remarkable to me is actually animal
brains. So, that red number is how many

126
00:07:22,120 --> 00:07:26,317
neurons are in the brain. And if you
kind of scale it linearly down to like a

127
00:07:26,400 --> 00:07:30,677
a monkey's brain, it runs on 1 watt.
To put that in perspective, the cell phone

128
00:07:30,760 --> 00:07:33,317
in your pocket runs on about 1 watt.

129
00:07:33,400 --> 00:07:36,117
And other animals like, you know, um

130
00:07:36,200 --> 00:07:40,717
rats and bats and things like this,
they run on milliwatts of energy. So, just

131
00:07:40,800 --> 00:07:43,437
something that's pretty relatable
to everyone is a squirrel. You probably

132
00:07:43,520 --> 00:07:47,077
watched how accurate they can be.
They jump between branches and they do it

133
00:07:47,160 --> 00:07:50,957
perfectly a thousand times out of a
thousand. Their brain runs on eight

134
00:07:51,040 --> 00:07:55,397
milliwatts of energy. You could run over
a hundred squirrel brains on your phone,

135
00:07:55,480 --> 00:07:59,437
and it it has very precise and accurate
behavior. So, biology's created

136
00:07:59,520 --> 00:08:02,677
something quite incredible.
In fact, it's the right kind of physical

137
00:08:02,760 --> 00:08:07,957
substrate for intelligence. So,
this is a quote I love. I I don't

138
00:08:08,040 --> 00:08:11,237
feel like we truly understand something
until we can create it. We've gotten a

139
00:08:11,320 --> 00:08:15,517
lot better at creating intelligent
systems. However, they do it in a kind

140
00:08:15,600 --> 00:08:19,717
of inefficient way. What kind
of inefficiency is there? Well, as I kind

141
00:08:19,800 --> 00:08:23,437
of hinted at the beginning, most of the
energy in a computing system goes into

142
00:08:23,520 --> 00:08:28,037
moving information around. Just to put
some numbers on it, a human cortex, the

143
00:08:28,120 --> 00:08:32,437
squiggly part of your brain, the outside
of it, um only moves about 16 billion

144
00:08:32,520 --> 00:08:35,797
bits per second. That's actually kind
of a small number if you think about it,

145
00:08:35,880 --> 00:08:37,957
cuz there's, you know, some 13 or 14

146
00:08:38,040 --> 00:08:40,997
billion neurons in that cortex. Um a a

147
00:08:41,080 --> 00:08:42,917
GPU or, you know, a high-end computing

148
00:08:43,000 --> 00:08:46,117
system moves nearly 30 trillion bits in

149
00:08:46,200 --> 00:08:49,957
and out of memory per second. That's
outside the chip. Inside the chip, it's

150
00:08:50,040 --> 00:08:54,037
probably, uh, you know,
10, 100x more than that. So,

151
00:08:54,120 --> 00:08:57,597
we're moving a lot more bits in these
synthetic systems than the brain does.

152
00:08:57,680 --> 00:09:00,997
And that's actually what
drives the energy demand.

153
00:09:01,080 --> 00:09:05,117
And, you know, how do we get here?
You know, computers have been around for

154
00:09:05,200 --> 00:09:08,437
hundreds of years, actually, in some form.
They were mechanical, they became

155
00:09:08,520 --> 00:09:11,517
analog around the turn of the century,
and they became digital back in the

156
00:09:11,600 --> 00:09:15,077
1930s, 1940s. So, this the operation of

157
00:09:15,160 --> 00:09:20,077
that computer in 1940, 1945, is actually
very similar to how they operate today.

158
00:09:20,160 --> 00:09:23,117
There's not a huge paradigm shift.
They actually have this memory on the

159
00:09:23,200 --> 00:09:25,637
outside, you have some kind of computing,
and you move bits back and

160
00:09:25,720 --> 00:09:30,237
forth. That operation creates a machine
that just requires a lot of movement,

161
00:09:30,320 --> 00:09:34,477
but it's very fast. We built computers
to be fast. They were always faster than

162
00:09:34,560 --> 00:09:38,117
the alternative. Uh,
incidentally, that computer in 1945

163
00:09:38,200 --> 00:09:43,077
that was built called ENIAC was built to
do artillery trajectory calculation. And

164
00:09:43,160 --> 00:09:46,397
it was built to do it faster than the
alternative. The alternative were humans

165
00:09:46,480 --> 00:09:49,957
who actually did the calculations. Now,
the alternative typically is some other

166
00:09:50,040 --> 00:09:52,957
computer. This This computer is twice
the speed of that computer. That's how

167
00:09:53,040 --> 00:09:56,357
we sell computers. But it's
ever doesn't contemplate energy

168
00:09:56,440 --> 00:09:59,517
efficiency.
And that's what we're changing.

169
00:09:59,600 --> 00:10:01,917
And you know, these have been trends
going on for a long time where the

170
00:10:02,000 --> 00:10:05,917
number of transistors kept going up, but
we couldn't keep scaling the frequency.

171
00:10:06,000 --> 00:10:09,957
We could couldn't keep scaling uh
single thread performance. And now we're

172
00:10:10,040 --> 00:10:13,677
actually not scaling efficiency any
longer. Moore's law, if you may have

173
00:10:13,760 --> 00:10:17,397
heard of this, it's like making
transistors smaller has largely ended.

174
00:10:17,480 --> 00:10:21,117
So, we're not seeing efficiency gains
just for making transistors smaller. So,

175
00:10:21,200 --> 00:10:23,797
we need to rethink the problem a bit.

176
00:10:23,880 --> 00:10:27,437
So, how do we do it? Well, a good
intuition is that we cut out the middle

177
00:10:27,520 --> 00:10:31,597
man. So, computers have been built up
with these abstractions. So, I mentioned

178
00:10:31,680 --> 00:10:34,197
digital. Digital means one and zero.

179
00:10:34,280 --> 00:10:38,397
That itself is an abstraction of the
physical world, right? We don't actually

180
00:10:38,480 --> 00:10:42,037
have systems that behave as one and zero.
A transistor actually has states

181
00:10:42,120 --> 00:10:45,637
in the middle, but we engineer it
to behave that way. And that that's an

182
00:10:45,720 --> 00:10:48,957
abstraction. So, we kept building
these abstractions up, and eventually we

183
00:10:49,040 --> 00:10:53,077
started creating neural networks and
learning machines on top of it. Each one

184
00:10:53,160 --> 00:10:55,317
of these abstractions actually is lossy.

185
00:10:55,400 --> 00:10:58,357
It means it's inefficient.
It It doesn't contemplate all the

186
00:10:58,440 --> 00:11:01,317
complexity underneath it.
That's why it's an abstraction.

187
00:11:01,400 --> 00:11:05,157
So, what we're doing is kind
of simplifying this in some sense. We find

188
00:11:05,240 --> 00:11:08,917
an abstraction of the physics of the
semiconductor and connect that to the

189
00:11:09,000 --> 00:11:12,157
neural network. If you think about
your brain for a moment, it has a bunch of

190
00:11:12,240 --> 00:11:16,397
neurons in it, but there's no linear
algebra. There's no uh floating point

191
00:11:16,480 --> 00:11:20,717
math. It's actually the physics of the
neurons give rise to intelligence. We

192
00:11:20,800 --> 00:11:25,800
want to kind of mimic some
of that uh with a semiconductor.

193
00:11:26,520 --> 00:11:28,117
And this is also not a new concept.

194
00:11:28,200 --> 00:11:32,477
There's actually computation all
throughout nature. Uh birds in a flock,

195
00:11:32,560 --> 00:11:35,637
you may have seen things like this where
you know, a bird does a simple behavior.

196
00:11:35,720 --> 00:11:39,197
It looks left, looks right, figures
out where the next guy's going. And when

197
00:11:39,280 --> 00:11:42,477
they do that, they actually create these
interesting flocking behaviors, this

198
00:11:42,560 --> 00:11:46,517
emergent behavior. And we see that with
ant colonies. Ant colonies actually do

199
00:11:46,600 --> 00:11:51,037
intelligent things just by very simple
rules. And uh this study is called

200
00:11:51,120 --> 00:11:55,037
dynamical systems theory. It's basically
how I get these emergent properties from

201
00:11:55,120 --> 00:11:58,717
very simple behaviors of individual
components. Our brain actually works

202
00:11:58,800 --> 00:12:02,477
this way as well. But we're taking these
ideas and actually starting to build

203
00:12:02,560 --> 00:12:06,277
circuits out of them. So,
let me give you an example of such a

204
00:12:06,360 --> 00:12:09,397
such a system. So, everyone here
probably knows what a metronome is.

205
00:12:09,480 --> 00:12:12,277
Just, you know, when you're learning
to play the piano or something like that,

206
00:12:12,360 --> 00:12:16,477
it's just tick-tock, tick-tock. Physical
thing moving back and forth. So, this is

207
00:12:16,560 --> 00:12:20,557
kind of cool. If you put multiple ones
of them on a rigid plank, and that plank

208
00:12:20,640 --> 00:12:23,917
can just roll back and forth, you'll
actually see them start to synchronize.

209
00:12:24,000 --> 00:12:27,517
And that's just due to the physics
of the system. They each push against the

210
00:12:27,600 --> 00:12:30,357
plank just a little bit. And even if
they're off by a little bit from each

211
00:12:30,440 --> 00:12:33,317
other, they're all synchronized to exactly
the same phase. You can actually

212
00:12:33,400 --> 00:12:36,317
scale this up to hundreds of of
metronomes on a physical system, and

213
00:12:36,400 --> 00:12:40,757
they'll all synchronize. So, this is
actually a form of a dynamical system

214
00:12:40,840 --> 00:12:43,877
that goes through, you know,
some starting point of all these different

215
00:12:43,960 --> 00:12:47,837
phases, and always synchronizes. You
can imagine a slightly more complicated

216
00:12:47,920 --> 00:12:50,877
version of this where maybe they don't
all synchronize, maybe half of them are

217
00:12:50,960 --> 00:12:53,477
synchronized with each other,
and another half are synchronized

218
00:12:53,560 --> 00:12:57,237
uh in an opposite pattern or something
like this. But, the idea is that this is

219
00:12:57,320 --> 00:13:01,637
a physical system that that kind of uh
behaves the way it does just by by the

220
00:13:01,720 --> 00:13:05,597
the inherent interconnection
of the system itself.

221
00:13:05,680 --> 00:13:10,317
Uh so, we asked the question,
can we actually use such a system

222
00:13:10,400 --> 00:13:14,117
like like that metronome system to do
computation? We want to connect that to

223
00:13:14,200 --> 00:13:17,637
generative AI. That's what we really
care about. So, we released a model we

224
00:13:17,720 --> 00:13:21,157
called Uno, which actually demonstrated
this. It's an image generation model

225
00:13:21,240 --> 00:13:25,797
that's built on a set of oscillators
like that. And uh we simulated it and

226
00:13:25,880 --> 00:13:29,197
made it open source so people can
play with it. But this was a the first

227
00:13:29,280 --> 00:13:30,837
demonstration that I can actually scale

228
00:13:30,920 --> 00:13:36,880
something up, train it, and actually
get useful output like image generation.

229
00:13:37,520 --> 00:13:41,037
And uh you know, this is kind of some
of the analysis that we uh we we provided

230
00:13:41,120 --> 00:13:44,757
in that uh in that write-up.
What you're seeing here is

231
00:13:44,840 --> 00:13:46,477
we call it a state space trajectory.

232
00:13:46,560 --> 00:13:49,757
It's basically how the system evolves
in time. You can imagine if you

233
00:13:49,840 --> 00:13:54,317
characterize the state of the system
is all of the phases of those oscillators,

234
00:13:54,400 --> 00:13:57,917
um and then look at how it evolves
through time, and you basically

235
00:13:58,000 --> 00:14:01,557
condition that on the output. You say,
"I want to generate an airplane or a car

236
00:14:01,640 --> 00:14:05,157
or a bird." It'll actually go through
different state space trajectories. And

237
00:14:05,240 --> 00:14:07,717
so that's what we're seeing here is just
an analysis of this. And these are

238
00:14:07,800 --> 00:14:11,357
actual images that were generated by it.

239
00:14:11,440 --> 00:14:14,997
Now, it turns out we actually started
to build a lot of that fundamental science

240
00:14:15,080 --> 00:14:19,037
up over the last couple of months, and we
found other things that enabled this

241
00:14:19,120 --> 00:14:22,677
to work even better. So, this is
a concept of what we call sparsity.

242
00:14:22,760 --> 00:14:27,077
Sparsity means that, you know, if I have
uh let's say I have a bunch of elements

243
00:14:27,160 --> 00:14:30,597
all connected to each other like we
have on the left-hand side there. So, you

244
00:14:30,680 --> 00:14:33,797
know, if I have uh 10 elements and I
want to connect them all to each other,

245
00:14:33,880 --> 00:14:37,557
I have 10 * 10 elements.
I have 100 connections.

246
00:14:37,640 --> 00:14:41,037
That's okay, but if I have a thousand,
now I have thousand * thousand, which

247
00:14:41,120 --> 00:14:44,957
becomes a million. So, this doesn't
scale very well. We call this n²

248
00:14:45,040 --> 00:14:48,877
scaling. So, the more I add, it actually
becomes way more hard to scale, right?

249
00:14:48,960 --> 00:14:52,557
So, that's not great. So, sparsity
allows us to actually say, "Well, can I

250
00:14:52,640 --> 00:14:54,437
throw away some of those connections?"

251
00:14:54,520 --> 00:14:58,597
If I throw them away, now can I actually
rescue the behavior of the whole system?

252
00:14:58,680 --> 00:15:02,557
And it turns out you can actually not
only throw away some of the connections,

253
00:15:02,640 --> 00:15:04,997
but you can actually get better behavior
out of the whole system. It actually

254
00:15:05,080 --> 00:15:09,037
becomes more trainable. Uh there's a lot
of theoretical reasons for this, but we

255
00:15:09,120 --> 00:15:12,757
actually are able to do this in um in not
only simulated systems, but actually

256
00:15:12,840 --> 00:15:16,237
real physical systems. So, it's one
of these rare things where you get

257
00:15:16,320 --> 00:15:19,717
something that's more efficient,
uh that's actually more scalable, and even

258
00:15:19,800 --> 00:15:23,277
gives you more performance. So, this is
kind of a holy grail. It's been a It's

259
00:15:23,360 --> 00:15:25,917
been a problem for a long time, but we
had to frame the problem the right way

260
00:15:26,000 --> 00:15:29,197
to actually find a solution.
And so, this is actually the first time

261
00:15:29,280 --> 00:15:33,517
I'm talking about this publicly. This is I
wanted to do it at this venue because

262
00:15:33,600 --> 00:15:37,077
I think it's a really big deal.
This is actually

263
00:15:37,160 --> 00:15:40,077
the first physical dynamical computer

264
00:15:40,160 --> 00:15:45,747
ever built. We did this in 5
months. This company started

265
00:15:45,830 --> 00:15:50,677
>> [applause] >> Thank you.
Started in earnest in January. We didn't

266
00:15:50,760 --> 00:15:53,397
even have a team, but we said we're
going to build this first physical

267
00:15:53,480 --> 00:15:56,637
prototype and do it, you know, this year.
And actually, we taped out the

268
00:15:56,720 --> 00:16:00,677
design, meaning we sent it to the fab
in June on June 1st. The chip is back in

269
00:16:00,760 --> 00:16:04,397
our lab, and we actually have results
from it. So, these are the first ever

270
00:16:04,480 --> 00:16:06,757
images generated from such a computer.

271
00:16:06,840 --> 00:16:10,997
Now, great, cool,
but does it do anything useful

272
00:16:11,080 --> 00:16:14,117
beyond just images?
You can actually do any kind of

273
00:16:14,200 --> 00:16:17,477
a task like sequence modeling
or language models, but the interesting

274
00:16:17,560 --> 00:16:20,597
thing here is it's only five 500 or so

275
00:16:20,680 --> 00:16:22,677
nanojoules per image.

276
00:16:22,760 --> 00:16:24,957
So, to put that in perspective,

277
00:16:25,040 --> 00:16:27,957
a normal computer like a GPU is on the

278
00:16:28,040 --> 00:16:32,157
order of millijoules. So, nanojoules
is 10 to the minus ninth. It's really,

279
00:16:32,240 --> 00:16:36,757
really small. So, it's many orders of
magnitude more efficient than a standard

280
00:16:36,840 --> 00:16:39,757
computer. And it's because it just
doesn't move information around. This is

281
00:16:39,840 --> 00:16:43,157
proof positive this works. So, this is
literally the first time we're talking

282
00:16:43,240 --> 00:16:47,317
about it publicly. So, um Yeah, thank you.

283
00:16:47,400 --> 00:16:49,420
>> [applause]

284
00:16:50,720 --> 00:16:53,277
>> So, what's cool here is this is really
the emergence of something new. So,

285
00:16:53,360 --> 00:16:58,517
computers have gone from CPUs to GPUs,
going more and more parallel, to compute

286
00:16:58,600 --> 00:17:01,517
in memory, which is even more
parallel and fine-grained.

287
00:17:01,600 --> 00:17:04,557
Um but all of these are what do we call
von Neumann architecture. They have a

288
00:17:04,640 --> 00:17:07,837
memory and compute, and we move
information back and forth. What we

289
00:17:07,920 --> 00:17:11,397
built is what we call a dynamical
computer, which actually has compute and

290
00:17:11,480 --> 00:17:14,837
memory in one thing. We don't have
a memory interface. Each individual

291
00:17:14,920 --> 00:17:18,197
computing element is a memory. It's
a completely different way to look at the

292
00:17:18,280 --> 00:17:22,077
problem. And so we call
this 4D computing where

293
00:17:22,160 --> 00:17:25,317
we use the time dimension and the
dynamics, and we actually use the

294
00:17:25,400 --> 00:17:28,917
physical three dimensions of die
stacking, putting things vertically as

295
00:17:29,000 --> 00:17:32,637
well as in a in a planar form.
So we have three dimensions from from the

296
00:17:32,720 --> 00:17:35,877
physical and we have one dimension
in time. So this is actually kind of a new

297
00:17:35,960 --> 00:17:39,477
way of thinking about a computer,
and it it really is proving to work uh for

298
00:17:39,560 --> 00:17:42,957
efficiency's sake. So now,
what are the implications if we

299
00:17:43,040 --> 00:17:46,237
build something that's a thousand times
more power efficient? I think this is

300
00:17:46,320 --> 00:17:50,437
actually pretty cool. So intelligence
per watt is what we care about. Can we

301
00:17:50,520 --> 00:17:54,437
optimize this and make it better over
time? There is actually a thermodynamic

302
00:17:54,520 --> 00:17:59,397
limit which you can never exceed. And uh
the mammalian brains, so animal brains,

303
00:17:59,480 --> 00:18:02,997
are somewhere within one or two orders
of magnitude of that. Today we're on the

304
00:18:03,080 --> 00:18:05,797
far left of this graph, and we're about

305
00:18:05,880 --> 00:18:08,597
10 billion times away. That's 10 one

306
00:18:08,680 --> 00:18:12,957
with 10 zeros after it uh
from that thermodynamic limit.

307
00:18:13,040 --> 00:18:16,637
We think in that three and a half
years we can hit the the limits of 2D

308
00:18:16,720 --> 00:18:21,117
lithography. And the the overarching
goal of this company is to beat biology.

309
00:18:21,200 --> 00:18:25,077
We want to make something better
and enable, you know, compute everywhere,

310
00:18:25,160 --> 00:18:27,917
including compute in new robotic um

311
00:18:28,000 --> 00:18:32,997
forms uh and things like
this in the next uh decade or so.

312
00:18:33,080 --> 00:18:36,317
I think what will be interesting is that
we'll see the shift going from like big

313
00:18:36,400 --> 00:18:40,757
big data centers with gigawatts to many
small data centers all over the place. I

314
00:18:40,840 --> 00:18:42,837
think this is a good thing.
It actually makes things that are more

315
00:18:42,920 --> 00:18:45,117
environmentally friendly, more local,

316
00:18:45,200 --> 00:18:47,677
more adaptive.

317
00:18:47,760 --> 00:18:50,517
And as I said, I think enabling the

318
00:18:50,600 --> 00:18:55,437
ability to build billions of robots
that kind of dynamically assemble to solve

319
00:18:55,520 --> 00:18:58,797
problems in our world is something
that's I think is actually really cool.

320
00:18:58,880 --> 00:19:01,717
This is something that will enable us
to think about bigger problems and do even

321
00:19:01,800 --> 00:19:05,277
more. And you know,
I talked about AI being a

322
00:19:05,360 --> 00:19:10,317
trillion-dollar market. Well, if we
disrupt it by a thousand X, it's there's

323
00:19:10,400 --> 00:19:13,517
a there's a concept called Jevons
paradox where when you drop the

324
00:19:13,600 --> 00:19:18,597
underlying um cost of an asset,
you actually consume more of more than the

325
00:19:18,680 --> 00:19:20,917
drop of that asset. So, if you make
something at half the price, you'll

326
00:19:21,000 --> 00:19:25,077
consume more than 2x. If you can you
make something 1/1000 the price, you'll

327
00:19:25,160 --> 00:19:29,077
consume more than 1/1000 of it. And uh
I think this will create the largest

328
00:19:29,160 --> 00:19:31,277
market that humanity's ever seen.

329
00:19:31,360 --> 00:19:35,437
>> I mean, that was extremely
unexpected. I got to say that.

330
00:19:35,520 --> 00:19:38,277
Amazing. Uh let me let
me start with actually

331
00:19:38,360 --> 00:19:41,317
probably the thing that's
on everybody's mind, which is

332
00:19:41,400 --> 00:19:45,197
to the extent that this works, Naveen,
you probably saw Jensen earlier. There

333
00:19:45,280 --> 00:19:48,757
just needs to be an entire ecosystem
of people that are beside you and around

334
00:19:48,840 --> 00:19:50,637
you, whether it's the fabs, packagers,

335
00:19:50,720 --> 00:19:53,317
etc., etc.

336
00:19:53,400 --> 00:19:56,557
What does it take to get from this early

337
00:19:56,640 --> 00:20:00,357
version to something
that sits in somebody's

338
00:20:00,440 --> 00:20:04,797
hand or that people use? And how do
you see that path in terms of time and

339
00:20:04,880 --> 00:20:06,317
complexity? What does that look like?

340
00:20:06,400 --> 00:20:09,917
>> Yeah, time-wise, we're
within 2 years of of getting it

341
00:20:10,000 --> 00:20:13,317
>> to a full product. I mean,
>> And And what is the product?

342
00:20:13,400 --> 00:20:16,277
>> Yeah, it's this >> It's a VM that sits
somewhere that you guys manage and

343
00:20:16,360 --> 00:20:19,757
>> Effectively, we're we're building a new
data center product first. So, it's a

344
00:20:19,840 --> 00:20:21,837
whole rack, it's a system. And the idea

345
00:20:21,920 --> 00:20:27,237
is that basically we'll run those um
models on it. And so, tokens in, tokens

346
00:20:27,320 --> 00:20:29,837
out through a network cable, but the inner
guts are completely different than

347
00:20:29,920 --> 00:20:32,517
an existing computer.
>> And do you expect that you'll have to

348
00:20:32,600 --> 00:20:36,877
move to support the existing model
families and existing architectures? And

349
00:20:36,960 --> 00:20:40,677
will this work in a world where,
you know, we've spent all of this time,

350
00:20:40,760 --> 00:20:45,197
like, okay, KV cache, and let's all like
this all just so mechanically reductive

351
00:20:45,280 --> 00:20:47,677
based on, as you said, these
abstractions that we've lived on.

352
00:20:47,760 --> 00:20:49,997
>> Right. >> So, how do
you expect the rest of us to

353
00:20:50,080 --> 00:20:53,157
kind of move towards this? Cuz I mean,
I think you see that efficiency curve,

354
00:20:53,240 --> 00:20:55,277
we'd all want it. So,
how do we take advantage of it?

355
00:20:55,360 --> 00:20:58,437
>> Yeah, so I think there's a there's
a sliding scale between, you know, how

356
00:20:58,520 --> 00:21:01,997
much better something is and how much
pain you'll you'll take to move to it.

357
00:21:02,080 --> 00:21:05,357
And you know, I I basically took the
the tack of like, let's make it really,

358
00:21:05,440 --> 00:21:09,197
really compelling to move. There
is going to be some work to port things

359
00:21:09,280 --> 00:21:12,477
over. We actually don't port at the
operations layer, you port at the model

360
00:21:12,560 --> 00:21:17,117
layer. So, yes, the existing models will
work, but there's a fair bit of compute

361
00:21:17,200 --> 00:21:19,277
required to make that transition happen.

362
00:21:19,360 --> 00:21:23,677
>> And very basic elements like map
will does this is that exist in your it

363
00:21:23,760 --> 00:21:25,877
doesn't exist. It's >> I mean,
you can completely you can

364
00:21:25,960 --> 00:21:28,717
characterize it as map will,
but it doesn't implement it as map will.

365
00:21:28,800 --> 00:21:31,997
>> Okay. >> It implements it as a sort
of time-varying behavior, but each one of

366
00:21:32,080 --> 00:21:35,597
those time steps you can analyze
as basically a matrix of the current state

367
00:21:35,680 --> 00:21:37,997
times a transition matrix.
>> And when you're building a team like

368
00:21:38,080 --> 00:21:42,357
this like who are these people? These
are biologists plus physicists plus what

369
00:21:42,440 --> 00:21:46,117
are these people? >> Yes.
Um like it's sort of like theorists that

370
00:21:46,200 --> 00:21:49,317
that come up with these uh like
dynamical systems theory has been around

371
00:21:49,400 --> 00:21:52,877
for 100 years. So, we got people from
that world and then we got people who

372
00:21:52,960 --> 00:21:54,797
actually build chips and they
don't talk to each other.

373
00:21:54,880 --> 00:21:56,797
>> talk to each other. >> So,
we had to facilitate that. That's

374
00:21:56,880 --> 00:21:59,477
actually one of the most challenging
things about this company is the span of

375
00:21:59,560 --> 00:22:03,357
talents that we have is so big that
getting them to kind of all coordinate

376
00:22:03,440 --> 00:22:04,957
and build one thing
is actually pretty hard.

377
00:22:05,040 --> 00:22:08,237
>> And what is this like CUDA-like
equivalent if you will just to use a bad

378
00:22:08,320 --> 00:22:11,997
analogy that allows these people up
here to talk to these people down there?

379
00:22:12,080 --> 00:22:15,437
>> Yeah, so we actually have
built a set of libraries in Python.

380
00:22:15,520 --> 00:22:17,357
>> In Python. >> So,
it's it's Python. It's not CUDA, but

381
00:22:17,440 --> 00:22:21,157
it's you know, it's it's a language of
sorts that allows you to kind of express

382
00:22:21,240 --> 00:22:24,757
time-varying elements
that have stochastic behavior.

383
00:22:24,840 --> 00:22:26,197
>> I mean,
it's incredibly impressive. It's

384
00:22:26,280 --> 00:22:28,917
so ambitious. Thank you very much.
It's great to see you. Great to see you.

385
00:22:29,000 --> 00:22:31,990
Amazing. >> [music]
