WEBVTT
Kind: captions
Language: en-US

00:00:00.257 --> 00:00:00.872
[MUSIC] 
 
 

00:00:00.872 --> 00:00:06.000
DOUG BURGER: This is The Shape of Things to Come,&nbsp;
a Microsoft Research Podcast. I’m your host,&nbsp;&nbsp;

00:00:06.000 --> 00:00:12.720
Doug Burger. In this series, we’re going to&nbsp;
venture to the bleeding edge of AI capabilities,&nbsp;&nbsp;

00:00:12.720 --> 00:00:15.440
dig down into the fundamentals,&nbsp;
really try to understand them,&nbsp;&nbsp;

00:00:15.440 --> 00:00:22.000
and think about how these capabilities are going&nbsp;
to change the world—for better and worse.   
 
 

00:00:23.760 --> 00:00:31.280
In today’s podcast, I’m bringing on&nbsp;
two AI researcher-experts: Nicolò Fusi,&nbsp;&nbsp;

00:00:31.280 --> 00:00:37.440
who is an expert in digital, transformer-based&nbsp;
large language model architectures and learning,&nbsp;&nbsp;

00:00:37.440 --> 00:00:42.160
and Subutai Ahmad, who is an expert in&nbsp;
biological architectures, specifically&nbsp;&nbsp;

00:00:42.160 --> 00:00:48.400
the human brain. And the question we’re going&nbsp;
to discuss is, are machines intelligent?  
 
 

00:00:48.400 --> 00:00:53.440
And what I mean by that: are digital&nbsp;
intelligence, large language models,&nbsp;&nbsp;

00:00:53.440 --> 00:00:57.680
on a path to surpass humans, or&nbsp;
are the architectures just so&nbsp;&nbsp;

00:00:57.680 --> 00:01:01.520
fundamentally different that one&nbsp;
will do one set of things well,&nbsp;&nbsp;

00:01:01.520 --> 00:01:07.200
the other will do something else very well?&nbsp;
And so we’ll be debating the architecture of&nbsp;&nbsp;

00:01:07.200 --> 00:01:11.760
intelligence across digital implementations&nbsp;
and biological implementations because the&nbsp;&nbsp;

00:01:11.760 --> 00:01:17.413
answer to that question, I think, really will&nbsp;
determine the shape of things to come. 
 
 

00:01:17.413 --> 00:01:21.840
[MUSIC FADES] 
 

00:01:21.840 --> 00:01:25.760
I'd like to ask each of my guests&nbsp;
to introduce themselves. Tell me a&nbsp;&nbsp;

00:01:25.760 --> 00:01:30.000
little bit about your background and&nbsp;
what you're currently working on—to&nbsp;&nbsp;

00:01:30.000 --> 00:01:35.266
the extent you can talk about it—in AI.&nbsp;
So, Nicolò, would you please start? 
 
 

00:01:35.266 --> 00:01:41.600
NICOLÒ FUSI: Yeah, thank you, Doug, for&nbsp;
having us and having me here. It's so much&nbsp;&nbsp;

00:01:41.600 --> 00:01:47.040
fun. So I'm Nicolò Fusi. I'm a researcher at&nbsp;
MSR [Microsoft Research]. So Doug is my boss,&nbsp;&nbsp;

00:01:47.040 --> 00:01:50.880
so I will be very, very, very&nbsp;
good to Doug in this podcast.  
 
 

00:01:50.880 --> 00:01:57.120
No, but jokes aside, my own background is in&nbsp;
Bayesian nonparametric. That's what I started&nbsp;&nbsp;

00:01:57.120 --> 00:02:03.440
studying. So Gaussian processes and things&nbsp;
like that. And then equally, I would say,&nbsp;&nbsp;

00:02:03.440 --> 00:02:08.960
in computational biology, because I found it,&nbsp;
like, one of the most interesting use cases&nbsp;&nbsp;

00:02:09.520 --> 00:02:14.720
for AI techniques. And that, kind of,&nbsp;
has been true throughout my career. And&nbsp;&nbsp;

00:02:14.720 --> 00:02:21.440
pretty much like everybody else, eventually,&nbsp;
I moved away from the kernel methods and the&nbsp;&nbsp;

00:02:21.440 --> 00:02:27.600
Bayesian nonparametrics and I started working&nbsp;
more on language models, transformer models,&nbsp;&nbsp;

00:02:27.600 --> 00:02:32.960
with a particular eye towards information theory&nbsp;
and the connection between information theory and&nbsp;&nbsp;

00:02:32.960 --> 00:02:40.080
generative modeling. And that's, kind of, one of&nbsp;
the main things I do today other than, kind of,&nbsp;&nbsp;

00:02:40.080 --> 00:02:45.760
managing the research of people who do much&nbsp;
more interesting work than I do. [LAUGHS]  
 

00:02:45.760 --> 00:02:53.040
BURGER: I have to interject there, Nicolò, because&nbsp;
you dragged a piece of bait across my path.  
 
 

00:02:53.040 --> 00:02:53.938
FUSI: I figured.  
 
 

00:02:53.938 --> 00:02:57.280
BURGER: You know, at Microsoft Research,&nbsp;
I have a management rule that I can't tell&nbsp;&nbsp;

00:02:57.280 --> 00:03:00.880
anyone what to do because we hire some of the&nbsp;
best people in the world. You have to trust&nbsp;&nbsp;

00:03:00.880 --> 00:03:07.280
them. And everyone is always completely free to&nbsp;
call BS on me. And so Nicolò was joking there;&nbsp;&nbsp;

00:03:07.280 --> 00:03:13.434
[LAUGHTER] he does not have to toe the party&nbsp;
line. In fact, I encourage him not to. So, so … 
 

00:03:13.434 --> 00:03:16.712
FUSI: I just have to be well-behaved. That's&nbsp;
the only thing I will say. [LAUGHS] 
 

00:03:16.712 --> 00:03:19.200
BURGER: Yeah. Thank you, thank&nbsp;
you for baiting me. [LAUGHS]&nbsp;&nbsp;

00:03:19.200 --> 00:03:22.800
Because he knew exactly what he was&nbsp;
doing. And I love him for it.  
 
 

00:03:22.800 --> 00:03:25.985
Subutai, can you tell us a&nbsp;
little bit about yourself? 
 

00:03:25.985 --> 00:03:27.760
SUBUTAI AHMAD: Sure. Thank you so much, Doug,&nbsp;&nbsp;

00:03:27.760 --> 00:03:31.760
for having me. I'm really looking forward&nbsp;
to the conversation between us all.  
 
 

00:03:31.760 --> 00:03:35.600
So I see myself fundamentally as&nbsp;
a computer scientist. You know,&nbsp;&nbsp;

00:03:35.600 --> 00:03:39.760
I've been studying computer science&nbsp;
for longer than I care to admit.&nbsp;&nbsp;

00:03:42.480 --> 00:03:46.880
But something changed for me during my&nbsp;
undergrad years. I decided to minor in&nbsp;&nbsp;

00:03:46.880 --> 00:03:51.760
cognitive psychology, and I started to get&nbsp;
really interested in how the brain works. 
 

00:03:51.760 --> 00:03:56.640
And to me, understanding intelligence and&nbsp;
implementing intelligence was the hardest&nbsp;&nbsp;

00:03:56.640 --> 00:04:02.160
problem a computer scientists could ever solve.&nbsp;
So I got very, very interested in that. You know,&nbsp;&nbsp;

00:04:02.160 --> 00:04:07.760
I couldn't see how to really commercialize that. I&nbsp;
was very interested in making products and stuff.&nbsp;&nbsp;

00:04:07.760 --> 00:04:11.760
So I stopped, you know, working on&nbsp;
that for a while. I did a number of&nbsp;&nbsp;

00:04:11.760 --> 00:04:17.520
startups doing computer vision, you know,&nbsp;
video processing, a lot of that stuff. 
 
 

00:04:17.520 --> 00:04:22.720
And then when Jeff Hawkins started Numenta&nbsp;
back in 2005 with the idea of really deeply&nbsp;&nbsp;

00:04:22.720 --> 00:04:26.800
understanding how the brain works and&nbsp;
figuring out how to apply that to AI,&nbsp;&nbsp;

00:04:26.800 --> 00:04:32.160
for me, it was like all my worlds coming together.&nbsp;
This, like, this is what I had to do. None of us&nbsp;&nbsp;

00:04:32.160 --> 00:04:36.880
thought [LAUGHS] it would take as long as it&nbsp;
did. We spent the last couple of decades really&nbsp;&nbsp;

00:04:36.880 --> 00:04:41.600
deeply trying to understand neuroscience from a&nbsp;
computer scientist—from a programmer's—standpoint,&nbsp;&nbsp;

00:04:41.600 --> 00:04:45.360
the underlying algorithms. And that's&nbsp;
really what I'm passionate about,&nbsp;&nbsp;

00:04:45.360 --> 00:04:50.960
just trying to translate what we understand&nbsp;
about the neuroscience to today's AI.  
 
 

00:04:50.960 --> 00:04:54.720
And in terms of what we're working on today,&nbsp;
it's, you know, the human—maybe we'll get into&nbsp;&nbsp;

00:04:54.720 --> 00:04:59.840
some of this—the brain is super efficient in how&nbsp;
it works—power efficient, energy efficient—and&nbsp;&nbsp;

00:04:59.840 --> 00:05:05.480
we're trying to embody those ideas and trying to&nbsp;
make AI a lot more efficient than it is today. 
 

00:05:05.480 --> 00:05:09.440
BURGER: Great. I think we'll get into&nbsp;
efficiency a little bit later in the&nbsp;&nbsp;

00:05:09.440 --> 00:05:13.120
podcast because that's a subject that's&nbsp;
near and dear to my heart, you know,&nbsp;&nbsp;

00:05:13.120 --> 00:05:17.040
being a computer architect&nbsp;
originally by training.  
 
 

00:05:17.040 --> 00:05:21.680
I want to go back to, you know, one&nbsp;
of the reasons I got involved with&nbsp;&nbsp;

00:05:21.680 --> 00:05:25.360
Numenta is, you know, Subutai and I&nbsp;
have been exchanging emails, like,&nbsp;&nbsp;

00:05:25.360 --> 00:05:30.240
discussing collaborations, you know,&nbsp;
visiting each other through the years,&nbsp;&nbsp;

00:05:30.240 --> 00:05:36.960
and the thing that really stuck with me was&nbsp;
when I read one of the earlier books from Jeff&nbsp;&nbsp;

00:05:36.960 --> 00:05:44.400
on intelligence. And there was an example in the&nbsp;
book that talked about how, you know, the human&nbsp;&nbsp;

00:05:44.400 --> 00:05:49.520
brain learns continuously. I think biological&nbsp;
organisms in general learn continuously.  
 
 

00:05:49.520 --> 00:05:54.960
And the anecdote that I remember was this anecdote&nbsp;
if you're walking down your basement steps,&nbsp;&nbsp;

00:05:54.960 --> 00:05:58.560
you know, you're walking down the stairs to your&nbsp;
basement and there's one step that’s always been&nbsp;&nbsp;

00:05:58.560 --> 00:06:03.120
a few inches off and you decide to fix it, and&nbsp;
so you raise it so it's even with the others,&nbsp;&nbsp;

00:06:03.120 --> 00:06:07.280
and then the next time you go down the stairs,&nbsp;
you don't remember and you're wildly off and,&nbsp;&nbsp;

00:06:07.280 --> 00:06:11.760
you know, you hit that step, you hit it earlier&nbsp;
or later than you anticipated, you go out of&nbsp;&nbsp;

00:06:11.760 --> 00:06:14.640
balance. You're flailing around. You know, you&nbsp;
get all this adrenaline. You think you're going&nbsp;&nbsp;

00:06:14.640 --> 00:06:19.760
to pitch headfirst down the stairs. Hopefully&nbsp;
you don't. And then the second time you do it,&nbsp;&nbsp;

00:06:19.760 --> 00:06:24.560
you're a little off balance, but it's not crazy.&nbsp;
And the third time you maybe notice a little bit,&nbsp;&nbsp;

00:06:24.560 --> 00:06:27.680
and the fourth time, it's, like,&nbsp;
it's your basement stairs. 
 
 

00:06:27.680 --> 00:06:33.280
And so somewhere between that first time&nbsp;
down and the third and fourth times down,&nbsp;&nbsp;

00:06:33.280 --> 00:06:38.080
there are molecular changes in your brain&nbsp;
that have learned the new timing of your&nbsp;&nbsp;

00:06:38.080 --> 00:06:43.360
basement steps. And I remember just that example&nbsp;
vividly from the book. And that got me thinking,&nbsp;&nbsp;

00:06:43.360 --> 00:06:50.400
wow, this is so different from the way&nbsp;
our digital AI works. I'll turn it over&nbsp;&nbsp;

00:06:50.400 --> 00:06:53.680
to you to comment for that and then&nbsp;
I think we'll go into the digital. 
 

00:06:53.680 --> 00:06:58.800
AHMAD: Yeah, no, that's a great example.&nbsp;
I think it's remarkable how our brain is&nbsp;&nbsp;

00:06:58.800 --> 00:07:02.800
constantly modeling our entire&nbsp;
world at such a granular level,&nbsp;&nbsp;

00:07:02.800 --> 00:07:05.760
and we're not even aware of it&nbsp;
perceptually. Like, you know,&nbsp;&nbsp;

00:07:05.760 --> 00:07:11.040
that example of the steps is probably not&nbsp;
… you wouldn't consciously be aware of it,&nbsp;&nbsp;

00:07:11.040 --> 00:07:16.400
yet if something is different about anything in&nbsp;
your world that you're very familiar with, you'll&nbsp;&nbsp;

00:07:16.400 --> 00:07:21.280
instantly notice it. And then you'll, you know,&nbsp;
you'll update your world model, you'll adjust,&nbsp;&nbsp;

00:07:21.280 --> 00:07:26.360
and you'll continue on. It's really remarkable&nbsp;
how the brain’s able to do that so seamlessly. 
 

00:07:26.360 --> 00:07:31.040
BURGER: And a lot of that is based on&nbsp;
neurotransmitters, right? Because there's just&nbsp;&nbsp;

00:07:31.040 --> 00:07:35.600
a … you know, when you have that physical reaction&nbsp;
to “I’m about to pitch down the stairs,” you get&nbsp;&nbsp;

00:07:35.600 --> 00:07:40.000
a flood of transmitters that actually changes the&nbsp;
way your brain's learning or at least the rate. 
 

00:07:40.000 --> 00:07:46.880
AHMAD: Yeah, there's a flood of neurotransmitters&nbsp;
and neuromodulators, as well, that invoke change,&nbsp;&nbsp;

00:07:46.880 --> 00:07:50.800
sometimes very rapidly. Another example,&nbsp;
you know, if you touch a hot stove—that's&nbsp;&nbsp;

00:07:50.800 --> 00:07:57.280
the canonical example—you will learn that very,&nbsp;
very quickly. So there's a lot of chemical changes&nbsp;&nbsp;

00:07:57.280 --> 00:08:03.040
that happen. But it's also really interesting&nbsp;
that we can update things and update our world&nbsp;&nbsp;

00:08:03.040 --> 00:08:07.920
knowledge without impacting everything else&nbsp;
that we know. This is something that's very,&nbsp;&nbsp;

00:08:07.920 --> 00:08:11.920
very different, again, from today's AI&nbsp;
models. We're able to make these changes&nbsp;&nbsp;

00:08:11.920 --> 00:08:15.680
in a very contextual and very,&nbsp;
sort of, fine-grained way.  
 
 

00:08:15.680 --> 00:08:21.600
BURGER: So, Nicolò, I want to go and talk a&nbsp;
little bit now to transformers. So I think,&nbsp;&nbsp;

00:08:21.600 --> 00:08:25.840
you know, you and I and Subutai&nbsp;
were all working in the AI field,&nbsp;&nbsp;

00:08:26.400 --> 00:08:31.040
you know, many years before 2017,&nbsp;
when the transformer hit. You know,&nbsp;&nbsp;

00:08:31.040 --> 00:08:36.960
I was building, you know, with my team hardware&nbsp;
to accelerate RNNs [recurrent neural networks],&nbsp;&nbsp;

00:08:36.960 --> 00:08:42.800
LSTMs [long short-term memory], you know, which&nbsp;
had this awful loop-carried dependence, you know,&nbsp;&nbsp;

00:08:42.800 --> 00:08:47.840
the bottlenecked computation, and then the&nbsp;
transformer was just much more parallelizable.  
 
 

00:08:47.840 --> 00:08:54.000
So what do you think's really going on in&nbsp;
these things? And maybe we could start—I know&nbsp;&nbsp;

00:08:54.000 --> 00:08:58.240
you and I have talked a lot about this—maybe&nbsp;
just start with the major blocks. You know,&nbsp;&nbsp;

00:08:58.240 --> 00:09:02.320
you've got the attention layer. You've got&nbsp;
the feedforward layer. You've got, you know,&nbsp;&nbsp;

00:09:02.320 --> 00:09:07.680
the encoder stack and the decoder stack and the&nbsp;
latent space in between. Can you just, kind of,&nbsp;&nbsp;

00:09:07.680 --> 00:09:12.320
walk us through those pieces at a high level&nbsp;
and tell us what you think is going on? 
 

00:09:12.320 --> 00:09:13.680
FUSI: Yeah. Yeah, I mean,&nbsp;&nbsp;

00:09:13.680 --> 00:09:19.548
I have a very opinionated view of&nbsp;
why transformers are so great.  
 
 

00:09:19.548 --> 00:09:20.111
BURGER: That's why you’re here. [LAUGHS] 
 
 

00:09:20.111 --> 00:09:22.640
FUSI: Maybe, like, yeah, maybe I’ll&nbsp;
inject it. I don't know. I don't&nbsp;&nbsp;

00:09:22.640 --> 00:09:27.360
think it's a super novel creative&nbsp;
opinion, but it is an opinion. So&nbsp;&nbsp;

00:09:27.360 --> 00:09:31.840
I guess the two principal … the two main&nbsp;
components you already described: the, you know,&nbsp;&nbsp;

00:09:31.840 --> 00:09:36.080
the transformer [read: attention] layers and&nbsp;
the feedforward layers. One way to think about&nbsp;&nbsp;

00:09:36.080 --> 00:09:42.320
them is, how does information in your context&nbsp;
relate to each other and what is every token&nbsp;&nbsp;

00:09:42.320 --> 00:09:46.880
referring to, for instance, in the case&nbsp;
of transformers in language models? 
 

00:09:46.880 --> 00:09:50.400
So by context, we mean, like, the&nbsp;
information you feed through the model,&nbsp;&nbsp;

00:09:50.400 --> 00:09:53.120
that the model keeps continuously&nbsp;
generating and appending to. 
 

00:09:53.120 --> 00:09:54.240
BURGER: So like your chat history. 
 

00:09:54.240 --> 00:09:59.824
FUSI: Your prompt. Your what? Your chat history&nbsp;
or your particular prompt in a chat session.  
 
 

00:09:59.824 --> 00:10:00.343
BURGER: OK.  
 
 

00:10:00.343 --> 00:10:04.960
FUSI: That prompt, which is a sequence of words,&nbsp;
gets discretized in a series of tokens. Tokens can&nbsp;&nbsp;

00:10:04.960 --> 00:10:09.920
be individual words, can be multiple words, kind&nbsp;
of, connected together. The way we go from words&nbsp;&nbsp;

00:10:09.920 --> 00:10:16.240
to tokens typically is through an algorithm that&nbsp;
tries to basically collapse as much as possible.&nbsp;&nbsp;

00:10:16.240 --> 00:10:22.000
Multiple words, like “the dog,” may be just one&nbsp;
token as a first, kind of, level of compression&nbsp;&nbsp;

00:10:22.000 --> 00:10:27.840
to feed into the model. So it just tries to bring&nbsp;
things together as efficiently as possible.  
 
 

00:10:27.840 --> 00:10:32.640
Then there is, you know, within these models,&nbsp;
there is a transformer layer. This transformer&nbsp;&nbsp;

00:10:32.640 --> 00:10:37.680
layer or this attention layer, sorry,&nbsp;
tries to basically figure out what the&nbsp;&nbsp;

00:10:38.960 --> 00:10:46.480
“the” refers to—the term “the” in “the&nbsp;
dog,” or “the dog jumps on the table,”&nbsp;&nbsp;

00:10:46.480 --> 00:10:51.600
“jumps” refers to the dog. So there is this&nbsp;
kind of, like, mapping that happens. 
 
 

00:10:51.600 --> 00:10:57.120
And then there is, like, feedforward layers,&nbsp;
which in modern large language models,&nbsp;&nbsp;

00:10:57.120 --> 00:10:59.840
they store a lot of information.&nbsp;
Like, that's kind of, like,&nbsp;&nbsp;

00:10:59.840 --> 00:11:05.840
where the knowledge typically kind of sits in,&nbsp;
the things that the model just knows. You know,&nbsp;&nbsp;

00:11:06.800 --> 00:11:15.360
that, I don't know, if you slam your arm&nbsp;
against [the] cup of water on your table,&nbsp;&nbsp;

00:11:15.360 --> 00:11:19.200
that cup of water falls off the table.&nbsp;
That's something that the model, kind of,&nbsp;&nbsp;

00:11:19.200 --> 00:11:24.800
has baked in through reading a lot about cups&nbsp;
falling off of tables when they’re hit. 
 

00:11:24.800 --> 00:11:31.360
So that's, kind of, those are, for me, the&nbsp;
two fundamental components, and the reason&nbsp;&nbsp;

00:11:31.360 --> 00:11:37.840
why I have an opinionated view is that, you&nbsp;
know, honestly, I do believe that RNNs and,&nbsp;&nbsp;

00:11:37.840 --> 00:11:43.280
you know, even state-space—modern incarnations&nbsp;
of state-space models—are good enough to learn&nbsp;&nbsp;

00:11:43.280 --> 00:11:48.800
over these, you know, language data or&nbsp;
whatever or vision data or audio data. 
 

00:11:48.800 --> 00:11:52.720
The good thing about transformers is that&nbsp;
they do two things very well. One is they&nbsp;&nbsp;

00:11:52.720 --> 00:11:56.240
get out of the way. They don't have this&nbsp;
notion of “everything has to be encoded&nbsp;&nbsp;

00:11:56.240 --> 00:12:00.640
through a state” like recurrent networks.&nbsp;
And two, they do that very computationally&nbsp;&nbsp;

00:12:00.640 --> 00:12:05.040
efficiently as you were saying. There&nbsp;
isn't a computational bottleneck. And&nbsp;&nbsp;

00:12:05.040 --> 00:12:09.520
so they created this nice overhang where&nbsp;
they happen to be the right architecture&nbsp;&nbsp;

00:12:09.520 --> 00:12:13.811
at the right time to unlock enough flow&nbsp;
of information through the model … 
 
 

00:12:13.811 --> 00:12:14.349
BURGER: Yeah.  
 

00:12:14.349 --> 00:12:16.120
FUSI: … that we could get&nbsp;
through these amazing things. 
 

00:12:16.120 --> 00:12:21.360
BURGER: Let me press you on one thing.&nbsp;
Like, you know, in the attention blocks,&nbsp;&nbsp;

00:12:21.360 --> 00:12:26.160
you can figure out which words or which tokens&nbsp;
relate to which tokens. So I put in the prompt and&nbsp;&nbsp;

00:12:26.160 --> 00:12:33.360
it's finding all the relations and then feeding&nbsp;
those relations up to, you know, the feedforward&nbsp;&nbsp;

00:12:33.360 --> 00:12:39.840
layer—well, the feedforward unit within a layer.&nbsp;
And you said that knowledge is encoded there,&nbsp;&nbsp;

00:12:39.840 --> 00:12:45.120
but then what does it really mean for those maps&nbsp;
to then access knowledge, but then you project it&nbsp;&nbsp;

00:12:45.120 --> 00:12:52.227
back into, you know, the output and then feed it&nbsp;
up to the attention block in the next layer?  
 
 

00:12:52.227 --> 00:12:52.879
FUSI: Again, yeah.  
 
 

00:12:52.879 --> 00:12:55.520
BURGER: So it seems kind of weird that&nbsp;
I’d be, like, accessing knowledge and then&nbsp;&nbsp;

00:12:55.520 --> 00:12:59.680
taking that knowledge, merging it, and&nbsp;
going back to another attention map. 
 
 

00:12:59.680 --> 00:13:04.960
FUSI: Well, you can see it as a mixing&nbsp;
operation that happens in the feedforward&nbsp;&nbsp;

00:13:04.960 --> 00:13:10.320
part of the layer. You know, like, you're&nbsp;
attending, then you're mixing, and, kind of, like,&nbsp;&nbsp;

00:13:10.320 --> 00:13:18.960
reprojecting to some space with higher-information&nbsp;
content or, like, a different level of information&nbsp;&nbsp;

00:13:18.960 --> 00:13:24.560
extraction. And then you're putting it back into,&nbsp;
“OK, so let me do another round of processing”&nbsp;&nbsp;

00:13:25.840 --> 00:13:30.720
and, kind of, attending and then a mix again. And&nbsp;
then I do it again and then I do it again.  
 

00:13:30.720 --> 00:13:34.560
So I think that the information that&nbsp;
is present in the prompt and in the,&nbsp;&nbsp;

00:13:34.560 --> 00:13:39.600
you know, that has been baked into the weights&nbsp;
gather further and further refined. Whether&nbsp;&nbsp;

00:13:39.600 --> 00:13:48.880
that refinement is extraction of structure or&nbsp;
aggregation into higher-level concepts, I'm not&nbsp;&nbsp;

00:13:48.880 --> 00:13:55.040
sure. I think it's just structure gets extracted&nbsp;
and things that are irrelevant get kind of pushed&nbsp;&nbsp;

00:13:55.040 --> 00:13:59.120
away. But that doesn't necessarily mean that it&nbsp;
gets aggregated through the architecture.  
 
 

00:13:59.120 --> 00:14:04.560
BURGER: So now I'm going to try to, like, restate&nbsp;
what I think I hear you saying. So, you know,&nbsp;&nbsp;

00:14:04.560 --> 00:14:09.920
we're adding information and we're kind of adding&nbsp;
information at a higher level but not necessarily&nbsp;&nbsp;

00:14:09.920 --> 00:14:13.577
throwing away the low-level information,&nbsp;
at least that's not relevant, right?  
 
 

00:14:13.577 --> 00:14:14.111
FUSI: Yeah. 
 
 

00:14:14.111 --> 00:14:16.880
BURGER: Because, you know, if the higher-level&nbsp;
stuff depends on the low-level stuff, I have to&nbsp;&nbsp;

00:14:16.880 --> 00:14:21.760
have that first. And so then you get to the top of&nbsp;
the encoder block and you're in the latent space&nbsp;&nbsp;

00:14:21.760 --> 00:14:27.840
with all of that information kind of maximized.&nbsp;
Is that a way to think about it? And if you agree,&nbsp;&nbsp;

00:14:27.840 --> 00:14:32.240
can you talk about what the encoder block&nbsp;
really is and what the latent space is? 
 

00:14:33.440 --> 00:14:40.640
FUSI: I tend to agree, yes. I mean, there&nbsp;
is … you're describing … I think you're&nbsp;&nbsp;

00:14:40.640 --> 00:14:48.800
describing what I think is happening, which is&nbsp;
there is given the context in your prompt and&nbsp;&nbsp;

00:14:48.800 --> 00:14:52.960
given the task that the model perceives&nbsp;
or, like, figures out that you're doing,&nbsp;&nbsp;

00:14:52.960 --> 00:14:59.760
it has to highlight and pull out the relevant&nbsp;
information. And it does that not by summarizing&nbsp;&nbsp;

00:14:59.760 --> 00:15:07.280
layer by layer, but it does it by, you know,&nbsp;
increasing the prominence of that information&nbsp;&nbsp;

00:15:07.280 --> 00:15:13.440
and suppressing other things. So I think that's&nbsp;
ultimately what happens up to the point where&nbsp;&nbsp;

00:15:13.440 --> 00:15:21.440
you reach this beautiful point in concept space,&nbsp;
which identifies both your intent and the things&nbsp;&nbsp;

00:15:21.440 --> 00:15:25.160
in the prompt and in the knowledge of the&nbsp;
model that are necessary to solve it. 
 
 

00:15:25.160 --> 00:15:28.800
BURGER: And so one last question, and then&nbsp;
I want to go to Subutai for a second.  
 
 

00:15:29.440 --> 00:15:35.120
So now when we go through the decoder stack, are&nbsp;
we just going the other way and stripping out the&nbsp;&nbsp;

00:15:35.120 --> 00:15:40.880
high-level concepts early and then getting down to&nbsp;
the granular tokens? Or, you know … because you go&nbsp;&nbsp;

00:15:40.880 --> 00:15:45.040
up through the encoder stack, those attention&nbsp;
blocks and feedforward layers, to get to that&nbsp;&nbsp;

00:15:45.040 --> 00:15:49.840
magical latent space. And now we're going to go&nbsp;
the other direction. How do you think about that&nbsp;&nbsp;

00:15:49.840 --> 00:15:55.520
other direction through the decoder stack, which&nbsp;
is the same primitives as the encoder stack? 
 

00:15:56.240 --> 00:16:04.080
FUSI: Same primitives. You can think of it&nbsp;
as kind of the reverse operation. Like you,&nbsp;&nbsp;

00:16:04.080 --> 00:16:09.360
you never lost information throughout. You just&nbsp;
kind of suppress or privileged different kinds of&nbsp;&nbsp;

00:16:09.360 --> 00:16:15.680
information. And now you're basically just&nbsp;
projecting it back out to a space that is,&nbsp;&nbsp;

00:16:15.680 --> 00:16:22.240
you know, intelligible. And it's, kind of,&nbsp;
where the model gets it's … I hesitate to&nbsp;&nbsp;

00:16:22.240 --> 00:16:25.200
use the term reward because it has a&nbsp;
particular implication, but that's,&nbsp;&nbsp;

00:16:25.200 --> 00:16:29.280
kind of, where the loss gets computed and&nbsp;
then gets pushed back through the model. 
 

00:16:29.280 --> 00:16:32.480
BURGER: Right, as you're trying&nbsp;
to evolve and train all those&nbsp;&nbsp;

00:16:32.480 --> 00:16:36.480
parameters—the relationship between words,&nbsp;
the information in the feedforward layers,&nbsp;&nbsp;

00:16:36.480 --> 00:16:41.240
the design of that latent space, and the&nbsp;
extraction of the knowledge from it. 
 

00:16:41.240 --> 00:16:46.480
FUSI: That's right. And so in encoder-decoder&nbsp;
model, you push through the whole thing,&nbsp;&nbsp;

00:16:46.480 --> 00:16:50.720
you decode back to a particular token,&nbsp;
which for people who don't know, it's,&nbsp;&nbsp;

00:16:50.720 --> 00:16:53.120
like, literally a number out of a vocabulary,&nbsp;&nbsp;

00:16:53.120 --> 00:17:01.642
like word No. 487. And if it was word&nbsp;
No. 1,500, you get, you know, like, … 
 

00:17:01.642 --> 00:17:02.177
BURGER: Something else. 
 

00:17:02.177 --> 00:17:05.840
FUSI: … a bad reward. Yeah. Yeah.&nbsp;
And then … and if you got it right,&nbsp;&nbsp;

00:17:05.840 --> 00:17:09.760
you get a positive signal that then&nbsp;
just flows back through the model. 
 
 

00:17:09.760 --> 00:17:14.320
BURGER: I'd like to go over to Subutai now. So&nbsp;
after hearing this, you've studied, you know,&nbsp;&nbsp;

00:17:14.320 --> 00:17:19.600
neuroscience and the neocortex and cortical&nbsp;
columns and all of this for a long time,&nbsp;&nbsp;

00:17:19.600 --> 00:17:24.960
and you and I have had lots of debates. Is&nbsp;
the human brain doing something different&nbsp;&nbsp;

00:17:24.960 --> 00:17:29.760
than that? You know, are we&nbsp;
just building latent spaces,&nbsp;&nbsp;

00:17:29.760 --> 00:17:34.240
then extracting? The architecture is very&nbsp;
different, but what's going on under the hood? 
 

00:17:34.240 --> 00:17:37.120
AHMAD: Yeah, the architecture&nbsp;
is very different. You know,&nbsp;&nbsp;

00:17:37.120 --> 00:17:41.200
as Nicolò was describing what happens&nbsp;
throughout a transformer stack,&nbsp;&nbsp;

00:17:41.200 --> 00:17:46.160
I was trying to relay and relate, you know,&nbsp;
what we know in the brain, as well.  
 
 

00:17:47.040 --> 00:17:52.320
In a typical, you know, transformer model, there&nbsp;
is, at the end of the day, there is a single&nbsp;&nbsp;

00:17:52.320 --> 00:17:58.480
latent space from which the next token is output.&nbsp;
That does not happen in the brain. There are&nbsp;&nbsp;

00:17:58.480 --> 00:18:03.360
thousands and thousands of latent spaces that are,&nbsp;
sort of, collaborating together, if you will.  
 
 

00:18:04.560 --> 00:18:09.520
You know, a lot of what we publish is under&nbsp;
the moniker the Thousand Brains Theory of&nbsp;&nbsp;

00:18:09.520 --> 00:18:15.280
Intelligence. And Jeff has published a&nbsp;
book a few years ago on that. And that,&nbsp;&nbsp;

00:18:15.280 --> 00:18:19.520
kind of, dates back to discoveries in&nbsp;
neuroscience from the ’60s and ’70s&nbsp;&nbsp;

00:18:19.520 --> 00:18:23.840
by the neuroscientist Vernon Mountcastle,&nbsp;
who was a professor at Johns Hopkins. 
 
 

00:18:23.840 --> 00:18:24.409
BURGER: Yup. 
 

00:18:24.409 --> 00:18:28.320
AHMAD: And what he discovered … he made&nbsp;
this remarkable discovery that, you know,&nbsp;&nbsp;

00:18:28.320 --> 00:18:32.240
our neocortex, which is the biggest part&nbsp;
of our brain—that's where all intelligent&nbsp;&nbsp;

00:18:32.240 --> 00:18:39.280
function happens—is actually composed of roughly&nbsp;
100,000 what he called cortical columns. 
 
 

00:18:39.280 --> 00:18:39.880
BURGER: Right.  
 
 

00:18:39.880 --> 00:18:46.240
AHMAD: And each cortical column is maybe&nbsp;
50,000 neurons. And there's a very complex&nbsp;&nbsp;

00:18:46.240 --> 00:18:50.640
microcircuit and microarchitecture between&nbsp;
the neurons in a cortical column.  
 

00:18:50.640 --> 00:18:56.480
But then there's 100,000 of them, and&nbsp;
every part of your brain—whether it's&nbsp;&nbsp;

00:18:56.480 --> 00:19:02.080
doing visual processing, auditory&nbsp;
processing, language, thought,&nbsp;&nbsp;

00:19:02.080 --> 00:19:07.760
motor actions—they're all composed of this,&nbsp;
essentially, this same microarchitecture. And&nbsp;&nbsp;

00:19:07.760 --> 00:19:12.720
this was a remarkable discovery. It says&nbsp;
that there's a universal architecture.&nbsp;&nbsp;

00:19:12.720 --> 00:19:17.840
It's not a simple one. It's complex. But&nbsp;
it's repeated throughout the brain. 
 
 

00:19:17.840 --> 00:19:21.280
And that's where this, you know, the&nbsp;
idea of the Thousand Brains … each of&nbsp;&nbsp;

00:19:21.280 --> 00:19:27.600
these cortical columns is actually a complete&nbsp;
sensory-motor processing system. It has inputs;&nbsp;&nbsp;

00:19:27.600 --> 00:19:32.320
it has outputs. It's getting sensory&nbsp;
input. It's sending outputs to motor&nbsp;&nbsp;

00:19:32.320 --> 00:19:37.920
systems. And it's building, in our&nbsp;
theory, complete world models. So there&nbsp;&nbsp;

00:19:37.920 --> 00:19:41.760
isn't a single latent space. There's&nbsp;
thousands of these latent spaces. 
 

00:19:41.760 --> 00:19:46.800
And each little cortical column is trying&nbsp;
to understand its little bit of the world.&nbsp;&nbsp;

00:19:46.800 --> 00:19:50.480
You know, one cortical column might&nbsp;
be getting, at the lowest level,&nbsp;&nbsp;

00:19:50.480 --> 00:19:55.440
maybe one degree of visual information from&nbsp;
the top right-hand corner of your retina.&nbsp;&nbsp;

00:19:55.440 --> 00:20:00.800
Another one might be focusing on specific&nbsp;
frequencies in the auditory range. You know,&nbsp;&nbsp;

00:20:00.800 --> 00:20:06.080
each one has its own little view of the world,&nbsp;
and it's building its own little world model. 
 

00:20:06.080 --> 00:20:12.240
And then they all collaborate together. There's no&nbsp;
top or bottom here. There's no homunculus in the&nbsp;&nbsp;

00:20:12.240 --> 00:20:19.440
brain. Everything is sort of equal. And they're&nbsp;
all simultaneously collaborating and voting and&nbsp;&nbsp;

00:20:19.440 --> 00:20:26.160
coming up to, you know, what is the, you know,&nbsp;
consistent interpretation of all of these sensory&nbsp;&nbsp;

00:20:27.040 --> 00:20:32.160
inputs that we're getting? What is the&nbsp;
single consistent, you know, concept,&nbsp;&nbsp;

00:20:32.160 --> 00:20:37.440
if you will, and, based on that, make the motor&nbsp;
actions that are most relevant to that. 
 

00:20:37.440 --> 00:20:42.640
So it's a sensory-motor loop. It's a, you&nbsp;
know, it's a constantly recurring system;&nbsp;&nbsp;

00:20:42.640 --> 00:20:48.160
we’re constantly making predictions.&nbsp;
As we discussed earlier, you know,&nbsp;&nbsp;

00:20:48.160 --> 00:20:53.520
we are constantly learning. Every cortical&nbsp;
column is constantly updating its connections,&nbsp;&nbsp;

00:20:53.520 --> 00:20:58.560
constantly updating its weights. It's building&nbsp;
and incrementally improving its world model&nbsp;&nbsp;

00:20:58.560 --> 00:21:05.520
constantly. So it's a massively distributed,&nbsp;
you know, set of processing elements that we&nbsp;&nbsp;

00:21:05.520 --> 00:21:10.240
call cortical columns that are, they're&nbsp;
all equal, operating in parallel. 
 

00:21:10.240 --> 00:21:14.000
So I think there are similarities, for&nbsp;
sure, between them. But at least the&nbsp;&nbsp;

00:21:14.000 --> 00:21:20.000
way I described it, I think it's very&nbsp;
different in its operation than what I&nbsp;&nbsp;

00:21:20.000 --> 00:21:24.880
understand today’s LLMs to be. I don't&nbsp;
know if you agree with that or not. 
 
 

00:21:24.880 --> 00:21:28.960
FUSI: Yeah, I … To better understand,&nbsp;
I had a question, which is,&nbsp;&nbsp;

00:21:28.960 --> 00:21:32.560
are these cortical columns relying on&nbsp;
the fact that these are essentially&nbsp;&nbsp;

00:21:32.560 --> 00:21:37.440
multiple views of the same process&nbsp;
and those multiple views, like,&nbsp;&nbsp;

00:21:37.440 --> 00:21:41.840
the, you know, the part of the sensory&nbsp;
input that gets allocated or subdivided,&nbsp;&nbsp;

00:21:41.840 --> 00:21:47.760
is it happening at the same time point? So in&nbsp;
other words, if you could artificially delay&nbsp;&nbsp;

00:21:47.760 --> 00:21:53.440
by some time t some cortical columns with respect&nbsp;
to the rest, would the learning suffer?  
 
 

00:21:53.440 --> 00:21:53.910
AHMAD: Yes, absolutely. Yeah.  
 
 

00:21:53.910 --> 00:21:57.744
FUSI: And so in other words, how important is&nbsp;
it that it's, kind of, on the same schedule? 
 
 

00:21:57.744 --> 00:22:01.920
AHMAD: [LAUGHS] Yeah, I mean, that's another … I&nbsp;
mean, LLMs today, you know, you get your input,&nbsp;&nbsp;

00:22:01.920 --> 00:22:05.040
one layer processes it, then the next,&nbsp;
then the next, and the other layers are&nbsp;&nbsp;

00:22:05.040 --> 00:22:09.440
not operating. In the brain, it’s not&nbsp;
like that. Everything is operating in&nbsp;&nbsp;

00:22:09.440 --> 00:22:14.720
parallel asynchronously. And this is important.&nbsp;
They're constantly trying to make predictions&nbsp;&nbsp;

00:22:14.720 --> 00:22:18.640
and so on. So if you were to artificially&nbsp;
slow down some of your cortical columns,&nbsp;&nbsp;

00:22:18.640 --> 00:22:22.000
you would absolutely suffer. Your&nbsp;
thinking would absolutely suffer. 
 
 

00:22:22.000 --> 00:22:28.560
BURGER: I wanted to interject here just because&nbsp;
this is where … this discussion is where,&nbsp;&nbsp;

00:22:28.560 --> 00:22:33.440
you know, I got super interested in the&nbsp;
difference and then spent a bunch of time&nbsp;&nbsp;

00:22:33.440 --> 00:22:41.120
with Subutai to learn from him. So if I think&nbsp;
about my skin, you know, which is an organ,&nbsp;&nbsp;

00:22:42.160 --> 00:22:48.400
you know, as I understand it, there's a&nbsp;
cortical column attached to each patch of&nbsp;&nbsp;

00:22:48.400 --> 00:22:55.760
my skin and the size of that patch, kind of,&nbsp;
corresponds to the nerve density there.  
 
 

00:22:55.760 --> 00:22:57.200
AHMAD: That’s right. Yeah. 
 
 

00:22:57.200 --> 00:23:01.920
BURGER: So in my brain, there is a set of&nbsp;
cortical columns that are skin sensors,&nbsp;&nbsp;

00:23:01.920 --> 00:23:04.800
and I could actually … if I numbered&nbsp;
all the cortical columns in the brain,&nbsp;&nbsp;

00:23:04.800 --> 00:23:08.880
I could draw a map on my skin and say,&nbsp;
“This is No. 72 in this patch. This&nbsp;&nbsp;

00:23:08.880 --> 00:23:14.240
is No. 73 in this patch.” Now are human&nbsp;
cortical columns, like, better than, say,&nbsp;&nbsp;

00:23:14.240 --> 00:23:19.824
what we see in a mouse? And, of course, this is&nbsp;
a leading question because I know the answer. 
 

00:23:19.824 --> 00:23:24.160
AHMAD: [LAUGHS] Yeah. So, yes, it, you know,&nbsp;
cortical columns in your sensory areas,&nbsp;&nbsp;

00:23:24.160 --> 00:23:29.760
primary sensory areas, each, you know,&nbsp;
pay attention to or get input from a,&nbsp;&nbsp;

00:23:29.760 --> 00:23:33.120
you know, some patch of your skin&nbsp;
somewhere on your body. And there's&nbsp;&nbsp;

00:23:33.120 --> 00:23:37.280
many more cortical columns associated&nbsp;
with your fingertips than, you know,&nbsp;&nbsp;

00:23:37.280 --> 00:23:43.680
a square centimeter of your back, for example.&nbsp;
So there's definitely, you know, areas of sensory&nbsp;&nbsp;

00:23:44.240 --> 00:23:49.120
information that we pay a lot more attention to&nbsp;
and devote a lot more physical resources to.  
 
 

00:23:49.760 --> 00:23:57.920
In terms of a mouse and humans, it's pretty&nbsp;
remarkable that the cortical columns … so all&nbsp;&nbsp;

00:23:57.920 --> 00:24:01.760
mammals have cortical columns; all mammals&nbsp;
have a neocortex. All mammals have cortical&nbsp;&nbsp;

00:24:01.760 --> 00:24:07.280
columns from a mouse all the way up to humans.&nbsp;
And mice have cortical columns that are very,&nbsp;&nbsp;

00:24:07.280 --> 00:24:13.040
very similar to what a human has. It's not&nbsp;
identical. There are differences. But by and&nbsp;&nbsp;

00:24:13.040 --> 00:24:18.480
large, the architecture of a cortical column&nbsp;
in a mouse is, you know, very, very similar&nbsp;&nbsp;

00:24:18.480 --> 00:24:23.200
to cortical columns in humans. Human cortical&nbsp;
columns are bigger. There are more neurons,&nbsp;&nbsp;

00:24:23.200 --> 00:24:27.432
and there's more detail there, but&nbsp;
essentially, it's the same. And …  
 

00:24:27.432 --> 00:24:29.200
BURGER: Maybe just scaled up a little bit.  
 
 

00:24:29.200 --> 00:24:34.400
AHMAD: Yeah. So evolution basically discovered&nbsp;
this structure—that it's really excellent&nbsp;&nbsp;

00:24:34.400 --> 00:24:40.560
for processing information and dealing with&nbsp;
it—and then through, you know, very fast in&nbsp;&nbsp;

00:24:40.560 --> 00:24:45.760
evolutionary time, basically figured out that if&nbsp;
you could scale up the number of cortical columns,&nbsp;&nbsp;

00:24:45.760 --> 00:24:53.280
you get more intelligent animals. And that's&nbsp;
what happened very, very fast evolutionarily. 
 

00:24:53.280 --> 00:24:57.280
FUSI: I didn't know about the unevenness&nbsp;
of cortical columns present. Like,&nbsp;&nbsp;

00:24:57.280 --> 00:25:02.880
this is not … I'm not a neuroscientist, and&nbsp;
so this is interesting because one of the&nbsp;&nbsp;

00:25:02.880 --> 00:25:08.640
biggest frustrations with many modern&nbsp;
architectures of models is that they&nbsp;&nbsp;

00:25:08.640 --> 00:25:12.480
deploy a constant amount of computation&nbsp;
no matter what the input is.  
 
 

00:25:13.120 --> 00:25:21.760
So I go through the same number of layers whether&nbsp;
I'm trying to predict the word “dog” after “the”&nbsp;&nbsp;

00:25:21.760 --> 00:25:26.960
or whether I'm trying to solve, like, give the&nbsp;
final answer to a very complicated math question&nbsp;&nbsp;

00:25:26.960 --> 00:25:33.840
or, you know, whether a theorem was proven or not&nbsp;
in the prompt. And so that's interesting because,&nbsp;&nbsp;

00:25:33.840 --> 00:25:38.800
like, some current instantiations of modern&nbsp;
architecture actually deploy … try to cluster&nbsp;&nbsp;

00:25:38.800 --> 00:25:43.520
things together such that you have a constant&nbsp;
amount of information that you then push&nbsp;&nbsp;

00:25:43.520 --> 00:25:48.640
together through the model. [LAUGHTER] And&nbsp;
so maybe like on my fingertips, I need more&nbsp;&nbsp;

00:25:48.640 --> 00:25:54.000
processing than I need on my elbow because, like,&nbsp;
you know … and so this, kind of, makes sense. 
 

00:25:54.000 --> 00:25:58.560
BURGER: Nicolò is being humble. He was working&nbsp;
on this problem two years ago and told me about&nbsp;&nbsp;

00:25:58.560 --> 00:26:02.880
it. It was one of the things I learned from&nbsp;
you that made me think differently. So … 
 

00:26:02.880 --> 00:26:05.520
FUSI: I just like to refer to people&nbsp;
are working on this ... [LAUGHS] 
 

00:26:05.520 --> 00:26:10.640
BURGER: Random average people who are not&nbsp;
all necessarily brilliant AI scientists.  
 
 

00:26:11.280 --> 00:26:15.120
So the prediction part of this, though,&nbsp;
is really what's fascinating to me,&nbsp;&nbsp;

00:26:15.120 --> 00:26:20.240
because, again, something else Subutai and I&nbsp;
discussed many years ago, you know, if I'm,&nbsp;&nbsp;

00:26:20.240 --> 00:26:25.920
like, moving my finger towards the table and …&nbsp;
my brain is making predictions because I have&nbsp;&nbsp;

00:26:25.920 --> 00:26:31.680
a world model. It knows a table is there. And the&nbsp;
cortical columns representing that patch of skin,&nbsp;&nbsp;

00:26:31.680 --> 00:26:35.600
as it's getting closer, they're starting&nbsp;
to predict that I'm going to feel something&nbsp;&nbsp;

00:26:35.600 --> 00:26:40.480
that feels like the table. And, yup,&nbsp;
there; I hit it. Prediction met.  
 
 

00:26:40.480 --> 00:26:46.160
But if I touched it and it felt really&nbsp;
icy cold or super hot or fluffy or not&nbsp;&nbsp;

00:26:46.160 --> 00:26:50.080
there—I pass through it—I'd get&nbsp;
a flurry of activity because the&nbsp;&nbsp;

00:26:50.080 --> 00:26:54.160
prediction wouldn't match the world model,&nbsp;
and that's where learning would happen.  
 
 

00:26:54.160 --> 00:26:57.760
Subutai, does that sound like the&nbsp;
right model and intuition?  
 

00:26:57.760 --> 00:27:01.920
AHMAD: Yeah, that's definitely a very important&nbsp;
component of it. We're constantly making&nbsp;&nbsp;

00:27:01.920 --> 00:27:06.640
predictions. And as you said, you know,&nbsp;
you're moving your right fingertip down;&nbsp;&nbsp;

00:27:07.760 --> 00:27:12.480
you know, perhaps you've never sat&nbsp;
in this room before or, you know,&nbsp;&nbsp;

00:27:12.480 --> 00:27:16.524
seen this table before, you would still have&nbsp;
a prediction, a very good prediction of it. 
 

00:27:16.524 --> 00:27:17.440
BURGER: Yeah. Because you know what a table is. 
 

00:27:17.440 --> 00:27:22.720
AHMAD: You know what a table is. And if&nbsp;
it was different, you would, you know,&nbsp;&nbsp;

00:27:22.720 --> 00:27:27.680
you would notice it right away. But if your left&nbsp;
hand, which you weren't paying attention to,&nbsp;&nbsp;

00:27:27.680 --> 00:27:33.600
also felt icy cold, then you would notice&nbsp;
that, as well. So you're actually making not&nbsp;&nbsp;

00:27:33.600 --> 00:27:39.512
just one prediction; you're making thousands and&nbsp;
thousands of predictions constantly about ... 
 

00:27:39.512 --> 00:27:40.416
BURGER: Every cortical column. 
 

00:27:40.416 --> 00:27:43.680
AHMAD: Every cortical column is making&nbsp;
predictions. And if something were anomalous,&nbsp;&nbsp;

00:27:43.680 --> 00:27:47.680
highly anomalous, you would notice&nbsp;
it. So this is something, you know,&nbsp;&nbsp;

00:27:47.680 --> 00:27:52.320
we don't often realize; we're making&nbsp;
very, very granular predictions&nbsp;&nbsp;

00:27:52.320 --> 00:27:56.400
constantly. And when things are&nbsp;
wrong, we do learn from it.  
 
 

00:27:57.280 --> 00:28:01.600
And the other interesting thing—and this&nbsp;
is, again, possibly different from how&nbsp;&nbsp;

00:28:01.600 --> 00:28:06.800
LLMs work— you know, if I were to tell&nbsp;
you to touch the, you know, the bottom&nbsp;&nbsp;

00:28:08.880 --> 00:28:13.840
surface of the table, you could without, again,&nbsp;
without looking at the table or opening your eyes,&nbsp;&nbsp;

00:28:13.840 --> 00:28:20.240
you would be able to move your finger and touch&nbsp;
the bottom of your table because you have a,&nbsp;&nbsp;

00:28:20.240 --> 00:28:23.654
you know, set of reference&nbsp;
frames that relate to ...  
 
 

00:28:23.654 --> 00:28:25.280
BURGER: Yup …
AHMAD: There you go. Yep. You're able to do it.

00:28:25.280 --> 00:28:27.080
BURGER: I did it! Yeah. Amazing. 
 

00:28:27.080 --> 00:28:29.520
AHMAD: Even though you maybe&nbsp;
never have been in this room;&nbsp;&nbsp;

00:28:29.520 --> 00:28:31.520
maybe you’ve never seen this table&nbsp;
before. It doesn't matter. 
 

00:28:31.520 --> 00:28:34.480
BURGER: I’ve been in this room because we&nbsp;
had to prep for the podcast series. But&nbsp;&nbsp;

00:28:34.480 --> 00:28:36.960
I didn't touch the underside of the&nbsp;
table, that's for sure. [LAUGHS] 
 

00:28:36.960 --> 00:28:41.920
AHMAD: Yeah, exactly. [LAUGHS] So, you know, we&nbsp;
know where things are in relation to each other,&nbsp;&nbsp;

00:28:41.920 --> 00:28:47.360
where our body is in relation to everything,&nbsp;
and we can very, very rapidly learn. And again,&nbsp;&nbsp;

00:28:47.360 --> 00:28:52.680
if the bottom part of the table was anomalous, you&nbsp;
would notice it and potentially remember that. 
 
 

00:28:52.680 --> 00:28:54.480
FUSI: I'm not going to lie. I was expecting you to&nbsp;&nbsp;

00:28:54.480 --> 00:28:58.360
find something under that table,&nbsp;
[LAUGHTER] like a talk show. 
 

00:28:58.360 --> 00:28:59.360
AHMAD: Or chewing gum or something. 
 

00:28:59.360 --> 00:29:03.920
FUSI: And if you reach under the table, you're&nbsp;
going to find a copy of my paper. [LAUGHS] 
 

00:29:03.920 --> 00:29:07.760
BURGER: [LAUGHS] You know, if I&nbsp;
was smarter and better prepared,&nbsp;&nbsp;

00:29:07.760 --> 00:29:11.520
that's exactly what would have&nbsp;
happened. But, sorry, guys.  
 
 

00:29:12.880 --> 00:29:18.880
I think you told me something, Subutai, you know,&nbsp;
that … and I'll give a little bit of preamble.  
 
 

00:29:18.880 --> 00:29:24.800
So, you know, the brain has these dendritic&nbsp;
networks in each neuron, and they form synapses.&nbsp;&nbsp;

00:29:24.800 --> 00:29:30.400
And so a neuron fires, and that, you know, the&nbsp;
axon of the neuron that's firing will propagate&nbsp;&nbsp;

00:29:30.400 --> 00:29:34.960
a signal through the synapses, which might do a&nbsp;
little signal processing to the dendrites of the&nbsp;&nbsp;

00:29:34.960 --> 00:29:39.840
downstream neurons, and those downstream—the&nbsp;
dendrites can then prime the neuron to fire.&nbsp;&nbsp;

00:29:39.840 --> 00:29:45.040
That's one of the fundamental mechanisms.&nbsp;
And it's the formation of those synapses,&nbsp;&nbsp;

00:29:45.040 --> 00:29:48.480
you know, between the upstream and&nbsp;
downstream neurons, the dendrites,&nbsp;&nbsp;

00:29:48.480 --> 00:29:59.520
that seem to be the basis of learning, and to me,&nbsp;
that feels a little bit like an attention map. 
 
 

00:29:59.520 --> 00:30:00.312
AHMAD: Yes.  
 
 

00:30:00.312 --> 00:30:03.520
BURGER: So maybe the dendritic network is&nbsp;
doing something akin to self-attention,&nbsp;&nbsp;

00:30:03.520 --> 00:30:10.640
and we have some work going on in that direction&nbsp;
at MSR. But the thing you told me was that your&nbsp;&nbsp;

00:30:10.640 --> 00:30:18.320
brain is actually forming an incredibly large&nbsp;
number of synapses speculatively. In some sense,&nbsp;&nbsp;

00:30:18.320 --> 00:30:26.160
sampling the world when something happens in case&nbsp;
it will recur. You know, it's a more … maybe it's&nbsp;&nbsp;

00:30:26.160 --> 00:30:30.040
a version of Hebbian learning, right? You know,&nbsp;
things that fire together, wire together. 
 

00:30:30.040 --> 00:30:30.560
AHMAD: Exactly. 
 

00:30:30.560 --> 00:30:33.520
BURGER: But then if that pattern doesn't recur,&nbsp;&nbsp;

00:30:33.520 --> 00:30:38.480
then they get pruned. And I’m just going&nbsp;
to, you know, what is the fraction of your&nbsp;&nbsp;

00:30:38.480 --> 00:30:42.000
synapses to get turned over every three&nbsp;
or four days, you know, ballpark? 
 

00:30:42.000 --> 00:30:45.280
AHMAD: OK. Yeah, I remember this.&nbsp;
This was an absolute mind-blowing&nbsp;&nbsp;

00:30:45.280 --> 00:30:49.760
study in [The Journal of]&nbsp;
Neuroscience. So, you know,&nbsp;&nbsp;

00:30:49.760 --> 00:30:54.960
the way a lot of learning happens in the brain&nbsp;
is by adding and dropping connections. 
 
 

00:30:54.960 --> 00:31:00.400
In AI models, it's usually strengthening, you&nbsp;
know, high-precision floating-point number,&nbsp;&nbsp;

00:31:00.400 --> 00:31:03.920
making it higher or lower. But you're not adding&nbsp;
and dropping connections. The connections are&nbsp;&nbsp;

00:31:03.920 --> 00:31:10.400
always—in fact, everything is fully connected,&nbsp;
right, between layers. And so in the brain,&nbsp;&nbsp;

00:31:10.400 --> 00:31:15.040
you're always adding and dropping connections.&nbsp;
That's a fundamental mechanism by which we learn,&nbsp;&nbsp;

00:31:17.520 --> 00:31:19.280
one of the fundamental mechanisms.  
 
 

00:31:19.840 --> 00:31:26.480
What I read in this study is that they&nbsp;
looked at adult mice and adult animals,&nbsp;&nbsp;

00:31:26.480 --> 00:31:32.160
and what they found is that they would look&nbsp;
at the number of synapses that were connected&nbsp;&nbsp;

00:31:32.160 --> 00:31:36.960
over the course of a couple of months—and they&nbsp;
were able to trace individual synapses in this&nbsp;&nbsp;

00:31:36.960 --> 00:31:44.560
particular part of the brain—and what they found&nbsp;
is that every four days, 30% of the synapses&nbsp;&nbsp;

00:31:44.560 --> 00:31:51.840
that were there were no longer there four days&nbsp;
from now. And there was a new 30%. And there's&nbsp;&nbsp;

00:31:51.840 --> 00:31:58.560
a huge number of connections that are constantly&nbsp;
being added and constantly being pruned. And my&nbsp;&nbsp;

00:31:58.560 --> 00:32:03.600
theory of what's going on there is that we're&nbsp;
always speculatively trying to learn things. 
 

00:32:03.600 --> 00:32:08.320
So, you know, there's all sorts of&nbsp;
random coincidences and things that&nbsp;&nbsp;

00:32:08.320 --> 00:32:13.760
we are exposed to on a day-to-day basis.&nbsp;
We're constantly forming connections there&nbsp;&nbsp;

00:32:13.760 --> 00:32:17.920
because we don't know what's actually&nbsp;
going to be required and what's real&nbsp;&nbsp;

00:32:17.920 --> 00:32:21.920
and what's random. Most of it's random;&nbsp;
most of it's not necessary. And the stuff&nbsp;&nbsp;

00:32:21.920 --> 00:32:26.720
that actually is necessary will stay on.&nbsp;
But we're constantly trying to learn. 
 

00:32:26.720 --> 00:32:32.160
This is a part of continuous learning that's often&nbsp;
not appreciated, I think, is that we're constantly&nbsp;&nbsp;

00:32:32.160 --> 00:32:36.640
forming new connections, and then we prune&nbsp;
the stuff that we don't need. In an AI model,&nbsp;&nbsp;

00:32:36.640 --> 00:32:40.720
if you were to do that, it would just go, I&nbsp;
don't know, it would go bananas. [LAUGHTER]  
 

00:32:40.720 --> 00:32:47.273
BURGER: Well, so let's double-click on that.&nbsp;
So when you told me that, the way I … 
 

00:32:47.273 --> 00:32:49.407
AHMAD: This is mind-blowing, this 30%.  
 
 

00:32:49.407 --> 00:32:49.961
BURGER: It’s crazy.  
 
 

00:32:49.961 --> 00:32:52.800
AHMAD: Your brain is going to be totally&nbsp;
different a few days from now. 
 
 

00:32:52.800 --> 00:33:00.000
BURGER: It's so mind-blowing. When you told&nbsp;
me that, I spent some time processing it,&nbsp;&nbsp;

00:33:00.000 --> 00:33:04.000
so a whole bunch of synapses were created&nbsp;
and destroyed during that time.  
 
 

00:33:04.000 --> 00:33:08.960
But it just made me think that we have,&nbsp;
you know, we have all of these columns&nbsp;&nbsp;

00:33:08.960 --> 00:33:16.800
getting all of this input continuously. You&nbsp;
know, eyes, hearing, smell, taste, skin,&nbsp;&nbsp;

00:33:16.800 --> 00:33:22.880
heat, and then, you know, interactions with&nbsp;
people, and then planning and experiences,&nbsp;&nbsp;

00:33:22.880 --> 00:33:30.160
just at every level. And they're constantly&nbsp;
sampling all this noise coming in and basically&nbsp;&nbsp;

00:33:30.160 --> 00:33:36.240
filtering out the noise. It's like, kind of,&nbsp;
like a low-pass filter. But when something&nbsp;&nbsp;

00:33:36.240 --> 00:33:42.040
statistically significant recurs, it's going&nbsp;
to lock and then become persistent.  
 
 

00:33:42.040 --> 00:33:46.800
AHMAD: Yeah, yeah, I think so. There's so much&nbsp;
that's happening, and you’re constantly learning,&nbsp;&nbsp;

00:33:46.800 --> 00:33:51.120
and, you know, when you touch a hot stove&nbsp;
or something, there's a flood of dopamine&nbsp;&nbsp;

00:33:51.120 --> 00:33:57.200
specific to those areas that caused these&nbsp;
synapses to strengthen very, very quickly.&nbsp;&nbsp;

00:33:58.720 --> 00:34:04.297
You know, most of these synapses that are&nbsp;
learned are very, very weak synapses.  
 

00:34:04.297 --> 00:34:04.821
BURGER: Yup. 
 
 

00:34:04.821 --> 00:34:07.120
AHMAD: And so, yeah, you know,&nbsp;
when you look … in this study,&nbsp;&nbsp;

00:34:07.120 --> 00:34:12.160
they also quantified the turnover in, kind&nbsp;
of, strong synapses versus weak synapses.&nbsp;&nbsp;

00:34:12.160 --> 00:34:16.800
And it's comforting to know that the strong&nbsp;
synapses stay there. It's really these weak&nbsp;&nbsp;

00:34:16.800 --> 00:34:20.440
synapses that are constantly added and dropped.&nbsp;
And then some of them will become strong. 
 

00:34:20.440 --> 00:34:25.280
BURGER: Now I want to go back … return&nbsp;
to Nicolò, but with an observation.   
 
 

00:34:25.280 --> 00:34:32.560
So when I'm training a transformer, it's&nbsp;
also a prediction-based system. You know,&nbsp;&nbsp;

00:34:32.560 --> 00:34:39.280
I'm running … I have my input in the training&nbsp;
set; I have my masked token or the next token&nbsp;&nbsp;

00:34:39.280 --> 00:34:45.920
I'm trying to predict. I run it through. I look&nbsp;
at how successfully did it make that prediction,&nbsp;&nbsp;

00:34:45.920 --> 00:34:51.040
and the worse it was, the, sort of,&nbsp;
the steeper the error, you know,&nbsp;&nbsp;

00:34:51.040 --> 00:34:56.720
I drive back through the network. So, you&nbsp;
know, if it's spot-on, I don't learn very&nbsp;&nbsp;

00:34:56.720 --> 00:35:00.960
much. But if the prediction is way off,&nbsp;
I've got to change a bunch of stuff. That&nbsp;&nbsp;

00:35:00.960 --> 00:35:05.680
sounds analogous to what Subutai was just&nbsp;
describing with the cortical columns. 
 

00:35:05.680 --> 00:35:14.720
FUSI: No, that's right. I mean, with,&nbsp;
I don't know, with one big pet peeve of&nbsp;&nbsp;

00:35:14.720 --> 00:35:18.219
mine in pretraining, in particular around&nbsp;
pretraining these language models.  
 
 

00:35:18.219 --> 00:35:18.746
BURGER: OK. 
 
 

00:35:18.746 --> 00:35:21.520
FUSI: So again, for context, like,&nbsp;
language models in particular, but,&nbsp;&nbsp;

00:35:21.520 --> 00:35:27.040
you know, many other instantiations of large&nbsp;
models, are trained in a few phases usually.&nbsp;&nbsp;

00:35:27.040 --> 00:35:33.120
One of them is pretraining, where you have&nbsp;
some ground truth text and you remove,&nbsp;&nbsp;

00:35:33.120 --> 00:35:37.360
let's say, just the last word, and then you&nbsp;
ask the model to predict the last word. And&nbsp;&nbsp;

00:35:37.360 --> 00:35:42.000
that's when you get that loss. Do you get the&nbsp;
word right? Do you get the word wrong?  
 
 

00:35:43.680 --> 00:35:46.080
One of the big problems that I have is that,&nbsp;&nbsp;

00:35:47.040 --> 00:35:52.400
you know, in human experience, we do not&nbsp;
get feedback every single thought.  
 
 

00:35:52.400 --> 00:35:56.000
The problem with language models, the way we&nbsp;
are training them, at least in pretraining,&nbsp;&nbsp;

00:35:56.000 --> 00:36:00.640
is that they do a thing called teacher forcing. So&nbsp;
they guess the word, then they get immediately the&nbsp;&nbsp;

00:36:00.640 --> 00:36:04.480
signal, and then the right word gets filled&nbsp;
in, and then they predict the next one. 
 

00:36:04.480 --> 00:36:09.360
So when you go through, like, a passage&nbsp;
of text, you constantly get this reward.&nbsp;&nbsp;

00:36:09.360 --> 00:36:13.360
And it's such a bizarre way to train a&nbsp;
model. It's necessary because you want&nbsp;&nbsp;

00:36:13.360 --> 00:36:20.320
a lot of flow of supervision. Like, you want,&nbsp;
like, a lot of supervision to essentially use&nbsp;&nbsp;

00:36:20.320 --> 00:36:25.040
all the computation available. But at the same&nbsp;
time, it actually makes the models arguably a&nbsp;&nbsp;

00:36:25.040 --> 00:36:29.840
little bit worse than what they would be if you&nbsp;
had enough compute to train them without this. 
 

00:36:29.840 --> 00:36:32.444
I went on a tangent just because&nbsp;
it's a pet peeve. [LAUGHS]  
 

00:36:32.444 --> 00:36:35.600
BURGER: It's a really important point,&nbsp;
though, because your goal when you're&nbsp;&nbsp;

00:36:35.600 --> 00:36:42.480
training a model is to get to your loss&nbsp;
target with the minimal cost and time.&nbsp;&nbsp;

00:36:42.480 --> 00:36:46.400
Or, of course, like, fixed budget&nbsp;
and, like, lowest loss target.  
 
 

00:36:46.400 --> 00:36:51.280
But, you know, biological systems, also, their&nbsp;
goal is survival with energy minimization. And so,&nbsp;&nbsp;

00:36:51.280 --> 00:36:55.200
like, once you've built a world model that&nbsp;
works, right, like touching the table,&nbsp;&nbsp;

00:36:55.200 --> 00:36:58.560
touching the underside of the table—nope,&nbsp;
still nothing exciting there—like,&nbsp;&nbsp;

00:36:59.120 --> 00:37:04.400
it takes very little energy to do that. And&nbsp;
I think a tragedy is that we all have these&nbsp;&nbsp;

00:37:04.400 --> 00:37:09.040
supercomputers in our heads. You know, the&nbsp;
neocortex is what, about 10 watts? And it's&nbsp;&nbsp;

00:37:09.040 --> 00:37:14.400
this amazing thing, right, that can compose&nbsp;
symphonies. But once we have a world model,&nbsp;&nbsp;

00:37:14.400 --> 00:37:19.360
a lot of us just stop learning because it's&nbsp;
comfortable, right. You don't have to perturb&nbsp;&nbsp;

00:37:19.360 --> 00:37:24.400
the state. You can go through … and, you know,&nbsp;
I mean, how many of us go through every day and&nbsp;&nbsp;

00:37:24.400 --> 00:37:28.480
all of our predictions succeed [LAUGHTER],&nbsp;
and there's no surprises, you know?  
 
 

00:37:28.480 --> 00:37:32.400
So all the new synapses get swept away,&nbsp;
right. That's not a goal of pretraining&nbsp;&nbsp;

00:37:32.400 --> 00:37:35.600
because then you're just wasting energy. But&nbsp;
we're trying to minimize energy consumption.&nbsp;&nbsp;

00:37:36.240 --> 00:37:40.080
So it does feel, kind of,&nbsp;
aligned to me in some sense. 
 

00:37:40.080 --> 00:37:44.000
So I've got a straw man I want to&nbsp;
hit you with, but before we do,&nbsp;&nbsp;

00:37:45.440 --> 00:37:51.280
Nicolò, I want you to talk about your view&nbsp;
on compression, like LLMs as compressors,&nbsp;&nbsp;

00:37:51.280 --> 00:37:53.600
because I know this is something&nbsp;
you're very passionate about and&nbsp;&nbsp;

00:37:53.600 --> 00:37:56.400
opinionated about. And I've learned&nbsp;
a lot from you on this, too. 
 
 

00:37:57.360 --> 00:38:02.240
And then, Subutai, after this, I'd like&nbsp;
to hear your biological response. I mean,&nbsp;&nbsp;

00:38:02.240 --> 00:38:07.120
your response from a biological&nbsp;
perspective. [LAUGHTER] And …  
 
 

00:38:07.120 --> 00:38:08.360
AHMAD: You'll get both.  
 
 

00:38:08.360 --> 00:38:11.200
BURGER: That's right, of course.&nbsp;
And then I want to try … I want&nbsp;&nbsp;

00:38:11.200 --> 00:38:15.600
to throw out this hybrid straw man. So,&nbsp;
Nicolò, tell us about compression. 
 

00:38:15.600 --> 00:38:22.400
FUSI: The view is that basically the&nbsp;
generative models are compressors in&nbsp;&nbsp;

00:38:22.400 --> 00:38:27.440
an information theoretic sense, and so&nbsp;
trying to come up with a better generative&nbsp;&nbsp;

00:38:27.440 --> 00:38:32.388
model is equivalent to trying to find the&nbsp;
best compressor for some data. And … 
 

00:38:32.388 --> 00:38:35.120
BURGER: Now when you say compressor,&nbsp;
do you mean lossless or lossy? 
 

00:38:35.120 --> 00:38:37.120
FUSI: I mean lossless.  
 
 

00:38:37.120 --> 00:38:37.840
BURGER: OK. 
 
 

00:38:37.840 --> 00:38:45.760
FUSI: You can basically look at literally my&nbsp;
much-maligned objective function that you use&nbsp;&nbsp;

00:38:45.760 --> 00:38:52.800
for pretraining, which is, you know, next-token&nbsp;
prediction, and you can basically draw a complete&nbsp;&nbsp;

00:38:53.440 --> 00:38:56.880
parallel to what you would do if&nbsp;
you were trying to come up with the,&nbsp;&nbsp;

00:38:56.880 --> 00:39:00.880
you know, try to do compression, which is&nbsp;
coming up with the shortest possible code&nbsp;&nbsp;

00:39:02.240 --> 00:39:03.920
for something that you're trying to compress. 
 

00:39:03.920 --> 00:39:10.080
And so the two things are the same, and it,&nbsp;
kind of, fits into a broader picture that,&nbsp;&nbsp;

00:39:10.080 --> 00:39:16.720
you know, like, goes back to Occam's razor&nbsp;
and Kolmogorov complexity and Solomonoff's&nbsp;&nbsp;

00:39:16.720 --> 00:39:23.520
principle of induction, which is, you want&nbsp;
short descriptions for likely things that&nbsp;&nbsp;

00:39:23.520 --> 00:39:27.440
happen in the world and you want your&nbsp;
algorithm that produces those short&nbsp;&nbsp;

00:39:27.440 --> 00:39:30.640
descriptions to be also short. That's the&nbsp;
minimum description length principle.  
 
 

00:39:30.640 --> 00:39:36.000
And I do feel like it fits in, kind&nbsp;
of, also what you were saying about&nbsp;&nbsp;

00:39:36.000 --> 00:39:41.280
the concept of you have a good world&nbsp;
model, why look for surprise? Because&nbsp;&nbsp;

00:39:41.280 --> 00:39:46.640
it simultaneously affects both terms, both&nbsp;
the algorithm, like your own world model,&nbsp;&nbsp;

00:39:46.640 --> 00:39:50.080
but also the loss that you incur&nbsp;
when something unexpected happens. 
 

00:39:50.080 --> 00:39:55.120
And so if I'm an agent in the&nbsp;
world trying to minimize the&nbsp;&nbsp;

00:39:55.120 --> 00:39:59.520
minimum description length of the&nbsp;
world, I’d like to go and seek some&nbsp;&nbsp;

00:39:59.520 --> 00:40:04.160
in-distribution data such that I don't&nbsp;
bump up my surprise term too much. 
 

00:40:04.160 --> 00:40:07.840
BURGER: Right. And I think you&nbsp;
said at some point that, you know,&nbsp;&nbsp;

00:40:07.840 --> 00:40:11.440
when I'm training a model, even though&nbsp;
you took the same loss point, you know,&nbsp;&nbsp;

00:40:11.440 --> 00:40:17.760
between Model A and Model B, if I have a&nbsp;
steeper loss curve in Model A than Model B,&nbsp;&nbsp;

00:40:17.760 --> 00:40:24.160
you know, it's getting to a better, sort&nbsp;
of, compressed-based vocabulary faster,&nbsp;&nbsp;

00:40:24.160 --> 00:40:28.360
which makes it more general. The shape of that&nbsp;
curve matters from a compression perspective. 
 

00:40:28.360 --> 00:40:32.160
FUSI: Yeah. I mean, I think&nbsp;
it would help here to expand&nbsp;&nbsp;

00:40:32.160 --> 00:40:33.672
on what I was talking about in terms of, … 
 
 

00:40:33.672 --> 00:40:34.262
BURGER: Yes. Please.  
 
 

00:40:34.262 --> 00:40:36.560
FUSI: … like, minimum description length&nbsp;
principle. The minimum description length&nbsp;&nbsp;

00:40:36.560 --> 00:40:40.320
principle is basically the loss&nbsp;
of the model you're training;&nbsp;&nbsp;

00:40:40.320 --> 00:40:46.640
that's one component. And so it's a sum over the&nbsp;
mistakes you make at predicting or, you know,&nbsp;&nbsp;

00:40:46.640 --> 00:40:52.560
the mistakes you make at predicting each word.&nbsp;
And that's one term. And the other term is how&nbsp;&nbsp;

00:40:52.560 --> 00:40:57.812
long it takes you in code to describe the&nbsp;
model and the training procedure, … 
 
 

00:40:57.812 --> 00:40:58.348
BURGER: Right. 
 
 

00:40:58.348 --> 00:41:01.342
FUSI: … to get to that training curve,&nbsp;
to produce that training curve.  
 
 

00:41:01.342 --> 00:41:01.884
BURGER: Right. 
 
 

00:41:01.884 --> 00:41:06.480
FUSI: So, yes, if you look at collectively,&nbsp;
one term is, kind of, fixed. It's an amount&nbsp;&nbsp;

00:41:06.480 --> 00:41:10.720
of code it would take you to write out a&nbsp;
language model, for instance, in code. Like,&nbsp;&nbsp;

00:41:10.720 --> 00:41:17.520
literally implement it, not the weights, just&nbsp;
implement the initialization of it and then the&nbsp;&nbsp;

00:41:17.520 --> 00:41:21.680
training loop. And then on the other side, you&nbsp;
have this training loss that gets generated as&nbsp;&nbsp;

00:41:21.680 --> 00:41:26.960
you start observing data. And, of course, because&nbsp;
it's a sum, you want to minimize really the area,&nbsp;&nbsp;

00:41:27.600 --> 00:41:34.640
like, you want to minimize the sum. And so,&nbsp;
like, a flatter curve is much better than,&nbsp;&nbsp;

00:41:34.640 --> 00:41:39.240
like, the steeper curve, you know, even if it&nbsp;
ends up at the end to be slightly better. 
 

00:41:39.240 --> 00:41:42.240
BURGER: Yeah. Concave is better than convex. 
 

00:41:42.240 --> 00:41:45.032
FUSI: Among other things, yes. [LAUGHTER] 
 

00:41:45.032 --> 00:41:52.880
BURGER: Sorry. So, you know, I think that we&nbsp;
could do a whole episode on this compression&nbsp;&nbsp;

00:41:52.880 --> 00:41:58.240
view because it's really fascinating. And&nbsp;
the lossless part of it is what blew my mind.&nbsp;&nbsp;

00:41:58.960 --> 00:42:03.200
And I think, you know, I'm guessing there&nbsp;
are multiple camps here, and you're squarely&nbsp;&nbsp;

00:42:03.200 --> 00:42:08.400
in one camp, so I'm guessing we'll get a&nbsp;
bunch of feedback from the other camps. 
 

00:42:08.400 --> 00:42:14.360
So, Subutai, you know, can I think of&nbsp;
cortical columns as compressors? 
 

00:42:14.360 --> 00:42:22.640
AHMAD: Yeah, it's a good question. You know,&nbsp;
I, you know, there's so much in the compression&nbsp;&nbsp;

00:42:22.640 --> 00:42:27.440
literature that you can draw insight from.&nbsp;
You know, if you look at the representations&nbsp;&nbsp;

00:42:27.440 --> 00:42:35.120
in cortical columns and that populations that&nbsp;
neurons have, you know, some of the things you&nbsp;&nbsp;

00:42:35.120 --> 00:42:41.040
have to deal with are that the brain doesn't have&nbsp;
a huge nuclear power plant attached to it. 
 
 

00:42:41.040 --> 00:42:47.680
You know, we only have 12 watts or so to process&nbsp;
everything we want to do, and the representations&nbsp;&nbsp;

00:42:48.240 --> 00:42:54.000
that evolution has discovered are incredibly&nbsp;
sparse. And what that means is that you may&nbsp;&nbsp;

00:42:54.000 --> 00:42:59.600
have thousands and thousands of neurons in a&nbsp;
layer, but only about 1% of them will actually&nbsp;&nbsp;

00:42:59.600 --> 00:43:06.960
be active at a time. And so it's a very small&nbsp;
subset of neurons that are actually active.  
 
 

00:43:07.920 --> 00:43:13.200
I don't know about this minimum description&nbsp;
length, whether that applies. I can say a&nbsp;&nbsp;

00:43:13.200 --> 00:43:17.920
couple of things about that. There's, you know,&nbsp;
by and large, the representations are very sparse&nbsp;&nbsp;

00:43:17.920 --> 00:43:23.737
when you're predicting well. When you see a&nbsp;
surprise, there's a burst of activity.  
 
 

00:43:23.737 --> 00:43:24.264
BURGER: Yup. 
 
 

00:43:24.264 --> 00:43:27.912
AHMAD: When there's something that's unusual,&nbsp;
there's a lot more neurons that fire, and ... 
 

00:43:27.912 --> 00:43:29.320
BURGER: That's why learning is tiring!  
 

00:43:29.320 --> 00:43:33.280
AHMAD: That's why learning [LAUGHTER] …&nbsp;
exactly. No, no, that's right, that's right.

00:43:33.920 --> 00:43:37.120
And so what we think is&nbsp;
happening is that, you know,&nbsp;&nbsp;

00:43:37.120 --> 00:43:41.280
the actual representation of something is a very&nbsp;
small number of neurons. When you're surprised,&nbsp;&nbsp;

00:43:41.280 --> 00:43:44.480
there may be many things that are&nbsp;
consistent with that surprise,&nbsp;&nbsp;

00:43:44.480 --> 00:43:49.200
and so your brain represents a union&nbsp;
of all of those things at once. 
 

00:43:49.200 --> 00:43:51.520
And when you have a very sparse representation,&nbsp;&nbsp;

00:43:51.520 --> 00:43:55.760
you can actually have a union of many, many&nbsp;
different things without getting confused. So&nbsp;&nbsp;

00:43:55.760 --> 00:44:01.520
that's what we think is going on there. So it is&nbsp;
a very compressed, very efficient representation.&nbsp;&nbsp;

00:44:02.480 --> 00:44:06.720
And because it's such a small percentage&nbsp;
of neurons that are firing, we are very,&nbsp;&nbsp;

00:44:06.720 --> 00:44:13.680
very parsimonious in how we represent things&nbsp;
and extremely energy efficient metabolically. 
 

00:44:13.680 --> 00:44:18.800
BURGER: I wanted to get to the&nbsp;
efficiency point, but before I do,&nbsp;&nbsp;

00:44:18.800 --> 00:44:24.000
you know, you talk about this 1, you know, 1 to&nbsp;
2% of the neurons firing. But it's, actually,&nbsp;&nbsp;

00:44:24.000 --> 00:44:27.331
the brain is actually much sparser&nbsp;
than that at a fine grain, right?  
 
 

00:44:27.331 --> 00:44:27.881
AHMAD: Yes, yes.  
 
 

00:44:27.881 --> 00:44:30.320
BURGER: Because, you know, you&nbsp;
have 1% of the neurons firing,&nbsp;&nbsp;

00:44:30.320 --> 00:44:34.160
but they aren't connected to all&nbsp;
the other neurons in the region. 
 

00:44:34.160 --> 00:44:35.280
AHMAD: That's right. Yeah. 
 

00:44:35.280 --> 00:44:39.040
BURGER: So really the sparsity should&nbsp;
be the product of the connectivity&nbsp;&nbsp;

00:44:39.040 --> 00:44:41.840
fraction times the activity factor. 
 
 

00:44:41.840 --> 00:44:43.040
AHMAD: Yeah. Yeah. 
 

00:44:43.040 --> 00:44:46.560
BURGER: Right. That's about one out&nbsp;
of 10,000. Something like that. 
 

00:44:46.560 --> 00:44:52.160
AHMAD: Exactly. Yeah. So something like maybe 1%&nbsp;
of the neurons are firing at any point in time,&nbsp;&nbsp;

00:44:52.160 --> 00:44:56.320
and maybe 1% of the connections that are possible&nbsp;&nbsp;

00:44:56.320 --> 00:45:02.480
are actually there at any point in&nbsp;
time. So it's a very, very small,&nbsp;&nbsp;

00:45:02.480 --> 00:45:06.560
you know, subnetwork through this massive&nbsp;
network that's actually being activated,&nbsp;&nbsp;

00:45:06.560 --> 00:45:11.360
a tiny percentage of neurons going through a&nbsp;
very, very tiny piece of the full network. 
 

00:45:12.160 --> 00:45:14.080
You know, it's common to,&nbsp;
you know, some people say,&nbsp;&nbsp;

00:45:14.080 --> 00:45:19.440
“Oh, we're only using 1% of our brain.” That's&nbsp;
not true. It just means at any point in time,&nbsp;&nbsp;

00:45:19.440 --> 00:45:24.880
you're only using 1%, but at other points in&nbsp;
time, a different 1% is being used. So, you know,&nbsp;&nbsp;

00:45:24.880 --> 00:45:30.280
the activity does move around quite a bit. But,&nbsp;
any point in time, it's extremely small. 
 

00:45:30.280 --> 00:45:33.200
BURGER: So, OK, the sparsity, I think, you know,&nbsp;&nbsp;

00:45:33.200 --> 00:45:38.720
the representation—how the brain is doing this&nbsp;
compression biologically—is super fascinating.&nbsp;&nbsp;

00:45:38.720 --> 00:45:43.440
And I want to go on a little bit of a&nbsp;
detour now to efficiency. So I remember&nbsp;&nbsp;

00:45:43.440 --> 00:45:50.640
in 2017 when in MSR we were building, you&nbsp;
know, hardware acceleration for RNNs. 
 

00:45:50.640 --> 00:45:54.960
And then the transformer hit, and they were&nbsp;
optimized, you know, to be highly parallelizable&nbsp;&nbsp;

00:45:54.960 --> 00:46:02.640
across this quadratic attention map for GPUs. The&nbsp;
way I would describe it is that that transition&nbsp;&nbsp;

00:46:04.080 --> 00:46:08.880
to semi-supervised training moved us from&nbsp;
an era when we were really data limited,&nbsp;&nbsp;

00:46:08.880 --> 00:46:15.360
like you had to have good high-quality&nbsp;
labeled data, to you were compute limited.  
 
 

00:46:15.360 --> 00:46:19.920
And when that transition happened, we&nbsp;
hockey-sticked from, “I'm building faster&nbsp;&nbsp;

00:46:19.920 --> 00:46:26.400
machines but I'm limited by data” to the bigger&nbsp;
machine I can build, as long as I have enough,&nbsp;&nbsp;

00:46:26.400 --> 00:46:31.200
you know, unlabeled data of high quality, the&nbsp;
better I can do with the model. And so we went&nbsp;&nbsp;

00:46:31.200 --> 00:46:38.000
on the supercomputing arms race, and now we're&nbsp;
building these, like, just gargantuan machines. 
 

00:46:41.040 --> 00:46:46.320
And really, we've kind of been brute-forcing it.&nbsp;
I mean, we've done a lot of things to optimize,&nbsp;&nbsp;

00:46:46.320 --> 00:46:51.440
like quantization, you know, and other&nbsp;
and, you know, a better process node,&nbsp;&nbsp;

00:46:51.440 --> 00:46:55.120
you know, a better, more efficient&nbsp;
tensor unit design. But to first order,&nbsp;&nbsp;

00:46:55.120 --> 00:46:59.520
we've been training bigger models&nbsp;
by building bigger systems.  
 
 

00:46:59.520 --> 00:47:08.880
And I just wonder, do you think that the brain&nbsp;
at this 10 to 12 watts in the neocortex just&nbsp;&nbsp;

00:47:08.880 --> 00:47:14.400
has a fundamentally more efficient learning&nbsp;
mechanism? Or do we think that, you know,&nbsp;&nbsp;

00:47:14.400 --> 00:47:18.880
what we're doing in transformers in the&nbsp;
most advanced silicon is as efficient,&nbsp;&nbsp;

00:47:18.880 --> 00:47:21.440
we're just building much&nbsp;
larger, more capable models? 
 

00:47:21.440 --> 00:47:25.920
AHMAD: Oh, I think without a doubt,&nbsp;
transformers are extremely inefficient&nbsp;&nbsp;

00:47:25.920 --> 00:47:31.200
and very, very brute force. We touched on this&nbsp;
a little bit earlier in the attention mechanism,&nbsp;&nbsp;

00:47:31.200 --> 00:47:35.680
where we’re, you know, transformers&nbsp;
are essentially comparing every token&nbsp;&nbsp;

00:47:35.680 --> 00:47:39.840
to every other token. I mean, there are&nbsp;
architectures which reduce that, for sure,&nbsp;&nbsp;

00:47:39.840 --> 00:47:44.960
but it's essentially an n-squared operation.&nbsp;
And we're doing this at every layer. 
 

00:47:44.960 --> 00:47:50.240
I mean, there's nothing like that in&nbsp;
the brain. Our processing, you know,&nbsp;&nbsp;

00:47:50.240 --> 00:47:56.080
in some sense, the context for the very next&nbsp;
word I'm about to say is my entire life,&nbsp;&nbsp;

00:47:56.080 --> 00:48:01.200
right? And the amount of time I take&nbsp;
to take the next word doesn't depend&nbsp;&nbsp;

00:48:01.200 --> 00:48:06.400
on the length of the context at all. It's&nbsp;
a constant time dependence on context. 
 

00:48:06.400 --> 00:48:13.600
So it's a significant, you know, reduction&nbsp;
in the compute that's required. You can kind&nbsp;&nbsp;

00:48:13.600 --> 00:48:18.480
of think about, like the brain—I think&nbsp;
has somewhere around maybe 70 trillion&nbsp;&nbsp;

00:48:18.480 --> 00:48:23.360
synapses. When I say the brain, I mean the&nbsp;
neocortex, has about 70 trillion synapses.&nbsp;&nbsp;

00:48:23.360 --> 00:48:28.560
And it's using only 12 watts. And a synapse&nbsp;
is roughly equivalent to a parameter. 
 

00:48:28.560 --> 00:48:35.520
And if you were to take the most efficient GPUs&nbsp;
today and try to run a 70 trillion parameter&nbsp;&nbsp;

00:48:35.520 --> 00:48:41.840
model, it would be something like a megawatt&nbsp;
of power. It's tens of thous ... it's orders&nbsp;&nbsp;

00:48:41.840 --> 00:48:49.200
of magnitude more inefficient than what our&nbsp;
brain is doing. So I absolutely believe that. 
 

00:48:49.200 --> 00:48:51.840
BURGER: The metric I use, to go&nbsp;
back to your point, you know,&nbsp;&nbsp;

00:48:51.840 --> 00:48:56.080
is, this is something, I think we talked&nbsp;
about this back in the day, right? When,&nbsp;&nbsp;

00:48:57.520 --> 00:49:00.800
you know, after this kicked off for a few&nbsp;
years, we were trying to project, like,&nbsp;&nbsp;

00:49:00.800 --> 00:49:04.800
how far would this go under the current model&nbsp;
to inform the research and the directions you&nbsp;&nbsp;

00:49:04.800 --> 00:49:07.600
took. Which is why I got so interested&nbsp;
in sparsity and working with you.  
 
 

00:49:07.600 --> 00:49:11.280
And we would look at a training run and just&nbsp;
say, how many joules did it take to train the&nbsp;&nbsp;

00:49:11.280 --> 00:49:17.840
whole model? How many parameters do we have? And&nbsp;
sort of what's our parameters per joule? And,&nbsp;&nbsp;

00:49:17.840 --> 00:49:22.480
if by that metric, you know, we were off by&nbsp;
many orders of magnitude where the brain is,&nbsp;&nbsp;

00:49:22.480 --> 00:49:26.720
but I don't know that that's the right&nbsp;
metric. So any thoughts on that? 
 

00:49:26.720 --> 00:49:31.120
AHMAD: Yeah. I mean, in some&nbsp;
ways, you know, transformers,&nbsp;&nbsp;

00:49:31.120 --> 00:49:35.894
you know, embody more knowledge&nbsp;
in them than any human has.  
 
 

00:49:35.894 --> 00:49:36.425
BURGER: Right.  
 
 

00:49:36.425 --> 00:49:41.112
AHMAD: It has memorized, you know, the entire&nbsp;
internet's worth of knowledge, essentially. 
 

00:49:41.112 --> 00:49:42.873
BURGER: All scientific papers ... 
 

00:49:42.873 --> 00:49:45.920
AHMAD: All scientific papers.&nbsp;
You know, good and bad, whatever,&nbsp;&nbsp;

00:49:45.920 --> 00:49:48.800
you know, it has memorized everything.&nbsp;
So that's something that, you know,&nbsp;&nbsp;

00:49:48.800 --> 00:49:54.800
humans just cannot do. So there's definitely stuff&nbsp;
that's better in transformers than humans.  
 
 

00:49:54.800 --> 00:50:00.000
But fundamentally, I think, you know,&nbsp;
we're extremely efficient in how we process&nbsp;&nbsp;

00:50:00.000 --> 00:50:05.600
the next token or the next bit of information&nbsp;
that's coming in. And I think there's a lot&nbsp;&nbsp;

00:50:05.600 --> 00:50:10.880
we can learn from the brain and apply&nbsp;
to LLMs and future AI models there. 
 

00:50:10.880 --> 00:50:14.160
FUSI: I was going to ask a question related&nbsp;
to that because ... forget memorizing the&nbsp;&nbsp;

00:50:14.160 --> 00:50:19.040
internet. But let me give you another example that&nbsp;
transformers do really well. And I'm wondering,&nbsp;&nbsp;

00:50:19.040 --> 00:50:25.440
like, you know, the human aspect of this or the&nbsp;
brain aspect of this because transformers, because&nbsp;&nbsp;

00:50:25.440 --> 00:50:28.640
of the n-square computation, they're really&nbsp;
good at stuff, like a needle in the haystack. 
 

00:50:28.640 --> 00:50:32.960
So I can tell you right now, I can speak, I can&nbsp;
talk to you, and I can tell you the password is&nbsp;&nbsp;

00:50:32.960 --> 00:50:37.600
something silly like “podcast microphone&nbsp;
blue,” whatever. That's the password. And&nbsp;&nbsp;

00:50:37.600 --> 00:50:43.200
then I can proceed and read the entire Odyssey&nbsp;
or a bunch of other books to you out loud for&nbsp;&nbsp;

00:50:43.200 --> 00:50:48.480
the next 5 or 6 hours. And then I can ask&nbsp;
the transformer, what was the password? And&nbsp;&nbsp;

00:50:48.480 --> 00:50:54.480
transformer will do this nice n-square computation&nbsp;
many times, and it will spit out the password.  
 
 

00:50:54.480 --> 00:50:59.040
A human, you know, there will be a decay&nbsp;
of that password. And then at some point,&nbsp;&nbsp;

00:50:59.040 --> 00:51:03.520
it won't remember, and depending on the human,&nbsp;
it may be in the first chapter of the Odyssey or&nbsp;&nbsp;

00:51:03.520 --> 00:51:08.960
like at the end, but … so fundamentally the type&nbsp;
of computation that is done is very different. So&nbsp;&nbsp;

00:51:10.000 --> 00:51:12.640
it always makes me wonder about the&nbsp;
efficiency because it's just, like,&nbsp;&nbsp;

00:51:12.640 --> 00:51:17.760
it's a different type of computation. So the&nbsp;
efficiency of … like, efficiency is kind of like,&nbsp;&nbsp;

00:51:17.760 --> 00:51:24.000
what are you doing divided by how good are&nbsp;
you at doing it. And so when the things we're&nbsp;&nbsp;

00:51:24.000 --> 00:51:28.320
doing are so incomparable in many ways,&nbsp;
that always makes me ... always troubles&nbsp;&nbsp;

00:51:28.320 --> 00:51:34.073
me a little bit. I don't know... I don't know&nbsp;
if there's any question in there. [LAUGHTER] 
 

00:51:34.073 --> 00:51:37.680
AHMAD: Yeah. I mean, transformers can&nbsp;
do the stuff that humans find very,&nbsp;&nbsp;

00:51:37.680 --> 00:51:43.520
very difficult to do. Absolutely. You know,&nbsp;
maybe there's a way to get the best of both.&nbsp;&nbsp;

00:51:43.520 --> 00:51:49.520
I don't know. You know, I don't know&nbsp;
that it's fundamentally necessary to&nbsp;&nbsp;

00:51:49.520 --> 00:51:54.707
have such brute-force computation&nbsp;
to get all of these features. 
 

00:51:54.707 --> 00:51:55.402
FUSI: That's right. 
 

00:51:55.402 --> 00:51:57.680
BURGER: Yeah. Yeah, it is a&nbsp;
weird thing because, you know,&nbsp;&nbsp;

00:51:57.680 --> 00:52:01.200
this is why memory palaces work so&nbsp;
well. Like, there is a way, though,&nbsp;&nbsp;

00:52:01.200 --> 00:52:05.840
for a human to remember that my microphone&nbsp;
is gray. It's not actually blue, Nicolò. 
 

00:52:05.840 --> 00:52:09.032
FUSI: Mine is blue. You don't see it. It's&nbsp;
off camera. You see, your world model …  
 

00:52:09.032 --> 00:52:11.600
BURGER: It’s off camera. Yeah, I&nbsp;
know. I was just teasing you.  
 
 

00:52:11.600 --> 00:52:14.800
But there's a way, like, if I can&nbsp;
just connect it to enough things,&nbsp;&nbsp;

00:52:14.800 --> 00:52:19.120
get that connectivity graph, then I'll&nbsp;
remember it because it's captured the&nbsp;&nbsp;

00:52:19.120 --> 00:52:23.040
signal out of the noise and connected&nbsp;
to enough things I can retrieve it.&nbsp;&nbsp;

00:52:23.680 --> 00:52:27.360
And retrieval would be a whole other topic&nbsp;
we don't have time to get into today.  
 
 

00:52:27.360 --> 00:52:32.560
But I do … now, I want to go to the straw&nbsp;
man. So let's take continual learning off&nbsp;&nbsp;

00:52:32.560 --> 00:52:39.200
the table. Let's imagine that, as I go&nbsp;
through my day, I'm just saving all of&nbsp;&nbsp;

00:52:39.200 --> 00:52:46.640
the sensory data to put in my training set.&nbsp;
And now imagine that I take 100,000 little&nbsp;&nbsp;

00:52:46.640 --> 00:52:52.160
transformer blocks, and I'm training&nbsp;
them each with what they're seeing. 
 

00:52:52.960 --> 00:52:57.360
OK, I replay the day so I don't have to, again,&nbsp;
I don't have to worry about continuous learning&nbsp;&nbsp;

00:52:57.360 --> 00:53:03.600
and whatever cross-cortical column, you&nbsp;
know, routing feature of the outputs,&nbsp;&nbsp;

00:53:03.600 --> 00:53:07.200
the inputs, and there's—Subutai, we've&nbsp;
talked about this—there’s a complex set&nbsp;&nbsp;

00:53:07.200 --> 00:53:13.600
of wiring there to bring features from here to&nbsp;
there that gets learned. If I replicated that,&nbsp;&nbsp;

00:53:13.600 --> 00:53:18.480
could a transformer block kind of do&nbsp;
what the cortical columns are doing? 
 

00:53:19.760 --> 00:53:23.680
Could I just instrument all my sensory&nbsp;
patches with little transformer blocks&nbsp;&nbsp;

00:53:23.680 --> 00:53:27.520
and then wire them up in the&nbsp;
right way and have it work? 
 

00:53:27.520 --> 00:53:33.040
AHMAD: I think there'll be … there's still a&nbsp;
couple of things we need. One is that cortical&nbsp;&nbsp;

00:53:33.040 --> 00:53:38.880
columns are fundamentally sensory motor. And so&nbsp;
they're actually, each one, each cortical column&nbsp;&nbsp;

00:53:38.880 --> 00:53:46.640
is initiating actions, as well. So you cannot have&nbsp;
a static dataset fundamentally ahead of time. It's&nbsp;&nbsp;

00:53:46.640 --> 00:53:53.752
always a dynamic because we're constantly making&nbsp;
movements to get the next bit of data. And so … 
 

00:53:53.752 --> 00:53:55.840
BURGER: Couldn’t I tokenize that, though? 
 

00:53:56.880 --> 00:54:02.480
AHMAD: I mean, you could tokenize the input&nbsp;
and you can tokenize the output, but, you know,&nbsp;&nbsp;

00:54:02.480 --> 00:54:07.120
if you were to play the same set of inputs&nbsp;
back again to a network that … a cortical&nbsp;&nbsp;

00:54:07.120 --> 00:54:11.920
column that’s randomly wired differently, it&nbsp;
may make a different set of actions. And so&nbsp;&nbsp;

00:54:11.920 --> 00:54:16.880
as soon as it makes the first action that's&nbsp;
different, that dataset is no longer valid,&nbsp;&nbsp;

00:54:16.880 --> 00:54:23.360
right? It's, you know, there is ... you can't&nbsp;
fundamentally … you have to have a simulation&nbsp;&nbsp;

00:54:23.360 --> 00:54:28.960
of an environment rather than a static&nbsp;
one-way dataset, if that makes sense.  
 
 

00:54:30.000 --> 00:54:36.160
So I think that's one piece that I think’s&nbsp;
missing in transformers today, is this,&nbsp;&nbsp;

00:54:36.160 --> 00:54:42.376
sort of, sensory-motor loop. And then the other&nbsp;
piece we talked about is continuous learning. 
 
 

00:54:42.376 --> 00:54:42.901
BURGER: Yeah. 
 
 

00:54:42.901 --> 00:54:44.952
AHMAD: I guess you said take&nbsp;
it off the table, but … 
 
 

00:54:44.952 --> 00:54:46.137
BURGER: It's fundamental.  
 
 

00:54:46.137 --> 00:54:49.920
AHMAD: Fundamental … different. Yeah, yeah.&nbsp;
And maybe one other difference. We talked,&nbsp;&nbsp;

00:54:49.920 --> 00:54:53.840
you know, much earlier about a single latent&nbsp;
space and the prediction that's being made&nbsp;&nbsp;

00:54:53.840 --> 00:54:58.800
at the top of the transformer that you compute&nbsp;
the loss function, and that's back-propagated&nbsp;&nbsp;

00:54:58.800 --> 00:55:04.880
through the transformer. That's not how&nbsp;
neurons learn. Neurons are making … every&nbsp;&nbsp;

00:55:04.880 --> 00:55:09.520
neuron is actually making predictions,&nbsp;
and every neuron is getting its input. 
 

00:55:09.520 --> 00:55:13.680
And it's learning independent of&nbsp;
anything that happens at the top.&nbsp;&nbsp;

00:55:13.680 --> 00:55:18.640
And so it's a much more granular learning&nbsp;
signal. And information does flow from the&nbsp;&nbsp;

00:55:18.640 --> 00:55:22.160
top to bottom. But there's also many,&nbsp;
many other sources of information that&nbsp;&nbsp;

00:55:22.160 --> 00:55:28.400
it's learning from. So it's different in&nbsp;
that sense, as well, mechanistically.  
 

00:55:28.400 --> 00:55:33.520
BURGER: The reason I ask, and now I'd like&nbsp;
to get into, you know, some of the ... the&nbsp;&nbsp;

00:55:33.520 --> 00:55:38.320
fun speculation because I've just ... it's been a&nbsp;
phenomenal discussion with the two. I think we've&nbsp;&nbsp;

00:55:38.320 --> 00:55:42.720
kind of elucidated the differences. Something I've&nbsp;
wondered after I've talked to both of you … and,&nbsp;&nbsp;

00:55:42.720 --> 00:55:47.680
you know, Nicolò, kind of learning about&nbsp;
this compression view of the world, lossless&nbsp;&nbsp;

00:55:47.680 --> 00:55:51.600
compression, and, Subutai, just, you know,&nbsp;
the Thousand Brains Theory and these cortical&nbsp;&nbsp;

00:55:51.600 --> 00:55:59.520
columns and the sampling of, you know, the world&nbsp;
to capture the signal that you can learn from. 
 

00:55:59.520 --> 00:56:05.280
So let's say that I was able to design a&nbsp;
really small, efficient digital cortical&nbsp;&nbsp;

00:56:05.280 --> 00:56:10.880
column. Maybe it's transformer-based with&nbsp;
some, you know, a sparse representation&nbsp;&nbsp;

00:56:10.880 --> 00:56:15.520
and some sensory-motor mechanism built in.&nbsp;
Maybe it's more dendritic-based, you know,&nbsp;&nbsp;

00:56:15.520 --> 00:56:26.400
mapped into digital hardware. And I put a cortical&nbsp;
column on every sensor I have in the world,&nbsp;&nbsp;

00:56:27.280 --> 00:56:33.120
associated with every person, and wire them&nbsp;
up together with some of this and then have a,&nbsp;&nbsp;

00:56:33.120 --> 00:56:37.120
you know, billions of them that can&nbsp;
form higher-level abstractions. Like,&nbsp;&nbsp;

00:56:37.120 --> 00:56:40.360
what do you think would&nbsp;
happen? What could we do? 
 

00:56:40.360 --> 00:56:47.120
AHMAD: That's a fantastic thought exercise,&nbsp;
I think [LAUGHS]. You know, again,&nbsp;&nbsp;

00:56:47.120 --> 00:56:52.640
assuming the cortical column is faithful and can&nbsp;
generate, you know, or suggest motor actions,&nbsp;&nbsp;

00:56:52.640 --> 00:57:00.240
as well. I mean, in some sense, you could&nbsp;
potentially have a super intelligent system,&nbsp;&nbsp;

00:57:00.240 --> 00:57:04.960
right, that's far more intelligent&nbsp;
than anything else on the planet.  
 
 

00:57:04.960 --> 00:57:11.511
Now we're scaling the number of cortical columns,&nbsp;
you know, not from a mouse, you know, to a hundred&nbsp;&nbsp;

00:57:11.511 --> 00:57:17.760
thousand columns that a human might have, but&nbsp;
potentially billions of cortical columns and way&nbsp;&nbsp;

00:57:17.760 --> 00:57:24.960
more. And there's no reason to think there's any&nbsp;
fundamental limit there. So this sort of a system&nbsp;&nbsp;

00:57:24.960 --> 00:57:29.680
is, I think, the way that superintelligent&nbsp;
systems will eventually be built.  
 

00:57:29.680 --> 00:57:33.149
BURGER: But this is a very&nbsp;
different direction … 
 

00:57:33.149 --> 00:57:33.703
AHMAD: It’s a very different ... 
 

00:57:33.703 --> 00:57:35.520
BURGER: … than the one we're&nbsp;
currently headed down with,&nbsp;&nbsp;

00:57:35.520 --> 00:57:41.200
like, these monolithic models where we're&nbsp;
doing tons of RL, you know, to capture,&nbsp;&nbsp;

00:57:41.920 --> 00:57:46.320
you know, to get high-value human&nbsp;
collaboration in distribution. 
 

00:57:46.320 --> 00:57:51.600
AHMAD: Yes. It's completely different&nbsp;
than the direction we're proceeding.  
 
 

00:57:51.600 --> 00:57:57.120
So I think they, you know, to go down that path,&nbsp;
there needs to be a fundamental rethinking of&nbsp;&nbsp;

00:57:57.120 --> 00:58:02.400
some of our assumptions, potentially even&nbsp;
down to the hardware architectures that are&nbsp;&nbsp;

00:58:02.400 --> 00:58:05.920
necessary to implement it. The, you&nbsp;
know, fundamental learning algorithms,&nbsp;&nbsp;

00:58:05.920 --> 00:58:09.600
the fundamental training paradigm. We talked&nbsp;
about, you know, you can't have a static&nbsp;&nbsp;

00:58:09.600 --> 00:58:13.600
dataset. You're constantly moving around in&nbsp;
the world and doing things. So it's a very,&nbsp;&nbsp;

00:58:13.600 --> 00:58:18.160
very different way of going about&nbsp;
AI than what we're doing today. 
 

00:58:18.160 --> 00:58:21.760
BURGER: Sounds like a great&nbsp;
time to be an AI researcher. 
 

00:58:21.760 --> 00:58:24.872
AHMAD: Absolutely. [LAUGHTER] 
 

00:58:24.872 --> 00:58:29.200
BURGER: Nicolò, what was your&nbsp;
reaction to that hypothesis? 
 

00:58:29.920 --> 00:58:33.840
FUSI: It sounds super interesting. I&nbsp;
mean, my brain was churning. You know,&nbsp;&nbsp;

00:58:34.960 --> 00:58:41.600
my background is very different. And so, like, I'm&nbsp;
in a much worse position to answer this question.&nbsp;&nbsp;

00:58:41.600 --> 00:58:46.640
But I was starting to think, OK, so let's say&nbsp;
I do this. What would be my loss function?&nbsp;&nbsp;

00:58:47.840 --> 00:58:51.680
What, you know, how would information&nbsp;
flow through the system? Like,&nbsp;&nbsp;

00:58:51.680 --> 00:58:54.400
sounds like cortical columns would&nbsp;
each have their own loss that then&nbsp;&nbsp;

00:58:54.400 --> 00:58:59.280
I would aggregate—and then I would add a&nbsp;
contribution that is, like, higher level. 
 

00:58:59.280 --> 00:59:01.120
And then back to my question. You know,&nbsp;&nbsp;

00:59:01.120 --> 00:59:08.560
how is the temporal information coordinated?&nbsp;
Because one way to see this is that, you know,&nbsp;&nbsp;

00:59:08.560 --> 00:59:12.320
the way I'm coming to understand this is that&nbsp;
it's kind of like a multi-view framework. 
 

00:59:12.320 --> 00:59:18.640
You have the same phenomena represented to&nbsp;
multiple independent, but at the same time, views.&nbsp;&nbsp;

00:59:18.640 --> 00:59:24.240
And so part of me is like it feels like that&nbsp;
you need to tie together these cortical columns&nbsp;&nbsp;

00:59:24.240 --> 00:59:28.160
in such a way that they all get that gradient&nbsp;
feedback if you're training with gradient-based&nbsp;&nbsp;

00:59:28.160 --> 00:59:34.000
methods, for instance. And so that's, kind&nbsp;
of, it feels super, super interesting. 
 
 

00:59:34.000 --> 00:59:38.160
It is related to a lot of, you know, very&nbsp;
superficially, to a lot of ideas in machine&nbsp;&nbsp;

00:59:38.160 --> 00:59:42.400
learning around, hey, is it better to have&nbsp;
one giant super deep network? Is it better&nbsp;&nbsp;

00:59:42.400 --> 00:59:46.880
to have a bunch of shallow networks? But the&nbsp;
difference is also in the way you train them,&nbsp;&nbsp;

00:59:46.880 --> 00:59:50.960
right? We typically train this bunch&nbsp;
of shallow networks on kind of the&nbsp;&nbsp;

00:59:50.960 --> 00:59:54.320
same objective and the same&nbsp;
data and not typically into&nbsp;&nbsp;

00:59:54.320 --> 01:00:01.200
an experiential cycle. Whereas this sounds&nbsp;
like this is a different way to do it.  
 
 

01:00:01.200 --> 01:00:08.480
BURGER: Right, right. I think … I want to&nbsp;
pull this back around to the title of the&nbsp;&nbsp;

01:00:08.480 --> 01:00:16.400
podcast. And so I'll share an observation.&nbsp;
You know, so I've been using some of the&nbsp;&nbsp;

01:00:16.400 --> 01:00:20.960
latest models to code. You know, they're&nbsp;
getting better really fast. I've been using&nbsp;&nbsp;

01:00:20.960 --> 01:00:24.880
them to kind of relearn some of the physics&nbsp;
that I never really understood deeply. 
 

01:00:24.880 --> 01:00:30.480
You know, especially in general relativity,&nbsp;
like E=MC2. Like, why is C in there at all,&nbsp;&nbsp;

01:00:30.480 --> 01:00:34.160
right? Just stuff like that. Because&nbsp;
now it can actually explain it to me,&nbsp;&nbsp;

01:00:34.160 --> 01:00:38.560
and I can keep beating at it until I&nbsp;
understand it, and then, of course, work. 
 
 

01:00:38.560 --> 01:00:45.040
And at some point, I asked the model, “Can you&nbsp;
describe how I think?” And I was just curious.&nbsp;&nbsp;

01:00:45.040 --> 01:00:51.840
And it, you know, it gave me a page description&nbsp;
that my jaw dropped because I said this, this&nbsp;&nbsp;

01:00:52.720 --> 01:00:58.240
thing knows me better than I know myself.&nbsp;
I don't think any human being, including me,&nbsp;&nbsp;

01:00:58.240 --> 01:01:03.040
could have captured kind of the way my approach&nbsp;
to learning and my brain works, and I just&nbsp;&nbsp;

01:01:03.040 --> 01:01:06.880
read it as, like, like, yep, that's right. And&nbsp;
I learned something about myself.  
 
 

01:01:08.240 --> 01:01:11.920
So I wouldn't say that it passed the Turing&nbsp;
test because this is way beyond Turing test.&nbsp;&nbsp;

01:01:11.920 --> 01:01:15.520
This was like, this thing knows me way&nbsp;
better, you know, than I thought any&nbsp;&nbsp;

01:01:15.520 --> 01:01:19.600
machine ever could. I mean, I'm having a&nbsp;
conversation with it. It could be human,&nbsp;&nbsp;

01:01:19.600 --> 01:01:30.880
but it's superhuman. So in some sense, it's&nbsp;
like intelligent beyond human capabilities&nbsp;&nbsp;

01:01:30.880 --> 01:01:36.720
with its ability to discern patterns in&nbsp;
how someone's interacting.  And yet it's a&nbsp;&nbsp;

01:01:36.720 --> 01:01:41.840
tool. You know, it's not conscious.&nbsp;
It doesn't have agency, embodiment,&nbsp;&nbsp;

01:01:41.840 --> 01:01:47.120
emotion. It understands a lot of that stuff&nbsp;
from the training data. But at the end of the&nbsp;&nbsp;

01:01:47.120 --> 01:01:51.600
day, it's a stochastic parrot, right? It's got,&nbsp;
you know, it's got the weights, and I give it&nbsp;&nbsp;

01:01:51.600 --> 01:01:55.440
a token, and it outputs a token. So, like,&nbsp;
are these machines intelligent or not?  
 
 

01:01:55.440 --> 01:02:01.513
FUSI: I’ll let Subutai answer first. [LAUGHS] 
 

01:02:01.513 --> 01:02:10.640
AHMAD: OK. You know, you know, it's definitely&nbsp;
a savant, right? It knows a huge amount about&nbsp;&nbsp;

01:02:10.640 --> 01:02:15.200
the world. It's absorbed a lot of stuff,&nbsp;
and it can articulate that in ways that&nbsp;&nbsp;

01:02:15.200 --> 01:02:21.360
are just amazing. And, you know, it's&nbsp;
taken your chat history with, you know,&nbsp;&nbsp;

01:02:21.360 --> 01:02:26.480
presumably thousands of chats and able to&nbsp;
summarize that in a way that's remarkable. 
 

01:02:27.200 --> 01:02:33.440
At the same time, I think, you know,&nbsp;
transformers are not intelligent in&nbsp;&nbsp;

01:02:33.440 --> 01:02:39.120
the way that a three-year-old is, right?&nbsp;
A three-year-old human is very curious, is&nbsp;&nbsp;

01:02:39.120 --> 01:02:46.000
constantly learning. It can learn almost anything.&nbsp;
And, you know, a three-year-old Einstein was able&nbsp;&nbsp;

01:02:46.000 --> 01:02:51.280
to learn and eventually come up with theories&nbsp;
that shook the world. That, you know, E=MC2. 
 

01:02:53.520 --> 01:02:58.240
And so, you know, could a transformer do&nbsp;
that? I don't think so. And so I think&nbsp;&nbsp;

01:02:58.240 --> 01:03:05.280
there's still a difference. There's things&nbsp;
it can do that are amazing. But there are&nbsp;&nbsp;

01:03:05.280 --> 01:03:10.560
still basic things that a child can do that&nbsp;
transformers cannot do. So I think there's&nbsp;&nbsp;

01:03:10.560 --> 01:03:17.040
still a gap there. Exactly how to articulate&nbsp;
it, and how to bridge that gap, is, of course,&nbsp;&nbsp;

01:03:17.040 --> 01:03:21.160
the trillion-dollar question. But it is&nbsp;
bridgeable. And there is a gap today. 
 

01:03:21.160 --> 01:03:24.080
BURGER: Right. Nicolò? 
 

01:03:24.080 --> 01:03:32.400
FUSI: You know, I think, from my perspective,&nbsp;
they are intelligent. And from my perspective,&nbsp;&nbsp;

01:03:32.400 --> 01:03:34.720
I go back to the definition&nbsp;
of intelligent, which is like,&nbsp;&nbsp;

01:03:34.720 --> 01:03:40.960
can you achieve your objectives in a variety&nbsp;
of environments? It's a very basic fundamental,&nbsp;&nbsp;

01:03:40.960 --> 01:03:45.200
but it's kind of, you know, it can be&nbsp;
embodied, a form of embodied intelligence,&nbsp;&nbsp;

01:03:45.200 --> 01:03:50.080
an agentic intelligence. If I plop you in&nbsp;
an environment, and I give you an objective,&nbsp;&nbsp;

01:03:50.080 --> 01:03:57.520
can you achieve it? And the wilder the&nbsp;
environment, the harder the task is.  
 
 

01:03:57.520 --> 01:03:59.680
And I do think … I agree with Subutai. Like,&nbsp;&nbsp;

01:03:59.680 --> 01:04:02.057
there is a jaggedness of&nbsp;
intelligence we keep describing.  
 
 

01:04:02.057 --> 01:04:02.578
BURGER: Yup. 
 

01:04:02.578 --> 01:04:07.680
FUSI: Like these things cannot be simultaneously&nbsp;
super good, you know, Olympiad-level&nbsp;&nbsp;

01:04:07.680 --> 01:04:14.240
mathematicians and still give you stupid answers&nbsp;
when you're trying to, I don't know, you know,&nbsp;&nbsp;

01:04:14.240 --> 01:04:19.983
figure out which cable goes where in your … in&nbsp;
your car's battery, you know, like, whatever. 
 

01:04:19.983 --> 01:04:22.720
BURGER: [LAUGHS] Well, then it's better than&nbsp;
me. I'm not an Olympiad-level mathematician,&nbsp;&nbsp;

01:04:22.720 --> 01:04:24.280
and I do stupid stuff all the time. 
 

01:04:24.280 --> 01:04:28.320
FUSI: I know exactly. Well, you know,&nbsp;
whatever that was, that was a bad example.&nbsp;&nbsp;

01:04:28.320 --> 01:04:34.320
But you get it. But part of it goes back to&nbsp;
the compression view. Like, I do believe that&nbsp;&nbsp;

01:04:35.360 --> 01:04:39.600
intelligence is compression. So the ability to&nbsp;
come up with succinct explanations for complex&nbsp;&nbsp;

01:04:39.600 --> 01:04:46.480
phenomena and even succinct explanations for&nbsp;
complex worlds, and then it implies or leads&nbsp;&nbsp;

01:04:46.480 --> 01:04:51.360
to your ability to operate within them, and the&nbsp;
fact that we have these things that they can&nbsp;&nbsp;

01:04:51.360 --> 01:04:59.040
prove crazy theorems but at the same time fail&nbsp;
at fairly rudimentary tasks is a sign that the,&nbsp;&nbsp;

01:04:59.040 --> 01:05:03.360
yes, transformers are great in terms of&nbsp;
inductive biases they put on the world and&nbsp;&nbsp;

01:05:03.360 --> 01:05:09.600
computation that are great, but we're ultimately&nbsp;
all subject to the No Free Lunch Theorem. 
 

01:05:09.600 --> 01:05:14.160
You know, across the world, the set&nbsp;
of tasks that you could be pursuing.&nbsp;&nbsp;

01:05:14.880 --> 01:05:20.320
You know, you have certain inductive biases that&nbsp;
kind of privilege certain tasks at the expense&nbsp;&nbsp;

01:05:20.320 --> 01:05:25.120
of others. And there isn't, like, a thing&nbsp;
yet that has expanded our set of tasks that&nbsp;&nbsp;

01:05:25.120 --> 01:05:31.920
are addressable. And so I do think that it's a&nbsp;
matter of rethinking our approach to a few things,&nbsp;&nbsp;

01:05:31.920 --> 01:05:39.040
whether I think likely both on the architecture&nbsp;
front and on the losses and the way we train these&nbsp;&nbsp;

01:05:39.040 --> 01:05:45.200
systems front. I think there is an opportunity&nbsp;
to expand the intelligent frontier of these&nbsp;&nbsp;

01:05:45.200 --> 01:05:51.680
models. But yeah, from my perspective, they are&nbsp;
intelligent already just in a jagged way. 
 

01:05:51.680 --> 01:05:55.120
BURGER: It's such an interesting question, and&nbsp;
I know a lot of people write a lot about this,&nbsp;&nbsp;

01:05:55.120 --> 01:05:59.120
so I don't think treading any&nbsp;
new ground here. But, you know,&nbsp;&nbsp;

01:05:59.120 --> 01:06:06.880
there's the diversity of the tasks you can excel&nbsp;
at. You know, are you able to handle nuance and&nbsp;&nbsp;

01:06:06.880 --> 01:06:12.560
understand things deeply? Are you able to learn&nbsp;
continuously? Right now, the systems can't,&nbsp;&nbsp;

01:06:12.560 --> 01:06:17.200
right. Are you embodied? I don't know if&nbsp;
that matters. Do you have an objective? Well,&nbsp;&nbsp;

01:06:17.200 --> 01:06:21.600
we could give them one. Are you conscious? Is&nbsp;
that … I mean, that's a whole other thing.  
 
 

01:06:21.600 --> 01:06:30.160
So it just feels like there's a bunch of check&nbsp;
boxes, and we've checked a bunch of them, and a&nbsp;&nbsp;

01:06:30.160 --> 01:06:36.480
bunch of them are unchecked. And maybe there's&nbsp;
no consensus on, like, where that threshold is&nbsp;&nbsp;

01:06:36.480 --> 01:06:41.200
because there are many dimensions of intelligence,&nbsp;
and some of which humans don't even have. 
 

01:06:41.200 --> 01:06:46.080
FUSI: And that's why we have the term AGI and ASI,&nbsp;&nbsp;

01:06:46.080 --> 01:06:51.360
and people are debating the G and the S—what&nbsp;
is general, what is specialized. So there is,&nbsp;&nbsp;

01:06:51.360 --> 01:06:56.800
like, it's a huge discourse, like, for sure.&nbsp;
But that's why we had to start characterizing.&nbsp;&nbsp;

01:06:56.800 --> 01:07:00.720
But if you go back in the definition,&nbsp;
going back to my schooling, go back to&nbsp;&nbsp;

01:07:00.720 --> 01:07:05.840
the definition of intelligence from Plato and&nbsp;
Aristotle and Descartes, like, in some sense,&nbsp;&nbsp;

01:07:05.840 --> 01:07:10.615
you see the goalpost moving through the centuries&nbsp;
around what we define as intelligent.  
 
 

01:07:10.615 --> 01:07:12.320
BURGER: Right.
FUSI: And I feel like we are still doing it.

01:07:12.320 --> 01:07:14.880
BURGER: Yeah. We’ll be doing&nbsp;
it for a long time, you know,&nbsp;&nbsp;

01:07:14.880 --> 01:07:18.880
which in AI velocity is probably&nbsp;
another like four or five years.  
 
 

01:07:18.880 --> 01:07:23.600
Hey, I just want to thank you&nbsp;
both for the dialogue. You know,&nbsp;&nbsp;

01:07:23.600 --> 01:07:30.080
I treasure both of you as, you know,&nbsp;
intellects and scholars and friends.&nbsp;&nbsp;

01:07:31.040 --> 01:07:37.480
It was just a joy to nerd out with you all.&nbsp;
So thank you both for taking the time. 
 

01:07:37.480 --> 01:07:40.000
AHMAD: Thank you so much, Doug, for having me.  
 

01:07:40.000 --> 01:07:45.811
FUSI: Thank you for having us. This was great. 
 
 

01:07:45.811 --> 01:07:45.824
[MUSIC] 
 

01:07:45.824 --> 01:07:46.355
STANDARD OUTRO: You’ve been listening&nbsp;
to The Shape of Things to Come,&nbsp;&nbsp;

01:07:48.357 --> 01:07:51.440
a Microsoft Research Podcast.&nbsp;
Check out more episodes of the&nbsp;&nbsp;

01:07:51.440 --> 01:08:05.840
podcast at aka.ms/researchpodcast or on&nbsp;
YouTube and major podcast platforms. 
 
 

01:08:05.840 --> 01:08:06.623
[MUSIC FADES]

