0:07
Okay, welcome to Alternative Compute Club.
Okay, welcome to Alternative Compute Club.
So, this night started with a dinner with a close friend of mine who um posed this prompt to me.
I thought it was a really interesting prompt that you guys can take with you to your dinner parties.
Um it was like, if you could only ask the aliens one question, these super intelligent aliens, one question, what would you ask?
And so we're thinking about, you know, maybe how do you solve power?
And I was kind of thinking, I think we kind of know Dyson spheres or I nuclear or something like that.
Um but I thought it was more interesting about how do you how do they compute their flops, their floating point operations.
And so, um, I think that we're so far down this rabbit hole of hardware, software co-adaptation, or more specifically, hardware architecture, and optimizer co-adaptation.
And I'm going to give I've been doing deep learning since roughly 2012.
I'll give a little bit of like my understanding of what happened in the last, you know, 14 years.
Um, and so roughly, let's say 2012, uh, Alex Net comes out.
Uh, CNN's probably had their era up until maybe 2019, 2020.
And at that time, I don't think per personally, you know, we Focal Systems, my company, went through this um, Nvidia, you know, um, kind of thing.
They gave us $125,000 for winning this like, you know, whatever competition.
We interfaced somewhat with Jensen.
My understanding is that he really didn't care about AI at all until 2020.
Um before then it was there was a really hot topic uh in Silicon Valley called crypto and that made a lot more money until SPF.
Um and you know and so uh I think he was most mostly focused on flops on getting more compute per jewel more flop flops per jewel.
um and specifically for CNN's since they were not memory bound and they're not uh bandwidth bound um and also for um uh not Bitcoin as Alok uh corrected me on but Ethereum. Okay.
So, and then after 2020, um, they started to figure it out and they, I think it was the GPT3 stuff, maybe GPT2 really put it on his radar and said, "These transformers demand more memory capacity and bandwidth than ever before."
And then they launched uh, Ampear was really the first one of of of this of this um, ilk.
And so what happened is they started to care a lot less about compute efficiency and a lot more about memory capacity and memory bandwidth.
and memory capacity and bandwidth are in this exponential. I wonder why.
Well, attention is n squ. So that's probably why.
And compute efficiency, well, as long as it produces the flops, it's going to, you know, uh, burn up all the oceans in the world, but that's okay.
Um, and so this is roughly what's been going on.
And so this is hardware architecture optimizer, code adaptation, um, uh, happening in real time.
So, and then even worse, this is like looking at uh from a G-flop G-lop G-flop perjule over time, we've kind of petered out.
In the last two years, there really hasn't been that much um improvement on this.
And so, we need something else.
And I spent my entire PhD thinking about one really, how does the brain work as as as number one, but two, what else? What are there?
There's no way that you know I I was around and doing deep learning before people knew about back prop.
Um, I think that in 10 years we'll stop using Backdrop.
And so I wanted to start working on something else.
And we'll get into my stuff in a little bit, but uh the the the premise is that the cost per gigaflop has to come down by significant amounts for to be human level.
The human brain draws like 20 watts.
Um, that's like a light bulb.
Um, and it can do insane amounts of of things that the presenters tonight are going to talk about.
Um, and there's some hope.
There are people now that we have this bifurcation of chips that are training chips and that are inference chips, then uh people are focused on different things and not trying to make everyone happy, but they're largely still focused on the transformer architecture.
So, they're going to have lots of um SRAMM uh maybe not HBM, but definitely SRAMs will have a lot of memory, which you know maybe isn't required.
Um, and so the brain is massly feed forward.
And so, uh, any argument that I've tried to that I've listened to that the brain does back propagation where the would require that these neurons that the this neuron here fires forward and then it would have to fire backwards.
It would have to turn off the brain and then fire backwards so that we can update some the syninnapse here.
Um, and it certainly doesn't do that.
Uh, it certainly does freed forward.
And to make Hebian learning work or any of these models, um there's something called the weight transfer problem where it does a forward pass and then in in uh one uh synaptic pair and then has to come back through a different one and you have the weight transfer problem where then weights have to be exactly equal, which is how is that going to be possible?
Um it turns out that if you put an activation thing in there, um and they both see all the same signals, then it kind of works, but it's still not back prop.
that's just a cyclic graph.
Um, which you know doesn't need doesn't require retention of activations and all those other things like that.
Um, there's some chemistry happening.
So there's some residual after the the activation and things and timings and things like that.
Um, the other cool thing that I'll reference later in that happens in the brain that people don't know about is that these things it's not just big cobweb of connections.
There's massive inhibition in the brain.
And so the these are these are cortical columns.
There's bunch many of them in the brain and they're basically these assemblies of that are repeated all throughout the brain with massive inhibitions all around to allow each um uh assembly to be uh to learn independently.
Um and so there's some hope here.
There's there's a lot of ideas to one use the brain that four and a half billion years of evolution already figured out to do with Tony Watts.
Um we have Sean who's who's going to be talking about that teaching a brain to to play Doom which is super super cool.
Um we have uh Ilker to talk about this opt optical computing.
I think if I make an argument why optics is probably the right answer.
Um we communicate to the data center in light then we hit a switch and then we switch it back to electrons to tell it where to go.
we switch it back to light and then we you know back to to the the DAG DAC is that the DAC uh digital analog converter and then it goes back to light uh and then uh we switch back with a a ADC and then we do burn 99.
99% of the of the jewels and then we switch it back to light.
Why don't we just do the whole thing in light and it would be instantaneous and it would draw no watts.
And so that's what I really think the aliens are probably doing.
But I'm open to a lot of different ideas and so neural computing is obviously really really exciting and cool as well.
Um and then neuromorphics we have a LO uh presenting on that as well which you know is is basically trying to mimic um what the brain is doing uh in analog uh circuitry um which is super cool as well.
And then we still aren't done yet because if we're trying to do back prop in any one of those substrates it wouldn't work because these things storing bits in light is actually very expensive. It's very hard.
Um and you know we'll have Ilker talk about that.
Um so the properties of this are really really interesting.
The cost per gigaflop is way way low to do computation in light is like free. You can do lots of it.
You just can't it's hard to like store stuff.
You have all these other non-ifferiability uh uh um activation issues and things like that.
Um but so the cost of VRAM of of VRAM of per per gig of VRAM is very expensive and bandwidth with light is pretty much instantaneous with neural compute.
No neuromorphics just copper wire probably could could do it as well.
These will have different properties.
So there's I spent my entire PhD so far working on alternative optimizers.
Um again I was around before people knew back prop.
back prop. there's a bunch of other learning procedures and I was curious around what are the other how far could we take um non-gradient-based methods and so I spent I grabbed a zero order to textbook and I literally implemented every single method in the my first six months of my PhD and this the setting was train a 1 billion
parameter LSTM to do next token prediction that's it all I want to do what will get the best results and the learning procedure the the update rule that worked the best is something called SPSA that pretty much no one knows about except James Ball and his students that invented this um learning update rule. And basically if you remember calc 1 you
And basically if you remember calc 1 you have two forward passes of some theta uh some some theta model that you're going to perturb in the random direction with some radius epsilon uh plus and then minus and you have a finite difference and you scale you update by that random direction by the finite difference and you do that number of pertabbation times.
If the finite difference is zero then you don't uh uh scale that direction at all.
If there's a huge finite difference, then you scale in that direction a lot.
And that's basically what happens here.
So it's called finite difference method, SPSA, central difference method.
These are all kind of the same names for the same thing.
And what's cool about this, the other property is there's no gradient.
So I can optimize 01 loss landscapes completely fine.
If this is a famous one called a function and you can look at back prop where it gets stuck in a local minima pretty quickly.
Um, and uh, zero order finds the the bottom of the of this very non-lipshits bowl very quickly.
Um, so there's some cool properties there.
So we have the optimizer, we have a good optimizer. Well, what do we try? What do we train? What's the architecture?
And then on what compute are we going to train it?
The architecture I've been playing around with a bunch. Why LSTMs?
LSTMs were designed to be to allow for gradients to flow through it.
So we don't care about gradients anymore.
We're just doing forward passes.
We don't there's no gradient flowing anywhere.
anywhere. Um and so maybe not uh LSCMs and transformers I've tried it you know transformer was really built for back propagation is you basically are have this parallelization in and during during training which is so beautiful um it was trained with back prop in mind it
was trained with GPUs in mind in fact um I may go farther to say that so um there might be and so I've been working on this is what I'll be submitting to ilear on Friday along with 60,000 other people Um, don't worry the reviewers they they always separate the wheat from the chaff. There's like, you know, perfect
There's like, you know, perfect F1 score. Um, so I trust that.
Anyway, um, this idea I call soma sharted optimization mixture of assemblies.
So with this cortical column idea in mind, what if we just take our common crawl data set, uh cluster it with TF and some SVD thing, cluster the data set and then ship a parquet file to a GP to n GPUs all throughout the world with n clusters.
So one uh parquet file per GPU and just train a little expert, a little LSCM, small one like 41K.
LSCM, small one like 41K. And then at test time you pull them all together and then you can route with that little router and you can it's like the same benefit of except it's mixture of models not not um uh just one layer feed feed
forward uh uh feed forward uh layers and so this is basically the the the premise how far can this go and the beauty in this is one of the things I didn't say here is this solver this optimizer the error in the gradient estimate uh becomes gets uh scales worse. It gets
It gets worse and worse, more less and less accurate as the model size grows and bigger and bigger and bigger um linearly so.
And so you can't train big models with this.
If you do a 10 billion parameter model, the number of pertabbations you'll have to do is extremely large, extremely infeasible, certainly more than a forward pass and a backward pass um in terms of flops.
So it gets really really uh compute inefficient if you shard it.
Then that gradient noise is capped to some amount by the size of each expert.
And so you only need maybe 64 pertabbations uh per step or maybe even less um and it works completely fine.
And the cool thing is we have these scaling laws that happen in all three dimensions.
We knew about pertabbations.
We knew about increasing batch size to reduce the grain estimate uh noise.
We didn't know that if you shard the model more and more, it continually gets better and better and better.
And so why does this lose the back prop?
Because the cost of forward pass is still really, really expensive. Well, what if it wasn't?
If we're doing this in light, forward passes are really, really cheap.
It's just beams of light.
It's just doing computation. It's like free.
Um, if we're doing this in neurons, it's 20 watts.
If we're doing this in in analog, it's whatever amps you put it through, I guess.
Um, and so I'm really excited about hosting this one uh because it will make my PhD worth something.
Um, so if none of these uh people are successful, then my PhD will be meaningless.
And we should I should have done back prop the whole time.
So I'm hoping on you to to solve it.
So uh I have three we have three great speakers.
Let me just introduce them a little bit.
Uh Ilker is a a PhD out of EPFL in Switzerland.
uh has worked on optical uh AI computing for how many years? Five, six years.
Um so probably one of the the forefront uh uh thinkers on this this concept um and is now doing his postto for the last year uh at Stanford uh where we got to meet.
Um Aloc is a PhD at a double E worked on semiconductors and nano optics um and has been a wildly successful uh venture capitalist and now runs a $ 1. 5 billion fund. no big deal.
Um, and uh, really focuses on the frontier and early stage.
Um, and just like hang out with them.
Uh, he's, you know, brilliant and he can gro any concept that we throw at him.
So, he offered to do neuromorphics.
Um, and then Sean just came through the bash NYC.
He got his masters in neuromorphics and now the CEO of a cool company called Parasma, which I asked him like, "What does that mean?
Is there some like deep meaning of like what uh, you know, maybe some neuroscience term that I didn't know?"
He's like, "No, it's just an item in Dota."
And that's kind of funny.
All right, Elker, you're up. Thank you so much.
>> So, as France um introduced very well, um we today do most of our data communication in optics and we know it's very efficient for that purpose.
But today I want to float the idea of how we can use light for computing. What are the advantages?
What are the challenges out there?
And as a use case um I'll give our study as an example in which we show that we can program the light propagation for doing diffusion based image generation.
So as we all know photons are great for carrying information and that's why we have a field called photonix or optics and communication is one of the biggest applications of it.
Nearly all of the transmission between um AI data center racks are now done with optical fibers.
So information is transmitted as photons and 99% of the intercontinental data transfer is also optical. Why?
And the answer is very simple and comes from the fundamental physics.
Photons has compared to uh electrons has about 10,000 times lower loss and 10,000 times larger bandwidth uh compared to electrons.
And that's why when we want to transmit some information today nearly more than one meter, we always convert it to photons and then send it over optical fibers. generally.
So if it provides such a low loss and high bandwidth, why don't we process information with photons as well?
And all uh what we have seen is that actually the technologies out there right now is at a convergence point because historically electronic processors were using very high precision 32bits 64 bits representations for doing information processing and then these symbols would be converted to lower bit depth uh two levels or four levels and then transmitted in optics.
But what we see recently is that with the um emergence of AI is that uh we can actually do our tasks with lower and lower bit precision uh to be able to do much more operations go faster and at the same time optical communications are able to support larger and larger bit depths so that we can see that there is an interface that is appearing between these two domains.
Another fundamental advantage of optics uh that exists for the computational purposes is that by definition it is massively paralyzable because photons don't interact with other photons.
So in at a given space um as many as optical beams can coexist without affecting each other.
On the other hand in electrical circuits we have to define different paths for each pet of signal.
So a quick comparison in terms of the principles of compute in electronics.
Generally what we do is that for for instance for matrix vector multiplier operations uh we charge wires uh we effect those wires with transistors and then discharge them and at each step of these computation we spend energy.
these computation we spend energy. So for n dimensional inputs and n dimensional outputs our energy consumption generally correlates with n squ on the optical end what we can do is and what's a general practice in the emergent studies is that we can modulate directly our input information on n many
modulators and let it propagate through a designed weight mask so that the weights of this operation can modulate the light beam passively without spending any extra active energy and then at the end the result of this computation can be summed on the detectors again through physics and the sum values can be read out directly on these detectors. So again the premise
So again the premise there is the number of input units or output units which is called n correlates with the energy consumption not n squ.
But if this is so great then why don't we have optical computing everywhere?
Why isn't widespread and already there are many companies out there trying to make it a reality and there are very competitive results coming out of them as well.
An example is light matter in which they show them electrical opto electronical PCIe board where they have the compute numbers comparable to GPUs and in terms of power efficiency as well.
And if you take a closer look into their energy budget, what we see is that actually what we spend on photonics is a small portion of all the uh complete budget of energy. Why?
And the answer to this will actually tell us what are the challenges out there for optical computing.
So uh one issue is that today our data or models are in digital electronic memories and if you want to transfer them to our optical um processing units, one of the challenges is that we pay this large price for converting digital information to analog and then optical domain.
And then once we do our computation, we have to go back to that those memories or transmit the information which requires us to convert the light to analog and then digital domain again.
and we pay a big price for it and we see that in the uh budget breakdown as well.
Another um expenditure item is that to be able to program these devices and keep them calibrated and precise, we also have to actively spend energy to uh put the parameters or control there.
And another bigger challenge especially for AI workloads is that the passive optics support linear operations in massive scales.
But when we want to do the nonlinear activation functions, they require additional mechanisms.
They are not as easily accessible as in the case of transistors or digital electronics.
One of our tries was that we used short pulses of very intense light in multim mode fibers and the light matter interactions there created optical nonlinearities.
But as you can see it adds another layer of complexity to have these nonlinear activations.
these nonlinear activations. So here the take-h home message that I would get from the first part is that for the first demonstrations of useful optical compute we have to carefully choose the use case and the algorithm that we would like to run on the optical machines so that we can create a tangible advantage
compared to electronics and as a try on this I'll talk about our ner paper um in my previous laboratory in EPFL and it was a collaboration with Google in which we try to take the task of image generation in the diffusion based models and we will want to do it by programming optical propagation as a den noising or generation uh inference unit. But first a very brief introduction to
But first a very brief introduction to diffusion based generative models just to set the scene.
They work by mapping after learning a uh train neural network they work by mapping a random distribution of data.
It could be pixels, it could be action items, it could be videos.
And these models take steps on these random distributions with this train neural networks to gradually transform the distribution to a desired data distribution.
This could be images, this could be other types of data sets.
And for the image case, what you can see is that we get a random set of pixels.
And by repeatedly applying this neural network on this random set of pixels, we can create clean and realistically looking images.
Our proposition here is that the using a general purpose electronic hardware for this task creates a big redundancy because we are using a single train neural network on a general purpose hardware again and again.
Sometimes it can go up to thousand times for creating a single sampling of a single image.
And what we propose here is that let's make an application specific and passive device so that the physics of the propagation of light inside this device takes a step on the diffusion procedure.
And by re uh iterating this architecture or reiterating on this optical system, we can start from a random distribution of pixels and by just letting the physic do its work, we can create uh new samples from the desired data set. How does it work?
Just giving a a bit of a closer look into the system.
It just depends on the definition and the main components that are required are that we know the how the physics work and we have a way of manipulation manipulating it and unfortunately still we are using back propagation to do that but maybe it will change.
But then the goal is that we define a target function as in the neural networks case and we let the physical system to be designed so that they perform the task at hand. Here it is.
We provide noisy samples to this physical system and we would like to have the predicted noise term in those images so that as the light propagates through the system, what we end up is a filtered version of the image and by using it over and over again, we can subtract the noise term and at the end create a clean image with the system.
And what you see as the layers are basically transparencies that modulate the phase of the light.
So that as they prop as the light beam propagates through this um set of layers uh they defract and only the information that we would like to end up with is transmitted.
Uh one key property of optics that I mentioned was that it's hard to program it or hard to change the weights and we designed our generation algorithm around it.
Instead of adding uh time dependent activation manipulations like the digital neural networks do in this case we created different subset of um schedule sub uh units let's say and if you are taking 1,000 steps over time we divided these 10,000 steps to 100 steps each uh subunits and for each of these 100 subunits we kept the parameters fixed.
So this means that if you have 10 of these devices, you could basically use one for 100 times, pass the activation to the other one, use the other one for 100 times and overall without changing any weights on the uh physical hardware, we could do all this uh procedure of different noises and uh denoising of over the process.
For now the because of the limited number of parameters we show small data sets like Mnest digits or fasion emnest.
But what we can clearly see is that this model or this optical device behaves as an diffusion model and by starting from a random distribution it can gradually generate images uh that are similar to the original data set and we can see that with uh image generation quality metrics.
So we have both the diversity and quality and that is shown by the FID going lower.
But maybe one more interesting property is that uh when we looked into the behavior of the system in terms of the number of parameters when we scale the number of defractive layers on this system or maybe change the number of pixels at each layer.
What we have seen is that it shows a similar power low scaling uh behavior when compared to the digital neural network.
So this means that we expect for these models to scale in a similar way as we add more parameters to it and as you remember the promise of optical computing is that as we add more par parameters actually the energy consumption scales more preferably and we could also show that by adding these extra parameters we can in a uh expectable way the performance improves in terms of the image quality.
So what is the bottom line?
is the bottom line? The bottom line is that when we take took this small model and compared with similarly performing models on the digital neural network on a commercial GPU what we have seen is that for generating a single image we had a significant energy performance
energy advantage when we used our optical system here there's an asterics we assumed our weights passive meaning that u just micr fabricating these weight layers and putting them there so that once they are trained they are fixed and they're eternally the same which is uh currently easily can be done. In our experimental implementation
In our experimental implementation uh we took an easier prototyping approach.
We used a device called special light modulator.
This is generally used also in uh light projectors and by coupling it with a mirror we could actually show that we can do a deep network on a single device and um looking at our camera images and input laser distributions we can do back propagation on spot and update our weights so that we don't need to go through different phases of micr fabrication and trying again.
So it gave us a useful means for prototyping this system.
But of course in the future for discussing a useful optical compute unit, we would want to produce these layers and put it in place and literature already shows that this is quite accessible.
Finally, um the points that I want to make in this presentation is that we could show that propagation of light can be programmed for an application specific relevant AI task which is in noising diffusion based image generation and I believe the opportunity in optical computing is through designing hardware and algorithms together.
In our case, we hand tuned our generation algorithm and we made sure that we repeatedly use the same weights over and over again so that we can really tap into the area where optics is useful.
optics is useful. And I think uh for making this technology more practical the next milestone is demonstrating an end-to-end system and to do it in a very large scale like a billion parameter scale model so that we can really claim
these energy advantages or latency advantages survive when the systems are scaled out and um I can say that we show the first premises that when we add more parameters we expect them to perform better as well as in the digital neural network case. Thank you very much. Thank you very much. Sorry. Nice.
>> I worked on NY light matters uh the first AI put silicon photonics accelerator.
One of the challenges I worked exact exactly on the the DAC and the ADC um subsystem that would do that for the their what is the photonic score.
uh that would basically be the offload engine for most tensor ops.
One of the problems was basically the ADC and DAC resolution uh like you know accuracy that would eventually end up in like you know prediction correctness over time basically in the operational you know in field.
Um there were there was a project that basically I cannot talk about but eventually that was observed and which was which was the reason why light matter eventually pivoted to you know interconnects for HBMs >> um and kind of like uh like kind of like stopped work on basically their multisspectral compute uh arm of research work. >> Yeah, I I agree. I totally agree.
>> Yeah, I I agree. I totally agree. uh when there is a large number of parameters or large large amount of data compared to the number of parameters has to be transmitted to the optical system then the advantages generally gets overtaken by the uh disadvantages that's
why I believe the first approach to this type of hardware is not a general purpose computing hardware like CPUs but it's more AIC type uh application specific approaches where we can really go and minimize the as you said the bit depth the precision or the data rate as well. >> So for the back prop, how do you
>> So for the back prop, how do you actually do that?
Because you're saying that I can do a forward pass where I keep hitting the same uh like F pass many times.
I do I do my outer loop and for my diffusion t times and so then but then I need to go backwards.
How how are you do you flip it and then send the beams backwards?
How do you how does it work?
>> There has been some work on that as well.
But in our case, we actually calibrate a fully digital twin and actually back propagation happens in digital electronics which is a normally the biggest bottleneck currently.
>> So um the other thing too that you could do so basically diffusion is an RNN that is teacher forced at every single step.
And so then if you remove that then there is no like L1, L2, L up to L10, LT.
Um you can just like do treat it as an RNN completely.
um have you tried that and and is there any advantage to that um from like the ADC DAC issue of reading things off and computing loss?
reading things off and computing loss? I haven't tried but I I agree that that's also another very interesting future direction because the usefulness of optics I think depends on how small interface that you do with electronics or digital environment and in that case
if we can find a way of looping the information in the optical system or in the analog system for longer uh time period or longer uh time steps then definitely I think the real advantage wouldn't be like seven times that we showed but it could scale scale up to hundreds or even more. >> I'm very curious if you try to do this
>> I'm very curious if you try to do this with much much much much much larger uh sort of uh weights uh what is the actual physical challenge with manufacturing kind of the relevant sort of components and kind of how does that scale and what kind of bottlenecks do we hit? >> Mhm.
So manufacturing wise I think it uses a very similar technology stack with electronics.
So people showed with etching and deposition that type of uh simple chemical operations the manufacturing could happen.
Um of course light has a longer wavelength compared to electrons.
So the unit size a bit larger but I think those are all um addressable challenges.
Maybe one limitation is that the nonlinearities because if you want to have a very large number of parameters what we benefit from in the digital case is making deep networks and in this case we need non.
So here the nonity is actually feeding back to the modulator again.
So we should come up with uh cheap ways of doing that nonity or maybe doing a massive operation and finding a model where it's useful.
I think uh that's the challenge making it very deep network for increasing the parameter count. Thank you. >> Thank you.
>> Just a quick question.
Electrons are much smaller than like the typical like wavelengths of visible light.
So are these kinds of photonic like chips limited by that size?
Like what is the theoretical like minimum unit size in optical compute?
And is there like are people trying to make that smaller? >> Yeah.
um people definitely try to look into different directions in terms of making better use of the space I would say.
So in electrons generally it's um as as we were discussing the bandwidth is smaller.
So one for instance one approach that uh many people took was that using the same unit size which is on the order of micrometers as opposed to nanometers in the electronic trans uh transistors but to use for instance something called com optical com.
So this means that you have hundreds of different wavelengths and there are ways of encoding different weights on each one of these wavelengths so that you could do multiplexing not in space but on the wavelength dimension.
Uh but essentially that's true that um yeah the feature size should be larger so that you can interact with the visible light. Thank you.
So I'm traditional computer architect.
Um so my question is looks like we have optics based uh processing units optics based interconnects.
What are your thoughts about optics based storage in traditional computing?
We are working on that from last like 70 80 years and we make significant progress but if we can have optics based storage we can bypass that ADC at all.
So what are your thoughts on that?
>> Definitely I think that's one of the biggest bottlenecks of the technology as well.
Um so actually optics as you know CDs, DVDs is a great storage medium for archival storage or for long-term storage.
Um it works quite well and there are u many different ways of doing that.
But of course I think the challenge is how can we quickly and cheaply write and read memory in the optical domain.
Um historically there has been lots of attempts on that.
has been lots of attempts on that. uh holographic memories is one like photorefractive materials there are attempts at it but I think so far none of them was as cheap and as accessible as the electronics so that's why it's one of the reasons why electronics
actually stayed as the main uh medium of compute but I would say um the reason why for instance we took application specific and fixed weight is that is exactly that because it's very expensive to rewrite or keep the weights in the optical domain Again u another line of research that is looking into this is phase change materials. They use them in
They use them in electronics but also in optics as well.
And uh this could be a way to solve that too.
So there are attempts at it but none of them is as major and um easy to use as NAND or like electronic based memories. Thank you. >> Hi.
Uh so one question that I had was uh have you guys considered I think non nonlinearity is something that is difficult with light but I was curious if you guys considered uh like you know polarized filters which almost have like a trigonometric relationship and like potentially composing uh several like polarized filters to do like you know like almost like a how forier is able to construct any signal with with trigonometry.
I was curious if that was something that you guys had considered uh working with.
I think that that's a good idea.
So polarization is another um dimensionality.
So light has different degrees of freedom.
Uh intensity is one, phase is one, those we used.
But polarization is another one.
And any manipulation there actually would manifest itself on the amplitude domain which we generally use with our sensors or detect with our sensors as a nonlinearity.
I think that's a good one.
I think one has to think about also how deep we can go with it because once we modulate the polarization uh we get a nonlinear response but if you want to do a deep neural network it has to act on each one of those layers each one of
those levels and I'm not sure once we filter it how we can uh keep making it nonlinear so >> but I I don't know at least personally I don't know a big research on that so that could be very interesting yeah >> thanks >> thank you >> all right thank you so Thanks. You guys can chat more AFTER
You guys can chat more AFTER I'm Aloque and um today we're going to talk about neworph computing.
I was told 10 minutes so I'm going to try and keep it under 10 minutes. Um a bit about me.
I am a co-founder and a partner at a venture firm called Standard.
We're one and a half billion in assets under management and we work on early stage technology companies.
Um, new types of computers have been a common thread in what I've worked on over the course of my career and I don't think there's been a more interesting time to build new computers than today.
Um, a little bit of street cred establishment.
um uh three degrees in double E undergrad at UT Austin, master's PhD here at Stanford.
And my focus in my in my research and work experience was semiconductors, fetonics, and nanomaterials.
Um and part of why I left was I I thought we would never find a use for this stuff.
And you know, here we are. Um, okay.
So, uh, to to start talking about neuromorphics, you always have to start by glowing up the brain a little bit. Um, it's miraculous.
It's a marvel of of evolutionary engineering.
Some of its properties, it's capable of language, reasoning, and creativity all at once. Um, it's multimodal. It's real time.
You can you can see it sees it hears it acts um visual audio physical embodiment all of these things just just come with it and it's real time.
You can have conversations with people and your brain is precomputing what to say back and you have you know millisecond hundreds of milliseconds response times back and forth.
Um it's just a remarkable computer.
Um it keeps learning all throughout your life. No context windows.
um it just it just keeps going and and adapts and learns new things.
Even things that evolution never really prepared it for.
Um cars, computers, all these things uh the brain didn't really have in its pre-train.
And and yet we can do it.
Um it can just learn from a couple examples.
Uh my favorite example here is if you show um a 16-year-old a stop sign once or twice, they'll recognize stop signs forever with several nines degrees of accuracy.
But how many different angles and pictures and shadows of a stop sign do you need to show to a autonomous car to get it to identify them reliably? Um, a lot.
And so again, it's it's a marvel of of engineering.
And um and and perhaps the most miraculous thing is that it just operates on about 20 watts, which is an iPhone charger.
is an iPhone charger. Um and so um in the in the late 1980s was we we basically had enough advances in neurobiology understanding how the brain itself works and enough advances in in electronics um we could start to make chips, we could
start to make circuits and and do them at scale where um um people started thinking hey we should we should these communities should intersect and and maybe there are things we can learn in designing our circuits and designing our chips um from how the brain works. And
And so, uh, Carver Meade at at Caltech in the late 80s, who's a one of the godfathers of of of the semiconductor industry, um, he coined the term neuromorphics and and his group did a bunch of the seminal work and and starting to really tease apart how the brain works and mapping it onto circuit models, um, for these things.
Um, so what is neuromorphic computing?
Um at at the highest level it is can we borrow the brain's organizing principles somehow in designing circuits and chips and and things that we use to do computation.
Um what you will find if you if you dig into the field more deeply is that there are many people have many different definitions for what it means and depending on your goal you kind of stretch the definition in a way that either includes or excludes what you're doing.
Um but we can abstract a handful of of highle principles um about how the brain works um that we can kind of use as as as concepts um to anchor on for neuromorphics.
What are those attributes?
Uh number one it is intertwined memory and computation.
So in the brain, the connections have memory, they have weights, they play a role in in computation itself rather than just simply relaying signals back and forth.
Um that's something that's really unique.
Uh that's something that we don't have in in in digital electronics, for example.
Um the way that communication occurs and and almost like state transition occurs is via events.
We call them spikes in the brain.
And um and it's it's a continuous machine that that takes sensory input and spikes get generated and propagated throughout and uh and and that's kind of the um the engine that drives things and and spikes themselves are these bizarre things that we don't really have any sort of perfect mirroring to and in and how we do computation otherwise.
Um it's it's this quasi it's you know mixed signalish if you were thinking about it like a circuit.
It has parts of it that are kind of digital, parts of it that are kind of analog.
Um and they all fit together in this in this unique way and and that's how the brain works.
Um and I'll get into that um a little bit in into more depth in in a toy model of how this works because it's a really fascinating concept.
And then um with experience uh with learning it changes both structure and function.
um like your brain is constantly rewiring itself um as you learn things.
And you know when we make chips, they're they're frozen.
And so this idea that this this this brain can can both logically and and even physically adapt itself over time in response to experiences um it's it's a it's a remarkable um it's a remarkable occurrence.
So people have tried to apply these learnings from the brain um onto onto designing chips and and of course like the neural network right the name neural network obviously um pays homage to this concept.
Um and uh oh sorry uh not there yet.
So the uh I wanted to go a little bit into the into a toy model of what a spiking neuron looks like.
And so this is what I mean by this kind of bizarre behavior that's not quite analog, not quite digital.
Um, and so if you're if you're a circuit bro like me, then uh you can you can think about it like a leaky capacitor is one way to think about it, but it's a leaky capacitor that when certain charge accumulates above a threshold, then it fires something that's kind of digital, a spike out.
Um, and so, uh, these spikes is basically like this this pulse train where pulses come in and charge accumulates on on the membrane of a neuron, but then over time it leaks, it dissipates.
So, you see these curves, it starts with a spike, but then it drops off.
But then if you have a certain number of spikes that come in coincidentally where they the the charge that you're accumulating on on on the on the membrane of a neuron sums up really quickly above a certain threshold then the neuron emits a spike.
Um but if not enough get there then it kind of relaxes and it resets.
And so we we don't have a circuit like that.
And it you know it's kind of analog in the sense that you have these decays and leaks.
Uh but it's also kind of digital in the sense that it it emits a spike in a in a discrete way.
Um you know the relationship in time between when these spikes occur matters a lot.
That's actually continuous variable.
The amplitude of the spike matters a lot.
And so it's it's just this this bizarre consequence.
But this is what's happening in your brain over and over and over again.
brain over and over and over again. and um and and and this these are one of these uh this is one of the core concepts that people have explored various iterations of you know can we do this in in electronics um okay so the point I was making earlier which is um you know people have tried to think
about what parts of this could actually apply to doing useful computations and doing things and the neural net is is the most famous of course it is um brain inspired but that's about it Um so when we think about like modern AIs and and modern neural networks um some concepts are absolutely neuro inpired and borrowed from biology. Um
Um those include having large networks of interconnected neurons.
Um as we learn we manifest those learnings as changes in connection strengths between these neurons.
Um and when we think about how information is encoded in these networks um it you know patterns across many neurons are how we contain pockets of knowledge and pockets of information.
of information. Um but one of the common threads in in neuromorphic computing um is that that it turns out that that digital electronics is is really good and we got really good at making them and scaling them and making them parallel and and it turns out that that going native and actually leaning into what can you do with this like highle
concept but its implementation in in digital silicon that's where all the real innovation has taken place and when you look at a modern AI you have these concepts like back propagation, transformer attention, how we train them at at massive scale um on digital hardware, like none of these map to what we have in the brain. So, we've already
So, we've already kind of gone off off course and and and we're doing things native to the manifestation of these neural networks in silicon rather than than kind of staying staying true to um you know, faithful to the brain.
And so this to me is is probably like the biggest takeaway I had in neuromorphics.
Um which is the brain is this ultimate model hardware co-design that's taken place like the hardware and the software are one and the same and um and they really have co-evolved in this really unique way.
And so if you are going to take pieces of it and try to apply it onto other computing substrates, onto other types of physics, you have to be very careful because maybe some concept will apply at a high level, but it turns out that whatever is native to that form factor is probably going to be the right way to take it to the next level and improve it.
Um, and so um, you know, where are we with neuromorphics today?
um in 2026 like my assessment is that we are still squarely in in R&D land and I think we are trying to figure out again this this kind of matchmaking problem of where can we take the right level of inspiration from what we find in the brain and where do we pair it with some sort of physical implementation and and figure out if we can actually do some magic with it.
Um so there there are kind of three highlevel ideas or premises that that that I see people exploring right now.
Um, some of them are we've got a problem that we're observing that we think that we have an opportunity to solve by leveraging some of these these biological insights.
Others are we're coming up with a bunch of new toys to play with and they have weird properties.
Um, maybe we can exploit these weird properties by by combining them with something that we derived from biology.
So, uh, the first one is is obviously people look at the brain.
I mentioned it 20 watts extraordinarily energy efficient.
um modern AIS are not quite energy efficient and so the idea is is there something we can take away from that property in order to um to to make AIS far more energy efficient.
Um and so some of the ideas that are being tried right now are co-mingling memory and computation.
Um um people are exploring analog ideas for even training and inference.
Um there you know and and I would say that this is where like you stress the definition of neuromorphics because you know this is you can purely digitally is is how people are thinking about co-mingling memory and computation right you can combine SRAMM and logic for example um you know Dmatrix is is a startup that's that's done some good work there.
They had a great hot chips presentation a couple weeks ago.
Um IBM has has done some work here.
Um but you know it's purely digital.
So is it is it neuromorphic?
Like yes if you if you have a very broadtent view of of you know intertwined memory and computation as being neomorphic.
We also have um you know on on the edge for thinking about different types of um chips and hardware they need to sense and react and adapt.
Um, can we can we take some inspiration from um the continuous sensing that the brain does and apply that to um how we have um a sensor on a drone, something that's doing something in real time?
And so um people are looking at at spiking networks.
Um can you train a neural net on on spikes rather than conventionally?
Um can you have devices that are ondevice learning?
Can can the network itself actually self-improve based on what it perceives?
Um, and and can we have like event- driven computation in the same way the brain works as there?
Um, you know, Intel has probably done the most work here.
They've they've got a a chip out of their R&D lab um that that they've done some some, you know, they've got a drone that can fly itself um that that can fly with it.
And then last is uh what I mentioned, we've got these new toys, right?
We've got we've got fetonics.
We've got a bunch of exotic electronic devices that have been cooked up in in semiconductor R&D labs over the years.
Things like resistive memories, things like meristers.
Um we've we've got circuit ideas like coupled oscillators.
And all of these systems exhibit very interesting new physics and new characteristics.
Um they have they have some analog properties um nonlinearities.
And so the question is can we can we start with these as as kind of a why now technological substrate and can we start to build models around the specific physical behavior of these systems um in order to co-design something that actually uh has a very interesting result or property.
This is the 30,000 ft view of the field.
Um I think it's been it's been a very interesting field.
it's been, you know, perpetually a source of inspiration, but then it's a little kernel of it here and there, but then for the most part, um, you know, we don't have anything that really resembles the brain that we're using in practice.
And, um, it's to me it's a very open question whether we will.
Um, but but I I think it's still a really interesting research direction to explore.
Um, and you know, think about something like neural nets, right?
something like neural nets, right? they were they were very sleepy for basically the first you know you know many many years of of their existence and then it took a while before we actually figured out how to make them useful and so I I there's just a gestation period for any
of this stuff and and so I think as a research direction it's it's it's super compelling um but you have to be patient and and you have to be a little bit lucky um so these are some of the resources um that that that I found really useful and helpful for directions to go forward Um but with that, thank you. >> Hi, thank you for coming. Um my question
>> Hi, thank you for coming.
Um my question is uh given uh all that you've seen in neuromorphic computing, what's the most promising startup or project or direction that you've seen uh take hold?
>> So I think I have to use the really broad tent definition of neuromorphic in order to answer that. Right.
I think a lot of the um a lot of the stuff that you would consider like very brain inspired is is still pretty early.
Um I think it's probably around the um either the um co-mingling of of memory and compute.
Um I think I think there there are bunch of startups that are doing that again purely digitally.
Um Dmatrix is one um that feels like it's it's they're kind of makeable now. It works.
um we can make we can manufacture them and and it feels like you can land the plane in terms of the market taking them up.
So that that's probably the answer there.
It's probably like a copout answer to some degree because it's none of like the real sci-fi stuff.
Um on the on the sci-fi stuff, I think that there there a couple very ambitious startups.
Um um Naveen Ralph's company, for example, is trying to do coupled oscillators in pure semos.
and and I think that at least like they're holding the constraint that we need to be able to make these things at scale.
Um, but very interesting types of dynamics and physics that you can come out of these systems.
So, I think that's a really interesting direction, but again, it's he's he's the kind of guy that can raise enough money to actually go after it. >> Thank you.
Thank you for that elegant presentation. I learned a lot. Thank you.
Um uh coming off of the hot chips uh presentation from Dmetrix I think the innovation around that is basically kind of like co- packaging and I'm bringing the SRAMM uh really close in the memory hierarchy closer to the compute on the decode side so that you know as we have uh we are hitting the decode ceiling they're trying to solve that problem.
So my question is that like um having that basically increases the surface area of the packaging and that basically means that you have you have uh you have lesser servers in the same data center uh amount of data centers.
My question is that do you think that there's innovation currently going on in neuromorphic uh world to kind of address you know packaging efficiency uh around around that?
I don't know if I consider that neuromorphic in terms of addressing that issue.
I think that there are a bunch of really interesting solutions for addressing that, right?
Like the core problem is at the end of the day um you know you have the capacitance of an interconnect is far greater than the capacitance of of the gate of a transistor.
And so the bulk of your of your energy expenditure is on moving bits, you know, across wires far and wide rather than like flipping the bit um of an individual transistor.
Um I I personally think optics is going to be how we address this.
>> If I just uh give a question here.
So I find it very uh like it seems like the the goal of neuromorphics is to develop circuitry or analog or optics or whatever that um outbrains the brain and like just becomes the brain.
And it's kind of like trying to estimate a sinoid with a bunch of lines.
It's like the best way to approximate a sinoid is just with a sinoid. Like just use a sinoid.
So like then if we're really trying to mimic the brain, then just use the brain, right? Just use neurons, right?
Is that is that the best hope we have for neuromorphics is neuro?
>> It's a it's a great question.
Um and and I think we've got a I think we've got a talk that might speak to that coming up soon. Um >> that was a segue. >> Yeah.
No, but it's a it's an interesting question.
I think that there are limitations um of the brain, right?
There are things that we can do in computers that that are way better than what we can do with the brain, right?
For example, like the clock speed um of a computer is so much faster than what we can do with a brain.
Um I think even in terms of memory, right?
We have like you know the retrieval and things like that are tougher but but you can you have infinite storage capacity with a computer whereas like you know we have limitations biologically.
Um, and so like it's possible, but then the question is like if we can get it, can we scale it?
Like maybe, but that's certainly like a direction of of of research that we that we've not really gone down yet.
Uh, thanks so much for the great talk. Uh, my name's Arun.
Um, you talked a lot about software hardware co-optimization when it comes to setting up next generation comput systems.
And I was wondering given that the cycles the iteration cycles for both uh are collapsing uh and getting shorter and shorter.
Do you think it's more important to iterate a model in software and then pick the hardware platform that best supports it or do you think it's more important to find the best hardware platform you can and then design a model around that?
>> Oh, that's a good question.
Um like the the copout answer is it depends, which is it does.
Um, and I think it depends on are you starting with like an insight around the physical property of some system and you know we're really good at making an optical computer and we deeply understand the constraints of the system but we we deeply understand where we can really lean into its advantages and and therefore we can't blindly apply a model architecture to it.
We need to work within the strengths and weaknesses of the system in order to design it.
I think similarly if you're if you're approaching it the other angle saying hey I've got a really unique architecture that I think can really work um then like now I need to pick what's the right like physical manifestation in in which to bring it to life and and and so I think it just kind of depends on on where your where your angle comes from. >> Yeah.
>> Thank you for this presentation.
So my question is u I I mean looks like all of our electronics are silicon based whereas brain are fundamentally carbon based.
Do you think like a lack of research in that direction is why we have to take uh like learnings from brain and simulate on silicon powered devices.
>> We're just really good at doing computation in silicon.
So I think it's more like if we're trying to mirror it as like an electronic circuit of some sort then like you're doing it in silicon.
Um, I think that like over time, if you look at our our the types of materials we've incorporated into chips, it's just gotten more and more exotic over time, and that should continue.
Um, you know, there's supposed to be there was a um, you know, there's some there's a team that was supposed to give a talk here tonight that's doing u they're working on diamond as a substrate.
And so there you go, that's carbon, but um, but but you know, so I think that like new materials, this and that, but I I wouldn't overly index on like silicon versus carbon.
And it's more about like manifestation as a as a um as electronics and silicon happens to be our our workhorse for doing that. >> All right, cool.
We got time for one more.
So, I'm actually really interested in another factor that I would associate with neuromorphic computing that maybe sort of wasn't mentioned, but then I'm interested what you think about it, which is the role of noise and reproducibility in computation.
Uh and because at anytime we're doing anything with sort of traditional silicon computing at the bit level, the intention is for the computation to either be reproducible or extremely close to reproducible because of parallelism sort of stuff like that like little edge cases.
Uh whereas if you think about the grain brain it's in some sense it's not trying to have a perfectly like low-level reproducible computation at the lower level.
And so I'm curious sort of what role that plays both in the field as you've described it um and sort of more broadly for the kind of ideas uh discussed.
>> Yeah, it's it's a great question.
Um even just even determinism within digital computation is very interesting, right?
As we get to the limits of floating point precision, then anyway, it's it's a very deep question.
it's it's a very deep question. Um I would say that there's one category of of of of attempt that I've seen um is thermodynamic computing which is again actually if you look back to the origins of neural nets people were even Jeff Hinton was looking at these things called Boltzman machines back in the day
and so I think this is almost like a reincarnation of some of those ideas um it's in analog circuitry but the idea is can you can you put noise to work for you in a useful way and the hard part is you want some of the noise not all the noise and and the noise introduced in the system itself is can can muddy the picture. Um but but to answer the
Um but but to answer the question I actually I think that's a really interesting um place to think about and um and and there are some directions that that are leveraging that and thinking very deeply about it but um I think it's a very profound question as per SPSA as long as you can make the noise reproducible and it's premise off the same seed then it actually is completely learnable.
Um you can do forward difference if you make it reproducible just twice if you do one to two the perturb and then one to scale the gradient and three times if you want to do central difference and then it's fine but the issue is making it reproducible.
I spent a lot of time with Shaw Drman at Stanford who's uh one of the leading neuroscientists over there and he's like I can make it reproduc I I can't make it reproducible is is really the issue.
I can't justify how you get this seed to a vector twice.
That's really the the the the main reason why zero order um is may perhaps not biologically plausible. It's a good question.
All right, we got to move on to neuromorph or to to to brain inspired. Thank you and luck. >> Hi everyone.
Um so guys, some of you may have seen human brain cells playing Doom.
So that's some of the work that I did in collaboration with Cortical Labs.
So, uh, they kind of put brain cells on chips and I wrote the algorithms for it.
Um, I think my goal here today is to kind of impart some of my thinkings of how we're treating this as a dynamical systems problem.
um rather than trying to force the brain to um or force the brain with innate calculations like matrix multiplications um we want to treat it as kind of like an input output machine where um some set of stimuli can you know produce some set of intelligence that allows us to get some form of output that's useful computationally.
So a a rough kind of example of this is you know how we talk to brain cells.
Um this is primarily through electrical stimuli.
So um in Doom in particular we have a game state.
have a game state. Uh so this is stuff like you can think of it as like scalar vectors and so in Doom for particular we have stuff like ammo health um but we also have you know kind of an image of the screen um so we we take all these variables we encode them into a set of
stimulus so this can be um which channel do we stimulate on this 2D multi-elerode array as well as stuff like ampl amplitude and frequency so we're controlling all these variables um of which you know I think we saw with the kind of leaky integrated fire neurons not exactly the same. Uh but these
Uh but these neurons do, you know, kind of accumulate charge and spike over time.
Uh what we do in Doom for this specific instance is we decode those spikes into a specific action in the game.
So that's kind of how we get useful compute out of these cells.
Um I'll talk a bit about some of the previous work.
So sorry if this is a bit small, but um so what Corticle did about four years ago was Pong.
So they got the brain cells to play Pong.
Um and Pong is a relatively simple game compared to Doom.
Um what this means is you can kind of hardcode how you want to do encoding and decoding.
Um so for Pong in particular um they figured out you know kind of which were the best encoding channels.
So if you stimulate these channels it would give you more diverse um outputs.
So more kind of a separable outputs for your decoder.
Um and then your decoder would kind of move the region or move the paddle to the region where it thinks the ball would be going.
Um so it's very much a handdesign thing and you can do that because it's a very low bit game.
Um, for Doom it's a lot harder because you know it's a 3D game for one kind of it's early Doom so it's like kind of 2D but you get it.
Um, so you know instead of one ball and a paddle you have you know multiple enemies in a field that it has to kill.
Um, and of course you know your action space is much larger in Doom.
So instead of pong where your paddle goes up or down uh we have a lot of joint actions.
So, you know, kind of strafing, turning, attacking, moving.
Um, these are things that we would have to account for.
Um, and this makes it kind of impossible to do encoding and decoding or at least a kind of a hand mapping for it because we only have 59 channels and becomes really really difficult um to get separable outputs for 54 different actions on 59 channels because some neurons might be on similar uh channels.
So, you might get very coupled behavior.
Um so what we realized was um for harder tasks you know if we want to scale this to GPT level intelligence encoding and decoding have to be learned end to end we need to learn some form of a way for us to stimulate the brain and we need to learn the decoding or the decoder as well for us to convert that into a useful action uh for compute.
Um so kind of the what we did initially was a closed loop PO architecture.
So PO is proximal policy optimization is RL kind of algorithm for it.
Um so we had Doom, we had the encoder uh you know we talked about encoder before.
So it would turn that into a stimulus for the neural culture.
Uh the spikes would then be sent to the decoder and then we'll convert that to action in game.
Um we'll talk a bit more about why we have a critic later.
Um but this is primarily to keep feedback to the cells stable.
So the thing with the cells is we need to provide feedback to it over time.
You know how good it's doing, how bad it's doing.
Um and we found that the critic for PO was particularly well aligned with how we wanted to do feedback.
Um so kind of how the percept looks like um you know we take the doom observations encode it neural hardware um put all the spikes to a decoder and then we update it every you know 2,000 steps or so.
So the encoder network looks something like this.
I think we talked about um image and scaler observations but this is kind of what it looks like.
Uh so we pass the image to a CNN and we you know kind of concatenate it with the scalar observations.
Um we pass it through MLP and then that kind of spits out the frequency and amplitude we should stimulate each channel with.
I think um if I'm correct in the code, so this is open source code.
You guys can take a look at this afterward.
Um I believe in this specific code it did not pick the channel.
Um all it did was pick the frequency and amplitude.
But I think in subsequent experiments we experimented with um having the MLP pick different channels to stimulate as well.
So a bit more freedom for the encoder network.
Um this is kind of what it ends up looking like.
So you know you can see the encoding electrics we have particularly the ones in green.
Um and you can see that we are sending kind of different frequencies and different amplitudes to these electrodes in hopes that uh this specific um setup would encompass the game state well enough for the neurons to actually act on it.
Um at least far better than we can kind of manually hand program it.
program it. uh so expect for you know further and further intelligence you know think GPD2 and beyond on these cells what you would have to do is something that's very similar you know we can't handcode 50,000 vocabulary words um this is something that we'll probably have to do and what we can
control instead is you know frequency and amplitude um so for the encoder you know for us to actually train the encoder because we can't propagate gradients uh through kind of the the multi-elerode array here um what ends up happening is we have to do do some form of stoastic stimulation. Uh so you know kind of beta sampling to
Uh so you know kind of beta sampling to figure out what's the best amplitude and frequency to stimulate it at.
Um I will say that this is fairly different from reservoir computing.
Um the the primary difference is your culture is changing.
Um so if you're sending something of this frequency and this amplitude uh that changes the culture in a different way every single time.
Uh so you don't really know how it's going to perform.
So the beta sampling in a way is also part of a policy learning algorithm for the cells.
This is changing the physical structure of the cells and changing the outputs of the cells over time.
So you can think of it as every single moving part in the system is is actively affecting how the cells process information and give us outputs.
Um decoder we use a linear readout.
So we had a few issues with the decoder initially.
I think uh what we did was we oversized the decoder bit too big and we we started realizing that oh it could play on itself if you just zeroed out the spikes.
The decoder could play itself if you had enough biases.
Um, so what we did was we kind of like massively undersized the decoder, ran a bunch of ablations to make sure that the spikes themselves actually carried information to play the game.
Um, and then we set the bias of the the kind of linear decoder at this point.
Kind of a linear readout, I would say, is a better way to put it, um, to zero.
And that kind of gave us better results where the decoder wasn't overfitting the task and it was able to start making progress on the game.
Um this is particularly important because if your silicon decoder overfits what ends up happening is the cells never learn.
They never have to learn.
They can keep their stochastic behavior and then the decoder itself can just play the game.
So this is particularly important and um kind of want to hammer down on it because I think this is something that you know startups and companies get wrong a lot of the time.
Um it's very easy to cheat performance out of these systems.
All you have to do is, you know, massively oversize your decoder, run back propagation on it, and suddenly you have brain cells that do anything, play anything.
Um, I think most recently we saw this with, you know, fly brains, right?
You know, you have these like connetos, and what people do is they just have these like massively oversized decoders.
It's never the fly doing anything.
It's just silicon, right?
You have this decoder that's doing everything for it.
Um, so I would say be very careful with that and kind of information that propagates that idea um over time.
Um so yeah this is kind of what the decoder looks like is a few sets of behavior.
So forward backward none um strafing left and right turning left and right and attacking.
Um so every kind of action step we do a softmax over the joint probabilities of the actions.
Um so feedback is something that um a lot of people got wrong.
Uh this was fairly widely covered by media.
Um but I think not a single outlet got it right.
They got very very confused over how it happened.
Um so to provide good and bad feedback to cells um how we define good and bad feedback is um through asynchronous and synchronous stimulation.
Um it is our kind of belief through the free energy principle Carl Fristen that um asynchronous stimulations would is something that the brain wants to avoid.
Um so you can think of it as if it does something bad give it asynchronous stimulation.
If it does something good give it a synchronous stimulation.
Um, however, because we're doing it in a kind of reinforcement learning style, it gives us a very unique problem, which is, um, it sucks.
You know, at the start of the game, it really sucks.
It does a lot of things that would give it negative stimulation.
And the cells that we have right now aren't good enough for, you know, long-term credit assignment.
Um, so what we had to do was, um, we had to predict surprise essentially.
So we use that critic um in the silicon decoder to figure out how far is it away from the silicon critic, how surprising is this action specifically and we scale our feedback based on the surprise.
So you know if it expected this action was going to be good and in reality it picked something and was really bad then would give it negative feedback.
So this allows us to modulate feedback in a much more I would say equal way where you know we don't keep stimulating with negative feedback.
we actually allowed it to build uh connections over time and the way in which it builds those connections is through minimizing surprise uh to the silicon uh critic.
So that's kind of what we do and of course you know the we have different stimulation patterns for each of these.
So um on our website we have this and it's kind of interactive in a way.
Uh so you actually see the channels light up and see the different simulation policies.
Um so you know synchronous obviously this would be synchronous and then asynchronous for negative feedback these would be firing at different times at random.
Um and then of course on top of that we scale it by surprise.
Um so we kind of take the um what I would say the on on top of the surprise that we had here uh we still scale it by surprise over time using the TD error and that specific scaling is on the frequency and amplitude.
So not only do we have kind of the reward sign that is modulated by the critic uh but we also change the frequency and amplitude of that feedback that we send to these cells.
Um so I think in the diagram here we have different channels for feedback and different channels for encoding.
Um I'm not sure the code if it's exactly the same.
Um but you know definitely check out the code on GitHub.
I think this is my last slide.
Um if you have any questions FOR >> YEAH.
I was uh curious about like the positive and negative signals.
So it seems like different frequency like uh like electricity is applied to the neurons or something. Yeah.
>> I was curious like what like does it feel does like it feel bad the like when you apply a 40 Hz thing like what exactly does it mean negative versus positive uh for the neuron? >> Yeah.
So the neurons themselves don't feel bad.
You know, they have no pain receptors.
There's no kind of like noxious poison we're introducing to these cells.
Um when I say, you know, positive and negative feedback, that's in relation to the game.
So, you know, if it does something bad in the game and we don't want it to do that action, we send it asynchronous feedback to reinforce that we don't want it to do that specific action.
Um so the brain itself, you know, at least this culture, it doesn't optimize for feeling good or feeling bad.
It optimizes for kind of reducing the entropy in the culture itself. >> Okay?
And so like the different like 40 Hz thing increases the entropy of the culture. >> Yeah.
I guess it it wants to reduce the amount of times that it gets kind of this high entropy stimulation basically like it wants to reduce the rate at which it gets this asynchronous stimulations.
Um and it would kind of self-organize to prevent that. >> Got it. That makes sense.
If you're scaling the TD error is what you mean by scaling the surprise, right? >> Yes.
uh aren't you like and you don't have a static reward model itself.
So you're changing the reward model itself.
How does it reach an equilibrium?
>> I think this is the interesting thing.
It never reaches an equilibrium.
Um I forgot to mention this in the slides actually but um you can think of it as you know you have these two dynamical systems that are almost fighting for control in some ways.
Um and your your enemy is entropy.
um entropy would disrupt these two systems kind of pushing themselves towards you know semiattractor states.
Um so you know in our code actually we uh did something called entropy penalties.
Um this is very very abnormal for PO.
So traditionally in reinforcement learning um you actually want a bit of an entropy bonus so you can start exploring the space around it.
Um we actually actively penalized entropy um because we found that it was helpful and the cells provided sufficient entropy for that.
Um so you can think of it as you know two dynamical systems that are pushing themselves uh to better and better outputs over time.
Um in terms of you know the sample efficiency of that I you know we never truly measured that down uh to the wire but that's uh that's what it looks like >> for the um stimulation you used.
How different is it to the real life electrical signals you might get from your eyes and ears?
>> I think it's very very different.
I think the way that I want to frame this is we're using this as a material and you know kind of going off on a tangent here.
I think that human intelligence is not optimal for compute.
Uh we're optimized for survival, you know, reproduction and you know the existence of our race.
Um but we're not optimized for compute.
It just happens that the fundamental substrate that we operate on is really good and is a good host for intelligence.
Um so I would say kind of it would be very different uh compared to you know traditional like I inputs you know it could evolve to a point where it starts looking similar and that might be a bit scary.
Um but my bet is that it will continue diverging and what we will see is a different form of intelligence arise on the same substrate that we see intelligence on ourselves.
>> Could I ask a followup?
Could you train on real life electrical signals from the brain or is that like just not >> like you know like grow some eyes and then >> yeah like record it and then maybe put it in a person.
>> I think I think some people have tried because you can grow like you know mini like eye organoids and attach it to the brain.
Um but I think that starts to get very scary because then you'd be kind of recreating a human um from first principles maybe. >> Yeah. >> Okay. Okay. Thank you.
How scalable is such a solution?
>> I think that's probably one of the hardest questions um that we have to answer for ourselves which is you know I see this you know solution of using biomp computing as having two main problems which is um you need to get frontier levels of intelligence on these cells and then you need to figure out how do you serve this intelligence to billions of people in the world.
the world. um we kind of know how to get intelligence on cells like we've we've seen it you know we're all intelligent beings on the same substrate um what we haven't seen is distributed computing right we we fundamentally don't know how to do this yet um so you know kind of a
shitty answer to that would be I don't know how scalable it is um but maybe a better kind of answer to that is um we're trying to find out um and we've kind of assembled a fantastic team uh toward finding and making steps toward scaling this technology >> thank Thank you. Thank you. Thank you. >> Thank you.