0:06
Hello everyone and welcome back.
Hello everyone and welcome back.
This is Jordan Anos here for sale media and semi analysis at UTC 2026.
I'm joined by Thomas Summers of Positron.
Uh we're going to chat a little bit about uh everything that's new, right?
You guys are moving on from um current generation with the Archer FPGA and into uh next generation chip with Azuma and the Titan system, so yeah, welcome and excited to have this chat. Yeah.
Yeah, thanks for having me.
All right, let's start with uh a little bit of an update.
You guys have had some uh some big announcements and I I have to say like for a chip startup, for a chip company to get to where you are right now where people can actually use it uh both from an API perspective and then also people can you know in some cases log in and run benchmarks for themselves is uh is is pretty great, right?
Haven't really raised that much money, haven't spent that much time actually doing development and you're making real progress on certain tokens.
Yeah, it's uh you know, been both wild.
Your company's almost 3 years old, but uh uh going from having, you know, first, you know, shipping pseudo product.
you know, shipping pseudo product. It It had a lot of uh bugs, but uh shipping our first product about 18 months in and then now just under 3 years in having the the announcement with Oracle um is uh uh I think unusual in in in the space to have that rapid of uh developments and and real market success, but uh
uh it's I think a credit to our team and and sort of the unique approach we took to starting with FPGAs and then making that into a true product and not just a prototyping vehicle, uh but having that and uh having that rapid iteration cycle with customers has really enabled us to to leapfrog uh folks that started a lot earlier had a lot more money uh etc. Yeah, cool. So, can we talk about the Yeah, cool.
So, can we talk about the feedback from customers like the big thing for me at least in our semi analysis where we use models and tokens as how uh you can pay more you know, more per token not just more uh for faster tokens right now, right? Opus 4.
6 fast is like the ratio is about six times more money for two and a half times faster, so you're paying like a super linear amount.
Do you think this is going to continue like we saw a Codex Spark announcements as well, still research preview.
Is that sort of the thesis of um the whole Positron business that it's going to be more expensive tokens, but faster or do you think there's a way to also save money and serve faster tokens?
Yeah, I would say the Positron thesis is more aligned to trying to reduce the cost of of tokens and still have them be high quality, fast etc.
Um I think, you know, while there's opportunity to um differentiate on on speed and there are other um uh providers or you know, be it Nvidia or self just reducing the the batch size is expected.
Like that's how Anthropic is on all of their hardware platforms delivering that higher level performance.
If you're paying, you know, three to six times more per token, they're willing to not cram that box with as many users.
Um so uh you know, anything from that approach to having architectures that are really solely designed for low batch, extremely high inference performance, but the the problem for service providers is that the actual aggregate, you know, throughput of that uh is significantly lower, so it's not as profitable unless you can charge a significant premium uh above that sort sort of super linear um uh scheme.
So uh for Positron, I'll say that, you know, we want to, you know, compete on both providing very high level of interactivity like that tokens per second per user that that individual user experience speed, uh but actually have that be driving cost down from even the the, you know, commodity uh you know, low to medium performance range that uh uh is the bulk of the actual inference taking place out there.
And then I would say just though that like using, you know, Opus 4. 6 fast uh is addicting.
Like uh my my personal sessions once I I switched that on uh I saw, you know, bill uh you know, grow very quickly, but uh um once you switch in that mode and have to then go back when you realize, oh, I don't actually need this result super fast it is painful. So, I I see that value.
Is it actually worth that huge price premium?
Probably not in most cases, but uh for me user's individual experience, it it is addicting. Yeah, yeah. Interesting.
So, uh maybe let's dig into the architecture a little bit.
So, you say there's you know, a scenario where Positron's used to reduce cost, right?
Let's say uh um serve cheaper tokens in general.
Uh can you talk about exactly how that's happening on the chip itself? Yeah.
happening on the chip itself? Yeah. So, uh in terms of, you know, what we've publicly disclosed about our architecture, you know, it is a systolic array of direct design, but really really focused on maximizing memory bandwidth utilization, so
uh you know, on our FPGAs we hit sustain about 93% of theoretical peak memory bandwidth and you know, what is the thing that is bound in in uh decode or in in inference overall especially when you look at the attention mechanism itself, be it in prefill or decode. Um
Um you know, that's a fundamentally memory bound matrix vector um uh operation.
And so that I would say architecturally we really focused on obviously getting uh as much of the the bottlenecks between the memory and compute out of the way so we can achieve, you know, as much of that that theoretical peak as possible, but then at the micro architectural level really treat matrix vector based math um as a a first class citizen which has been horribly neglected on on Nvidia, you know, Google's TPUs etc.
Like we've actually had regression in that over recent generations.
Going from Hopper to Blackwall actually had the ratio of your um matrix matrix, your your gem performance um uh that's when when you go to matrix vector on on Hopper it was a 16 to to one ratio of matrix matrix matrix vector, which is your your scaled dot product attention.
Um that went to to 32 to one.
So, uh even though they increased the the flops overall for the device while also, you know, doubling the amount of silicon um uh they actually had the ratio get worse for for, you know, the main constraint when you're scaling um uh tokens.
Like it you have your uh attention mechanism growing quadratically in computational complexity and they actually made that ratio worse.
And so our focus is really solving those bottlenecks which have been neglected in the era of training and when people are just focused on big, dense matrix matrix uh workloads.
Yeah, and maybe just uh be clear, are you guys talking about the ratio for your chip when compared to those? A one to one. One to one.
So, we we have most importantly the thing that is uh scaling the most in the compute, we get to use that, you know, orders magnitude more efficiently and and actual more throughput than uh than Nvidia and others.
Yeah, and can you talk about like the high level specs?
I mean, others have made decisions at the high level on the chip to focus on let's say, you know, SRAM or HBM or, you know, specific stuff in the actual uh logic.
Is there a high level summary that you can provide there?
Yeah, so with Azimov, our our next generation chip, uh we're taking a pretty uh uh counterintuitive or just against the the trend approach of uh actually leveraging LPDDR um as the the main memory technology.
And with that, we're able to offer drastically higher memory capacity.
So, uh you know, Azimov we've announced will be shipping uh you know, up to 2.
3 terabytes of memory capacity per chip um [clears throat] and in about 400 watts uh TDP profile.
So, uh you know, over an order of magnitude more more memory than, you know, a B200 uh solution today.
You know, even Rubin will be shipping with 288 to start, eventually go to 384.
Uh so, we're in that 68x, you know, greater memory capacity per chip while being, you know, less than a quarter of the power.
Um and we really think that scaling that memory capacity is is critical for longer context lengths beyond support large models.
You know, when we take four of those Azimov chips and put that in Titan Titan, which is the name of the server, um you know, that's design we want to be able to support 16 trillion parameter models running on a single box and think that uh if you are able to have that much memory capacity and that high realized memory bandwidth per server you know, we still have great chip to chip networking.
We've got 72 112 gig lanes dedicated for for interchip communication.
Um but you you rely less on that if you can do more per chip.
Um so, uh we think, you know, when Azimov is shipping next year, uh be able to support 16 trillion parameter models, you know, millions of tokens context length is uh very exciting, differentiated capability. Cool.
All right, 16 trillion parameters millions of tokens of context good token efficiency no HBM, no advanced packaging. No advanced packaging.
Regular organic substrates.
That's another big thing from a supply chain perspective.
Like, if we want to be able to support, deliver, you know, hundreds of megawatts, gigawatts of compute capacity to to the world, um having the same supply chain that Nvidia, AMD, Google, all the big guys are fighting over um is is a losing proposition as a startup.
And so we want to embrace, you know, commodity memory, packaging technologies, etc.
And enable um you know, our advancements through architecture and through, you know, clever uh scaling. Cool. Awesome.
Well, I think that was a great place to wrap.
Thanks for taking the time.
I got lots more questions for you, but yeah, we're going to sign off on this one for now.
Thanks everyone for watching. Thank you.