Why next-gen AI scale-up needs CPO

0:00

Semiconductors run on copper.

0:00

From the tiniest  metal layers that connect individual transistor and the traces that run through your motherboard,  all the way to the massive “spine” that allows 72 GPUs inside a single Nvidia NVL72  rack to communicate with each other. All of them use copper.

0:19

Without  copper, microchips wouldn’t work. But copper has a limit.

0:23

And a very short one for  that.

0:23

Your internet cable might be copper inside your house or to the curb, but that’s about  the maximum if you want really fast internet.

0:36

Beyond that, it’s all fiber, optics and lasers.

0:36

The same is true for datacenters.

0:36

But here, the demand for ultra-high bandwidth limits the  reach of copper even more.

0:42

The faster the speed that data is transmitted over a copper channel,  the shorter the distance the transmission can reach.

0:54

Two meters is about the maximum for  speeds of 200 Gigabits per second per lane, which is the speed that the latest AI Chips  communicate at.

0:59

That’s not a lot of reach.

1:06

Beyond that distance, copper cannot support the  immense bandwidth needs of modern AI servers.

1:13

So, why use copper in the first place, if  it’s so limited?

1:13

Optical must be better, how can you beat a laser?

1:18

Well, there’s a  reason why they say “use copper when you can and optical when you must”.

1:25

And that’s exactly  what we will figure out in this video.

1:30

We will talk about the advantages  and limits of copper, explore the current state of pluggable transceivers and  explain the future of Co-Packaged Optics. To understand copper vs.

1:43

optics,  we have to understand the different networking tiers a modern AI server is using.

1:51

To start, there’s the Front-End network.

1:51

That’s what every server has been using, long before AI became a thing.

1:55

It’s used for  loading data, SSH access and user requests.

2:02

Where it gets interesting is the  Scale-up and Scale-out Networks.

2:07

First, there’s the scale-up network.

2:07

Scale-up  connects all the compute and networking trays inside a single rack.

2:13

It’s basically  rack internal communication.

2:18

The scale-up network requires extremely high  bandwidth and super low latency, because it’s used to connect multiple GPUs, or other  AI accelerators, in such a way that they behave almost like a single GPU.

2:29

That means they have  to be able to communicate with each other in an instant and share extremely large amounts of data.

2:35

The most famous example for a scale-up network is probably Nvidia’s NVLink inside the NVL72 rack.

2:41

Scale-up networks are copper based and because the networking is limited to a  single, or maybe a double-wide rack, we are talking about distances of up to two  meters.

2:52

Well within the domain of copper.

2:57

Second, there’s the scale-out network.

2:57

Scale-out  handles the networking that goes outside the rack.

3:03

It basically covers the entire datacenter and  connects all the individual racks and servers with each other.

3:08

Scale-out is rack-to-rack, or to be  more precise, server-to-server communication.

3:14

Scale-out isn’t trying to turn multiple chips into  a single one, at least not in the extreme way that scale-up does, so it doesn’t have the same  extreme bandwidth and latency requirements.

3:25

But you still want it to be as fast as possible.

3:25

And because a modern AI datacenter contains a lot of racks, spread out over a pretty large  area, the scale-out network also has to cover a pretty large area.

3:36

One rack to the  next one might still be in reach of copper, but the next row of racks certainly isn’t.

3:41

That’s  why scale-out networks have to be optical.

3:47

And just to put the different bandwidth  requirements into perspective, the scale-out network needs 8-10x more than  the front-end and the scale-up network uses 10x again that of the scale-out.

3:57

So, now we know that generally, scale-up is rack internal and copper based and  scale-out is rack-to-rack using optics.

4:01

But how does that help us and where do co-packaged  optics come in? I’m so glad you asked!

4:12

Optical networking isn’t new by any definition,  no matter if ultra-long range fiber cables that literally cross entire oceans or optical  networks inside datacenters.

4:18

They have been around for a while and are tried and tested technology.

4:24

The most common form of optical interconnects used in datacenters are so called “pluggable  transceivers”.

4:29

They are called pluggable because, well, they are plugged right into the back of  a server tray and they are called transceivers, because they both transmit and receive signals.

4:39

If you have been inside a datacenter there’s a good chance you’ve seen one of those before.

4:44

Pluggable transceivers are not only a proven and widely used technology, they are  also standardized.

4:49

The latest and most common ones are OSFP and QSFP-DD.

4:54

That’s  why they look the same.

4:54

And there are a lot of companies competing in the pluggable  transceiver space.

5:00

Lots of different choices.

5:05

A standard pluggable transceiver contains four  main components.

5:05

First, the standardized physical connector that plugs electrically into the  server tray interface.

5:11

Second, a DSP, short for “Digital Signal Processor”.

5:17

The DSP has the very  important function of boosting and cleaning the incoming electrical signal before it’s translated  into the outgoing optical signal and vice versa.

5:29

Component number three is the transmitter, also  called “Transmitter Optical Sub-Assembly” or TOSA and includes the laser plus the modulation  function and number four is the “Receiver Optical Sub-Assembly”, or ROSA, a sensor  that catches the incoming light signals.

5:45

If you had to guess, which of these four  components is using the most amount of energy, what would your guess be?

5:50

The interface, the DSP,  the transmitter laser or the receiver sensor?

5:57

I’m pretty sure most of you would answer  the same as I would: obviously the laser.

6:02

I mean, it’s a freaking laser, right? But no.

6:02

The laser only uses about 15% of the typical power of a pluggable transceiver.

6:08

Only a few more percent than the receiver sensor.

6:15

The vast majority of the energy is consumed  by the Digital Signal Processor.

6:15

And with vast majority I really mean 60% or more. But that’s not all.

6:20

Every system, every interconnect always adds some kind of  latency into the network.

6:25

A pluggable transceiver usually takes about 150 to 200 nanoseconds to  translate the signal from electric to optical.

6:36

I’m not going to let you guess again, because  it’s even more extreme.

6:36

Over 90% of the entire latency delay is because of the DSP.

6:41

Looking at all the components inside a pluggable transceiver, the digital  signal processor is responsible for 60% or more of the entire energy consumption  and 90% or more of the entire added latency.

6:56

And that’s where co-packaged optics come in.

6:56

The  entire reason CPO even exists is to eliminate the need for a DSP in an optical transceiver. And how do you do that?

7:02

By placing the optical engine closer to the source, which means closer  to the networking switch or even the GPU.

7:14

The reason a pluggable transceiver needs  a DSP in the first place is because until the signal reaches the transceiver and  is translated into an optical signal, it’s an electrical signal traveling over copper.

7:22

The signal originates from the GPU or the switch, travels through the metal layers of the silicon  onto the package, from there to the motherboard and then finally to the transceiver plugged  into the very end of the server tray.

7:39

We are talking about up to 30 centimeters here.

7:39

And while it doesn’t sound like much, remember that for high-speed interconnects, copper caps out  at 2 meters or less.

7:44

So, 30 centimeters is a lot for copper. Can it handle it? For sure!

7:51

That’s how  pluggable transceivers have worked for many years.

7:57

But if you want to translate a signal from  electrical to optical, you need a super loud and clean signal.

8:02

30 centimeters of copper are enough  to degrade a modern high-speed signal to a point, where you need a DSP to boost and clean up  the signal.

8:08

And that takes time and energy.

8:14

Hence the power and latency penalty.

8:14

That’s what  CPO is trying to eliminate.

8:14

Get so close to the source, that you don’t need a DSP anymore.

8:20

There are different ways to go about it.

8:24

One interesting approach is called LPO or  “Linear Pluggable Optics”.

8:24

The idea is a mix of crazy and genius: it’s still a standard  pluggable transceiver, but without a DSP.

8:38

If you are confused right now, I get it.

8:38

Isn’t a  DSP required to clean up the signal? Yes, it is.

8:45

But LPO basically says “screw it”, takes the  still distorted electrical signal, translates that into an optical signal, and hopes for the  best.

8:51

And you know what, it actually works.

8:51

But only for a much shorter distance than a typical  optical network.

8:58

Because sending an already distorted signal drastically decreases the reach.

9:03

The first attempt towards actual CPO was called “On-board Optics” or OBO.

9:10

The idea is simple,  move the optical transceiver closer to the signal source to reduce DSP requirements.

9:16

A  good start, but it never got enough interest, because it wasn’t close enough to get rid  of the DSP entirely and at the same time it lost the easy access and exchangeability that  pluggable transceivers offered.

9:27

It was a good idea but ultimately combined the worst of  both worlds.

9:33

More complex integration and less repairability while still needing a DSP.

9:39

And then there’s NPO, or Near-packaged Optics.

9:46

NPO moves the optical transceiver even closer to  the ASIC, usually on a special high-performance substrate.

9:52

Much closer than OBO, NPO can  been seen as an intermediate step towards CPO and is actually being deployed  right now.

9:57

In the end the question is, if you are using NPO, why not go all the way?

10:02

And all the way is only co-packaged optics.

10:09

It’s already in the name.

10:09

Co-packaged, as  in on the same package. That’s super close.

10:14

AMD’s Infinity Fabric On-Package for example,  it’s already in the name, also uses on-package interconnects.

10:20

That how Zen 1 to Zen 5 scales  and connects compute and IO-dies. But we can get closer.

10:27

Let’s talk about CPO tiers.

10:27

The first tier is the minimum, what we just talked about.

10:32

The optical engine  is placed on the same package as the switch.

10:38

Both are connected via copper traces that run over  the shared packaging substrate.

10:38

It’s close enough to get rid of the DSP entirely.

10:43

But on-package  still requires a high-speed interconnect that runs over the package, to connect the ASIC  with the optical engine, which means you need SerDes that translate the electrical signals  from parallel into serial and back again.

11:00

The second CPO tier is introducing an interposer.

11:00

This interposer can be silicon based or organic and the ASIC and the optical engine are sitting  on the same interposer.

11:06

They are still packaged together, but the interconnects aren’t routed via  the package substrate, but via the interposer.

11:17

And because a interposer allows for much higher  interconnect density, this design doesn’t require SerDes anymore.

11:23

ASIC and optical engine are  connected via a wide fabric that allows for full parallel integration.

11:29

This is the final  boss of CPO.

11:29

Placing the optical engine so close to the ASIC, it not only doesn’t require a DSP  anymore, but also gets rid of SerDes entirely.

11:43

Using more advanced packaging technology,  like hybrid bonding for example, could in theory result in an even closer and lower  power integration.

11:46

No matter if true 3D stacking, with the optical engine above or below the AISC,  or 2.

11:52

5D stacking like TSMCs latest SoIC-mH.

12:00

But once the optical engine is so closely  packaged that it doesn’t require SerDes anymore, you have truly mastered CPO.

12:06

But there’s one more thing.

12:09

Did you notice how we’ve talked about a  networking switch or a GPU or other AISCs?

12:15

What we are seeing right now with Nvidia’s  Quantum- and Spectrum-X chips for example is CPO for the networking switch.

12:20

But the  final destination isn’t the switch, it’s the GPU or AI AISC.

12:25

Co-packaging the optical  engine with the GPU instead of the switch isn’t a different tier, but an entirely different level.

12:33

Now that we know basically everything about CPO, let’s talk about implementation.

12:39

At the  beginning we talked about “use copper when you can and optical when you must”.

12:43

How does that  relate to CPO and where will CPO be used first?

12:50

The first target of CPO is  actually the scale-out network.

12:54

It’s to replace the pluggable transceivers.

12:54

And  with all the hype around CPO it seems like an easy choice.

12:59

But interestingly it isn’t.

12:59

Scale-out networks connect multiple racks in a datacenter, they have been optical for a  while, because they cover a distance that’s out of the reach of copper.

13:10

As I said in the beginning,  pluggable transceivers have been the standard for many years now.

13:15

And we just learned how the only  reason CPO exists in the first place is to get rid of the latency inducing and power-hungry DSPs  that are a necessity for pluggable transceivers.

13:28

But as always, everything comes with pros & cons.

13:28

Yes, pluggable transceivers need a DSP.

13:35

Yes, they use a lot of energy.

13:35

And yes, the  add quite a bit of latency into the network.

13:41

But they are also super easy to  handle.

13:41

They are literally pluggable.

13:45

If one fails, you just change it out.

13:45

Something  any datacenter technician can do in in no time.

13:51

And then you should never underestimate working  with technology you know.

13:51

Pluggable transceivers have been around for so long, every datacenter  knows how they work and what their flaws are.

14:01

Plus, because they are standardized and there are  many suppliers, large datacenters never have to worry about supply constraints or high prices.

14:07

If  your current supplier gets too expensive, another one will gladly step in.

14:14

Because of that, moving  to CPO isn’t as easy of a choice as it seems.

14:19

With CPO, you are buying the optical  transceiver as a part of the sever hardware.

14:25

It’s literally packaged right next to  the switch or maybe even the GPU itself.

14:29

Which means if you buy Nvidia hardware,  you have to buy the Nvidia solution.

14:33

And if you buy Broadcom, you obviously  have to buy the Broadcom CPO solution.

14:38

It also means, if one optical interface fails,  you have to replace the entire switch.

14:38

Because you can’t repair something at package level.

14:44

But  at the same time, you do get an already tested and working system that is more reliable over all.

14:50

In the end it all comes down to pain points.

14:56

Where are the datacenter providers and the  hyperscalers feeling the most amount of pain?

15:01

What is most important to them?

15:01

And  when it comes to AI datacenters, the most important aspects, the largest “pain  points”, are energy and system utilization.

15:12

When you spend billions of dollars on AI hardware,  you have to make sure that it doesn’t sit idle.

15:18

That means you want to reduce latency and increase  bandwidth, even for the scale-out network.

15:24

And because power is such a massive issue  for AI datacenters, you want to reduce energy consumption for everything that’s  not AI compute to the absolute minimum.

15:34

And these two areas are where CPO shines.

15:34

So, AI datacenters would seem to benefit a lot from CPO and in fact, many are  considering or already preparing a switch, given that CPO promises better efficiency,  reliability and operational simplicity.

15:52

Rack-to-rack communication has a lower latency  and the networking layers consume less energy.

15:57

With these obvious benefits, the entire  industry should be moving in unison, right?

16:03

But looking at the industry right now shows  an interesting deviation.

16:03

Many hyperscalers seem to value repairability and especially  vendor diversity, with the important added factor of price control, more than a faster and  more efficient scale-out network.

16:14

Hyperscalers can see the technical benefits, but they also  want to avoid a vendor lock-in at all cost.

16:26

That’s why many hyperscalers are still  cautious about fully committing to CPO.

16:31

Because of this, there are even efforts to develop  a NPO and pluggable transceiver hybrid.

16:31

The entire industry around optical networking is moving  quickly and all possibilities are explored.

16:37

NPO could become a real middle-ground or might just  turn out to be an intermediate step towards CPO.

16:49

Neoclouds on the other hand are much more keen  on CPO.

16:49

They like buying a turn-key solution and the idea of an Nvidia CPO switch  is super appealing to many of them.

17:00

In any way, because scale-up  always has been optical, the rest of the infrastructure is already there.

17:03

So, scale-out networks in AI datacenters are starting to switch to CPO.

17:09

But what about  scale-up?

17:09

An entirely different beast and still fully dominated by copper. But why copper?

17:15

Because copper speaks the native language of semiconductors, it uses electrical signals.

17:20

Every microchip works with electrical signals, starting with the smallest layers,  deep inside the chip itself.

17:30

By using copper for the rack internal scale-up  network, you don’t have to translate signals at all.

17:35

Copper has a latency of about 5  nanoseconds per meter and at distances of up to 2 meters or less, we are talking about  maybe 10 nanoseconds of latency.

17:41

Remember, the DSP inside a pluggable transceiver added  about 150 to 200 nanoseconds of latency alone.

17:53

Copper is fast, because you don’t need to  translate the signal.

17:53

Copper is easy, because you don’t need to translate the signal.

17:58

And copper  is tried and tested, because it has been used since the very beginning of modern networking.

18:03

And because there’s no translation done at all, it’s still faster than co-packaged optics.

18:08

Yes,  CPO gets rid of the DSP and with that it removes the majority of the latency in legacy pluggable  transceivers.

18:15

But there’s still added latency.

18:20

And yes, CPO also uses considerably less energy to  translate the signals from electrical to optical and back.

18:27

But even without a DSP that’s more than  what you need for copper.

18:27

Because with copper, you don’t have to translate the signal at all.

18:33

For scale-out networks, CPO has obvious advantages, even if there are some downsides.

18:38

But it’s much more difficult when it comes to scale-up.

18:43

What are the actual  benefits of CPO over copper?

18:47

The benefits of CPO for scale-up start where  copper ends.

18:47

And I mean that literally.

18:54

Copper is a great material for high-speed  interconnects, but like we said at the beginning, it has a very short limit.

18:58

And we are  getting awfully close to that limit.

19:03

While current-gen 224G copper interconnects still  work with PAM4 and can reach up to 2 meters, next-gen 448G interconnects won’t have it  that easy.

19:10

PAM, or Pulse Amplitude Modulation, is a technique that allows you to transmit two  bits of data by using different voltage levels.

19:23

But for faster interconnect speeds, PAM4 isn’t  cutting it anymore.

19:23

PAM6 or maybe even PAM8 will be needed.

19:30

The problem is, using higher levels  of Pulse Amplitude Modulation creates a more unstable signal, which further reduces the reach  of copper.

19:35

Suddenly, 2 meters might be too much.

19:42

At some point, the signal to noise  ration becomes too much of a challenge.

19:46

Another angel is doubling the baud rate, basically  the signaling speed.

19:46

But in the end, it has the same problem as using higher PAM levels.

19:51

Both approaches shrink the reach of copper.

19:57

The argument for CPO for scale-up networks isn’t  really CPO itself, but it’s the limits of copper.

20:05

And that means, as long as copper scales for rack  internal communication, be it with PAM8 or beyond, copper will still be the go-to solution.

20:12

That’s  why Nvidia is still planning copper based NVLink networks for Rubin, Feymann and beyond.

20:17

But copper  can’t scale forever, the end is already in sight.

20:25

Nvidia Blackwell generation was a massive  leap for AI performance and efficiency.

20:30

And while part of that was definitely  due to the advanced GPU architecture, the real breakthrough was Nvidia’s NVL72 rack.

20:35

Before Blackwell, the scale-up domain was limited to eight GPUs.

20:42

The 9x increase to 72 GPUs was the  real performance boon.

20:42

Suddenly, 72 GPUs could act and work like a single GPU.

20:51

That’s what unlocked  the true performance advantage of Blackwell.

20:57

And it’s exactly this concept, where CPO  will be able to show its true advantage.

21:02

Once CPO is integrated at the GPU level, it will  unlock massive scale-up domains that directly translates to a larger world size.

21:09

And once  you compare a CPO based scale-up cluster with potentially thousands of GPUs to a copper-based  scale-up that might be able to connect only a few hundred GPUs, the choice will be obvious.

21:21

And that’s exactly what Jensen announced at GTC 2026.

21:27

A mixed scale-up network, that combines  copper and optical, to achieve a multiple rack world size.

21:35

Vera-Rubin Ultra NVL576 will be  the first system using this combines approach, with a total of eight Rubin Ultra NVL72 racks.

21:42

Internally, each rack still uses a copper-based scale-up network, but this time, the scale-up  network also connects up to eight NVL72 racks using optics.

21:55

And because eight times  72 is 576, Nvidia calls it 576.

22:03

And the next-gen Nvidia Kyber  NVL1152 is already on the horizon.

22:10

At this point it doesn’t matter that CPO adds  a tiny bit of latency and uses more energy.

22:15

The sheer scale of CPO scale-up will dominate  everything.

22:15

And because CPO offers a lot more scaling vectors than copper does, there’s  plenty of future proof network scaling left.

22:26

It might be time to change the principle  of “use copper when you can and optics when you must” to “use copper as long as you  can”, because the wall is approaching fast.

22:36

Co-packaged optics aren’t the holy grail,  just like any other technology they come with their own drawbacks and challenges.

22:41

For scale-out networks, the advantages are already clearly visible and we will see a  steady adoption of CPO over the coming years.

22:51

For scale-up it’s a little bit more difficult.

22:51

Copper still has a few tricks its sleeve and we will see every last bit pushed out of it.

22:56

But make  no mistake, the end of scaling is in sight and once copper has reached its literal limit, CPO  will take over scale-up in the blink of an eye.

23:09

Because no one can outcompete a larger  scale-up world size that connects more GPUs.

23:15

What we just covered in this video is just a  tiny part of the SemiAnalysis deep-dive into everything CPO.

23:20

If you want to understand how CPO  really works, I highly recommend checking it out.

23:28

And for everyone that not only wants to know  how CPO works, but where the entire networking industry is headed, including granular visibility  into the hardware, like Switches, Transceivers, Cables plus a top-down analysis of total  market conditions and vendor market shares, the SemiAnalysis AI Networking Model  offers industry leading insights in the fast-paced market.

23:49

As always, you can find the links in the video description below and  I hope I see you in the next one.