Training a 400B Model on 2,048 Blackwell GPUs for $20M | Researcher Conversations at GTC

0:07

Hi Lucas, thank you for joining us for this interview.

0:11

It's uh personally I'm a big fan of RC's work.

0:13

I'm very honored to meet you.

0:13

So yeah, can you first uh introduce uh yourself, please?

0:20

>> Yeah, my name is Lucas Atkins.

0:20

I'm the CTO and head of research over at RCAI.

0:24

You know, have worked in the open-source AI space for a long time.

0:26

you know, previously on the post-training side almost entirely and then, you know, we as RC made the move into pre-training our own models, you know, started work on it early 2025, but it didn't materialize till July 2025 and then since then we've been pre-training a lot and moving further down the stack.

0:45

So, >> first of all, some question about RC.

0:50

>> So, why a pre-trained lab or like an AGI model lab right now?

0:57

and um despite all the challenges like a competing in talent and everything >> I I've thought on that AGI lab question uh is like would I would I consider ourselves an AGI lab certainly our ambitions aren't currently AGI in that I don't think that as I go into the next release >> or planning out the next models were putting it against a road map of how is this going to help us reach AGI now certainly you know never say ever.

1:28

>> But it's really how could we be like the most economically viable company for developers, startups mainly, and then enterprises.

1:38

The why is a little more nuanced.

1:38

I think that that there's a few answers to it and sometimes the answer changes depending on who the audience is, but I I I'll try to give you, you know, the short version of each.

1:52

for us when we were largely doing just post training on top of other open models.

1:57

So Alama >> Mistrol back in the day and then later on Quen was a very popular one cuz we were trying to stay sub you know 20 billion parameters.

2:07

We reached a point where the ceiling of how well we could do on a particular uh job or a particular contract was largely dictated by how well the lab had pre-trained the base model.

2:24

>> And that was fine when you had kind of really competitive options coming out of the west, >> you know, when meta was kind of leading that frontier >> that became a little more limiting.

2:34

>> that became a little more limiting. um when they stopped and and you have most of if not all of the models sub 20B out of China and we had a lot of customers who uh stopped wanting us to work off of a a Chinese pre-trained base model which whether or not I agreed that that was

2:57

like a viable take based on the current risks didn't matter you know their their legal teams and compliance teams reigned supreme >> so that was one reason that we felt that if we were going to be able to offer the best services to customers, build the best products, um that owning the whole stack was going to be needed. There is

3:15

There is also the aspect of the West lacks a company whose sole existence relies on how good their base models are or just their pre their pre-trained models.

3:30

So you'll have like OpenAI did GBOSS which was amazing and then Nvidia has really spurred along with Neatron. >> Mhm.

3:38

>> Um but those are all sub projects of a much larger company whose revenue is not dependent on that.

3:44

And as long as that's the case, I do think that uh we're going to struggle to have like a reliable western openweight player and that it's continually kind of pushing, you know, what the frontier is in the west and or open models in general.

4:00

Um and then lastly, it's it's it enables us to do a lot more in that customization world for like downstream customers when we control the whole stack.

4:12

like when we know what data goes into it in pre-training that can very strongly inform what data you put into post- training >> um especially in the mid-training phase like I'd say that that last 10% of pre-training is some of the more consequential data for how well you do in post training.

4:30

Um, so those are the main factors that we did it as to why we did it and decided to jump in there with all of the risks and difficulties was uh we thought we could I mean we've always we made it to where we are as a company by picking like super ambitious things that we have no business trying to do and then like not stopping until we do it and we felt that pre-training was a similar uh vertical.

4:53

So yeah, those are those are the main reasons and then you know I could wax poetic all day about how I think open weight models are extremely important and that you know having alternatives outside of the frontier that you can own and have complete control of the data that goes in and out the way that it speaks and and the way that it interacts the way that you can you know interpretability research is super important there.

5:19

um that unless we have competitive base models or just like competitive open models outside of anthropic and open AI, you're never like unless you work at anthropic or open AAI, you're never going to get a chance to do interpretability research at that level.

5:35

So >> yeah, actually that's uh like my next question like why open weights?

5:38

Among all the reasons that you just briefly mentioned, could you pick maybe like one or two and explain why is it so important for us to have open weights?

5:47

I mean, if I'd pick one, it would be, you know, sovereignty.

5:53

I think the ability for someone to completely own the pipeline is extremely important.

5:59

I think that enterprise aside, which like is definitely the most strict on compliance and like data sovereignty, the ability to be in complete control of your product and development process is extremely important.

6:14

So, if you're building, let's say, an invoice processing app, right?

6:18

You're going to take in parsed uh receipts and then you want the model to then take all of those uh you know, charges and the the the prices associated with them and put them in an Excel sheet and then you want it to calculate kind of like monthly spend or something uh or you know, set it aside for your accounting team to do taxes on later.

6:42

If you reach a point where margins are too low, meaning like you're actually losing money, uh, or you have no margins because you're, you know, the the models coming out of those labs are so expensive. Yeah.

6:55

>> Yet improvement on your task is not increasing over time.

6:58

So the models are getting more expensive, they're taking longer, they're doing more tokens for the same task, but they're scaling like frontier mathematics, but you need them to scale like tool use, speed, cost.

7:12

um the ability for you to then take control and say, "No, actually, I'm going to take this four to eight billion parameter model.

7:17

I'm going to, you know, collect a bunch of my data.

7:20

I'm going to format it in a way that is conducive to fine-tuning and then I'm going to fine-tune it.

7:26

And then when I run into an issue in the future, say I want to build a new feature that requires a new skill set, I'm going to go and and make sure that my model's good at that, too."

7:34

I think it gives you uh a lot more control than you have when you're kind of reliant on the whims of you know uh these kind of frontier labs which ultimately are in a you know arms race for profitability and so they're going to make sacrifices or choices that lean towards that which aren't always in the best interest of consumers, developers, startups and then ultimately enterprises.

7:59

So I think control is the number one.

8:02

I think the second is from a research perspective.

8:07

Anything with kind of an unlimited cap on the potential for abundance and good which I think AI has necessarily kind of has the opposite potential as well, right?

8:19

Like it it you know it could be uh there could be bad side effects.

8:24

side effects. I think an example of that is like the mobile boom and phones and social media where like it really had this upside of like connecting the world >> in a way that's unprecedented but at the same time we're seeing the tail effects of that which is like reduced attention span um you know depression and anxiety

8:45

at like the highest levels and I think it's a similar thing that can happen with AI there's obviously going to be negative side effects and if the people who are able to monitor and see at a fundamental level and what those risks are and to analyze and try to mitigate them are only 200 safety researchers at OpenAI Anthropic. I think that

9:05

I think that uh from a transparency perspective, a trust perspective and from an effectiveness at combating those risks perspective that open models are unbelievably important.

9:17

I would rather the risks be very transparent.

9:19

I would rather the mitigations be extremely easy and and well, you know, um >> well understood >> understood and and and uh properly um what do you you know properly >> maybe like diffused. >> Yeah.

9:39

like the ability is, you know, easier for everyone to implement and then I would like the smartest minds in the world able to analyze those and I think that having that out in the open is extremely important. >> Yeah, totally.

9:52

I think uh open- source community is great evidence of how um we can do uh greater good by collaborating by being transparent with each other and uh I think the early progress of AI al definitely also benefited from oh yeah open source um if uh AlexNet they don't open source anything or >> um if pietorch isn't open source I I don't know where we will be >> that's scary idea yeah where where we are because of open weights and open research.

10:25

I mean open weights aside like open models aside, open research in academia and then you know having that then applied at a lab level and then I mean all of we are only where we are because stuff was done in the open and and uh done for other people to judge and and build on top of.

10:45

Um, and I think that by commercializing everything to the extent that we are, um, and having everything be this arms race of capital and talent especially, but compute as well, uh, you run the risk of of not having those same compounding breakthroughs, right? Mhm.

11:09

>> You know, if if OpenAI is, you know, able to solve continu, you know, they have some path to continuous learning and Anthropic has their own and Google has their own, um, it's going to, you know, saturate to the wider world much slower than if it were to happen, you know, in in the open with artifacts for others to build on top of.

11:30

Now, >> there's no going back.

11:33

I mean there's no world in which um no matter how much you know whining and and advocating is done there's no world in which these players are going to open up their research to the degree that they once did.

11:50

>> But not allowing us to ignore it entirely in the open world and for other players to be pushing for it I think is is going to be extremely important.

12:00

>> Definitely not going back.

12:00

Uh they definitely have some uh um from a company's perspective they need to guard their secrets and everything. >> Yeah. >> Uh yeah.

12:09

Another question is um so how do you how do you plan to compete with uh the west um AI labs in resources uh i. e.

12:21

compute and also more uh importantly actually talent.

12:27

Talent's actually easier than compute because compute comes down to funding which comes down to money which comes down to you know revenue which is um if if it's not an easy problem to solve.

12:43

But talent is um easier in that uh not that it's easy to find the talent but retaining it and and and you know convincing someone to join a lab of of our size is been surprisingly um not as hard as I thought it was going to be.

13:02

Well, and I think there's a few reasons for that.

13:04

The first is that people want to work on open weights.

13:08

Like people want to work on models that are going to be out in the public where people can appreciate the research that went into it, the engineering, the infrastructure >> which um at you know labs where they're serving it via APIs or downstream products, we only see the symptoms of their work instead of the actual like work itself, >> right?

13:34

Like we see how quickly OpenAI and Anthropic are improving their models monthtomonth >> and you can infer that wow they have some amazing infrastructure engineers and wow their researchers have really built this amazing feedback loop >> because we can infer that that's what it takes to have that kind of progress at that kind of speed.

13:55

>> But we don't actually get to see it.

13:55

We don't actually get to appreciate the elegance that goes into it.

13:58

We can only um again kind of intuit it based on the what needed to go into it in order to have that outcome.

14:07

>> Um, >> and then also when we're small, >> uh, and I think regardless of if we went out and raised or somehow got access to$2 trillion dollars, I think I would still like to keep our research team under 30 people.

14:20

I think that the way that you make breakthroughs and momentum be a side effect of the way you work is by having really opinionated, talented, and um opinionated people >> uh in a room debating in good faith.

14:35

I think that is the way that you make these things become I think you make breakthroughs and the ability to solve these hard problems that you have no business trying to solve uh possible and consistent.

14:48

And and by that I mean is If I have 30 researchers who are all involved in every step of the training process >> instead of just one specific niche within a much larger training loop like you would get at the labs, right?

15:01

Like you know there might be one or two or three people who are aware of and involved in every step from pre- you know data collection and synthesis to pre-training to architecture to mid-training to SFT to you know but you not everyone isn't and therefore it can be hard to with that lack of context to make the best decision for your you know niche.

15:28

Um whereas at RC and other you know labs of our size, everyone's working on every single step of the process.

15:35

Uh which makes this very you know malleable.

15:38

I can have five people working on pre-training that then I need three of them to move over to mid-training because you know that team is struggling with a certain tricky challenge.

15:48

I can move them around and they're able to jump in without having to like go through a big >> onboarding process.

15:55

>> Um and people like that.

15:55

I think that you know for people who are intellectually and uh curious and unbelievably like love solving really hard problems being able to understand the whole the whole process is extremely enticing.

16:10

So the talent part is is one that I'm less concerned about.

16:12

Um now when it comes to compute and if you know you want to use compute as an aggregate for intelligence right like the smarter a model is likely the more compute went into it.

16:26

I think that framing openw weight models or just open source in general as needing to match or exceed the capability of these frontier labs is the wrong viewpoint.

16:43

>> You hear often like, oh, Chinese labs are six months behind or they're four months or three months. >> Yeah.

16:49

>> And that kind of implies that they're catching up or that they're trying to, you know, exceed.

16:53

And maybe in some cases they are.

16:54

Um, but I think that just over time the the distance between what Frontier Labs are focusing on and what open model labs are focusing on is going to be very different. >> Mhm.

17:09

>> Um, again, for instance, I'm not currently developing a team and a model to do like theorem proving or, you know, frontier math, >> right?

17:20

Um but opening I is they're hill climbing that very heavily.

17:28

>> Um so while their next model might go up 30% on frontier math the ability for it to generate front end and therefore like slide generation for you know apps that are building like PowerPoint um might that percentage might only go up 5. 5%. >> Right.

17:46

Um, and so I think that our goal at RC is to be as reliable as possible, as fast, as inexpensive as possible, and as performant in economically viable uh the 80% of economically viable tasks that don't require that level of insane uh kind of frontier intelligence.

18:06

>> So therefore, our goal isn't to catch up.

18:08

Our goal is to undercut, I guess, if that's a better way to put it.

18:11

So therefore it's a just a a completely different you know compute spend and and allocation than what they're doing.

18:20

>> Um so I think that you know raising the funds and being able to execute down for the next couple years is entirely feasible and well within motion.

18:33

motion. What becomes tricky is how do we navigate the kind of es and flows of this industry and and ensure that as compute demand goes up and down and the compute required to do any kind of new training paradigm shifts right is significantly more in you know compute and infrastructure intensive >> um than what post- trainining used to be

18:55

um you can expect that that those kind of paradigm shifts are going to continue to happen over the next months and years >> that we have a loose enough and modular enough kind of research structure that we can kind of move with those and ensure that we're not, you know, spending 10 times more than what we make. >> Mhm. I think uh RC has um the Trinity >> Mhm.

19:12

I think uh RC has um the Trinity series models.

19:16

So uh first of all please uh let the people who still don't know what Trinity stands for appreciate to know why it's called Trinity and then uh maybe talk about what is it like to uh work with the other companies and um how is it like to train the first model with uh uh big a big cluster of B300s.

19:36

>> Yeah, Trinity was named first and foremost because I thought it was a cool name.

19:40

Um it ended up being there nice reasons for it and that we're you know we work with a company called Daytology for helping us curate our pre-training data.

19:50

We work with prime intellect on infrastructure scale up and kind of our GPU management and um having three companies involved in making one model felt apt for that name and we were also making three models in the first generation.

20:03

So >> there's several different, you know, reasons for the name, but yeah, Trinity is is mainly just like, you know, three, you know, uh, three companies all focusing on different areas of the stack working in tandem to build something special that we think is better than one company of our size trying to own it all.

20:20

>> The B300's were uh, interesting.

20:23

It was it was a a a decision made out of a practicality and that they were available and we wanted to train as fast as possible.

20:31

We wanted this pre-training to take a month, not three.

20:37

And so B300's became the next, you know, the obvious choice.

20:42

It was definitely intimidating because there wasn't a lot of outside of what Nvidia had shared themselves or um PyTorch benchmarking.

20:55

there wasn't a lot of like atcale B300 benchmarks but also like tools to take best advantage of them.

21:02

Kernels weren't readily available especially like super sparse kernels like what we have.

21:07

You know up until that point we had largely been utilizing and building off of the work that had come out of other labs that had done uh models and shared their their artifacts and how they' done that.

21:21

Like a lot of it was, you know, this the ecosystem that built around the deepseeek uh like super sparse >> helped us on Hopper because a lot of those models were trained on Hopper.

21:30

So we could use that as like a a golden example of what throughput looks like for a model of this size and this scale.

21:40

>> Um and we could use that as a benchmark for how well we were doing in our own stack.

21:43

Um you know based on how how well others had trained models.

21:46

Thank you for letting us uh interview you and uh looking forward to uh the new uh maybe post-trained um Trinity models and um yeah, thank you again. >> I appreciate it. Thank you.