searchlore

Back to Resource

All Segments

OpenAI vs. Deepseek vs. Qwen: Comparing Open Source LLM Architectures

OpenAI vs. Deepseek vs. Qwen: Comparing Open Source LLM Architectures

15 segments available

OpenAI recently released its first open-weights model since GPT-2, entering a field led by DeepSeek and Alibaba's Qwen. YC's Ankit Gupta breaks down everything you need to know about these top OSS models, including what sets them apart under the hood. He’ll compare their approaches to mixture-of-experts, long-context training, and post-training techniques that shape reasoning and alignment—and explore how different design choices lead to surprisingly similar performance. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs 00:00 – OpenAI OSS Launch 01:00 – Comparing Open Source LLM Architectures 01:46 – GPT OSS Overview 02:37 – Under The Hood of GPT OSS 03:25 – Qwen-3 Architecture 04:17 – Qwen-3 Training 05:12 – Qwen-3 Post-Training 06:08 – Qwen-3 Reasoning & RL Innovations 06:52 – DeepSeek V3 Overview 07:40 – DeepSeek V3.1 Updates 08:39 – Attention Mechanism (MLA) 09:39 – Comparing Model Sizes 10:35 – Long Context Strategies 11:25 – Reflections on Methods 12:00 – Takeaways

Segments Timeline

1
0:00 - 1:02
1:02 duration176 words

OpenAI OSS Launch

"OpenAI recently dropped GPT OSS, its first open weights model since GPT2 in 2019. It's one of the highest profile open source model launches since DeepSeek R1 made waves back in January. But how does ..."

2
1:02 - 1:50
0:47 duration147 words

Comparing Open Source LLM Architectures

"share the same key value pairs to reduce memory use and speed up inference. It also includes swiggloo activations in the feed forward network layers which allow for more nuance transformations than si..."

3
1:50 - 2:38
0:48 duration156 words

GPT OSS Overview

"builds on the O200K tokenizer used in models like GPT40. As for the data set GPT OSS was trained on, OpenAI has only disclosed the broad strokes. The model was trained on a textonly corpus in the tril..."

4
2:38 - 3:29
0:51 duration161 words

Under The Hood of GPT OSS

"of open source AI, GPToss arrives as a fully equipped long context model ready for immediate use. As impressive as it is, however, it's just one of several models in a rapidly expanding field of open ..."

5
3:29 - 4:18
0:48 duration152 words

Qwen-3 Architecture

"3 incorporates features like group query attention, swiggloo, rope, and RMS norm. Quinn 3's sparse models share the same fundamental architecture as its dense models, but add a mixture of experts laye..."

6
4:18 - 5:12
0:54 duration169 words

Qwen-3 Training

"3 was trained on 36 trillion pre-training tokens, twice as many as the Quen 2.5 models. In addition to pulling data from multilingual texts, STEM and coding sources, and reasoning tasks, Quen 3 also u..."

7
5:12 - 6:09
0:56 duration174 words

Qwen-3 Post-Training

"Together, all of these optimizations allow the model to reason over much longer inputs at inference. Finally, Quen uses a four-step post-training pipeline with two goals. Giving users more control ove..."

8
6:09 - 6:53
0:44 duration150 words

Qwen-3 Reasoning & RL Innovations

"what developers did in this step was fine-tune the model on a mix of thinking data, which includes intermediate reasoning steps, and non-thinking data, which omits them, and then build a chat interfac..."

9
6:53 - 7:43
0:49 duration159 words

DeepSeek V3 Overview

">> The chatbot developed in China called Deep Seek. >> Deepseek is such a fundamental change to the economics of what's going on. >> The most downloaded free app in the US. This is an update in what p..."

10
7:43 - 8:40
0:57 duration184 words

DeepSeek V3.1 Updates

"on the original V3based checkpoint, extending it with a two-phase long context training approach and adding a hybrid thinking mode that lets the same model switch between reasoning heavy and lightweig..."

11
8:40 - 9:41
1:01 duration212 words

Attention Mechanism (MLA)

"let's take a step back from V3 to Quen to GPDoss. How should we think about at a high level the differences between these models? One big difference is size. The Quen 3 model family is the only one of..."

12
9:41 - 10:35
0:53 duration162 words

Comparing Model Sizes

"rotary positional embeddings so that it can handle far longer sequences than it was originally trained on. Normally rope starts to break down when you feed it more tokens than its base frequency was s..."

13
10:35 - 11:27
0:51 duration184 words

Long Context Strategies

"train model can do without more long context training. Personally, I think one of the most interesting things about these papers and the state-of-the-art in deep learning more generally is that a lot ..."

14
11:27 - 12:03
0:35 duration134 words

Reflections on Methods

"training efforts. And it's fascinating and pretty surprising how some of these RL efforts require very little amounts of data. just 4,000 data pairs in the case of Quinn. Another point here is that it..."

15
12:03 - 12:23
0:20 duration79 words

Takeaways

"are tons of high performing open source models that we didn't discuss in this video, like Kim K2 or Google Gemma 3. But when you peek under the hood of many of these, you'll find nuance differences th..."