We Asked AI and Humans the Same Question — Here's The Result

0:00

You're watching TVP.

0:03

>> Sam Alman says something interesting here.

0:05

He says, "I don't want to quibble on the exact definition of the touring test.

0:08

By the popular conception, we kind of have passed it."

0:12

And most people in 2020 did not think that was going to happen.

0:16

So now, if in the next 5 years we really we really deliver on AI, there's discovery and important new science and we were hyping our progress and our excitement a little bit too much.

0:26

I think we deserve some grace there, but it's annoying and we should do less.

0:29

So he says he says by the popular conception we have passed the touring test.

0:33

And so >> we were having a little bit of a debate. >> Yeah.

0:38

>> A couple nights ago talking about >> have we have we and so and so I ran I ran the touring test and so I asked chat GPT GPT5. >> We haven't.

0:49

>> Do you think we passed the touring test?

0:50

And I asked Tyler the same question.

0:50

And so one of these answers is from Tyler the human and the other is from chat GPT.

0:57

So um I'll give you answer A and then I'll give you answer B and you'll have to detect which one is AI. Is A or B AI?

1:08

>> So answer A to the question, do you think we passed the touring test?

1:10

That's a fascinating question and one that doesn't doesn't have a single yes or no answer because it depends on how you define passing the touring test or whether you think that test is still a useful benchmark.

1:24

But I'll walk you through the terrain then give you my verdict.

1:28

What is the touring test and what does passing mean?

1:30

Alan Turing proposed in 1950 that instead of trying to define thinking, we ask, can a machine imitate human conversational behavior well enough that a human evaluator can't reliably tell the machine from a human in a textual dialogue?

1:44

Over time, people have realized many problems.

1:46

The test doesn't measure understanding, consciousness, creativity, or reasoning depth, only imitation of conversation behavior and emotional stylistic cues. So that is answer A.

1:56

And then answer B is I asked uh do you think we passed the touring test? And the answer yes.

2:03

We did that a long time ago.

2:08

So which one do you think is AI?

2:08

Which one do you think is Tyler?

2:13

>> Wait, what was the Sorry.

2:13

The first >> the first one was that's a fascinating question and one that doesn't have a single yes or no answer because it depends on how you define passing the touring test or whether and then the second one is just yes.

2:23

I asked Tyler, do you think we passed the touring test? Yes, of course.

2:27

Uh, and it was just so so funny to hit him with that.

2:29

And, uh, it does it does clearly illustrate like the difference in how the that even though you can't tell the it it's speaking English, well, it still has like a very specific style to it.

2:41

Uh, whereas Tyler just says like, "Yeah, yeah, we passed it."

2:46

>> But in doing so, Yeah.

2:46

What's your reaction to this?

2:48

>> Okay, Tyler, you should have answered uh what what Bobby in the chat is saying. You're absolutely right. >> Just >> Yeah. >> I mean, yeah.

2:56

So like obviously chatbt has like a very specific style of speaking.

3:01

>> Um >> you can see um you know there's some uh it's like pokey uh it like they train they train the model or poke is that what it's called? >> Oh. Oh the the AI app. com.

3:14

>> You can train the model to speak with a different style. Yeah.

3:17

>> And then I think it's much harder to tell.

3:18

Um, like you I I think even if I just if you just prompted the model to say answer it very succinctly, very concisely, um, I think it'd be much harder to tell.

3:28

>> But I didn't have to prompt you for that.

3:29

I didn't have to tell you answer succinctly.

3:31

You just did because you're confident about that.

3:34

>> Well, sure, but it's like um, how do you define like prompting the AI?

3:39

If it's part of the system prompt, is that like part of the prompt or is that part of the model?

3:43

If you like bake it into the weights, then like where where does the where's the line, right?

3:47

does the where's the line, right? like >> um if you have one of those like poke model or interaction like all these things um >> like does it really matter if it's in the prompt or if it's like a RL step at the end like you know >> I don't know I I I think there's just like as I look through our text back and

4:05

forth there are sometimes when you drop a whole paragraph sometimes you ask a question sometimes you feed it back to me and it feels very different from the back and forth with GPT5 but it it's like yeah it does I think I think one correct if you showed these outputs to two people at a Walmart super center in Nebraska. >> Sure. >> Sure.

4:26

>> How many people would clock it >> or how many people would prefer the nuanced answer that GPT5 gave because a lot of people would I mean you didn't tell me about when Allan Touring proposed the touring test.

4:37

You didn't give me any backstory.

4:38

You just you just you just ripped an answer.

4:41

Some people might prefer that chatbot would give this like sort of short to the point answer without a lot of nuance, >> but I think it's I I think it is clear that we we did pass one definition of the touring test.

4:54

There's still something else going on.

4:56

And it's a little bit uh it's it's obviously a nuance question, but but it does uh but I think the point holds that that OpenAI has uh underhyped a few key things that have just like blown everyone away.

5:09

And people have been very very impressed by that.

5:13

And so you do have to give them a few uh you have to give them more credit.

5:17

Like you have to give them the benefit of the doubt on a lot of these things um because they've made so much progress.