
93 segments available
Noam Brown is a research scientist at FAIR, Meta AI, co-creator of AI that achieved superhuman level performance in games of No-Limit Texas Hold'em and Diplomacy. Please support this podcast by checking out our sponsors: - True Classic Tees: https://trueclassictees.com/lex and use code LEX to get 25% off - Audible: https://audible.com/lex to get 30-day free trial - InsideTracker: https://insidetracker.com/lex to get 20% off - ExpressVPN: https://expressvpn.com/lexpod to get 3 months free EPISODE LINKS: Noam's Twitter: https://twitter.com/polynoamial Noam's LinkedIn: https://www.linkedin.com/in/noam-brown-8b785b62/ webDiplomacy: https://webdiplomacy.net/ Noam's papers: Superhuman AI for multiplayer poker: https://par.nsf.gov/servlets/purl/10119653 Superhuman AI for heads-up no-limit poker: https://par.nsf.gov/servlets/purl/10077416 Human-level play in the game of Diplomacy: https://www.science.org/doi/10.1126/science.ade9097 PODCAST INFO: Podcast website: https://lexfridman.com/podcast Apple Podcasts: https://apple.co/2lwqZIr Spotify: https://spoti.fi/2nEwCF8 RSS: https://lexfridman.com/feed/podcast/ Full episodes playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 Clips playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOeciFP3CBCIEElOJeitOr41 OUTLINE: 0:00 - Introduction 1:09 - No Limit Texas Hold 'em 5:02 - Solving poker 18:12 - Poker vs Chess 24:50 - AI playing poker 58:18 - Heads-up vs Multi-way poker 1:09:08 - Greatest poker player of all time 1:12:42 - Diplomacy game 1:22:33 - AI negotiating with humans 2:04:58 - AI in geopolitics 2:09:43 - Human-like AI for games 2:15:44 - Ethics of AI 2:19:57 - AGI 2:23:57 - Advice to beginners SOCIAL: - Twitter: https://twitter.com/lexfridman - LinkedIn: https://www.linkedin.com/in/lexfridman - Facebook: https://www.facebook.com/lexfridman - Instagram: https://www.instagram.com/lexfridman - Medium: https://medium.com/@lexfridman - Reddit: https://reddit.com/r/lexfridman - Support on Patreon: https://www.patreon.com/lexfridman
Noam Brown discusses the skepticism surrounding Game Theory in poker, contrasting it with the belief that reading opponents' emotions is key to success. He shares insights from a match where an AI bot, using Nash equilibrium strategies, outperformed top human players without trying to exploit their weaknesses.
"a lot of people were saying like oh this whole idea of Game Theory it's just nonsense and if you really want to make money you got to like look into the other person's eyes and read their soul and fig..."
Noam Brown outlines his contributions to three groundbreaking AI systems: Libratus, Pluribus, and Cicero. Each system achieved superhuman performance in poker and the game of Diplomacy, showcasing the advancements in AI's strategic capabilities in complex games.
"you've been a lead on three amazing AI projects so we've got libratus that solved or at least achieved human level performance on No Limit Texas Hold'em poker with two players heads up you got pluribu..."
Brown explains the fundamentals of No Limit Texas Hold'em, the most popular poker variant. He highlights its differences from other poker types, particularly the absence of betting limits, which leads to rapid escalation of stakes and strategic complexity.
"hold 'em poker is the most popular variant of Poker in the world so you know you go to a casino you play sit down at the poker table the game that you're playing is no limit Texas Hold'em if you watch..."
In this segment, Noam Brown discusses the psychological aspects of playing No Limit Texas Hold'em. He describes how the potential for large bets can make players 'jumpy' and affect their decision-making, emphasizing the importance of strategy in high-stakes situations.
"other ones rewards more kind of calculated strategy or or no because you're sort of looking from an from an analytic perspective is is strategy also rewarded in No Limit taxes hold on I think both var..."
Noam shares his passion for poker, explaining how the game can be solved through strategic play. He reflects on the allure of discovering the optimal way to play, comparing it to solving chess and other games, and the thrill of mastering such a complex challenge.
"between No Limit and limit what about on the action side when you're actually making that big bet that's what I mean by crazy I was I was trying to refer to the technical the the technical term of cra..."
Brown introduces the concept of Nash equilibrium, explaining its significance in finite two-player zero-sum games like poker. He illustrates how optimal strategies can ensure players do not lose in expectation, using relatable examples like rock-paper-scissors.
"that play these games we'll talk about poker we'll talk about diplomacy are you um are you drawn in in part by the beauty of the game itself AI aside or is it to you primarily a fascinating problem se..."
In this segment, Noam discusses the high variance nature of poker and what it means to play in expectation. He clarifies that while players may not win every hand, a long-term strategy can lead to breaking even or profit, emphasizing the theoretical aspects of poker.
"this gets into the concept of an ash equilibrium yeah so it is a Nash equilibrium Okay so and any finite two-player zero-sum game there is an optimal strategy that if you play it you are guaranteed to..."
Noam elaborates on the constraints of zero-sum games and the existence of Nash equilibria in various game types. He explains how these concepts apply to poker and other strategic games, highlighting the complexities involved in multi-player scenarios.
"your stack generally speaking now that doesn't include anything about the fact that you can go broke it doesn't include any of those kinds of normal real world limitations you're talking you know in t..."
Brown discusses the human elements in poker, such as the emotional and psychological aspects that influence gameplay. He suggests that understanding these factors could enhance AI systems, making them not only effective players but also engaging opponents.
"some games there's no guarantee that you're going to win by playing a national equilibrium have you ever tried to model in the other aspects of the game which is like the pleasure you draw from playin..."
In this segment, Noam reflects on the difference between creating an AI that wins and one that is enjoyable to play against. He emphasizes the importance of personality and engagement in AI design, particularly in recreational games.
"fun to play with and fun to watch yeah and I think you know I've I've heard uh talks from like game designers and and they say like you know people that work on AI for actual recreational games that p..."
Noam and Lex discuss the potential of large language models to revolutionize NPC interactions in video games. They explore how this technology could lead to more dynamic and engaging gameplay experiences, moving beyond traditional scripted dialogues.
"they're just releasing the star feel good they do one game at a time yeah and so uh whatever it is whenever the date is I don't know what the data is calm down uh but it would be I don't know like uh ..."
Brown explains the process of finding Nash equilibrium through self-play in AI systems. He describes how AI learns from counterfactual reasoning, simulating games against itself to optimize decision-making and converge on optimal strategies.
"crazy stuff AI does we have some flexibility to play just like with the game of diplomacy it's a game this is not real geopolitics not real war it's a it's a game so you could you can have a little bi..."
In this final segment, Noam highlights the immense complexity of poker, particularly No Limit Texas Hold'em, which has an astronomical number of decision points. He discusses the challenges this presents for AI and the strategies used to navigate such a vast game space.
"it regret having not played that action in the past and when it encounters that same situation again it's going to pick actions that have higher regret with higher probability now it'll just keep simu..."
In this segment, Noam Brown compares the complexities of poker and chess, arguing that poker is more challenging due to its imperfect information aspect. He highlights how hidden information affects decision-making and introduces the concept of bluffing as a strategic element unique to poker.
"uh it's a simpler setting sure so I kind of in my brain the word self-play has mapped in you all networks but we're speaking something bigger than just neural networks it could be anything the self-pl..."
Noam explains how neural networks enhance poker strategies by allowing players to generalize from similar situations rather than relying on exact scenarios. He discusses the vast number of decision points in no-limit Texas Hold'em and how AI can adapt to various situations using learned experiences.
"neural nets come in I said like okay if it's in that situation again then it will choose the action that has high regret now the problem is that poker is such a huge game you know I think no limit Tex..."
This segment delves into the hidden information in Texas Hold'em, where players hold two face-down cards. Noam Brown explains how players must reason about their opponents' hands and the implications of bluffing, emphasizing the strategic depth that arises from imperfect information.
"complex game chess or poker or go or poker do you know that is a controversial question okay um I'm gonna it's like somebody screaming on Reddit right now it depends on which subreddit you're on is it..."
Noam discusses the importance of probability in poker, particularly in estimating opponents' hands and making strategic decisions. He uses rock-paper-scissors as an analogy to illustrate how players must balance their actions to remain unpredictable and maximize their chances of winning.
"they think I have what do they have what do they think I think they have that kind of stuff and um that's that's kind of where bluffing comes into play right because the fact that you can Bluff the fa..."
In this segment, Noam Brown explains how a player's reputation affects the value of bluffing in poker. He discusses the balance between bluffing and playing strong hands, highlighting the psychological aspects of poker that influence decision-making and outcomes.
"how what are the different approaches to the imperfect information game so the key thing to understand about why in perfect information makes things difficult is that you have to worry not just about ..."
Noam contrasts Game Theory Optimal (GTO) play with exploitative strategies in poker. He discusses how expert players and AI aim to approximate Nash equilibrium to remain competitive, emphasizing the importance of balancing strategies to avoid being predictable.
"now you take that to Poker what that means is the value of bluffing for example if you're the kind of person that never Bluffs and you have this reputation as somebody that never Bluffs and suddenly y..."
This segment explores the dynamic interplay between playing cards and reading opponents in poker. Noam discusses the ongoing debate in the poker community regarding the effectiveness of GTO strategies versus adapting to opponents' behaviors.
"the balance of how often in the key sort of branching is the bluff or not the bluff is that a is that a good crude simplification of the major decision in poker it's a good simplification I think that..."
Noam Brown elaborates on the psychological aspects of poker, including how players can manipulate their opponents' perceptions. He discusses the importance of understanding common knowledge beliefs and how expert players navigate the complexities of human behavior in poker.
"to think about what would I do if I had this different set of cards is there explicit estimation of like a theory of mind that the other person has about you or is that just a emergent thing that happ..."
In this segment, Noam discusses Phil Hellmuth's unconventional playing style and its effectiveness. He analyzes how Hellmuth's chaotic approach can confuse opponents and lead to success, highlighting the balance between unpredictability and strategic play.
"start by playing the Nash equilibrium and then maybe if they spot weaknesses in the way you're playing then they can deviate a little bit to take advantage of that they aim to be unbeatable in expecta..."
Noam shares insights from the 2017 competition where AI faced top professional poker players. He discusses the advancements in AI strategies and how they have evolved to approximate Nash equilibrium, leading to significant victories against human opponents.
"deviate from a Nash equilibrium style of play to try to take advantage of those perceived weaknesses and then counter exploit them so you kind of get into the Mind Games there so you think about these..."
Noam reflects on the evolution of poker AI from earlier competitions to the present. He explains the shift from pre-computed strategies to real-time adaptive algorithms that enhance decision-making and improve performance against human players.
"I think you know it we've we're playing for 50 100 blinds and over the course of about 120 000 hands it made close to two million dollars 120 000 hands 120 000 hands against humans yeah and this was t..."
In this segment, Noam discusses the search space in poker, focusing on the actions players can take and the probabilities associated with those actions. He explains how AI algorithms aim to maximize unpredictability and create challenging situations for opponents.
"money at stake where it would basically be divided among them depending on how well they did relative to each other so we wanted to have some incentive for them to play their best did you have a confi..."
Noam concludes by discussing the ultimate goal of poker strategies: maximizing expected value. He explains how players must consider their opponents' optimal responses and the importance of maintaining a balanced approach to achieve success in the game.
"in a game like chess the the search is like okay I'm in this chess position and I can like you know move these different pieces and see where things end up in poker what you're searching over is the a..."
Noam shares insights from a 2015 competition where their AI was defeated. He contrasts the human approach to poker with that of the bot, noting how humans take time to think through decisions, while the bot acted instantly. This segment explores the lessons learned and the evolution of AI strategies in poker.
"Nash equilibrium approach because if they're not then you're just going to make more money right like anything that deviates like by definition the national equilibrium is the strategy that does the b..."
Noam discusses the significance of search algorithms in enhancing AI performance in poker. He explains how incorporating search can dramatically improve decision-making, likening it to expanding the neural network's capabilities. This segment underscores the critical role of computational resources in developing effective poker strategies.
"lot from that competition and in particular you know what became clear to me is that the way the humans were approaching the game was very different from how the bot was approaching the game the bot w..."
In this segment, Noam Brown breaks down the concept of hand combinations in No-Limit Texas Hold'em. He explains how the AI evaluates potential hands during gameplay, emphasizing the complexity of decision-making based on the numerous possible combinations. This discussion highlights the intricacies of poker strategy and AI's analytical capabilities.
"outcome and we did some experiments in small scale versions of Poker and what we what we found was that if you do a little bit of search even just a little bit it was the equivalent of making your you..."
Noam explains how the AI integrates neural networks with search algorithms to enhance its decision-making process. He discusses the balance between pre-computed strategies and real-time adjustments, showcasing the AI's ability to adapt and optimize its gameplay. This segment reveals the sophisticated mechanics behind AI poker strategies.
"hands Okay so that search how do you fuse that with what the neural net is telling you or what the the the train system is telling you yeah so you kind of like where the train system comes in is is th..."
Noam reflects on the evolution of AI strategies in poker, particularly how the bot's ability to search and adapt has changed over time. He discusses the importance of searching multiple moves ahead and how this capability has led to significant improvements in AI performance. This segment highlights the ongoing advancements in AI technology.
"all these other hands as well okay but you are you in the search always going to the end of the game in liberatis we did uh so we only use search starting on the turn and then we searched all the way ..."
In this segment, Noam discusses the concept of overbetting in poker, a strategy that emerged from AI play. He explains how the AI's ability to place large bets created difficult decisions for human players, leading to a shift in high-level poker strategies. This discussion illustrates the impact of AI on traditional poker gameplay.
"the case of liberatis so for liberatis what's the the most number of re-racists have you ever seen uh you probably cap out at like five or something because at that point you're basically all in you k..."
Noam shares insights on the future of poker in light of AI advancements. He discusses how top human players are adapting to AI strategies and the ongoing debate about the capabilities of AI in various poker variants. This segment explores the evolving relationship between human intuition and AI precision in the game.
"second best hand like now you get a really tough choice to make and so the humans would sometimes think like five or ten minutes about like what do you do should I call should I fold and um and when I..."
Noam emphasizes the critical role of search in game AI, comparing it to historical AI achievements like Deep Blue and AlphaGo. He explains how search enhances decision-making and the importance of planning ahead in complex games. This segment highlights the foundational principles that drive AI success in strategic gameplay.
"um so he wasn't scared he was excited he was excited and uh and he all he honestly he wanted to play against the bot he thought he thought he had a decent chance of beating it um I I think he's you kn..."
In this segment, Noam discusses the differences between human intuition and AI calculation in strategic games. He explores how humans leverage their experience and intuition while AI relies on computational power and search algorithms. This discussion sheds light on the unique strengths of both humans and AI in competitive environments.
"um we've focused on the most popular variants so heads up no limit Texas Hold'em and then we followed it up with um with uh six player poker as well where we managed to uh make a bot that beat expert ..."
Noam Brown discusses the critical role of search algorithms in AI performance, particularly in games like Go and poker. He explains how without search techniques, AI systems struggle to achieve superhuman levels, emphasizing the difference between human intuition and AI's computational methods.
"moves ahead and you see like what the board state looks like um that's what I mean by search if you take out the search that's done during the game the ELO rating drops to around three thousand so eve..."
In this segment, Brown contrasts human decision-making in games with AI's search capabilities. He highlights how humans utilize intuition and experience, while AI relies on structured search methods like Monte Carlo tree search, which may not fully replicate human reasoning.
"network is doing the searching and I wonder what the human brain is doing in terms of searching because you're doing that like computation the human is Computing they have intuition they've got they h..."
Noam Brown elaborates on the unique challenges poker presents for AI, particularly due to hidden information and the need for strategic betting. He explains how traditional search methods fail in poker, necessitating a different approach to AI development in this domain.
"think it's a really important missing piece the ability to plan and reason more generally across a wide variety of different settings in a way where the general reasoning makes you better at each one ..."
Brown shares insights into the development of Liberatus, the AI poker bot that competed against top human players. He discusses the programming challenges, resource allocation, and the innovative techniques used to enhance the bot's performance in a competitive environment.
"kinds of of planning that we could do so when the broadest actually beat the poker plays what did I feel like what was that I mean actually on that day what were you feeling like were you were you ner..."
In this segment, Brown reflects on the strategies employed by human players to exploit weaknesses in the AI poker bot. He discusses the importance of understanding betting patterns and how the bot's limitations were revealed during the competition.
"but still a lot for even any gratitude today it's still tough to to get or even to allow yourself to think in that in terms of scale at CMU at MIT anything like that yeah and you know talking about te..."
Brown describes the dynamics of the competition between humans and the AI bot, emphasizing the collaborative strategies humans used to identify and exploit the bot's weaknesses. He shares the psychological aspects of the competition and the stress involved in the process.
"that I you know had to make uh just a fun question what what id did you use what uh for for C plus plus I think I used a visual studio actually yeah okay yeah is that still carried through to today vs..."
This segment covers the decision to provide human competitors with detailed logs of the bot's hands during the competition. Brown discusses the implications of this transparency and how it influenced the strategies employed by human players.
"be able to distinguish you know like having a king High flush versus an ace high flush and in some situations that really matters a lot and so they could put the bot into those situations and then the..."
Brown reflects on the emotional highs and lows experienced during the competition, sharing his thoughts on the significance of the AI's success in poker. He discusses the culmination of years of work and the impact of achieving a milestone in AI development.
"and so then they would review the hands and try to see like okay could they find patterns in the bot the weaknesses and could they then then they would coordinate and study together and try to figure ..."
In this segment, Brown discusses the challenges of extending AI strategies from heads-up poker to multi-way games. He explains the complexities involved and how depth-limited search techniques were adapted to handle the larger game space.
"in the long run how did it uh feel at the end like as a human being what it as a person who loves appreciates the beauty of the Game of Poker and the person who appreciates the beauty of AI is there d..."
Brown delves into the concept of Nash equilibrium in the context of multi-player poker, discussing the implications of strategy selection among multiple players. He highlights the challenges of achieving coordinated strategies in non-zero-sum games.
"it's a different it's different than chess and that aspect like people get that's why you want to look at Betty Marcus if you want to actually understand what people really think in the same sense pok..."
In this segment, Noam discusses the unique dynamics of six-player poker, emphasizing its adversarial nature. He explains how techniques from two-player poker can still be effective in this larger format due to the lack of cooperation among players.
"and so there was this big debate about whether Nash equilibrium and all these techniques that compute it are even useful once you go outside of two player zero some games now I think for many games th..."
Brown shares insights on the algorithmic advancements that led to the development of Pluribus, a poker AI that achieved superhuman performance. He contrasts the computational costs of Pluribus with its predecessor, Libratus, highlighting the significance of depth-limited search techniques.
"that's true more more broadly in extremely adverse serial games in general but that's sort of in practice versus being able to prove something that's right nobody has a proof that that's the case and ..."
Noam Brown discusses the role of neural networks in poker AI, revealing that neither Libratus nor Pluribus utilized them. He explains how the challenges in poker differ from games like Go, where neural networks excelled in feature extraction.
"so what are some interesting things about uh pluribus that was able to achieve human level performance on this or superhuman level performance on the six player version of Poker I personally I think t..."
In this segment, Brown elaborates on how modern poker AIs incorporate beliefs into their value functions. He contrasts this approach with traditional methods used in chess and Go, emphasizing the importance of understanding opponents' perceptions in poker.
"matter how would you describe the more General case of limited Dev search so it's basically constraining the scale a temporal or in some other way of the computation you're doing in some clever way so..."
Noam Brown tackles the challenging question of who the greatest poker player of all time is. He discusses the evolution of the game and how modern players have surpassed the skills of earlier generations, ultimately naming Daniel Negreanu as a standout player who has adapted to AI advancements.
"um but it wasn't the main challenge for poker like I think what neural Nets are really good for if you're in a situation where finding features for a value function is really difficult then neural Net..."
Transitioning to the game of Diplomacy, Brown describes its cooperative elements and strategic negotiations among seven players. He explains how the game differs from adversarial games like poker and chess, focusing on alliances and communication.
"the skill with which you avoided the question of the greatest of all time was impressive so my feeling is that it's a difficult it's a difficult question because just like in chess where you can't rea..."
Brown elaborates on the negotiation mechanics in Diplomacy, where players must form alliances to succeed. He highlights the importance of private communication and the strategic depth involved in making deals with other players.
"diplomacy yeah so I talked a lot about two player zero some games and what's interesting about diplomacy is that it's very different from these like adversarial uh games like chess go poker even Starc..."
In this segment, Noam discusses the unstructured communication style in Diplomacy, where players can freely negotiate and make deals. He compares the game to a mix of Risk, poker, and Survivor, emphasizing the social dynamics at play.
"period is done all the players simultaneously submit their moves and they're all executed at the same time and so you can tell people like hey I'm going to support you this turn um but then you don't ..."
Brown explains the mechanics of Diplomacy, detailing how players control units and issue move orders. He describes the objective of gaining control of territories and the strategic necessity of collaboration with other players.
"talk about anything you could say like hey let's have a long-term alliance against this guy you can say like hey can you support me this turn and in return I'll do this other thing for you next turn o..."
In this segment, Noam explains the mechanics of Diplomacy, including how players control units and issue move orders. He outlines the objective of gaining control of the majority of the map and emphasizes the importance of collaboration and negotiation with other players to achieve victory.
"by working with other players so on every turn you can issue a move order so for each of your units you can move them to an adjacent territory or you can keep them where they are or you can support a ..."
Noam elaborates on the concept of 'support' in Diplomacy, where players can assist each other's units. He highlights the strategic tension of making promises that may not be kept, illustrating the game's core dynamics of trust and betrayal among players.
"they'll have to retreat somewhere what does support mean support is like it's it's an action that you can issue in the game so you can say this unit you write down this unit is supporting this other u..."
This segment dives into the history of Diplomacy, including its creation in the 1950s and its association with notable figures like JFK and Henry Kissinger. Noam discusses the game's intention to teach diplomacy through its mechanics, reflecting on the failures of historical diplomacy that led to World War I.
"general is it true that Henry Kissinger loved the game and JFK and all those I've heard like a bunch of different people that or is that just one of those things that the cool kids say they do but the..."
Noam explains the balance of power in Diplomacy, noting that while France is considered the strongest power, the game's self-balancing nature prevents any one player from dominating. He discusses the unique starting positions of different nations and how they influence gameplay.
"know it kind of has a nice like wholesome take-home message then that you know war war is ultimately futile and uh and that optimal that feudal optimal could be achieved through great diplomacy yeah s..."
In this segment, Noam outlines the victory conditions in Diplomacy, where a player must control a majority of the map to win. He explains how draws are common among experienced players and discusses the scoring systems used to evaluate performance in the game.
"more vulnerable position because they have to like um they have a lot more neighbors as well got it larger territory more uh yeah right more border to defend okay uh what else is what else is importan..."
Noam shares insights into the history of AI research in Diplomacy, noting that efforts have been ongoing since the 1980s. He contrasts past rule-based approaches with modern strategies, emphasizing the complexity of the game compared to others like chess and poker.
"um there's a lot of different scoring systems the one that we used in our research um basically um gives a score relative to how much control you have of the map so the more that you control the highe..."
This segment focuses on the unique challenges AI faces in Diplomacy, particularly the natural language components and the need for effective communication. Noam discusses how the breadth of conversation topics complicates AI development and the necessity of understanding human interactions.
"sure um and you know it's understandable I mean the game is so incredibly different and so so much more complicated than the kinds of games that people were working on like chess and go uh and poker t..."
Noam emphasizes the importance of negotiation in Diplomacy, explaining that AI must learn to communicate effectively with humans. He discusses the need for AI to understand human expectations and social conventions to succeed in a game that relies heavily on alliances and trust.
"um the the depth and breadth of these conversations is is really complicated and it's all being done in natural language um now you could approach it and we actually consider doing this like you you k..."
In this thought-provoking segment, Noam compares the challenges of AI in Diplomacy to the Turing Test. He explains that while traditional Turing Tests focus on distinguishing humans from machines, Diplomacy requires AI to effectively collaborate with humans, highlighting the need for human-like behavior.
"like hey you might be getting attacked by by this other power Okay so what how we're supposed to think about okay so that's the natural language how do you even begin trying to solve this game it seem..."
Noam discusses the motivation behind pursuing AI development for Diplomacy, reflecting on the rapid advancements in AI and the desire to tackle complex challenges. He shares insights into the progress made and the goal of achieving human-level performance in this intricate game.
"to beat humans so how do you integrate the human into the loop of this so what you have to do is incorporate human data and to kind of give you some intuition for why this is the case like imagine you..."
Brown explains what it means for an AI to win in Diplomacy, highlighting the importance of convincing multiple players and achieving a high average score. He shares insights from their AI's performance, which reached human-level play but did not claim the top spot, sparking a discussion on measuring success in complex games.
"considering just how much progress there there was in Ai and that that progress has continued in the years since so winning in diplomacy what does that really look like it means talking to six other p..."
In this segment, Brown elaborates on how to measure AI performance in Diplomacy, arguing against testing solely against expert players. He draws parallels to self-driving car testing, emphasizing the need for a diverse skill level in opponents to accurately assess the AI's capabilities.
"about 40 games with with real humans online uh the bot came in second out of all players that played five or more games and um so not like number one but way way higher than well what was the expertis..."
Brown discusses the complexities of playing against human opponents in Diplomacy, noting that expert players are often more predictable than beginners. He highlights the unique challenges posed by human behavior and the need for AI to adapt to varying strategies and communication styles.
"that's quite brilliant because I played a lot of sports in my life like as a tennis Judo whatever and it's it's somehow almost easier to go against experts almost always I don't I think they're more p..."
Brown explains the process of incorporating human play data into AI training for Diplomacy. He describes how they trained a language model to generate dialogue with intent, ensuring that the AI's communication aligns with strategic goals while remaining human-compatible.
"um the really good diplomacy players are able to to take advantage of the fact that there is that there are some weak players in the game okay so if you have to incorporate human play data how do you ..."
In this segment, Brown details how the AI connects language to intent in Diplomacy. He discusses the combination of reinforcement learning and planning used to generate messages that align with strategic actions, emphasizing the importance of effective communication in achieving game objectives.
"how's that done just so as a starting point is that with reinforcement learning or is that just optimal determining what the optimal is for intents It's a combination of reinforcement learning and pla..."
Brown addresses the challenges of ensuring that the AI's communication is effective and appropriate. He explains the filters in place to prevent the AI from sending harmful or nonsensical messages, highlighting the importance of maintaining trust and strategic advantage in Diplomacy.
"planning is done is actually not using language so we're coming up with this like plan for the action uh that we're gonna play and the other person's gonna play and then we feed that action into the d..."
Brown discusses the critical role of trust in Diplomacy, arguing that successful players build relationships rather than relying solely on deception. He reflects on how the AI's approach to communication can inform our understanding of trust dynamics in human interactions.
"you have like an extra function that does the uh estimates the value of that message yeah so we have these kinds of filters that like so it's a filter so there's a there's a good and is that filter in..."
In this segment, Brown explores the broader implications of AI in Diplomacy for understanding human psychology and trust. He emphasizes the potential for this research to inform human-robot interactions and the importance of trust in intelligent entities.
"the the the uh the goal you want yeah and we're filtering for several things we're filtering like is this a sensible message you know so sometimes language models will send will generate messages that..."
Brown shares exciting news about open-sourcing their AI models and data from Diplomacy games. He discusses the potential for researchers to explore trust, negotiation, and persuasion using the extensive dataset, highlighting the significance of this research for the academic community.
"when it is telling you that hey I'm actually going to support you this turn is there some sense I don't know if you step back and think that this process well indirectly help us study human psychology..."
Brown concludes by affirming the richness of Diplomacy as a domain for studying human-AI interaction. He highlights the unique aspects of the game that make it an ideal setting for investigating negotiation and trust, positioning it as a valuable resource for future research.
"use to um we're hoping that the the academic Community the research Community is able to use it for for all sorts of interesting research questions so do you from having studied this game is this a su..."
Noam Brown discusses the importance of maintaining politeness in AI negotiations, particularly in games like Diplomacy. He shares insights on how researchers monitored AI behavior to prevent it from making inappropriate threats, emphasizing the need for AI to understand human-like strategies and interactions.
"you're threaten somebody you're supposed to do it politely yeah politely you know keep it in character um we actually had a researcher watching the bot 24 7 for well whenever we play a game we had a b..."
Brown explains the challenges of modeling human irrationality in Diplomacy. He contrasts the performance of AI trained in a two-player version of the game with its failure in a seven-player setting, highlighting the necessity for AI to adapt to human behaviors and emotions to succeed.
"um so what's required to win like what um what does it mean to mess up or to exploit the sub-optimal behavior of a player like uh is there is there optimally rational behavior and irrational behavior ..."
In this segment, Brown elaborates on the dynamics of cooperation in Diplomacy, likening it to geopolitical scenarios. He discusses how players must unite against a leading competitor, and how AI struggles to navigate human emotions and alliances, often leading to suboptimal outcomes.
"and be able to to work with that can you just Linger on that meeting like there's an individual there's an individual personality each player and then you're supposed to remember that but Woody means ..."
Brown highlights the difficulty of programming AI to understand human emotions and frustrations in Diplomacy. He illustrates how AI's rational strategies can clash with human emotional responses, leading to unexpected game outcomes and the need for AI to model human behavior more effectively.
"equilibrium would would change things is if it helped you so I what's the dynamic of cooperation that's effective in diplomacy do you always have to to have one friend in the game you always want to m..."
This segment focuses on the challenges of training AI with limited human data in Diplomacy. Brown discusses the balance between self-play and human data, explaining how they regularize AI behavior to better reflect human-like decision-making in complex game scenarios.
"and the bot will do this like the bot will work with the other players to stop the superpower from winning but if it doesn't really if it's trained from scratch or it doesn't really have a good ground..."
Brown explores the potential future applications of AI in dialogue systems, drawing parallels between Diplomacy and general communication. He discusses how the techniques developed for AI in games could enhance chatbot interactions and NPC behavior in video games, emphasizing the importance of intent in communication.
"don't have unlimited data we don't have unlimited neural net capacity um but it gives us some approximation uh is there some data on the internet that's useful besides just diplomacy so on the languag..."
In this concluding segment, Brown reflects on the implications of AI in real-world diplomacy and geopolitical decision-making. He argues that AI could help prevent conflicts by promoting cooperative strategies, drawing lessons from the game of Diplomacy to inform better diplomatic practices.
"is um and and then the language can correspond to that intent now I'm not saying that this is you know happening imminently but um I'm saying that this is like a future application potentially of this..."
Brown explores the potential for AI to simulate diplomatic scenarios, allowing leaders to anticipate the consequences of their actions. He discusses the importance of human data in these simulations and how they could help leaders navigate complex negotiations, ultimately aiming for peaceful resolutions in international relations.
"yeah I mean I just came back from Ukraine I'm going back there on deep personal levels think a lot about how peace can be achieved and I'm a big believer in conversation or leaders getting together an..."
This segment delves into the broader applicability of AI techniques developed for diplomacy to other forms of negotiation, such as legal disputes. Brown highlights the challenges of collecting data in real-world scenarios and the need for well-defined action spaces to effectively implement reinforcement learning in various negotiation contexts.
"but then you have to have human data right you really because it's like the game of diplomacy is fundamentally different than geopolitics you need data you need like I guess that's the question I have..."
Brown discusses the potential of using AI techniques from diplomacy to develop more human-like players in games like chess and Go. He explains how balancing human-like play styles with strong performance can enhance the gaming experience, making AI opponents more relatable and enjoyable for human players.
"like because it's natural language right you're operating in a space of intense and in a space of natural language that feels very close to the real world and it also feels like you could get data on ..."
In this segment, Brown addresses the ethical implications of creating AI that mimics human behavior in games. He discusses the challenges of cheat detection as AI becomes more human-like, raising questions about trust and fairness in competitive environments. The conversation touches on the broader societal impacts of integrating AI into human activities.
"um to elaborate on this a bit like one way to approach making a human-like AI for chess is to collect a bunch of human games like a bunch of human Grand Master games and just do supervised learning on..."
Brown reflects on the dual nature of AI technology, highlighting both its potential benefits and risks. He emphasizes the importance of ethical considerations in AI development, particularly regarding deception and trust in AI systems. The segment underscores the need for careful design to ensure AI serves humanity positively.
"where he needs to improve his strategy um and so I can Envision this future where data on specific chess and go players becomes extremely valuable because you can use that data to create specific mode..."
This segment explores the ethical dilemmas posed by AI systems capable of deception. Brown discusses the implications of developing AI that can lie and the potential societal consequences. The conversation raises critical questions about the moral responsibilities of AI developers and the future of human-AI interactions.
"right now it's really hard to learn how to get better in games like chess and poker and go because the way that the AI plays is so foreign and incomprehensible but if you have these AIS that are playi..."
Brown concludes by discussing how AI systems can reflect and challenge our understanding of human behavior and ethics. He emphasizes the importance of addressing deep philosophical questions as we design AI technologies, particularly in the context of diplomacy and conflict resolution, and how these systems can impact society.
"after somehow find where's the ethics in that and we're back to discussions inside relationships anyway what were you saying oh yeah I was getting like yeah this yeah that's kind of going to the quest..."
In this final segment, Brown addresses the challenges of data efficiency in AI development. He discusses the need for AI systems to learn from fewer examples, particularly in real-world applications. The conversation highlights the ongoing advancements in AI and the potential for future breakthroughs in creating more efficient and capable systems.
"creating AGI systems you've been a part of creating um by the way we should say a part of great teams that do this of creating systems that achieve breakthrough performances on before thought unsolvab..."
In this segment, Noam Brown explains how humans utilize extensive background knowledge to learn games like poker more efficiently than AI, which often starts from scratch. He suggests that allowing AI to leverage general knowledge across various domains could help address the sample complexity problem.
"gigantic background model language model and then you do um the training becomes like prompting that model to uh to essentially do a kind of querying a search into the space of the things it's learned..."
Brown shares valuable advice for beginners interested in machine learning, emphasizing the importance of building a strong foundation in mathematics, computer science, and statistics. He encourages newcomers to explore diverse perspectives and tackle challenging problems, drawing from his own atypical journey in AI research.
"um and then shifting more towards reinforcement learning as time went on and that actually had a lot of benefits I think because it allowed me to look at these problems in a very different way from th..."
The conversation shifts to the philosophical implications of AI in life optimization. Brown discusses the complexities of defining a reward function for living optimally, comparing it to the challenges of specifying objectives in AI. He reflects on the potential for AI to help clarify these concepts in the future.
"learning do you think one day we'll be able to since you're taking steps from poker to diplomacy one day we'll be able to uh figure out how to live life optimally well what is it like in in poker and ..."
In the closing segment, Noam Brown expresses his appreciation for the conversation and the work being done in AI. He quotes Sun Tzu from 'The Art of War,' emphasizing the importance of strategy and deception in both warfare and AI, leaving listeners with profound insights into the nature of intelligence and competition.
"function is sometimes it's pretty hard to specify the same way that you know trying to handcraft the optimal policy in a game like chess is really difficult it's not so clear-cut what the reward funct..."