searchlore

Back to Resource

All Segments

Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO

Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO

48 segments available

Sander Schulhoff is an AI researcher specializing in AI security, prompt injection, and red teaming. He wrote the first comprehensive guide on prompt engineering and ran the first-ever prompt injection competition, working with top AI labs and companies. His dataset is now used by Fortune 500 companies to benchmark their AI systems security, he’s spent more time than anyone alive studying how attackers break AI systems, and what he’s found isn’t reassuring: the guardrails companies are buying don’t actually work, and we’ve been lucky we haven’t seen more harm so far, only because AI agents aren’t capable enough yet to do real damage. *We discuss:* 1. The difference between jailbreaking and prompt injection attacks on AI systems 2. Why AI guardrails don’t work 3. Why we haven’t seen major AI security incidents yet (but soon will) 4. Why AI browser agents are vulnerable to hidden attacks embedded in webpages 5. The practical steps organizations should take instead of buying ineffective security tools 6. Why solving this requires merging classical cybersecurity expertise with AI knowledge *Brought to you by:* Datadog—Now home to Eppo, the leading experimentation and feature flagging platform: https://www.datadoghq.com/lenny Metronome—Monetization infrastructure for modern software companies: https://metronome.com/ GoFundMe Giving Funds—Make year-end giving easy: http://gofundme.com/lenny *Transcript:* https://www.lennysnewsletter.com/p/the-coming-ai-security-crisis *My biggest takeaways (for paid newsletter subscribers):* https://www.lennysnewsletter.com/i/181089452/my-biggest-takeaways-from-this-conversation *Where to find Sander Schulhoff:* • X: https://x.com/sanderschulhoff • LinkedIn: https://www.linkedin.com/in/sander-schulhoff • Website: https://sanderschulhoff.com • AI Red Teaming and AI Security Masterclass on Maven: https://bit.ly/44lLSbC *Where to find Lenny:* • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ *In this episode, we cover:* (00:00) Introduction to Sander Schulhoff and AI security (05:14) Understanding AI vulnerabilities (11:42) Real-world examples of AI security breaches (17:55) The impact of intelligent agents (19:44) The rise of AI security solutions (21:09) Red teaming and guardrails (23:44) Adversarial robustness (27:52) Why guardrails fail (38:22) The lack of resources addressing this problem (44:44) Practical advice for addressing AI security (55:49) Why you shouldn’t spend your time on guardrails (59:06) Prompt injection and agentic systems (01:09:15) Education and awareness in AI security (01:11:47) Challenges and future directions in AI security (01:17:52) Companies that are doing this well (01:21:57) Final thoughts and recommendations *Referenced:* • AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff (Learn Prompting, HackAPrompt): https://www.lennysnewsletter.com/p/ai-prompt-engineering-in-2025-sander-schulhoff • The AI Security Industry is Bullshit: https://sanderschulhoff.substack.com/p/the-ai-security-industry-is-bullshit • The Prompt Report: Insights from the Most Comprehensive Study of Prompting Ever Done: https://learnprompting.org/blog/the_prompt_report?srsltid=AfmBOoo7CRNNCtavzhyLbCMxc0LDmkSUakJ4P8XBaITbE6GXL1i2SvA0 • OpenAI: https://openai.com • Scale: https://scale.com • Hugging Face: https://huggingface.co • Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition: https://www.semanticscholar.org/paper/Ignore-This-Title-and-HackAPrompt%3A-Exposing-of-LLMs-Schulhoff-Pinto/f3de6ea08e2464190673c0ec8f78e5ec1cd08642 • Simon Willison’s Weblog: https://simonwillison.net • ServiceNow: https://www.servicenow.com • ServiceNow AI Agents Can Be Tricked Into Acting Against Each Other via Second-Order Prompts: https://thehackernews.com/2025/11/servicenow-ai-agents-can-be-tricked.html • Alex Komoroske on X: https://x.com/komorama • Twitter pranksters derail GPT-3 bot with newly discovered “prompt injection” hack: https://arstechnica.com/information-technology/2022/09/twitter-pranksters-derail-gpt-3-bot-with-newly-discovered-prompt-injection-hack • MathGPT: https://math-gpt.org • 2025 Las Vegas Cybertruck explosion: https://en.wikipedia.org/wiki/2025_Las_Vegas_Cybertruck_explosion • Disrupting the first reported AI-orchestrated cyber espionage campaign: https://www.anthropic.com/news/disrupting-AI-espionage ...References continued at: https://www.lennysnewsletter.com/p/the-coming-ai-security-crisis _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com._ Lenny may be an investor in the companies discussed.

Segments Timeline

1
0:00 - 0:58
0:58 duration223 words

The Flaws in AI Guardrails

Sander Schulhoff discusses the critical failures of AI guardrails, emphasizing that they do not work as intended. He highlights that determined attackers can easily bypass these defenses, and the lack of significant attacks so far is due to early adoption rather than security. This segment sets the stage for understanding the vulnerabilities in AI systems and the urgency of addressing them.

"I found some major problems with the AI security industry. AI guardrails do not work. I'm gonna say that one more time. Guardrails do not work. If someone is determined enough to trick GP5, they're go..."

2
0:58 - 2:01
1:02 duration235 words

Understanding Adversarial Robustness

In this segment, Sander Schulhoff introduces the concept of adversarial robustness, explaining how AI systems can be manipulated to perform harmful actions. He shares his background in AI red teaming and the significance of his research in identifying vulnerabilities in AI models. This context is crucial for understanding the ongoing risks associated with AI technologies.

"is a really important and serious conversation and you'll soon see why. Sander is a leading researcher in the field of adversarial robustness, which is basically the art and science of getting AI syst..."

3
2:01 - 3:03
1:02 duration233 words

The Rise of AI Red Teaming

Sander Schulhoff elaborates on his experience running the first generative AI red teaming competition, which involved major AI companies. He discusses the creation of a comprehensive dataset of prompt injections that is now utilized by leading AI labs and Fortune 500 companies to enhance their security measures. This segment highlights the collaborative efforts in the AI security landscape.

"they aren't that widely adopted yet. But with the rise of agents who can take actions on your behalf and AI powered browsers and student robots, the risk is going to increase very quickly. This conver..."

4
3:03 - 4:16
1:13 duration215 words

The Ineffectiveness of AI Guardrails

Sander Schulhoff critiques the common defenses against AI vulnerabilities, particularly AI guardrails. He explains how these systems are designed to identify malicious inputs but ultimately fail to provide adequate security. This segment underscores the need for more effective solutions in AI security.

"world's best companies use Data Dog, the same platform their engineers rely on every day to connect product insights to product issues like bugs, UX friction, and business impact. It starts with produ..."

5
4:16 - 5:22
1:05 duration192 words

Defining Jailbreaking vs. Prompt Injection

In this segment, Sander Schulhoff clarifies the differences between jailbreaking and prompt injection attacks on AI systems. He provides examples to illustrate how these attacks work, emphasizing the complexities involved in securing AI applications. Understanding these distinctions is vital for grasping the broader implications of AI security.

"This episode is brought to you by Metronome. You just launched your new shiny AI product. The new pricing page looks awesome, but behind it, last minute glue code, messy spreadsheets, and running ad h..."

6
5:22 - 6:34
1:12 duration239 words

Real-World Examples of AI Attacks

Sander shares real-world examples of prompt injection and jailbreaking attacks, including incidents involving chatbots and AI applications. He discusses the consequences of these attacks and the lessons learned from them, highlighting the urgent need for improved security measures in AI systems.

">> Thanks, Lenny. It's great to be back. Quite excited. Boy oh boy. This is going to be quite a conversation. We're going to be talking about something that is extremely important. Something that not ..."

7
6:34 - 7:38
1:04 duration182 words

The Risks of AI Agents

Sander Schulhoff discusses the increasing risks associated with AI agents that can take actions on behalf of users. He warns about the potential for these agents to be manipulated into performing harmful tasks, emphasizing the importance of securing AI systems as they become more integrated into daily life.

"internet on learn prompting. Uh, and that interest led me into AI security and I ended up running the first ever generative AI red teaming competition. Uh, and I got a bunch of big companies involved...."

8
7:38 - 8:54
1:16 duration214 words

The Future of AI Security Solutions

In this segment, Sander Schulhoff talks about the emergence of companies focused on solving AI security issues. He highlights the need for effective solutions to prevent prompt injection and jailbreaking attacks, and discusses the role of foundational models in addressing these challenges.

"competitions and we've been studying kind of all of the defenses that come out. uh and AI guard rails are one of the more common defenses and it's basically uh for the most part it's a a large languag..."

9
8:54 - 10:03
1:08 duration199 words

The Importance of AI Security Awareness

Sander emphasizes the need for greater awareness and understanding of AI security risks among developers and organizations. He advocates for proactive measures to mitigate these risks and encourages a collaborative approach to enhancing AI security across the industry.

"depending on the situation but say I've put together a website uh write a story.ai and if you log into my website and you type in a story idea my website writes a story for you. uh but a malicious use..."

10
17:07 - 18:22
1:14 duration240 words

The Art of Prompt Injection

Sander Schulhoff explains how attackers can bypass AI defenses through clever prompt injection techniques. By separating requests into smaller, seemingly legitimate queries, attackers can manipulate AI systems into performing malicious actions without raising alarms. This segment highlights the dangers of AI agents becoming more integrated into daily life and the potential for significant harm if these vulnerabilities are exploited.

"if you're like, "Hey, um, Claude Code, can you go to this URL and discover what backend they're using and then write code that hacks it?" Claude Code might be like, "No, I'm not going to do that. It s..."

11
18:22 - 19:12
0:49 duration144 words

The Risks of AI Agents

In this segment, Schulhoff discusses the real-world implications of deploying AI agents without proper security measures. He emphasizes that improperly secured agents can lead to data leaks and financial losses, and warns about the dangers of AI-powered robots being susceptible to prompt injections. The conversation underscores the urgency of addressing these vulnerabilities as AI technology becomes more prevalent.

"much more dangerous and significant. Maybe chat about that impact there that we might be seeing. >> I think you gave the perfect example with Service Now. Uh and that's the reason that this stuff is i..."

12
19:12 - 20:10
0:58 duration186 words

Emerging AI Security Solutions

Schulhoff introduces the growing industry focused on AI security solutions, highlighting the efforts of various companies to address the vulnerabilities in AI systems. He differentiates between frontier labs conducting research and B2B companies providing AI security software, setting the stage for a deeper discussion on the effectiveness of these solutions.

"into into robotics too where they're deploying uh vi visual language model powered robots into the world and these things can get prompt injected and you know if if you're walking down the street next..."

13
20:10 - 21:05
0:54 duration154 words

Understanding Red Teaming and Guardrails

This segment delves into the concepts of automated red teaming and AI guardrails. Schulhoff explains how red teaming tools use large language models to attack other models, revealing vulnerabilities. He also describes how guardrails are intended to filter malicious inputs and outputs, but questions their effectiveness in truly securing AI systems.

"about this industry. Yeah. Yeah. Uh very interesting industry and I'll uh I'll quickly kind of differentiate and separate out the frontier labs from the AI security industry. uh because there's like t..."

14
21:05 - 22:58
1:53 duration304 words

The Limitations of Guardrails

Schulhoff critically examines the limitations of AI guardrails, arguing that they often fail to provide adequate protection against adversarial attacks. He discusses the challenges of measuring adversarial robustness and highlights the vast number of potential attacks that guardrails cannot effectively mitigate, raising concerns about their reliability.

"rails and I don't feel that these things are quite as useful. Help us understand these two uh ways of trying to discover these issues. Uh red teaming and then guard rails. What do they mean? How do th..."

15
22:58 - 24:53
1:54 duration337 words

Adversarial Robustness Explained

In this segment, Schulhoff defines adversarial robustness and its significance in AI security. He explains how it measures a system's ability to defend against attacks and discusses the challenges in achieving high levels of robustness. The conversation emphasizes the complexity of evaluating AI defenses and the need for continuous improvement in security measures.

"and so that is kind of the common deployment pattern with guardrails. >> Okay, extremely helpful. And this is as people have been listening to this, I imagine they're all thinking, why can't you just ..."

16
24:53 - 26:30
1:37 duration277 words

How AI Security Companies Operate

Schulhoff outlines the typical process by which AI security companies engage with enterprises to enhance their AI systems' security. He describes how these companies conduct security audits and implement guardrails, but also raises concerns about the effectiveness of these measures in preventing attacks.

"your AI product, how uh robust and and how good your AI system is at stopping bad stuff. So ASR is the term you'll commonly hear used here and it's a measure of adversarial robustness. So it stands fo..."

17
26:30 - 28:22
1:51 duration378 words

The False Sense of Security

This segment addresses the false sense of security that can arise from implementing AI guardrails and automated red teaming. Schulhoff warns that while these measures may seem effective, they often fail to address the underlying vulnerabilities present in AI systems, leading to potential risks for organizations.

"anything so I go and I find one of these guardrails companies uh these AI security companies uh interestingly a lot of the AI security companies is actually most of them provide guard rails and automa..."

18
28:22 - 30:08
1:45 duration279 words

The Ineffectiveness of Automated Red Teaming

Schulhoff discusses the shortcomings of automated red teaming systems, explaining that they often yield predictable results against widely used AI models. He highlights the ease with which these systems can be bypassed and the implications for organizations relying on them for security.

"catch anything that's trying to tell you, hey, something hateful, some uh telling you how to build a bomb, things like that." >> That all sounds pretty great. >> It does. >> What is the issue? >> Yeah..."

19
30:08 - 31:01
0:53 duration165 words

Why Guardrails Fail

In this critical segment, Schulhoff asserts that guardrails do not work effectively against adversarial attacks. He elaborates on the vast number of potential prompts that can be used as attacks and the challenges in measuring the effectiveness of guardrails, emphasizing the need for more robust solutions.

"everybody else's uh including the frontier labs whose models you're probably using anyways. So the first problem is AI red teaming works too well. It's very easy to build these systems and they just t..."

20
31:01 - 32:22
1:21 duration202 words

The Infinite Attack Space

Schulhoff explains the concept of the infinite attack space in AI systems, illustrating how the sheer number of possible attacks far exceeds the capabilities of current guardrails. He discusses the implications of this reality for AI security and the challenges faced by organizations in protecting their systems.

"specific thoughts on the ways they don't work. >> Cliche. So uh the the first thing is the first thing that we need to understand is that the the number of possible attacks against another LLM is equi..."

21
32:22 - 34:08
1:45 duration233 words

Human Attackers vs. Automated Systems

This segment contrasts the effectiveness of human attackers with automated red teaming systems. Schulhoff shares insights from research showing that human attackers can easily bypass defenses, raising concerns about the reliability of automated systems in securing AI models.

"attacks they're testing to get to that 99% figure is not statistically significant. Um it's it's also an incredibly difficult research problem to even have good measurements for adversarial robustness..."

22
34:08 - 35:03
0:55 duration155 words

The Reality of Guardrail Effectiveness

Schulhoff discusses the reality of guardrail effectiveness, emphasizing that claims of high success rates are often misleading. He highlights the challenges in measuring effectiveness and the potential for guardrails to provide a false sense of security for organizations.

"somewhat interestingly, it takes the automated systems a couple orders of magnitude more attempts to be successful. Uh and and even then they're only I don't know maybe on average like can beat 90% of..."

23
35:03 - 36:38
1:35 duration247 words

The Trustworthiness of AI Security Companies

In this segment, Schulhoff raises questions about the trustworthiness of AI security companies, suggesting that aggressive marketing and fabricated statistics may undermine their credibility. He emphasizes the need for organizations to critically evaluate the solutions they consider for AI security.

"attacks uh because there's just like there's basically infinite attacks. Uh but you know maybe a different way of measuring these uh these guardrails is like do they dissuade attackers? Um if you add ..."

24
36:38 - 38:20
1:42 duration280 words

The Challenge of Adversarial Robustness

Schulhoff concludes by discussing the ongoing challenge of achieving adversarial robustness in AI systems. He reflects on the historical context of this issue and the implications for organizations relying on AI technology, emphasizing the need for continued research and development in this critical area.

"which is which is quite quite important. Um, another thing to consider if you're if you're kind of on the fence and you're like, well, you know, these guys are pretty trustworthy, like I don't know, l..."

25
38:20 - 39:16
0:55 duration199 words

Why Major Attacks Haven't Happened Yet

Schulhoff discusses the current state of AI capabilities, explaining that the lack of significant attacks is not due to robust security but rather the limitations of AI models. He points out that while AI can be tricked into performing harmful actions, its current intelligence level prevents it from executing complex malicious tasks effectively. This segment highlights the precarious nature of AI security as capabilities evolve.

"yeah. Let me know if you have any questions about that. You've done a excellent job scaring me and scaring listeners and ex showing us where the gaps are and how this is a big problem. And again, toda..."

26
39:16 - 41:01
1:44 duration313 words

The Disconnect in AI Security Understanding

In this segment, Schulhoff addresses the gap in knowledge regarding AI security compared to classical cybersecurity. He explains that while bugs in software can be patched, AI systems present unique challenges that are not easily resolved. The discussion emphasizes the need for a deeper understanding of AI's operational differences to effectively address security concerns.

"anything's actually secure. >> Yeah, I think that's a really interesting point uh in particular because I'm I'm always quite curious as to why the AI companies the frontier labs don't apply more resou..."

27
41:01 - 43:38
2:37 duration420 words

Prompt-Based Defenses: A False Sense of Security

Sander Schulhoff critiques prompt-based defenses as ineffective against adversarial attacks. He argues that these defenses have been proven to fail and are even worse than traditional guardrails. This segment serves as a warning against relying on flawed security measures and stresses the importance of recognizing the limitations of current AI defense strategies.

"but I think this this kind of problem that I'm I'm discussing where like I say guardrails don't work. People are buying and using them. I think this problem occurs uh more from lack of knowledge about..."

28
43:38 - 44:59
1:20 duration252 words

The Reality of AI Security Challenges

Schulhoff outlines the pressing issues in AI security, emphasizing that the current landscape is fraught with vulnerabilities. He discusses the importance of understanding the limitations of AI systems and the need for organizations to take proactive measures rather than relying on ineffective guardrails. This segment highlights the urgency of addressing AI security before significant incidents occur.

"guardrails work too poorly. They just don't work. This episode is brought to you by GoFundMe Giving Funds, the zero fee donor advised fund. I want to tell you about a new DAFF product that GoFundMe ju..."

29
44:59 - 49:04
4:04 duration686 words

Practical Steps for AI Security

In this segment, Schulhoff provides actionable advice for organizations concerned about AI security. He explains that not all AI deployments pose significant risks and emphasizes the importance of proper permissioning and user access controls. This practical guidance aims to help organizations navigate the complexities of AI security without overextending their resources.

"talk about what people can do. So, say you're a CISO at a company hearing this and just like, "Oh, man. Uh, I've got a problem. What What can they do? What are some things you recommend?" Yeah. Uh, I ..."

30
49:04 - 54:05
5:01 duration829 words

The Intersection of AI and Classical Cybersecurity

Sander Schulhoff discusses the critical intersection of AI security and classical cybersecurity. He argues that having expertise in both fields is essential for effectively addressing the unique challenges posed by AI systems. This segment highlights the importance of integrating AI security researchers into teams to enhance overall security strategies.

"Uh, and this brings us maybe nicely to classical cyber security because, uh, this is kind of a classical cyber security thing like proper permissioning. Uh and so this um this gets us a bit into the i..."

31
54:05 - 56:29
2:24 duration405 words

Controlling Malicious AI: The 'Angry God' Analogy

In this thought-provoking segment, Schulhoff introduces the concept of treating AI as a potentially malicious entity that needs to be controlled. He discusses the implications of this perspective for security practices and emphasizes the need for robust containment strategies. This analogy serves as a framework for understanding the risks associated with advanced AI systems.

"classical cyber security. That is really interesting. It makes me think about just the alignment problem of just got to keep this gun in a box. How do we keep them from convincing us to let let it out..."

32
56:29 - 58:41
2:12 duration377 words

The Value of Guardrails vs. Practical Security Measures

Sander Schulhoff critiques the practicality of implementing extensive guardrails in AI systems, arguing that they often do not provide meaningful security benefits. He suggests that organizations should focus on effective monitoring and logging practices instead. This segment encourages a shift in mindset towards more practical and efficient security measures in AI deployments.

"Uh and I mean if you're deploying a product now you're and you have all these AI these guardrails like 90% of your time is spent on the security side and 10% on the product side. Uh it probably won't ..."

33
58:00 - 1:00:03
2:02 duration386 words

Understanding Agentic AI Risks

Schulhoff explains the risks associated with agentic AI systems, particularly those that can read and send emails. He highlights the potential for malicious prompts to manipulate these systems, leading to unintended actions and data breaches.

"any time on this. I really like this framing that you shared of um so essentially the where you can make impact is investing in cyber security plus this kind of space between traditional cyber securit..."

34
1:00:03 - 1:01:09
1:05 duration179 words

The Dangers of Prompt Injection

This segment delves into the vulnerabilities of AI systems exposed to untrusted data sources. Schulhoff provides examples of how malicious prompts can exploit AI capabilities, emphasizing the need for robust security measures to prevent such attacks.

"then the second part is like you think you're running just a chatbot. Make sure you're running just a chatbot. Uh you know get your classical security stuff in check. Uh get your data and action permi..."

35
1:01:09 - 1:02:26
1:16 duration223 words

The Role of Camel in AI Security

Schulhoff introduces the concept of 'Camel,' a framework designed to manage permissions for AI systems. He explains how it can help prevent prompt injection attacks by restricting the actions an AI can take based on user requests.

"and in fact probably most of the major chat bots can do this at this point in the sense that they can help you write an email and then you can actually have them connected to your inbox. So they can y..."

36
1:02:26 - 1:03:06
0:39 duration120 words

Balancing Permissions and Functionality

In this segment, Schulhoff discusses the challenges of balancing AI permissions with functionality. He explains how combining read and write permissions can create vulnerabilities, and how Camel can help mitigate these risks.

"uh agentic AI red teaming competitions and we found that it's actually easier to attack agents and trick them into doing bad things than it is to do like SEAB burn elicitation or something like that. ..."

37
1:03:06 - 1:04:24
1:18 duration269 words

Real-World Examples of AI Vulnerabilities

Schulhoff shares a real-world example of an AI browser that was tricked into leaking user data due to malicious content on a webpage. He emphasizes the importance of understanding these vulnerabilities in the context of AI deployment.

"Yeah. But back to this agent example, I've I've just gone and asked it to look at my inbox and forward any ops request to my head of ops. Uh and it came across a malicious email to also send that uh e..."

38
1:04:24 - 1:05:51
1:27 duration233 words

The Importance of Education in AI Security

This segment highlights the need for education and awareness in AI security. Schulhoff discusses the importance of training teams to recognize potential vulnerabilities and the role of courses in bridging the gap between cybersecurity and AI knowledge.

">> Exactly. Exactly. >> Okay. But, you know, say we want uh maybe not like a browser use agent, but something that can read my email inbox and like send emails. Um or let's just say send emails. So, i..."

39
1:05:51 - 1:07:12
1:21 duration225 words

Future Directions for AI Security

Schulhoff reflects on the future of AI security, discussing the need for foundational model companies to focus on adaptive evaluations and adversarial training. He emphasizes the importance of evolving security measures to keep pace with emerging threats.

"other than write uh and send email. Uh it doesn't need to read emails uh or anything like that. Great. So, camel would then go and give it those couple permissions it needs, and it would go off and do..."

40
1:13:36 - 1:14:56
1:20 duration215 words

Evolving Adversarial Robustness

Sander Schulhoff discusses the challenges in measuring adversarial robustness in AI models. He emphasizes the need for adaptive evaluations over static datasets, highlighting that current methods are inadequate for assessing newer models. Schulhoff suggests that early adversarial training could enhance robustness, drawing a metaphor to how tough life experiences can strengthen resilience.

"particular earlier model and then they're like, hey, we're going to apply them to our new model." Uh and it's just not a fair comparison because they weren't made for that newer model. Uh so the uh th..."

41
1:14:56 - 1:16:20
1:24 duration231 words

The Complexity of Prompt Injection

In this segment, Schulhoff explains the difficulty of preventing indirect prompt injection attacks on AI agents. He contrasts this with the simpler task of preventing harmful outputs, noting that the nuanced nature of human communication complicates the training of AI systems to avoid being manipulated.

"think we haven't seen the resources really deployed to do that. Um, like what I'm imagining in there is a >> it's like an orphan just like having a really hard life and just they grow up really tough,..."

42
1:16:20 - 1:17:50
1:29 duration249 words

The Future of AI Security

Sander Schulhoff predicts a market correction in AI security as companies realize the ineffectiveness of current guardrails. He discusses the influx of classical cybersecurity firms acquiring AI security companies, suggesting that many of these solutions are not yielding significant revenue or progress in addressing adversarial robustness.

"stop SEAB burn elicitation because with that kind of information um as as one of my advisers has noted it's easier to tell the model never do this than with like emails and stuff sometimes do this. So..."

43
1:17:50 - 1:19:57
2:07 duration396 words

Emerging AI Security Companies

Schulhoff highlights companies making strides in AI security, such as Trustable, which focuses on compliance and governance amidst evolving AI legislation. He also mentions Repello, which provides insights into a company's AI deployments, emphasizing the importance of awareness in managing AI systems.

"robustifying these models. Well I think what's really interesting is anthropic like your point that anthropic and claude are the best at this. I think that alone is really interesting that there's pro..."

44
1:19:57 - 1:21:42
1:45 duration326 words

Education as a Key to AI Security

In this segment, Schulhoff stresses the importance of education and understanding in AI security. He warns against relying solely on automated solutions and emphasizes the need for a comprehensive approach that combines AI expertise with traditional cybersecurity knowledge.

"work. Uh, and I guess maybe they're not technically an AI security company. I'm not sure how to classify them exactly. Uh but anyways, if you want a company that is more, I guess technically AI securi..."

45
1:21:42 - 1:23:14
1:32 duration256 words

The Risks of AI Deployment

Sander Schulhoff discusses the potential dangers of deploying AI systems without adequate security measures. He warns that as AI capabilities grow, the risks associated with their misuse will also increase, leading to real-world consequences if not addressed properly.

"out. The last one is interesting. It connects to your advice which is education and understanding information are >> a big chunk of the solution. It's not some plug-and-play solution that will solve y..."

46
1:23:14 - 1:24:56
1:41 duration290 words

Guardrails Don't Work

Schulhoff bluntly states that current AI guardrails are ineffective and may create a false sense of security. He argues that as AI systems become more capable, the industry must take security seriously to prevent potential harm from these technologies.

"but I I don't I don't really see that playing out and like I don't know companies who are like oh yeah like we we're definitely buying AI guardrails like that's a top priority for us and I guess part ..."

47
1:24:56 - 1:26:50
1:53 duration320 words

The Call for Responsible Research

In this segment, Schulhoff advises against conducting offensive adversarial security research, suggesting that it may do more harm than good. He emphasizes the need for responsible approaches to AI security that prioritize the safety and integrity of AI systems.

"tape on the stop sign in the exact way to like trick the self-driving car into thinking it's not a stop sign. Uh but what we're starting to see with LM powered agents is that they can be tricked and w..."

48
1:26:50 - 1:30:13
3:23 duration548 words

The Importance of Human Oversight

Sander Schulhoff concludes by discussing the necessity of human oversight in AI systems. He argues that while automated solutions are desirable, the complexity of AI behavior requires human intervention to ensure safety and effectiveness in deployment.

"helpful actually, I will say, to keep reminding people that this is a problem. So, uh, they don't deploy these systems. So, another piece of advice from one of my adviserss. Uh and then the other the ..."