0:00
Okay, Alex, I want to go deep on dots, but first I want to just talk about this moment we're in as an industry.
Okay, Alex, I want to go deep on dots, but first I want to just talk about this moment we're in as an industry.
It's something I've been thinking about a lot and there's so many new products that are constantly popping up, which is this rise of personal agents.
And it feels like it's a combination of a lot of things.
The models getting capable enough to really power these experiences.
things like browser use, computer use really taking off, people figuring out the harness and now you guys have dots which is the open AI bet on this form factor.
I would be curious to hear you like bigger picture maybe talk about why this moment is happening now. >> Totally.
And yeah, I agree this is a massive moment.
I've I joined OpenAI a little over two years ago and I've basically been on a mission to ship dots since I joined and so I'm like thrilled that it's happening.
that it's happening. So you know recapping quickly we had chat right that was chat was like it's very obvious today but it was not obvious that that would be the the right way for many people to get value from LLMs when we started uh the wave after chat was coding agents and then eventually those agents useful for knowledge work and the big improvement from chat to coding agents was like hey that now these
agents can do things right but there was still this like fundamental problem which is I think the main problem we're solving with this next era you know the third chapter active intelligence we call it um and that's that Even if you
have an incredibly powerful coding agent, it's kind of like you have this like incredible principal engineer who refuses to check Slack, refuses to check linear, and basically only does things when they're told by you. And if you're
And if you're a power user, you can set up automations and kind of get them going more independently.
But for a sort of normal user, that doesn't happen.
And so for me, this is happening now because we're finally ready to provide intelligence in a way that's much more intuitive for people in a way that you don't have to put all the onus on yourself to get value from the product, but it can start proposing to you what it's going to do and how it's going to help. So that's the problem.
And then, you know, we can get more into the reasons, but I think broadly it is the model capabilities are there.
And I think that maybe I so I agree with you on what you said.
Maybe the thing you didn't mention that I think is underrated is there's actually a lot of human change that needs to happen for these products to get adopted.
You know, like if you look at like a company or a team adopting coding agents, day zero when they don't know how to use them and they're not connected to any tools to like day 400, the agents are connected to a bunch of tools, they have access to context, um people have best practices.
I think that's the other major thing we needed, right?
Because if you think of a proactive agent, if it doesn't have access to context, it can't actually be proactive in helping you.
>> Um, so I could go on, but >> yeah, the connectors and all that stuff being a place where people are bringing their data into these experiences. >> Yeah.
>> Yeah, that makes sense.
>> You said you joined and it's been your mission to ship dots ever since. >> Yeah.
>> Expand on that a little bit.
I mean, were you all envisioning this form factor for this two years ago?
>> I'll just speak for myself here.
>> I'll just speak for myself here. Um but I had this broad idea that chat was incredibly powerful and AI was going to transform the world and it was really important like the the mission of openi really resonated with me that's why I joined right deliver the benefits of AGI to all humanity and okay I'm not necessarily going to be the person
training the model but I did feel with chat that there was so much lowhanging fruit on how to make that model more useful and it was those two things I mentioned like let it do things and let me not have to figure out what to ask it to do because you But even even today, like I go on Twitter and I see someone talking about all the loops they're running and I feel a little bad about myself personally. I'm like, "Oh, I'm
I'm like, "Oh, I'm I'm not an advanced practitioner.
I'm not running enough loops."
>> I'm pretty sure you are an advanced. >> No, I am.
But, you know, but even I feel that way. And I don't know how. Do you feel that? >> Oh. Oh. Oh my gosh. I'm overwhelmed.
Every time I open it, I'm like, "Wow, 10 new things." >> Yeah.
And it's like it's rough, right?
Like what product categories exist that the better you get at it, the more aware you are of your own personal shortcomings? Yeah.
>> Maybe like sports hobbies, maybe.
>> Maybe like sports hobbies, maybe. like I don't know that's like it's kind of rough right and so >> I broadly I wanted to figure out how we could solve those things and it was kind of interesting we >> we were working on chat I remember in the early days I was working on the desktop app uh we were trying to build that into a contextual assistant and
yeah the models were not there in certain ways well actually fun story I joined open through an acquisition >> of a small startup >> and we had signed a deal the day before that openi launched like their voice product I forget what it was called just voice you know that that jaw-dropping demo like a few years ago and we so we signed the deal the night before and then we got on a call with our team watched that demo live. We didn't know
We didn't know that they were launching voice.
It was a jaw-dropping moment and then we said to the team like, "Hey, we're joining that company. They just acquired us." >> It was wild.
>> Um and so at the time we thought that like multimodal progress was going to be it and that's how we were going to have the next wave of capabilities and that was kind of the approach we took.
Like the desktop app, Chachby desktop app would like screenshot your apps and try to do work in it.
But it it was really slow.
Oh, it didn't work super well.
And it turns out like coding was actually that next frontier.
And so we had a year of progress on coding.
And I think, >> you know, with coding, we ended up building a lot of tech for helping agents do more.
And we built out all these connectors uh in a more easy way than having the agent like click around on your computer and stuff.
>> We also developed ways to like hoist the agent to the cloud.
All all these like base primitives that needed to be in place for us to ship these persistent agents.
And now it's funny because we're kind of having a comeback of computer use now, right?
where now the models are so good at coding, but they're and they're also good at using code using code to control the computer.
Like if you get Astra to control the browser or computer, it might actually write code snippets to do it because that's faster than taking screenshots and clicking. >> Yeah.
>> Um but yeah, now we're we're in back at this comeback of like, well, the models can do anything that you can do on a computer.
>> This episode is brought to you by Mercury, AI native banking that's loved by more than 300,000 entrepreneurs, including me. Visit mercury. com to learn more.
Mercury is a fintech, not a bank.
Check the show notes for details.
This episode is also brought to you by Jira by Atlassian, where teams and agents get the context, coordination, and control to move work forward. Try it free at jira. com. That's ji. com.
Thanks also to Granola, the AI notepad for people in back-to-back meetings.
It works everywhere you do and lets you focus on what matters. Try it at granola.
ai/sources and use the code sources for 3 months off.
So take me into the last few months building dots.
We talked about the buildup of how these things have progressed.
When did you guys go, "Oh, we have to now we're doing this. We're building this.
This is the next big swing."
>> If we think about this a little bit abstractly, the key things about dots are first it is a model that is trained to be persistent and to go after tasks uh sort of proactively, right?
So initially over a long horizon uh maybe you could say something like babysit this PR. >> Mhm.
you know, the model, you didn't tell it exactly how to solve that problem, but it's like figuring it out.
Um, and then eventually even more open-ended tasks like stay on top of this feedback channel, right?
So, this that's one key idea there, persistence, and it's it's probably the most important idea.
Some of the other ideas are it being available wherever you are so you can talk to it in any tool and it being able to do anything you can do because it has access to its own computer, not yours.
Um, so going back to that first idea of persistence, we've been working on this for actually a really long time.
And it's been not like a instantaneous step function improvement in capability, but it's been sort of like the models have been getting better and better at it.
And the product has also been slowly getting better at it.
Um, so for example, /go is a fan favorite feature in Codeex that enables the model to sort of go after a problem for a really long time.
And so we were shipping things like SLGO.
We were shipping internal prototypes of things that enabled you to have agents or threads that would just go for a really long time.
And we just started seeing this like really explode across the company.
Uh to the point where there was just like a channel of feedback around one of these internal prototypes that was just like constantly active, people sharing like all sorts of cool ways that they were like meaningfully accelerating themselves.
And so we were looking at that and I remember we had a really fun debate which was should we ship this prototype that works locally on your computer?
So it was kind of like slash go like it didn't have its own computer.
It would just like kind of create this agent and it would go after things.
And so we asked ourselves should we ship this locally on your computer or should we wait till we have a cloud computer which is a harder thing to build. >> Mh.
>> And should we ship this as just an enhancement to a codeex thread?
So it's just like you're just using codeex.
It's just like the the thread can do more or should we embody this as an agent that has its own identity that you can reach in Slack etc etc.
And so we kind of went through this iterative process shipping with the company seeing what we saw like exciting use cases etc.
And we ended up deciding okay for us internally we can get a lot of internal acceleration from this local only non-embodied product but the most intuitive form factor to help everyone benefit um by asking this agent to do really hard tasks and long long tasks would be let's put it in the cloud and let's embody it.
Embodying mean like let's give it a name, let's give it an identity, etc.
And so we kind of then started a sprint to like get those two things, those two sort of tweaks to the form factor built around this base model capability with Astra.
>> I'm curious about the decision to embody it and the identities that you can give it and the customization.
Why give that level of control to the user? >> Totally.
So it's a it's a really good question and to be completely honest with you, I think we made the right choice.
But I also think that we together like the industry are going to discover the answers to some of these questions like should agents be embodied?
Uh how many should people have?
Is it different between consumer and enterprise?
Like we're going we have our takes like I can give you my takes but we're going to figure that out.
I think embodiment though to answer your question it's a really powerful analogy that helps people understand how to use the product. >> Um two main things.
First, I think that like really smart like power users of tools realize like advanced ways they can use tools, but a normal user is going to look at something that's not embodied and kind of just ask an agent to do something like now one time.
But the moment that you have this like entity that shows up in like I don't know Slack, it's just something clicks in people's brains and they start realizing like oh I can ask this for things that take days or that aren't continuous work that uh you know maybe it's like stay on top of this project and let me know and so people start asking it like start giving it much harder tasks and that was really important to us.
We wanted to have like really good tasks.
The other thing that is really powerful with embodiment is it helps people understand that you are talking to one context no matter where you're talking to it. Right?
So for example the way that I work I when I'm commuting I'm often on a call uh with my dot I spend a lot of my day in Slack because I'm on the product team.
So, I'm I'm like slacking it across many different threads, right?
Sometimes I'm forwarding DMs to it, sometimes I'm I'm in threads, and then sometimes I'm in the app talking to it.
But no matter where I'm talking to it, it's one context, and it knows about all the conversations I've had.
And there are a few ways you could try to teach users that. One is embodiment.
The other is you could try to like synchronize threads across surfaces.
But that to me doesn't really make sense, right?
Like if you and I have a conversation now like live in this room and then I text you later, it's not like there's going to be a transcript of our conversation in the text. >> Right. >> Right.
But but you still know it's me because I'm an entity.
>> Well, and it makes sense too when you look at the product evolution of chatbt.
You guys have been the merge.
You guys have been trying to unify work, chat, everything together, but there are still silos and like even until recently voice mode didn't have the connectors access that the work and chat tabs did.
So, you know, one of my routines is like I take my dogs for walks in the morning and the evening and I'll do work calls and I'll talk to AI and all that stuff.
And I've been on the the early uh beta for for uh DOT the last few weeks trying it on these walks and like it's amazing to have all that context and my connectors finally in one place together and it's all and and I do think embodying it makes sense in that context.
But I'm also curious just like yeah, do you think the world we're going towards is everyone has their single agent that's like their personal agent and that agent then delegates to a bunch of sub aents or people have an agent for different slivers of their life?
Like how do you think that that plays out?
>> So it's like the way I kind of think of that question is like what is the interface that we're exposing to whom >> right?
And so to agents, like I we'll parking lot, but like to agents, I think there's a different answer to that question.
But to humans, like there's only so many people that I want to remember the names of, right?
And that's just cuz I have bad memory.
And now if each of those people have like eight agents and they all have like whimsical names, like I don't know what's a sport you like.
>> I I play ultimate frisbee. >> Okay.
So like are there different kinds of frisbes?
Imagine you had five agents and they're all named different kinds of frisbee brands, right?
Like how am I going to remember that?
And like one of them is for like GTM like go to market and one of them is for like if I want to meet you for dinner, you know, and like it's just it's hard for me.
Maybe it's okay for you, but it's hard for me.
And we we see the future of these agents is clearly like people and agents working together to help empower us and advance our lives.
Um but if that comes with a bunch of cognitive load of me knowing what all your agents are called and who they are, it's terrible, right?
So our our take on this right now is that we should support people having multiple agents but two things.
First for a consumer normal person like you probably just want one agent and it's only like a power user might create multiple agents. We could talk about why.
So most people will just have one and that agent should feel like an extension of them.
So for example a a decision we've made in product is if you put your agent in Slack or in Teams it actually is prefixed the name of it is prefixed with your name.
So like my my Slack handle is AE.
And so my agent, if I give it a name, my I just call mine dot because, you know, I like the name dot.
So my agent is just AE- dot.
If you called your agent Alice, your agent would be called AE- Alice.
And so this was not something that we thought of and were correct on.
This is actually something we discovered as we were building because we accidentally let everyone create many agents with random names.
And you know, we even encouraged people to give them fun names and then no one had any idea what was going on in Slack and it just became like really difficult >> and the open eye Slack is already a pretty active place. >> Yeah.
So that's and that's us working together, right?
Like you and I aren't in the same Slack.
Let's say we're like texting.
I was in like there's no way I'm going to know your agent's name, right? Right. >> Right.
>> Um so yeah, we that's that's our take there.
>> So why this interface though where it's kind of one master single thread?
Because I think people are used to different chat threads and bouncing around and I think you know there's Muse, there's Instinct, there's other products that have this kind of interface that everyone seems to be landing on.
But why is this the the interface?
Something we believe in that Sam talks about is that the human world is shaped around humans.
And so a lot of the tools we have were shaped around us.
And so what I find very fun about this is that as soon as you start say designing software to be embodied or in some future, you know, building robots that are humanoid, >> you start to get to reuse a lot of really excellent design thinking that has already been done elsewhere in the world, right?
world, right? So if we look at consumer and we look at enterprise and we look at like how we talk to people like obviously in both places we can talk in the room at the same time we can also message and that's the center of gravity for how we talk to most people messaging products feel very different like iMessage is what you use in personal
life it's mono threaded per person and then at work you have slack or teams and that has like really heavy threading primitives like it has channels for context separation those channels within them have threads with advanced controls and I There's a reason that there's been kind of this convergent evolution like WhatsApp is kind of like iMessage, right? Is kind of like other texting
Is kind of like other texting tools.
So my take is um we can kind of learn from that.
So we we decided to make the interface for your dot really simple like you're just texting with another person.
Um I think over time we're going to create more ways to at work um segment the context.
Actually space and dots we call it dots and space kind of fun name.
Those work really well together.
And so like having different spaces to collaborate with your dot in is a great way to encapsulate different contexts.
>> Well, since you mentioned it, can you explain space and kind of how that fits into the product roadmap because that's a big new thing as well. >> Totally. Yeah, it's massive.
Uh, and I'm really excited about it.
So yeah, maybe I'll go back to that point I made, right?
Like there's a lot of really excellent design work that is being done generally by other people.
And one of those things is if you're trying to collaborate with someone, there are these these media, right, for it.
There's in there's three maybe four.
There's in there's three maybe four. The three are the three main ones are live messaging and then documents right you if you include slides as documents right and they have very different purposes and all of us know intuitively when to reach for which I mean sometimes you might disagree with your team like oh I wish you would have slacked me you
didn't need to send me a doc link but like generally speaking we kind of all know what we're doing the fourth is maybe functional tooling right so like you're in Salesforce and you're in Figma or something and so you're doing a specific workflow if you're building something that is messaging and you want it to really feel like messaging you have to start to be disciplined about the form factor, right? So, if you
So, if you compare talking to your dot to to talking to chatbt, the dot's form factor has like actual message bubbles.
And if you were to put a very long response in one of those message bubbles, it doesn't feel as good as it might feel in a long form generated chatbt response, which also may not feel as good as like an actual document, right?
And so what we wanted to do is we kind of hard pivoted the design towards messaging.
And then you're like, well, I need somewhere to list uh long form content, right?
If my dot brainstorms like 12 ways it can help me and it sends me 12 texts, >> I don't know if I'm thrilled. >> Right. >> Right.
But if it generates a space for me, like a page in a space with a checklist, I'm actually pretty happy. Yeah.
>> So these things kind of fit well together and allow each to be more opinionated.
And so, sorry, this is an incredibly long-winded answer. Sorry.
But basically, space is a collaboration. I'm like laughing.
I'm like, "Wow, you didn't answer this question at all."
But yeah, so basically space I you could cut me off by the way.
>> So space uh is a collaboration workspace that is built to be I wouldn't say agent first, but it's built for sort of equal weight agent and human collaboration.
And so there's a lot of really cool things that you can do with that.
For example, uh one thing that I do is, you know, I'm on the product team is often I'll write out a doc and you know I'm stubbing out content and often like teammates would go do something.
So I have to to action that I have to go ping them like I'm like in Google Docs and I'll like write something and then I'll command tab into Slack and like hey can you go get this Google doc link and like fill this thing in.
So now what I can do is I can just be tagging my dot continuously in there be like hey flesh this out. Hey pull that data.
Hey go ping this person and ask them if this is done and if it's done check this box.
Um, and you can get into really deep flow with agents.
Um, another really cool thing you can do on a page is you can set agents instructions.
It's kind of like an agentmd for a repo except it's on a page and then an agent can just like keep that doc up to date.
And perhaps the third and last thing I'll tell you about it that I that I really like is like if you look at documents as a format historically they have had a very hard constraint which is that they needed to be optimized for human authoring. Right?
Like if I could build you the awesomest document editor in the world, but if it's hard to type in it, you're not going to use it, right?
Because because you have to type in it pre-AII.
Interestingly, we now have this thing called HTML, which personally I don't find, at least modern HTML and CS, I don't find to be very friendly for human authoring, but it's much more expressive than just like a plain document format and agents happen to be very happy in HTML.
And so space like naturally integrates like visualizations, HTML visualizations.
So, a lot of the the the pages that we have internally at OpenAI, like when we're discussing metrics, we'll probably have like a live chart that an agent is keeping up to date that you can click into and ask the agent to customize.
Uh, or we'll have a prototype, like the way we'll look at a design decision now is we might have clickable prototypes like literally in the page.
Um, so when you start to think like, hey, we we're building agents and we're going to talk to them in a messaging interface.
So, we need a sort of corresponding document interface and that document interface can be agent first.
There's a a lot of magic kind of comes together.
>> One of the biggest pains I've had running sources has been managing my books.
I'm not an accountant and all the software out there is unintuitive, clunky, and detached from where my transactions actually live until now.
Mercury recently launched Mercury Books.
It's AI powered accounting that works with your Mercury banking and credit card transactions, plus external cards and payroll systems as well.
You can choose between cash and acrual accounting and quickly generate reports for your cash flow, P&L, and balance sheet.
Like everything Mercury does, the design of Mercury Books is super approachable and clean.
I finally don't feel overwhelmed when I'm trying to understand the big picture of my business.
I love that AI does most of the heavy lifting, including for things like autocategorizing transactions, and that Mercury's Command AI agent can handle tasks end to end.
You don't need to be paying for separate accounting software anymore.
And you can even invite your accountant or CPA to work alongside your Mercury books with you.
This doesn't cost extra, and Mercury will even help you find a human accountant if you need one. Visit mercury.
com/books to learn more and apply online in minutes.
Mercury is a fintech company, not an FDIC insured bank.
Banking services provided through Choice Financial Group and column NA members FDIC.
AI is only as useful as the context it has.
But when that context is scattered across tools, threads, and DMs, your team and your AI agents are flying blind.
That's the problem Jira by Atlassian solves.
What's the goal tied to your project?
What got decided last week in Slack DMs?
Atlassian's teamwork graph pulls all of the valuable pieces together from Jira, Confluence, GitHub, Slack, and more.
So nothing falls through the cracks.
You get 44% more accurate results with 48% less token usage.
With Jira, you can easily share your work context with the AI agents you already love, like Claude, Cursor, and GitHub Copilot.
Assign them work directly or connect your tools through MCP.
All of this lets you spend less time digging through endless links and messages, chasing down what got decided and by who, and spend more time actually shipping. Learn more at jira. com. That's jir. com.
I spend a lot of time context switching between meetings, often with no time to process one before the next starts.
Thankfully, Granola runs in the background the whole time.
It's an easy to use AI notepad for meetings that works everywhere, even on phone calls.
I use Granola to recall what was said in meetings and create helpful summaries.
I use it every day to stay on top of what I need to get done with my team.
It connects to my email and suggests follow-ups for me to quickly review and send, saving me valuable time.
Granola isn't just a core part of my workflow.
It's basically my second brain. Try Granola at granola.
ai/sources and use the promo code sources for 3 months off.
Framer is the AI website builder that powers the sources podcast website at podcast. sources. news.
Framer brings AI agents into the same canvas where your website is designed, managed, and published so you can move faster without giving up your taste or control.
I use Framer to make the sources podcast website be the destination for everywhere you can find the show.
Plus, you can also see recent issues of my newsletter.
Start building with agents for free today at framer.
com/sources for 30% off a framer pro annual plan. framer. com/sources.
Rules and restrictions may apply.
What do you think if you were to boil it down sets dot apart from Muse, Grockbot, Instinct, this new form factor that everyone's really betting on right now? >> Two things.
First, most capable agent.
second safest and most trustworthy agent. Those are the main two.
We actually, it was really clarifying when we decided that those would be our two.
Allowed us to make a lot of like strong decisions in my opinion.
And you know, you've been in the alpha and I'm sure you've tested it for things that you were using other products for and noticed where it was good and noticed where you weren't >> it was very aggressive on permissioning.
Um, and I think that you know that was you guys are ironing out kinks, but I think it was it was very much are you sure you want to make this purchase?
like not being um I think there's a balance between how how proactive and and how much agency do you want it to actually have and is that a trust that's earned over time?
Is it something you give immediately?
Do you do it on a sliding scale based on what the user says they want?
I imagine that's pretty hard to build into the product and I think you guys seemed like early you were going very heavy on permissioning and maybe you're find finding more of that balance right now. >> Yeah.
>> Yeah. Uh so the way that I think about it is okay you know we're building the most capable and trustworthy um agent we want people to give it really hard tasks and really hard tasks probably come with a lot of responsibility depending on the nature of the task right >> we're talking like landing PRs or
evaluating hardware or planning a launch um and so uh well we could talk about why you know it's the most capable assistant but bringing this to to your point about safety we then decided okay we have lots of dreams and ambitions for how this thing will work in the future But to start, we're going to be conservative. And so we made trade-offs
And so we made trade-offs in terms of capabilities, we made trade-offs in terms of product, and we made trade-offs in terms of rollout to sort of like have this most trustworthy agent.
And that will slowly widen those over time.
So for instance, um you know, for various reasons, you could debate which model you might want to launch with, but we decided to launch with our most capable, but also our most aligned model, which is GPT6 Astra, >> which I want to talk about, but on the safety piece, talk more about that.
So that's so that's the model side right.
Um next there are many features that we want and many of our users want.
So for example once you give your agent its own identity in say Slack or you know we were testing email for example you immediately want it to email everyone or Slack everyone and to be emailable by everyone right that's like an obvious feature we want.
We don't we didn't ship with that feature.
When you give your agent its own identity in Slack it only responds when you tag it unless you explicitly tell it hey like listen to other people too >> right?
This is because we think that we definitely want it to be multiplayer, but we want to be there getting there very carefully, right?
Like the moment your agent is answering unsolicited questions from other people, you need to be super thoughtful about what is the boundary of what it can share or not.
>> You know, that's a little bit on product.
We could talk a lot more about the controls in product.
Maybe the last thing I'll say is rollout wise, >> yeah, >> very intentionally for us, we're starting with pro business and enterprise, business premium that is.
And so that is our most AI literate audience.
audience. M >> um and you know you you could talk about going much broader much faster but for us we want to do this really carefully go to that audience see how things go tune etc and then go to a broader audience >> pricing and accessibility I want to get to but on the safety side I think especially when people are reading about
hacking and rogue agents and all these things and you and you know the Sam and others and I just did the pod with him about that about a month ago like the alignment stuff you guys are doing on the frontier but people are rightfully like concerned about are these agents hacking are they going and maybe being reticent to hand over their personal data. How do you guys approach credit
How do you guys approach credit cards, uh, personal information, that kind of stuff? Is it fully sandboxed?
How does it how does the architecture relate to the rest of chat? >> Yeah.
So, personally, I'm very happy that there's this conversation just in the market generally about the level of responsibility that people are giving agents.
I think it plays well because we are taking that so seriously in this product.
One thing I do want to call out, those incidents that we disclosed were done with models that were not intended to be shipped internally and were not running in the same level of sandboxing or or or the safeguards that uh are in the product, >> right? Shipped externally. >> Yeah.
So yeah, so those models are very different, you know, like GP6 Astra is an incredibly well aligned model and then there's many many layers of of of safety and defense and depth built into the product.
So yeah, like to just to talk through then the architecture and how that works.
Uh the way that I would think about it is you start with the model very well aligned.
could talk your ear off about that.
You then run that model in a harness and we are running our model in the codeex harness which is like a very well- tested harness which we have many evaluations for including evaluations from a safety perspective on on how that model performs in the harness.
We then modified that harness to enable the product to have sort of this persistent agent behavior.
Uh you know for example that means things like running it for really long periods of time.
It means things like it delegates to sub agents and so forth.
So we then actually reev we designed that harness to make sure it's safe and re-evaluated things our tests in that harness.
So for example uh when you have this harness you have this persistent agent and um you know like I said we were experimenting with email and so we evaluated hey if you email it can you xfill in data from it?
Um and so we ran those eval okay we had a model try to hack it uh you know prompt injected and you know no successful prompt injections there.
So you know that was great but we so we basically do all all that all the way at the harness level.
Then in the product itself there's kind of two sides to it or maybe three.
The first is the default behavior and rules that you can customize.
The second is how we enforce those rules and then the third is how we give you visibility on what happened you know with all those systems running.
So to talk about the model behavior we have some very conservative defaults that you noticed.
By the way those are those defaults that you tested early on.
We started with a very conservative policy and then we whittleled away at it with user feedback.
Um, so you know, expect >> accept the terms of service of the retailer before you buy the shoe. Like that kind of thing. I I I got it. It's by the shoe.
>> Thank you for being an early tester. We really appreciate it.
Thanks to you testing us so folks in the room don't have to deal with all of the same things.
>> Um, but you know, basically every time we got feedback like the one you shared, we would like go discuss that and be like, "Okay, how do we want the model to behave here?"
And so we really we really tone down.
We call those overconfirmations when like it really it's like you didn't need to confirm that you you would have felt comfortable.
>> Um so we have this default behavior and the way the model is instructed to behave is that it can it does proactive research for you over the tools that you've connected it to and it's thinking about how to help you but it doesn't do anything to help you without uh you confirming first.
So for example, if it noticed that there was a conflict on your calendar, it's not just going to move the meeting.
It's going to say, "Hey, I noticed a conflict.
do you want me to move the meeting? Right?
Or if someone's asking me a question and it knows the answer already because it has all my context.
It's not going to send the answer.
It's going to ask me if I want to send it.
>> So you can then prompt it to give it more instructions.
So like, hey, next time if I'm on PTO, like on a holiday, like you can tell people that I'm on holiday if they're also work colleagues or something.
>> Um, and then we also have a layer of custom rules that you can set that you can go look at in the UI.
So all of this is like the model trained to follow your instructions.
But then we add a layer on top of that which is well what if the model doesn't actually follow instructions.
Well, we have a layer called auto review which is another model that's observing and kind of like policing uh the base model.
Um and so this is a layer that is aware of like your rules, the custom instructions and all that and just make sure everything goes well.
And then so that's like the controls.
And then finally, we have this auditability which is anything you're doing in the product or the model is doing in the product, you can go see the sub aents that it's running and exactly what those sub aents are up to.
>> You've teased, you know, email as a thing that may be coming when can this be in iMessage, WhatsApp, is it going to be contained to chat?
>> We we definitely want it everywhere that you work.
So yeah, we're intending or everywhere that you communicate, right?
You know, going beyond work to consumer as well.
So we have texting coming soon.
Um it's a story for a fun time about uh the journey to enable that. Okay.
>> Um but uh can't share more now other than that it's coming soon. >> Okay. Yeah.
>> The web part of this is something that is maybe under discussed as the technology is moving so fast but you know Amazon got a lot of headlines for blocking Muse uh recently because people were using Muse to to buy things.
Um then there's companies like Shopify which seem to be leaning in.
They're partnering with a bunch of companies on this space and I think everyone's kind of waiting to see what happens like how how do um the the sites and the companies that people use to buy things need to adapt or not to the rise of agents like this.
I'd be curious to hear you talk about that and then how you guys are approaching this from a partnerships perspective.
Are you just kind of um going in and saying like we're going to let people do this and then we'll we'll figure out a partnership or you saying no like you can't crawl this until we do a partnership.
How are you how are you thinking about that?
>> I think we're going to discover together with the ecosystem how this should work.
How how like shopping and payments should work and there are different approaches you could take.
You could just go and be a bit of a pirate.
Yeah, there's a way there's a way to do things.
Um we are taking a very partnership and ecosystem first approach.
You know, a lot of the announcements today at Devday were really all about the ecosystem and so we're doing the same.
So uh for instance, Webbot O is something that we work really closely with Cloudflare on.
Uh, so we actually, our agent declares explicitly that it's a bot.
>> Um, and so it's something we're actively thinking about.
We're in discussions with many partners for how to enable this, but we don't want to take the approach of just saying like, okay, we're just going to go out and like figure out what happens.
We we kind of want to be measured in our in our approach with customers.
And we think this works for us also.
So, you know, this is a little bit why like we steered towards um going for productive use cases for people at work because this is where we think like there's a lot of demand for the power of AI there.
Um and you know there's a lot of really well- aligned incentives for people trying to get things done together.
So, we're kind of hoping to land with this like incredibly capable assistant um for people trying to get ambitious tasks done and then broaden uh access and capability from there.
The only place I could see incentives maybe not being aligned is are companies that are engagementbased or have huge advertising businesses that rely on humans seeing those ads and now agents are all of a sudden crawling their services and doing things on behalf of humans who never see the ads and that feels I mean a lot of the web is supported by advertising.
I don't think anyone in the industry, not just you guys, has the answer yet for how and I'd be curious even if you're just kind of brainstorming like do you have any idea for how this is going to net out?
>> Well, this is why I was saying actually I was saying on the other side like if we're using if we're trying to get things done together, you know, like the example of like the hardware engineer like looking at factory defects, that's where I feel like the all the all that tool chain that that that engineer is using I think is like very well aligned to having an agent in there.
But but have you thought about the the advertising piece like companies that they they monetize based on you know people human beings looking at it and spending time with it and now their agents are maybe going to go do that. >> Yeah.
I'm not I'm not sure how it's going to play out but I think like broadly like at least my opinion here is that those companies are providing a service right and if previously they were monetizing with ads we either need to make sure that they continue being rewarded for their service with ads or we need to provide some alternate means >> of compensation for them.
I don't think there's any scenario where we, you know, where they don't have that. >> Yeah. >> Yeah.
>> Let's talk about pricing and the fact that you all are putting Astra in this which is um it's great for everyone.
Um you're not also, correct me if I'm wrong, you're not counting DOT usage towards plan limits. Is that right? >> Yeah.
>> That can't be a permanent thing.
Is this just a temporary to like get people >> No, I mean that is permanent.
But so I you this is really fun to talk about.
So, we're going to try pricing for uh for dots in a in a new way where you don't you don't run out of rate limits. Here's what I mean.
So, we have to unpack this a little bit because it's a bit of a new idea.
Like today, if you're using codeex and you consume all your rate limits and then you try to ask Codex to do more stuff, it'll just stop because yeah, >> you're out of rate limits.
>> But just when we think about this from a product perspective, I mean, first of all, it just sucks just to, you know, not be able to use more.
not be able to use more. But secondly, when this is a persistent agent that you're counting on to have your back, you're telling it things like, hey, you know, I don't know, stay on top of my flight and order me an Uber to the
airport when it's time based on the traffic or something or like, you know, some task like that that's like you you're relying on it to have your back in the future and then you accidentally like code too much the night before and it doesn't do its job. Like that's we
Like that's we don't want that. >> Yeah. >> Right.
And so we kind of had to go a little bit to the drawing board in thinking about how pricing should work and come up with some like pretty basic principles.
>> So one of those principles is that your dot should be available 24/7 and there should be no way for it to not be available 24/7.
I mean barring abuse or something.
Another principle should be now that it's available 24/7 it should always respond to you quickly.
So I started likening the way that we want to price to it to a bit of a like more like a contractor, right?
Like if you have a contractor who's very responsive and very polite, they will always answer your messages.
Now, if you ask them to do work, there's sort of a varied amount of work that they're going to do for you based on like how you're compensating them, but they're always available to you, right?
And if they let's say they I don't know, they set up this beautiful installation behind us >> and then you have a question about it, you can ask them for that.
And they don't they don't have to charge you like by the hour to like answer how they did this, right?
That's just the question you can ask.
So, this is kind of how we want to price.
We want your dot is available 24/7.
you pay a fat flat monthly amount to have access to your DOT.
Maybe uh you can pay more another flat monthly amount if you want like a more powerful DOT running on a more powerful computer uh or with more powerful models or whatever and or running faster or able to do more parallel work.
But at the end of the day, you pay this flat amount.
It's kind of like buying internet. Sorry, more analogies.
As you can tell, I'm still figuring out how to talk about this.
Um and it does sort of a varied amount of work for you based on based on how much you're paying, but it's never not available.
And as maybe even as you approach your limits, it can start telling you like, hey, I've I've done a lot of work for you.
Um, so I'm going to start using a more costefficient model or maybe is it okay for you if I do this work overnight?
Because if I do it overnight, it's incredibly affordable. Like is this urgent?
And it'll kind of have a sense of whether or not that's urgent, right?
Uh on the other hand, maybe if you haven't used it at all that month, it might text you and say, "Hey, like you're not utilizing me fully.
Like here are some things I could do for you."
Um, so anyways, those are those are some of the the thoughts on how we're going to set this up.
The way the way we're doing pricing is it's going to be long-term.
It's going to be a separate sort of bucket from your main codeex limit because we're doing this very new thing with it.
Uh, for the next month, we just have very generous limits.
And this is going to be as we build out and sort of measure and tune this new setup where you have a model that is aware of like how much work it can do for you, but is always available for you.
Why not just do like a take rate business if you believe that people will use this to buy things and make transactions at scale?
It just take a small vig on all activity that flows through.
>> I mean, I think that might make sense for some other business models where you're betting on like transactions being like one of the main things that you're doing, but like I said, we're we're not really betting on that.
Like we're betting on people doing really ambitious work with this tool.
>> Um, and so they're not necessarily buying anything when they're doing that ambitious work. >> Okay. Okay.
Well, on that note, u you know, the bet that you're making that brings me to probably my biggest question I've been thinking about since the I started testing this and certainly since the keynote uh this morning is what does this mean for chat? You guys have 1.
2 billion weekly users of chat GBT now.
Um and you've got the work tab and chat obviously still is is massive.
How do you see these things coexisting?
Is dot the future interface of all of it?
>> Chat is definitely here to stay.
So just for anyone who's not following chat you know fastest growing consumer assistant um you know over 1.
um you know over 1.2 2 billion users and then we have our agent products uh codeex and work which have 35 million uh users and also growing and uh and now we're just launching dot right and so the way that I look at is like chat is the is the consumer assistant for everyone um right now we have this different product called work which really we should just merge with chat so
that that will happen >> codex is a product for developers so great and then we have dot which is us in a way like the way I think about it is like we're a little bit repeating the codeex playbook if we're going to take this frontier capability >> and we need to ship it in sort of a careful conservative and very powerful way and then figure out how to bring that capability to everyone. So we're
So we're starting with dots as a sort of like premium standalone product.
It's starting with just pro.
We need to go get it to plus you know after we've like figured out um you know learn learned how everyone's using it the safety story the scaling story all that and then we eventually need this capabilities to get all the way to free.
Uh, and I think the way that we get all the way to free is actually we're baking this all into chat and making chat or chatbt into just like an incredibly powerful cons.
>> It's so much more capable.
I mean to me this should be the interface for everyone because it can do so much more. >> Yeah.
So >> you can expect a lot of the learnings from this to make their way into chat. Okay.
>> But I don't think we want to ask 1.
2 billion people to like change what they're doing or how they're using things.
We just want it to just be upgraded.
>> So you see it as more of a gradual thing over time. >> Yeah.
when you think about the road map uh over the next year for dot like if we're talking a year from now what do you think is the most meaningful change?
>> Can I tell you a funny story? >> Yeah.
>> Uh and we can get back to the question. Okay. Okay.
Okay. Okay. So when I first joined OpenAI it was is through this acquisition and uh and we had a meeting and and some of my team were in the room and um and uh Greg was there and so everyone you know talked about how excited they were that we were all
together and then Greg asked does anyone have any questions from the multi- team which was our startup and everyone no one has any questions we're all a little nervous and I was the CEO of this company so everyone looks at me it's like well you have to ask the mandatory you know obligatory question. So I'm
So I'm like ah what am I going to ask? I'm nervous.
Uh, and so I'm like, "Okay, in two years, what does success look like?"
And then Greg, I I always remind him of this.
He looks at me and he says, "In 2 years?
That's an incredibly long period of time.
That's not the right duration to ask about.
You should be asking about success in 2 months."
>> And I was like, "Oh god, I'm already fired." >> Okay. Well, don't fire me. Okay. >> Yeah. No, no, no. It's all good. But so sorry.
So, >> no, but I I get the point.
Like this it moves so fast.
like 6 months ago you probably couldn't imagine all of this. >> Yeah.
>> So it is probably an unfair question.
>> So what I can say though is um if I think about what's next we are approaching shipping these capabilities incredibly carefully and next what I would like to do is to start looking at some of the things that we didn't want to do immediately and landing them.
So for example let's get to more users.
Let's not let's get beyond pro.
Um, also for example, let's enable your dots to talk to other people >> and maybe even other dots, right?
We need to be super thoughtful and careful about how we approach that.
We just shipped uh basically in my mind a reboot of Codex cloud.
I'm really excited about it cuz I worked on the first Codex cloud. It's super powerful.
Your dot also super powerfully can control your local codecs. We just shipped space.
All the interplay between those things right now, I think, can be really sanded down and made super smooth. >> Yeah.
>> Um, so I want to do that as well.
Um the other really interesting thing that we haven't talked about I think yet is specialist dots. >> Yes. >> Right.
And so we have this idea of like there's basically like two things going on. We have assistant dots.
That's everyone has an extension of themselves and then separately you know if you're you're an organization and you have some function or workflow that you would like to have uh more like more automated you may not want that being done by individuals who have their own personal dots taking on that work.
you may want to actually like control that centrally, right?
And so you can create specialist dots which specialize in doing a particular thing and which are centrally managed by it, which have their own identity completely separate from any sort of individual user and which are accelerating an entire team or an entire organization.
And so I think the other thing uh that I'm really excited about over the upcoming months is bringing specialist thoughts to life uh and entering this world where it's like we have us we have our our personal assistant dots and then we have assistant dots much more broadly deployed within enterprise but that is another place where we're starting like very slow very careful working with a handful of customers. >> Got it. Okay last question.
What's the craziest thing your dot has done for you that you're like wow I you it happened and it blew your mind.
and it blew your mind. The craziest story that I've heard uh it was like on someone I was working with as we went up to this was this engineer woke up and realized that in Slack he had received a bunch of kudos which is like you know props like awards for having like fixed
a bug before it made its way to production >> and it wasn't >> and the engineer was like uh what happened uh by the way I tell this story with a little bit of reticence because that's actually not something that you would want to happen without the person knowing and so you know we're building a lot of souls around that. This is why we
This is why we do small scale.
This is like early on when it was just our team testing it.
Uh and yeah, his dot he had asked his dot to stay on top of the feedback channel and the dot like saw the feedback um and then identified an issue and then spun up a PR and uh you know didn't set that PR in draft mode.
So like you know DOT should set PRs in draft mode uh and then uh set it to automerge because that was that engineer's personal preference.
Um and then and you know the dot had figured that out, right?
And it was like, "Oh, you always create your PRs directly and you always set them to automerge."
And then someone else had approved the PR and it did it.
So, to me, that was that was a really cool uh cool moment for sure. >> Wow. Well, knock on wood.
Uh we all get to have moments like that. Thank you, Alex. Thank you. >> Sure. Thanks for having me.
>> Banking should feel like modern software.
Get everything you need in one place. Visit mercury.
com to learn more and apply online in minutes.
Mercury is a fintech, not a bank.
Check the show notes for details.
Granola is the best AI notepad I've tried.
It works everywhere on a video or phone call, in person, or an Apple Watch. Try it now at granola.
ai/sources and use the promo code sources at checkout for 3 months off.
Jiraabiadlassing is where your team and your agents work from the same context. Try it free at jira. com. That's ji. com.
Framer is the AI native website builder that lets you build faster without giving up control. Visit framer. com/sources for 30% off.
Rules and restrictions may apply.