Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
May 05, 2026
·
Seattle
Nate B. Jones - AI News & Updates: Agents Inside and Outside the Enterprise
Explore how AI agents, both internal and external to the enterprise, are shaping the future of AI news and updates in this mainstage presentation.
Overview
AI News & Updates: Agents Inside and Outside the Enterprise
Video
Transcript
Generated 3 months ago
Summary
Generating a talk summary...
View full transcript
Speaker 0: So. We have a dead mic. Still dead. Well, it, like, it came to life. Oh, there it is.
Speaker 0: You gotta just gotta Yes. Bring it up. Yes. From here. I can also talk loud.
Speaker 0: Well, it's great to be with all
Speaker 1: of you.
Speaker 0: I feel like I'm more excited to see what you're building than anything One, and so I'm very excited for the rest of the talks that are gonna come after me. But in terms of the news and updates that that came to mind as I was asked to sort of talk a little bit about what's going on, I think I wanna talk about this world that we're all building for that exists both inside and outside the enterprise or inside and outside the Staff up at the same time. And One world is the world of agents. Right? We talk about agents all the time.
Speaker 0: But I think that there are different dynamics when we are, building for the market versus when we are building internally for our own tooling. And I wanna just take, you know, 10 minutes, explore that a little bit, kick it around a little Agent, maybe we can have 1 or 2 questions, and and get going. So we'll we'll start with our customers first. I'm former Amazonian, customer obsession sort of got grilled into us. If you're building for the Internet now, the the way I like to put it is the Internet was the attention economy for a long time since, I don't know, 2000, since Google.
Speaker 0: How do you get attention? How do you keep human attention? How do you drive attention through the funnel? I, you know, I I was in marketing for a while. I built Partner tools.
Speaker 0: And it was always about how you connect with a person who's making a buying decision, whether they're buying a pair of Nike shoes or, you know, whether they're making an account decision for a SaaS. And you need to connect with them 1 way or the other. And the products you build have to be for that human world. And the thing that's changing underneath all of it like, we talk about agents, but if you think about the underlying dynamics that are shifting, is the Internet is shifting from an attention economy to an interpretation economy where you have to assume that whatever you are putting on the Internet is going to be filtered and interpreted through an agent, through an LLM before it ever gets done. And so a lot of the question then becomes, what kind of signal are you putting out there in your product and what you're building that survives effectively that compression action from the LLL that pops out and still offers genuine signal.
Speaker 0: And I'll give you a I'll I'll give you an example. I like to make things concrete. So, I had, until very recently, an ancient sound, and it was terrible. The wires were not good. My receiver wasn't good.
Speaker 0: I was limping along on it, and I finally decided I'm gonna bite the bullet and I'm gonna buy a new sound system. I did not go to the Internet into Google and just say, hey. Can I have a sound system? Please show me my options. That was not the the the ad selling exercise that Sergei wanted me to do.
Speaker 0: I did not do that. I went and said to my One, and I said, One. And I tried this with both Claude and Chad Chiboutier. I said, hey. I want a sound system, and I want you to give me options.
Speaker 0: And so already, I'm filtering the Internet through. I never looked at a web page until I clicked buy. And after, frankly, after stretch Agent sessions, I don't even know what I have to do that before. And so I went through I went through a process for over a week where I actually gave it the dimensions of the room. We talked about speaker placement.
Speaker 0: We talked about budget. We talked about what kind of wires they would need to get. All of that happened inside the LN out of view of the marketer, out of view of anybody else. And it was depending on the product availability being something that the that the LLM could filter through and that the LLM could actually pass signal on. Now I have no idea if I actually had what I would describe as the best consideration set.
Speaker 0: I don't know. Right? I just know that I'm in the habit of using an LLM, and I'm gonna use the LLM to do my purchase Intelligence, and the LLM is gonna provide me a consideration set. And so if you're building for that world, the question then becomes, how do you present product in a way that allows the product information to map reliably to user intent? I'll give you another example.
Speaker 0: I'm a former coffee person. I love coffee. And this is a wonderfully vague coffee request that I think you can meet with an agent that illustrates the data problem we face as builders. I want authentic coffee. Now what does authentic mean?
Speaker 0: Right? It could mean anything to anybody. But an agent who is working with you, who understands your history knows that when I say authentic coffee, what I mean is it naturally processes the Ethiopia that is sort of lightly roasted. I want it roasted within the last 2 weeks, and I want it on my door in 2 days. I have lots of requirements.
Speaker 0: And the agent knows seed well enough that it knows that I need that. And so it's going to go and interpret that. It's gonna go and look for that product. And if the product metadata is there, then it's gonna be able to find that and include that in the interpretive center. And I'll go farther.
Speaker 0: If you have an agent who doesn't have that background with you, you don't talk about coffee with your agent. You have much more important things to build with the agent. The agent then has to go out and infer and map the meaning of your intent. And the way it does that is by looking to see what products Sn the competitive set actually offer that kind of mapping that maps and says this is what authentic means. Right?
Speaker 0: This is the this is the authentic interpretation we have for coffee. And it's like, oh, thank god. Like, someone finally helps me, the agent, understand what's going on here. And that is like a tiny micro consumer example. But that's also what's happening if you talk to people who are marketing for enterprise right now.
Speaker 0: If they're marketing SaaS right now. If you're marketing tools right now. So much of it is getting compressed through the the interpretation layer that we're all bolting on to our experience of the Internet. Google even does this. Right?
Speaker 0: You go to Google. Are you leaving the Google search results page? Are you reading the Google answer and say, yeah. Good enough. I'm moving on.
Speaker 0: And and this is big enough that it's affecting organic traffic to sites because people are reading the AI summaries that Google provides. And so increasingly, the interpreted Internet is becoming the way we experience the digital world, the way we buy digital products and services. And so if we are building, what that implies to me is that we need to be more opinionated. Because if you're not more opinionated about your software, you're going to get flattened into the Internet average for your category. You're going to get sort of compressed into the middle part of the bell curve and to be like, well, you know, 1 of 3 other SaaS tools in this particular category and you don't have fun with that.
Speaker 0: Right? What of, you know, 18,000,000 AI tools that you still have right now? And there's no opinions there. So that's the outside look. That's the customer obsession look.
Speaker 0: That's that's where I think we are going, what we need to think about when we don't. Like, at that group right when we're building. If we come inside the firm for a minute, if we look at agents inside the firm, a lot of what I see is people obsessing over what I call sort of fancy chains of meaning. Right? Like, they're thinking about so I put the element here and the memory here One then I put my tool registry over here and this is my cool like stack of Lego bricks that allows me to build a really neat internal data tool or a really neat tool that's a pipeline for marketing materials or whatever we're building inside.
Speaker 0: But I think at roots, 1 of the things that I see that's missing is we are missing the understanding that the agents that we are working with need to be treated as if they are growing up quickly One the harnesses we are putting around them are very modular and very at least need to be ready to be very temporary. Boris Trini talked about this. He had gave a talk on I think earlier this week, and he said they effectively have to rebuild their harness at Anthropic for every model because the models have different strengths and weaknesses and model capability to scale. And that got me thinking how often are we thinking about not just modularity and what we put out for customers, but enough modularity internally that when we have the next decimal point model drop, instead of saying, oh, gosh. What do we do?
Speaker 0: We say, let's swap in and assume that we don't have to have the planner component of our workflow anymore because the new model seems really good at it. We'll just drop it out. It shouldn't take too long, and we'll be back at it in a regular pipeline. And I think about that a lot because I think that a lot of the nature of how we build is basically enabling constructor sets internally that open the door to a wider range of inputs to the build process as long as we maintain good code hygiene standards as long as we are guarding the value of the architecture that we are putting together. Like, I was talking to a designer, I think yesterday.
Speaker 0: And and we were talking about beaming and he'd lost his job and sort of how how he changes and all of that. And 1 of the things I called out is just like, I'm not technical. I'm never gonna learn to code. And I said, it increasingly does not matter. And it doesn't matter because you're gonna write code.
Speaker 0: That's not the point anymore. It's 2026. It doesn't matter because your engineers should be setting up a pipeline that allows you to submit clear intents, drive the design polish you want as long as the code you're writing passes their emails. And you should have a patient that is driving to pass those emails. And then you don't have to know if it works because the emails are so good.
Speaker 0: And your job is basically to put the design polish into the email, so that's really clean. And the engineer's job is to make sure the non functional requirements aren't shipped. Pardon my French. And so often I see less attention on effectively setting up our modular pipelines so that it it enables us to build for agents that are getting smarter faster. And so I want you to think about the idea that by Christmas time your agents are gonna have 5 or 10 x more blast radius to connect inside the enterprise and you have to build your pipelines now Sn that you're thinking about that modularly One you can extend that surface area.
Speaker 0: And I think that's something that we get so wrapped up in what the tools can do now, and it's super cool. But if we're responsible, we have to think about where they're gonna go and how we build now so that we suffer less come November, come December when, you know, CHI GPT 6 drops or whatever number it is drops. And suddenly, everyone's breathing down our necks saying, well, you gotta go for that. Right? You make sure you have this thing done by whatever.
Speaker 0: Build modularly and we're not gonna have that issue the same way. And I know that that's easy to say, and I know that it is hard to do. And I know it's increasingly hard to do because when we built modularly for developers before, it was building modularly for everyone inside the engineering team. And now when we're building modularly, like Stripe, this got kind of covered up, but Stripe built a platform that ties their designers and product managers into the build flow, and it's more visual. It's less sort of in the terminal.
Speaker 0: And they took the time to do that to wrap more people in. And so you have to think across many more teams as you're putting these pipelines together. So I get that it's harder, but I think that's where we're going. So that's my thought.
Speaker 1: Okay.
Speaker 0: Cool. You mentioned evolves and practically evolves are very hard. And, what what would you recommend in addition to modularity, on on evolves to be ready for this MR how do I wanna put this? I'll divide it up very neatly. So I think One, 1 of the patterns we're learning as of Agent a roll is that the idea of having a separate agent at roughly the same horsepower, roughly the same caliber that checks the work.
Speaker 0: That's the only job that's checked the execution and the intent of the original agent and make sure it stays on guard seems to be working very well. And so this is what auto review in 5.5 in codex does. The the auto review essentially instantiates a separate 5.5 agent and its only job is to make sure your original intent was followed. And the and the agent that's like happy go lucky and gonna try and happy his way around and get whatever he wants can cross certain lines. So when you think of that model, what it suggests to me is that we are getting to a point where you want to use evals as a constitution for your intents.
Speaker 0: And so the evals should speak to what are the functional requirements of this product that we are putting into production. How can we how can we shape those requirements Sn it is past fail? So it is extremely crisp and clear to a guard agent, a review agent, whatever you wanna name the thing. I think that, CJ, you called it a watch tower agent or something. But you you pick your name.
Speaker 0: It has to be able to go through, run a test, and basically say, is this software passing this particular unit test in our joint eval or not? Yes or not? It's binary. It should not have room for ambiguity. And that's part of the art of writing eval is to make sure that you minimize the interpretable ambiguity.
Speaker 0: And where possible, you run scripts that are deterministic to do the testing rather than using an element to do it all. And so I've seen cases where people are running something all weekend. And, yeah, there is an LLM that acts as a watchdog that's doing some of the eval but there are a bunch of scripts that are running that are like linting the code, whatever else they're doing, you know, checking for accessibility and contrast and all of that is deterministic and if it doesn't pass, there's no way for the element to gain that. It just gets told to fail and go back and do it again. Alright.
Speaker 0: Probably have time for 1 very short other question. Yeah. I know I can do an hour.
Speaker 1: Hey. So it's not working. Hello? Hello? I can't
Speaker 0: hear you.
Speaker 1: Hello? How's that?
Speaker 0: No. No. No. No. Just a few questions.
Speaker 0: Okay.
Speaker 1: So I just I just have a comment on that. I do a a lot of Agent ID work. I'm I'm gonna present it from later. Every single thing CLI everything anything I do, every time there's an epic or a break point, I always build in code reviews One code reviews against the plan. You have to do both of those, review against what you're planning to do, and then just do a general code review.
Speaker 1: Every single epic, every single breakpoint, every time you're doing a development, or else things just the agents just go wild. You have to check on it.
Speaker 0: That's a great point from Leon, building, reviews for Agent, but he does actually get to speak later. So you can talk about that later when you run steady. That sounds sounds sounds good. Okay. Alright.
Speaker 0: Thank you very much, David. Very happy to hear you.
Tech stack
Compose Email
Sending...
Email preview
Loading recent emails...