AI Engineer

Give Your Agent a Computer — Nico Albanese, Vercel

878 summary words 4 min summary Watch video

Start with the signal

4 min read

Summary

Give Your Agent a Computer — Nico Albanese, Vercel

Main Topics

  • Building AI Agents with Vercel and AI SDK 6
  • Tool loop agents architecture
  • Setting up Next.js development environment with Vercel CLI
  • Integration with AI gateway for inference
  • Three Core Building Blocks for Agents in 2026
  • Agent runtime (managing the loop and context)
  • System instructions/prompts (behavioral guidance)
  • Agent tools and capabilities
  • Agent Tool Types
  • Custom tools (developer-defined functions)
  • Provider-defined tools (optimized by LLM providers)
  • Provider-executed tools (handled by provider infrastructure)
  • Persistent Sandboxes and File Systems
  • Named, persistent sandboxes by Vercel
  • Bash tool implementation
  • File-based memory systems
  • Advanced Patterns
  • Context management across agent steps
  • Persistent memory using markdown files
  • Self-extending agents that build tools over time
  • Sub-agents for distributed work

Key Points

AI SDK 6 Fundamentals

  • Simplified API: Two-line installation with pnpm add ai, ai-sdk react, and zod
  • Global Provider Concept: Attach providers to all AI SDK functions by default; AI gateway is the default
  • Tool Loop Agent: Purpose-built for agents, keeps agent logic separate from streaming concerns

Web Search Integration

  • Add OpenAI provider to enable provider-executed web search tool
  • Zero additional code required—tool result automatically added to message state
  • Type-safe tool parameter definitions with location context support

Type Safety and Developer Experience

  • End-to-end typing: Infer types from agent definition across UI components
  • Benefits: Compile-time checks for tool existence and parameter types
  • Eliminates runtime type assertion issues

Context Management Between Steps

  • Prepare step callback: Runs before each agent step, allowing message manipulation
  • Enables sliding window context (e.g., keep only last 5 messages)
  • Preference for sub-agents over aggressive compaction to preserve input cache efficiency

Call Options Schema

  • Define structured inputs that modify agent behavior at call time
  • Example: Different model tiers for different customer types
  • Type-safe options passed at invocation time

Persistent Sandbox Architecture

  • Named sandboxes: Identified by string name, with multiple sessions per sandbox
  • Ephemeral sessions: Individual instances spin up/down automatically
  • State persistence: Snapshot file system between invocations
  • Feels like a continuous machine despite session turnover

Bash Tool Pattern

  • Single bash tool with description, input schema, and execute function
  • Agents excel at generating bash commands (ls, grep, find, etc.)
  • Runtime context provides sandbox access within tool execution

Memory Systems

  • File-based memory: Store memories in memories.md within sandbox
  • Inject at call time: prepare call reads memory file and injects into system prompt
  • Structured file system: Separate files for core memories, conversation history, scripts
  • Deterministic loading: Consistent, predictable memory retrieval

Self-Improving Agents

  • Agents create Python scripts for repeatable tasks
  • Scripts stored in sandbox and referenced in future interactions
  • Builds own toolkit over conversation lifespan
  • Perfect REPL loop environment for feedback and iteration

Notable Quotes

> "The big assumption that goes through every single AI SDK API decision is we want to have the agent definition being the source of truth that everything inherits from and derives from."

> "I would say memory is a file that you store in your sandbox...The file system becomes this playground, this environment for you to store in a structured way a lot of this information."

> "The fact that we just have that ID, that name, and it's just persisting whether it's alive or actively running or not. But we have that persistent state is really cool."

> "These are the building blocks of this agent that learns, that builds on itself as we go."

> "Bash is all you need. In a lot of cases, these agents are really good at writing bash commands."

> "I don't think having used this coding agent for 14 hours a day...compaction has not been an issue to me. And I was showing you earlier, I have a 95% cache token read ratio on that as well."

Takeaways

For Developers

  • Start simple: Basic agent with instructions works immediately
  • Leverage typing: End-to-end type safety prevents runtime errors
  • Use web search: Add context without custom implementation
  • Implement sandboxes: Persistent file systems enable stateful agents
  • Keep bash as primary tool: Simple, powerful, and model-efficient

For Agent Design

  • Instructions matter: System prompts significantly influence behavior
  • Provide file system: Enables natural memory and state management
  • Enable self-improvement: Let agents create reusable tools
  • Avoid aggressive compaction: Modern token windows and caching make full history preferable
  • Use sub-agents: Offload subtasks to maintain main thread efficiency

Architectural Patterns

  • Context injection at call time: More flexible than static prompts
  • Prepare callbacks: Powerful hook for message manipulation before steps
  • Persistent naming: Named sandboxes simplify lifecycle management
  • Runtime context: Tool access to shared state via context objects
  • Resumable streams: Enable background work with periodic result returns

Production Considerations

  • Cache token read ratios become critical metric (91-95% achievable)
  • Token efficiency improves with persistent state vs. compaction
  • Workflow integration adds durability (retries on failures)
  • Sub-agent pattern keeps main thread within 7,000 tokens
  • Real-world system serves 23 users, processes billions of tokens monthly
Full transcript 10396 words · 58 min read
0:14

SPEAKER_00

I now have my live deployment here and I can continue to dashboard and I've got my project. One out of five things for the production checklist. Is that all? Were you able to get that? Yeah. Cool. So we're gonna jump over to the repository on GitHub, clicking that big repo button and then I'm gonna quit out of Slack and all these things and I'm gonna jump into a terminal, head over to a directory that I'm happy with and I'm happy to pop this in and I'm gonna get clone that, oops, that's not it. I need to copy it, obviously. Click the big green code button, copy the clone and then git clone AIE London demo. Awesome.

0:54

SPEAKER_00

I'm gonna jump into that repo there and then if we jump back to the docs site, this is at nicoalbanese.com slash AIE and then that will redirect you here. If you don't have the Vercel CLI installed, you can install it with. What if you don't have the Vercel? So you can install that with npm i -g Vercel. That will probably take you to a Vercel login the first time you use it, which will pop up this signup process that you can allow and then you should be signed in and you can check that with Vercel whoami. Take, yes. Good thing is that in a small room this, the WiFi shouldn't be an issue. Were you able to log in? Yeah. Cool. [SPEAKER_03] Yep. Perfect.

1:59

SPEAKER_00

[SPEAKER_03] I think I have to, I think you never have to close.

2:10

SPEAKER_00

Our IM team is doing awesome work with the Vercel CLI. I'm a big fan. I am biased, but I am a big fan.

2:14

SPEAKER_03

[SPEAKER_00] So once we are all set up, logged in, we go into our repository that we just cloned and we run Vercel link.

2:16

SPEAKER_00

We say yes.

2:17

SPEAKER_03

[SPEAKER_00] Should be able to find it if you have the same name.

2:20

SPEAKER_00

And then we want to pull down our environment variables. And that should give us V, which I probably shouldn't show, but it gets invalidated. That should give us an OIDC token, which we'll be using to authenticate with both the AI gateway for our inference and for self sandbox, which is... Yeah, we get that a lot as well. We're internally, you can imagine we have a ton of teams. I think I'm on seven or eight different Vercel teams. And so authenticating with the right one can be... Or finding the projects can sometimes be a challenge, but making it easier every single time. Yeah, you may have already had a token in there from... A work one. [SPEAKER_03] Nice.

3:10

SPEAKER_00

Thanks. Oh, the Vercel CLI? [SPEAKER_03] Oh, yeah. Yeah, that is if you have, if it's installed with npm, you can do npm install -g vercel at latest. And that should get you on to the latest one. I had that recently as well because we just invalidated the old authentication method in favor of a new one. You can also look which vercel, and that should tell you what you installed it with. See, mine is with pmp instead of npm.

3:36

SPEAKER_00

Which, that classic frustrating problem of which package manager owns this. [SPEAKER_00] Which vercel. Were you able to set things up? Yeah.

3:42

SPEAKER_03

[SPEAKER_01] I love your shirt, by the way.

3:51

[SPEAKER_03] Homemade?

4:11

SPEAKER_03

[SPEAKER_00] It looks very professional.

4:12

SPEAKER_00

[SPEAKER_03] It's a robot.

4:17

SPEAKER_03

[SPEAKER_00] For those watching, it's TypeScript shirt. It's very cool.

4:18

SPEAKER_00

TypeScript till I die. I die on that hill. Is it getting set up or it's installing new?

4:48

SPEAKER_00

Great. Yeah. Which is the vercel link. By the way, if you haven't seen, we did ship a new landing page. Finally, after two years, we replaced my terrible landing page. Which, this was definitely not done by me. This was design engineers on our team. Very, very cool. [SPEAKER_01] Live download figures.

5:09

SPEAKER_00

I love showing this work, because it looks so good.

5:19

SPEAKER_00

Oh, and if you didn't know, if you've used open code, open code is all built on AISDK as well. And this is the most DAX quote ever.

5:26

SPEAKER_01

[SPEAKER_00] Open code uses AISDK.

5:30

SPEAKER_00

All right.

5:38

SPEAKER_00

We all set up?

5:40

SPEAKER_03

[SPEAKER_00] Cool. [SPEAKER_00] So we're good to go.

5:42

SPEAKER_00

So first thing we're gonna run is pnpm install or npm install. If you're using npm in 2026, use pnpm or bun. I love bun. And then we can just check that it is working. This is a fresh Next.js app, so we shouldn't have any issues. But if you run the dev server and then head to localhost 3000, you should see the wonderful, pretty new actually, starting screen.

5:43

SPEAKER_03

[SPEAKER_00] Is that all set up for everyone?

5:44

SPEAKER_00

Yes.

5:47

SPEAKER_03

[SPEAKER_00] Amazing. [SPEAKER_00] Good start.

5:48

SPEAKER_00

So we're gonna start with our agent.

5:48

SPEAKER_03

[SPEAKER_00] We are gonna build an agent.

5:52

SPEAKER_00

First thing we're gonna have to do is install a few dependencies because we are working with JavaScript, right?

6:00

SPEAKER_00

[SPEAKER_03] We love installing dependencies.

6:08

SPEAKER_00

So we're gonna jump into the terminal and we're gonna run pnpm add AI, add AI SDK react, and Zod. Don't worry. We're right at the beginning, so you're not missing anything. So we are, if you wanna follow along with your computer, I should write it, is there a white board or something? [SPEAKER_00] No, there isn't. This site that we're on is NicoAlbanese.com slash AIE. And that will get you to the project setup, which you can follow right here, and then we're just on, we're literally just starting off with building out the agent. [SPEAKER_00] So, first thing we're doing, I'm saying we are building with JavaScript, so we gotta install some dependencies.

6:22

SPEAKER_03

[SPEAKER_00] We are gonna be installing the AI SDK, which is this very nice two-line npm package right here. [SPEAKER_00] So we are, if you want to follow along with your computer, I should write it. Is there a whiteboard or something? [SPEAKER_00] No, there isn't. [SPEAKER_00] This site that we're on is NicoAlbanese.com slash AIE. And that will get you to the project setup, which you can follow right here, and then we're literally just starting off with building out the agent.

6:26

SPEAKER_00

So, first thing we're doing, I'm saying we are building with JavaScript, so we gotta install some dependencies. We are gonna be installing the AI SDK, which is this very nice two-line npm package right here. Our React adapter, which we'll see in a little bit. And Zod, which we'll be using for defining a few schemas that we'll be using later on in the project.

6:31

[SPEAKER_00] So, last time I did an AI workshop was in AI SDK 4, and back then everything was built along pretty much four primitives, which were generate text, stream text, generate object, stream object. Those still exist, although we've pushed as much into the text generation functions as possible, so you can now do structured outputs with just those generate text and stream text. But we've also been working on providing a more object-oriented approach to building agents with the SDK. And part of that was how much they were ballooning by having all of the LLM logic in the call site. So you had your API slash chat slash route dot TS and XJS route handler, and in there you only had 2,000 lines of code because you had the tools defined in line, and you had the system prompt defined in line. And so, part of the beauty of the AI SDK is it is lightweight JavaScript. And so you can define this once in code in a monorepo and then use it anywhere, whether it's a Next.js app, it's a simple button server, whatever it is, it's plain JavaScript. And so this was our first foray, I would say, into that, which is what we call a tool loop agent.

6:41

[SPEAKER_00] Now, we have the mastermind behind the AI SDK is German, and so naming is something we hold quite dearly. This is one of our shorter APIs, but we find that this is quite obvious as to what it does. It is a tool loop using agent. Does what it says on the tin. [SPEAKER_00] So, to build our first agent, all we need to do is define this, import this tool loop agent, and then specify a model. So we're gonna do this, we're gonna create a new file called agent.ts in a lib folder, and I'm actually gonna open this up in an editor, so it's a little easier for me, and I'm gonna bump up the font so everyone can see. Let's do... [SPEAKER_00] Is that better?

6:52

SPEAKER_01

[SPEAKER_00] Can we read? [SPEAKER_00] Yeah, cool.

6:58

SPEAKER_00

So we're gonna head into the app, actually, we can do it in the top level, and we're gonna create a new file, and it's gonna be in a lib folder, and it's gonna be called agent.ts.

7:02

SPEAKER_00

Now, in here, we'll copy this snippet, and as I said before, we're importing tool loop agent, specifying our model. The cool thing here, if you've used the AI SDK before, you may have seen this syntax where you import a provider, so you'd import openAI from at AIS, if I can type, at AISdk slash openAI, and then you could specify, this would create a model provider instance that you could then specify a model ID in here. One of the cool things that we ship with in AISdk 6 is the concept of a global provider, which allows you to effectively attach a provider to every single AISdk function that you have in your application, and by default, that is the AI gateway, and so by specifying just plain strings, you can access any model in the AI gateway. If you want, you can override this and use any provider as that global provider here, but this just makes it really easy to get started.

7:02

SPEAKER_00

So we are going to be using, what did we say, GPT-54-mini. Always to keep these things up to date. I did this last night, so this is pretty up to date, but probably when this video goes out, this will not be. This model will give us, I haven't used 5.4-mini a lot, I've used 5.4 a ton, but this gives us a nice way of, a cheap and fast way to experiment with our agent as we go. So this is our agent definition. It's surprisingly how little it is, but it literally is just this reusable agent using GPT-54-mini.

7:08

SPEAKER_00

We're then going to have to create a way to call our model, and we're going to do that by building a route handler. So that's going to be in a folder called API, which has a folder called chat, and then in there, with Next.js's conventions, we can define a post handler in a route.ts file. So again, that's app, API, chat, route.ts.

7:14

SPEAKER_00

And so one of the cool things, going back to why you might want to use this instead of just stream text out of the box, is that this allowed us to keep all of our agent definition in one place, while having just the concerns for streaming from that agent in another place. And we've abstracted a lot of the complexities for that streaming into really that one line of code, this create agent UIS stream response.

7:17

SPEAKER_00

Again, if you've used the AI SDK before, what's happening here is literally the same as const result equals stream text. You have your stream text call, and then you return result.to UI message stream response. So it's effectively the same thing. And on your agent, you also have myagent.stream, which is your agent's streaming function. So that gives us our way of, and I'm going to get rid of this, that gives us our way of actually calling our agent. You can see that our post handler is taking in some messages and just passing that in alongside our agent.

7:20

SPEAKER_00

And now we need to actually create a page and a client to call that agent or call that endpoint. So we can head to our page.tsx, our homepage, and replace it with the text that's here.

7:21

SPEAKER_00

And what we've got here is the infamous now use chat hook, which was actually the first component of the AI SDK when we launched whatever three years ago. And this is managing all of our message state and then helping us send those messages off to our route handler. So we have, we destructure messages, which will be our message store. It will all be on the client today. Error and the send message function. We check if there is an error. And then otherwise, in our markup, in the UI, we will map through the messages in our message state. So, we can head to our page.tsx, our homepage, and replace it with the text that's here.

7:24

SPEAKER_00

And what we've got here is the infamous useChat hook, which was actually the first component of the AI SDK when we launched whatever three years ago. And this is managing all of our message state and then helping us send those messages off to our route handler. So, we have, we destructure messages, which will be our message store. It will all be on the client today. Error and the send message function. We check if there is an error. [SPEAKER_00] And then otherwise, in our markup, in the UI, we will map through the messages in our message state. [SPEAKER_00] And then for each message, we will map through the parts. [SPEAKER_00] And we can, have you used the AI SDK before?

7:49

SPEAKER_00

Is this first? Okay. So, and you have as well. So, or you haven't. This is the first time seeing. Okay, great. So, most of you guys have seen this, so I can fly through this. But yeah, this is how we're rendering on the client. So, we can now run pnpm run dev.

8:16

SPEAKER_03

[SPEAKER_00] And we should now, I think I, we should now be able to say hi.

8:17

SPEAKER_00

And get back our response. [SPEAKER_00] So, we've got our chat bot, or our agent, as we should say, it's 2026, ready to go. Now, the first and most basic thing that we'll be wanting to update our agent is to change its behavior. One very core way of doing that is altering its system prompt or its instructions. And you can do that with the AI SDK in your agent definition by passing in an instructions parameter.

8:33

SPEAKER_00

So, you can do that here. I can say respond like a cowboy. And now, when I go back to our agent and I say hi, we should see howdy partner, how can I help you all today? With a nice cowboy emoji as well. So, this goes back to, I mentioned to Casper right at the beginning. What I wanted to do today is try and communicate what I believe to be the core building blocks for building agents in 2026.

8:59

SPEAKER_00

Which are an agent runtime being a way to build your harness effectively. So, how do you manage the loop? How do you manage context between the loop? That kind of thing. How do you manage the report. Folks kind of scoff their head at system prompt being something of 2023 when we were first really starting, but we'll see today how we're going to use the instructions alongside those two other components to really get the most out of an agent and really influence its behavior. So that's the starting point. We've built our basic agent here.

9:34

SPEAKER_00

We're going to do the first thing that I think a lot of folks will want to do is give the agent the ability to pull in more context in some kind of way. Basic thing that we can do today because we don't have a great use case is just bring in web search. So the way that we can do that is we're going to jump back to the terminal. I'm going to start my dev server here and I'm going to run this pnpm add ai sdk openai. And the reason we're going to add the openai provider is to be able to get the tool definition for something that's called a provider tool. And there are roughly three types of tools.

10:04

SPEAKER_00

There are custom tools, and I can actually find this right here because this is a good help for me. So there are three types of tools. You've got custom tools, which are tools that you define yourself. You provide a description, you provide an input schema, and you provide an execute function. This is effectively giving the agent the ability to run any kind of arbitrary function that you want to give it based on the context of the conversation. We then have provider defined tools. These are quite interesting. These are tools that an LLM provider will usually post train their models to use effectively.

10:44

SPEAKER_00

So Anthropic has a bash tool, which alongside all that cloud code usage that is subsidized, they are training to use a lot better. They also have a tool for computer use as well. And so the idea here that you define what the agent should do when it calls it, but they've worked really hard on the description and the input schema to make sure that the agent is very effective at calling it. The final type is provider executed. And these are tools that exist in the LLM provider's infrastructure. And so the classic example of this is web search. Oh, it's a web search tool. Anthropic has a web search tool as well.

11:15

SPEAKER_00

You don't actually provide the tool or what should happen when the tool is called, but you kind of opt yourself into the LLM provider being able to use it. And if the agent decides to use it, they will literally execute that on their server and add the tool result to the message state and return all of that back to you. So the really nice thing about it is that we don't have to write any more code.

11:37

SPEAKER_00

You get it out of the box. The bad thing about these provider executed tools is you're obviously tied to a single provider. But for today, for this demo, not so much of an issue and allows us to move quite quickly. And so what's happening when we import the OpenAI provider here, is that we're literally just using this to augment our request.

11:42

SPEAKER_00

They will literally execute that on their server and add the tool result to the message state and return all of that back to you. So the really nice thing about it is that we don't have to write any more code. You get it out of the box. The bad thing about these provider executed tools is you're obviously tied to a single provider. But for today, for this demo, not so much of an issue and allows us to move quite quickly. And so what's happening when we import the open AI provider here, is that we're literally just using this to augment our request that's going off to open AI ultimately to include that opt-in flag. I want the web search tool included. So we will add that into our agent definition and you'll see we remove our instructions as well. And actually that is literally it for us to, well, I need to run the dev server as well. PNPM run dev. But I can now jump back to my agent here and say, when is AIE London, AI engineer summit London? I think that's what this is called. And we'll see a long pause, terrible UI. But eventually we should see our response come back in here. Our augmented context, our rag bot, if you will, using the old archaic language. But our agent now has the ability to fetch in and pull in relevant context via web search. And there's some cool things that we can do here. If you pass in an object into this tool, what is this, factory function here, you can actually specify things like the user location here. And so you can say type. I just use the language server to help me here. But I can say that we're in London. And that allows you to maybe get like, when is AI engineer summit? So we don't include London. And in theory, we'll get more London oriented. Yeah, there you go. So that extra context or those parameters are not being passed into OpenAI, into the request that's going off to OpenAI ultimately. So cool. We've got a way now of augmenting our context. And we've provided our agent with its first tool. But the problem that we have, if we show this again, when is AI engineer summit? You guys already saw this. I think, what did I do? When is AI engineer summit? Is that there is nothing showing up in the UI right now. So user has no idea what's going on. And while the agent is actually going through multiple steps there and calling a tool. So what we really want to do is actually render in our markup in this page.tsx. We want to render for different, we want to render what should, we want to describe what should be rendered in the UI when different tools are being used. Specifically, when our tool web search. So we could go in here and we could say case. And then the AI SDK follows this convention of a prefix of tool. And then you would specify your tool name. So we'd go and we'd go to our agent and we'd see it's called web search. So we'd say tool web search. And then we'd define some kind of markup here. So we'd say we'd say div called web search. Oh my god, I'm typing. I haven't typed in months. And you can see now that our agent has called web search. But this process is terrible. What if we wanted to see what is the p.input. There's probably a query here that's going in. But it's typed as unknown. That's ugly. We don't want to do that. So this is where I'm going to delete all of that. That's why I typed it and it's not on the page. This is where we are going to leverage another really awesome component of the AI SDK, which is the end-to-end type system that we've built. So the big assumption goes through every single AI SDK API decision is we want to have the agent definition being the source of truth that everything inherits from and derives from. So we spend a lot of time across tool calls, across the UI library, making sure that it can flow nicely. And so where we start, where we have to start with that is the atomic unit of state in our application, which is the message. And so we can get a typed message by using this type helper infer agent UI message. So we're going to jump back into our agent definition and pass in the type of our agent. Now this agent UI message is fully type safe based on the tools that we pass. And so we can now, if we head to the route handler, we can update the route handler to define that the messages coming in are of our custom UI message type. But most importantly now, we can go back into the page and we can type use chat as having my agent UI message. And so what you'll notice once you've done that, and that's imported from our lib agent file. So when we scroll back down and we check for the different part types, we'll see that we should have our web search tool here now. And if you were to go in here and check for the input and output types, you'll see that they are typed, which is very cool. So I'm not going to bore you with trying to type out and build a component on the spot because I can't design at all. And I don't think any of us design anything really by hand anymore. But you'll see if we jump back, we now have a nice component in line for our web search. And we can try that out again and say, who is Nico Albanese? Sorry, egoistic. But now you can see our component has that pending state, which is really cool in line. So what do we have next? That's augmented our context. Now onto what I think is the most interesting component also in 2026, which is providing our agent with a computer, with a sandbox to interact with. And to take a second here, we saw this really take off where internally we have agent called D zero, which has access to all of our, like the chat with your data that we're talking about has access to pretty much all of the Vercel backend, our entire admin panel, all of Salesforce, all of these kinds of things. And the idea here was, I think it was our head of data who wanted, was destroyed from getting tagged in everything. And he wanted naturally a replacement Slack bot that could do all of that and help scale his team as well.

11:49

SPEAKER_00

That's augmented our context. Now onto what I think is the most interesting component also in 2026, which is providing our agent with a computer, with a sandbox to interact with. And to take a second here, we saw this really take off where internally we have an agent called D zero, which has access to all of our chat with your data that we're talking about has access to pretty much all of the Vercel backend, our entire admin panel, all of Salesforce, all of these kinds of things. And the idea here was, I think it was our head of data who wanted to be stopped from getting tagged in everything.

12:10

SPEAKER_00

And he wanted a replacement Slack bot that could do all of that and help scale his team as well. And there was a really interesting point where it went from, let's have all of this agent that has all of these tools and it would use maybe five or ten of them at a time and return an answer that was somewhat hallucinated to when he added a file system to it. And some instructions there to say, okay, every single piece of work, every session that you do, you have a scratch pad in the file system, which is where you're storing an initial plan that becomes your reference for exactly what you should be doing in each step.

12:30

SPEAKER_00

And then you have a directory that's for your research and you collect everything in there. And there were two things that were happening there. For one, the agent started actually following through and going through entire tasks because things weren't just being layered in this insanely long context window where the initial thing was getting thrown away. Like the instructions there were create this plan file and in that plan file was the objective right at the top.

12:45

SPEAKER_00

And then right below the instructions were follow this plan file to a T, check things off as you go. And all of a sudden now you have this fascinating thing where the agent is reading and pulling in and reminding itself. Step. Okay. This is my objective. And so it's staying on track more. And at the end, you get this really great artifact that shows exactly what the agent did, what the work went into it. And so it's this emergent behavior that we saw where they were actually very good pairing this file system, whether it was a full computer or just a virtual file system worked very, very well.

12:55

SPEAKER_00

So we've now seen it across pretty much every agent that we build internally. We've got an agent for the GTM team. We've got an agent, like I said, the data one. We have customer support one that pushed our customer support tickets down by 90% with I think a 95%. Like there are people that were actually saying, thank you. Thank you. Which we haven't really seen before. All of these are backed and underpinned by this file system backed agent approach. And so that's what I wanted to show off today. We are going to be using a Vercel product for this Vercel sandbox. And I have built a lot with different sandboxes over the years.

13:10

SPEAKER_00

And there's something really cool that is just shipping now that's in beta, which is called named sandboxes, persistent sandboxes. And I think I have them up right here. And that's what we're going to be taking advantage of today. So one of the issues that you have with sandboxes in general is that unless you're using it on your computer, they are fundamentally ephemeral, right? Like whether a provider has a sandbox that can live for 30 days or five hours at the end, it's done. And so we were thinking about how we could work around those challenges. And what the team ended up coming up with is this really cool concept of every sandbox has a name.

13:29

SPEAKER_00

And then every sandbox can have sessions and sessions are just instances of that underlying sandbox item, we can call it. And every time you in your product or in your code want to use a sandbox, you reference it by its name. And under the scenes, Vercel will either see if there is an active instance for that session and route you to that, or it will spin up a new one optimistically and route you to that. So in your code, you simplify all of this lifecycle management of is my computer running? Which one should I point to?

13:55

SPEAKER_00

I've written, I'll show if I have time at the end, something I've been working on that I built out an entire lifecycle management system where I was literally tarring the entire file system after every single request and then storing in blob and spinning up new. It was terrible. And this is just so much easier. So where we'll be getting to today is almost this kind of, I hate to drop the open file thing, but we will be able to have this specific computer that we'll be able to record.

14:06

SPEAKER_00

So this is a lot of reference throughout every single request and behind the scenes what will be happening is that Vercel will actually spin it down after a timeout after inactivity, but it will snapshot the file system. So the state remains the same. And when you make subsequent requests, it will spin up a new one with that snapshot and state. So it effectively feels like the exact same machine, which is really cool. So this is in beta right now, but I'm already using it a ton and it's working really, really well. And we've got a lot more that's coming here. So the first thing that we're going to do is we're going to jump. By the way, any questions so far?

14:31

SPEAKER_00

Yeah, of course. I'm wondering if you treat them because you have access to these sources. Yeah. Do you treat the messages that go from the user to the user, et cetera, so that it doesn't contain too much because you might not be there with the whole conversation? Yeah. Do you use Google or Google, et cetera? Yeah, great question. So the question, because I don't know if the mic picks it up, is, do you, you're saying for every agent request, do you send the entire history? I think that's the default.

15:28

SPEAKER_00

With use chat. Yes. So that is the default, is right here. When we make every single request with send message, this is going to send the entire message history in the route and send that over to your agent. Now you'll always want to send, or that's a bold claim. Your message history is your context, right? Going into your agent. And so that will always want to be something bigger, right? Like as much of your conversation history as you would like, but sending over the wire, all of those messages every single time is not really something that you'll do.

15:44

SPEAKER_00

And so the more classic pattern is sending just the most recent message here. And you can do that in your client with prepare. I think it's in the, I don't write code anymore. So I can't remember what the name is. It's probably on the transport, which is new. And then we have, in here, where is it? Oh, we're not importing anything from AI. From AI. I'll just show you very quickly. Actually, you know what I should be doing? Is just show you the docs. That's even better. I don't know why I'm trying to type this all out. What we will do here is you can specify to just send the most recent message. And then you fetch all of the messages on the server in your endpoint.

16:03

SPEAKER_00

And you put those together and send those to the model. But where I think you're leaning is more around the context thing, right? You don't want to send old irrelevant things to the model, right? Like the context engineering side of things of pruning stuff and making things go away. Okay. Is that where you're going? Yeah. So that's where things get harder to do, more from, they're simple to do with SDK, but it requires trade-offs. I don't know why I'm trying to type this all out. What we will do here is you can specify to just send the most recent message. And then you fetch all of the messages on the server in your endpoint.

16:20

SPEAKER_00

And you put those together and send those to the model. But where I think you're leaning is more around the context thing, right? You don't want to send old irrelevant things to the model, right? The context engineering side of things of pruning stuff and making things go away. Okay. Is that where you're going? Yeah. So that's where things get harder to do, more from they're simple to do with SDK, but it requires trade-offs. And it requires decisions from the developer that we don't necessarily impose from the framework.

16:24

SPEAKER_00

So we have a few approaches that you can do here. They're all based on using, or most of them are based on using two things for one, in behind the scenes, what's happening in our endpoint here is that we're actually taking our UI messages, these messages. And we use a function under the hood called convert to model messages, which strip away a lot of the UI specific timestamps, IDs and all that kind of stuff. But under the hood as well, you can define this, I believe to ignore certain things, convert data parts to different structures to alter your context that's coming in.

16:29

SPEAKER_00

So in general, you own the messages that are coming in. So if you wanted to messages.map and go through and say, if m.if message role is assistant and m.parts.includes. What is it if there is a tool call. This is where the typing starts getting fun. What is the type here? Basically you can go through and define it. I want to strip out this type of tool call. You can totally do that. The way that we see bigger code bases doing it is actually using a function that we have called prepare step. And this runs before if we think about everything that was happening in our previous web search is an agent, we send a message that spins up a new assistant message.

16:44

SPEAKER_00

In the first step of that assistant message, it called web search. And then it sent back. Basically the agent runtime took all of that context and then decided what to do next. And it could have used web search again, could have used the computer or whatever it might be. But each of those are individual steps. And with the AI SDK, you can listen in, you have this callback that runs before every single step.

16:59

SPEAKER_00

And then this allows you to modify any of the parameters going into that specific step. So what do you get in here? You get the messages, the context, the model that you're using, the step number and any of the steps data, what has happened. And then in here you can literally return any of those top level parameters. [SPEAKER_02] So you could return a completely different message state for if, let's say if step number is greater than or equal to 20, you could basically take in just, where is it?

17:16

SPEAKER_00

These are the messages that are actively running. You could say messages.slice and take just the last five. And so in this way, you now have running at the beginning of every single step. You've got this sliding filter that is only taking the most recent five messages. [SPEAKER_02] Now, this is a very interesting long rabbit hole that we can go down. Because I've been thinking about this a lot. I was telling Casper before, I've been working on a coding agent for a few months now. [SPEAKER_02] I have it like yesterday. It ran for, I wonder if I can find the exact photo, but it ran for 104 minutes in one single turn. Let's see if we have the exact.

17:27

SPEAKER_00

[SPEAKER_02] So I can prove to you, these are all my, I think it was right here. Yeah. 104 minutes. It ran for used 316 tool calls, changed 29 files. And it used only 32% of GPT 5.4's context window. I have zero compaction running on this. And I think this is a hot take today. Because everyone's got their own systems where they're, I previously had, I was stripping away tool calls that were earlier on in the message history.

17:35

SPEAKER_00

So I was, I'm not sure if I had a certain threshold, but I would stay between 40% and 60% of the token window or the context window. But you have a really big problem that comes up when you do that, which is you invalidate the input cache every single time you do that. Which I think is, oh shit. Yeah. [SPEAKER_02] Like, of course. But it was important when the token windows were 400 K. So when I was first building this main models, I was using was Opus with the 400 K token window and GPT 5.3 codex, both 400 K token windows. And so you would hit that 40% pretty quickly, but with these million token windows, I'm not seeing as much of an issue.

17:53

SPEAKER_00

So what I'm building instead for it is more of two things. One, more comprehensive, I have sub agents implemented here and we can talk about that a lot as well. Taking dedicated pieces of work that can be somewhat independent and putting them off of the main context thread.

18:03

SPEAKER_00

And then returning just a thousand tokens, which is the summary of whatever the objective was and bringing that back in. So that's just efficient use of the context window in a session. But also having a dedicated tool, the folks at Amp did a really cool implementation of this is giving the agent a handoff tool, which at a certain point agent can call generate some kind of context that it wants to pass into a fully new thread and that becomes the main thread to start off with. And so I feel as much as you can push into sub agent territory is better because these, summarization, this compaction by agent summers or by LLM summarization is lossy.

18:27

SPEAKER_00

There was the infamous Twitter thread of, I think the head of one of the heads of AI at Meta AI who asked in an open call, can you archive emails from yesterday? And it just deleted her entire inbox. And she was asking, telling it to stop. And it was like, no, you asked me to do this. I'm going to keep going. And the reason why was because there were so many emails that were brought in by an inefficient tool, that triggered auto compaction and the emails overwhelmed her initial instructions.

18:35

SPEAKER_00

And so her initial instruction that was don't do this was gone. And so that's what scares me a lot about. I don't have a great system. I don't know. I don't know. But that's my long tangent here. I can go into that more. My whole point here trying to say is I don't think having used this, my coding agent for 14 hours a day for every single piece of work that I've done in the last four months, compaction has not been an issue to me. And I was showing you earlier, I have a 95% cache token read ratio on that as well. And that to me is more valuable both in terms of speed, performance and cost. But that long winded rabbit hole for that. And we can go into that more as well.

18:46

SPEAKER_00

But this is how you would manipulate context between steps. And the last thing I'll mention here is a lot, we were very intentional with this API being functional, right? My whole point here trying to say is I don't think having used this coding agent for 14 hours a day for every single piece of work that I've done in the last four months, compaction has not been an issue to me. And I was showing you earlier, I have a 95% cache token read ratio on that as well. And that to me is more valuable both in terms of speed, performance and cost. But that's a long winded rabbit hole for that. And we can go into that more as well.

18:56

SPEAKER_00

But this is how you would manipulate context between steps. And the last thing I'll mention here is a lot of, we were very intentional with this API being functional, right? This is a function that runs on every single step rather than returning a new message array here and persisting that to the steps afterwards. This is very easy to reason about because this is an empty function that starts at every single implication. Your messages coming in here are the aggregated set of messages throughout. And it also gives you that when you do persist at the very end, you know that you're getting the full message history at the end. Sorry, is that helpful? Cool. All right.

19:13

SPEAKER_00

And this is good as well because I know I'm doing basic stuff and you guys have already used it. So please do stop me with interesting questions like that. I'm happy to blabber. So we're going to do, we're getting to the fun part now. We have, I think we installed the beta, right? No, we didn't. So we're going to go through and we're going to run this command pnpm add versel sandbox at beta. And then we are going to create a new lib file called sandbox.ts. And like I was explaining before with these persistent sandboxes, we're just going to create a function, and it leverages this functionality. So we pass in a name and then it's going to check if that sandbox exists.

19:35

SPEAKER_00

If it does, it will return that. Otherwise it will create it and return that sandbox instance. So we've got that set up. Now this is some new stuff that I don't think any of you would see if you haven't used aistk6 before. So we've been using the tool loop agent and we've been calling it just like this, passing our agent passing our messages. But very rarely is the only thing dynamic in your agent your context. A lot of the time there are different structured inputs that should be changing the way that your agent behaves.

19:52

SPEAKER_00

A lot of the time, the classic example I give here, if you have a customer support agent, one of the classic things that you'll want coming into that is the customer ID. And maybe their customer type, whether they're enterprise or hobby, for example. And then the behavior of that model might change based on that. If you're running a system for hundreds of thousands of people, maybe you're an airline, I think you're in travel, right? Or no? Yeah, I was in construction. In construction. So travel works. We'll use that for now. Imagine you're British Airways and you've got a chat bot and you are serving hundreds of thousands of customers every day with your chat bot.

20:19

SPEAKER_00

And maybe a blue member who's at the bottom of the tier, you want to give them GPT-4 mini, but for gold member, you want to give them GPT-4 pro or something like that. These are structured inputs that fundamentally alter and augment the behavior. And the way that we used to do that before is functional style. You'd say const create agent, and you'd have some kind of input that would change. You'd have your input here that would ultimately change the behavior and you're doing if statements and all that. And so we saw that and thought these are call options. They change the way that they're options that you pass at call time that change the way that the model behaves.

20:40

SPEAKER_00

And so the way that you can define those is with a call option schema that you pass in to your agent definition. So I'm going to copy over all of this code right here. And I may strip away some of it so it's easier to see. So maybe we'll close this off and focus just on the call option schema. So what we want for our agent that is going to be like this agent with a computer is it should take in a sandbox in each invocation. Right? And the agent is going to interact with that sandbox. So we define a ZOD object for our call options expected. And it's going to be an instance of our sandbox class from the Vercel sandbox.

20:59

SPEAKER_00

We can then pass in that call option schema into our tool loop agent definition. And if we were to go to our route, you can see in our create agent UI stream response, there is an options key. And in there we now have this type safe options object that we need to pass in. You can see we've got an error here. Our options expect a sandbox. Howdy. No worries. I didn't know you have capacity with us. I hope you can wait. Okay. Yeah. So we have our call option schema. Now we actually need to do something with our call options. And what we're going to do here, this isn't going to make a ton of sense right now, but it will as we go into it.

21:41

SPEAKER_00

We want to have our sandbox available within our agent runtime for any tool to use. And the way that we're going to do it is with something called context within the AISDK agent runtime. And what the context is, is this is more similar to React context than it is agent context. And so if you're familiar with React context, you have a provider and any component within that nested component tree can access arbitrary values within that. And so this is the same idea is that you can pass in any kind of arbitrary data, variables, functions, whatever they might be into this object, and then you can access it in your tool runtime and your tool execute functions.

21:57

SPEAKER_00

So what we're doing here is we're saying our agent now expects a sandbox to come in every single time it's used. And then the first time it's called, this prepare call function will run, it runs only once the first time that our agent is called. And then it's going to take that sandbox off of the call options and inject it into our runtime context, this runtime state that you have across the entire agent run within steps. So that's the big thing. And I know this is a lot right now. We'll see it in a second. So what we're doing here is we're saying our agent now expects a sandbox to come in every single time it's used.

22:22

SPEAKER_00

And then the first time it's called, this prepare call function will run, it runs only once the first time that our agent is called. And then it's going to take that sandbox off of the call options and inject it into our runtime context, this runtime state that you have across steps. So that's the big thing. We'll see it in a second. It will make a lot more sense. Because we are going to define our first tool. So we're going to define a tool called, in a file called tools.ts. And this is going to be a bash tool. Now, we feel very strongly that bash is all you need. In a lot of cases, these agents are really good at writing bash commands.

22:54

SPEAKER_00

And so for this whole session, we are going to have just this one tool, bash, that is going to do everything for us. So, how do you define a tool with the AISDK? You guys have all used AISDK before, but there are three components. Description, super important. This is what the model uses to decide whether to use your tool and can also influence how it uses your tool. Then input schema, obviously, this is what that tool needs in order to run. And then we have the execute function. And this is the code that will be executed every single time the agent uses it. And so, now you'll see where the context comes in.

23:19

SPEAKER_00

This is an argument on the second argument of the execute function that you can pull off in your tool runtime, access any of that runtime state, and then run commands. So, this is that way of your tool being effectively wholly independent from the agent it's interacting with, but expects this input to come in, and then can use it. And we have a update coming to AISDK 7, which actually adds typing for the context. And it types the main context on the top-level agent. So, if you were to pull in a tool that expects some kind of context, that will then throw an error in your agent definition for saying you need to provide this.

23:38

SPEAKER_00

Which is really cool and directly as a result of me pestering Lars for three months as I've been using this pattern a ton. But so, what we have under the scenes is we're literally just pulling off the sandbox from the context, running our bash command that the agent generated, and then returning the standard out, any errors, and then the exit code. So, we could jump back and head back to our agent, which I don't think is running. So, we'll run PN dev. And we could say run ls, run ls-la. Nope. Why aren't we... Oh! See? It's not working because we haven't provided the sandbox, of course. If I followed instructions, I would have known that.

24:31

SPEAKER_00

The next step, which is... Well, there are two steps that are quite important. One, we need to actually give our tool to our agent. So, we're going to head back to our agent definition.

24:44

SPEAKER_00

I'm going to replace it all. And you'll see that now, on line 21, we are passing in our bash tool that we just created. And then, most importantly, we actually need to update our route handler to get said sandbox and pass it in to the agent run. And so, we're going to do that with our helper function that we created before. [SPEAKER_02] Create or get sandbox. Specifying a sandbox name. So, this is just a random ID. In an application, you'd have these names tied probably to a user's name or to a session. And that becomes this persistent sandbox that you can use across invocations. So, we get our sandbox here. You can see. And then, we pass it in to our call options.

25:18

SPEAKER_00

You see, again, if you didn't pass this in, we would get an error because it's end-to-end type safe. And who doesn't love end-to-end type safety? So, the final thing that we're going to want to do... It will work right now, but you obviously need to wait and you don't see the terminal actually. The command's running in line. So, we can actually define a component for our bash tool. Which we'll do here by providing a case for tool-bash. So, I'm going to jump back to the page. I'm going to copy this in and we can look at this again. So, before we have what should be rendered if the model generates text or the user has some text.

25:53

SPEAKER_00

What should be rendered in the UI if the model uses web search. And then, finally, what should be rendered in the UI if the model uses bash. And this, obviously, I didn't create because it both looks good and it works. So, we can say run ls-la. And now, we should see our terminal that's spun up. [SPEAKER_04] It's running the command. [SPEAKER_04] And we'll see returned that we are in a Brutel sandbox. So, this is a small leap for us today, but a huge jump for our agent to be able to now have this. But, you'll notice... And I go back to instructions and behavior. If I said, what do you see... The agent is oblivious.

27:07

SPEAKER_00

It now knows that it's got a bash tool, but it doesn't really know when to use it, how to use it, what to use. And so, this is when we could literally say... We could jump into our agent's instructions here and say, you are an agent with a computer. You can access with bash. If the user asks what you can see, use ls, something like that. This is a terrible system prompt, but you can be like, what do you see? Naturally, it's going to start using the computer a little bit more. This is not hugely helpful right now. The agent is oblivious. It now knows that it's got a bash tool. But it doesn't really know when to use it, how to use it, what to use.

27:51

SPEAKER_00

And so, this is when we could literally say... We could jump into our agent's instructions here and say, you are an agent with a computer. You can access with bash. If the user asks what you can see, use ls, something like that. This is a terrible system prompt.

29:04

SPEAKER_00

But you can be, what do you see? Naturally, it's going to start using the computer a little bit more. This is not hugely helpful right now. I think a natural progression to having access to a file system is, let's store stuff in it. Right? And the pretty natural next step for storing things with an agent is memory.

30:19

SPEAKER_00

I think a lot of people are thinking about what's the ideal memory. My hot take here is that memory is a file that you store in your sandbox. And you have some kind of actual deterministic code for pulling that in, injecting it into the system prompt. And then potentially having some kind of structure in your file system for different memory types. So, you'll have a core memories.md, which is the stuff that gets sent in every single turn. And in there are probably some information for...

31:25

SPEAKER_00

To search conversation history, it's stored in conversations.jsonl or something like that. But the point being the file system becomes this playground, this environment for you to store in a structured way a lot of this information. And the beauty here is that these agents are so good at generating bash commands. That they can use things like find, ls, grep, glob, all of these kinds of things. So, we're going to... Our final thing that we're going to play around with here is adding this concept of persistent memory.

32:28

SPEAKER_00

So, I'm going to copy this in quickly in here. I explained the basic idea of what I was thinking about. Which is let's have a file in our file system called memories.md. The agent will know to put new memories inside there. And then we'll actually in that prepare call, which runs once before every single agent run, let's fetch that file and just inject it into the system prompt, into the instructions with some kind of context around it. So, that's what I'm doing here. I'm fetching the memories.md. I'm getting the string itself. And then I have the instructions from before. You're a coding agent with access to a computer.

33:40

SPEAKER_00

You have a memories.md file that you can read and write to. You should always add any facts the user shares to memories.md. And then if we have any memories, here are your memories. Otherwise, no memories yet. And we could even ask something... No, actually, that's fine. We can go like this. And so now I could say, hey, my name is Nico.

34:46

SPEAKER_00

Again, really dumb examples here. But we'll see the beginning of it. The user's name is Nico. And now if I refresh and I say hi. We'll see that it's doing some behavior that we don't want. Which is it's adding some arbitrary like user greeted with hi.

36:00

SPEAKER_00

But you can see that it did regardless. It had Nico in the file system. I could say something here, don't report every... Only record important memories. You say always add any facts.

37:09

SPEAKER_00

And it was a fact that you... You're absolutely right. Important facts the user shares. Only record important memories. Let's try it.

38:18

SPEAKER_00

You see this explains what's broken with most agents. And the difference is that even as a native English speaker, I write conflicting instructions that an agent will obviously take very literally.

38:54

SPEAKER_02

[SPEAKER_00] So I say hi.

39:12

SPEAKER_00

And now hopefully...

39:31

SPEAKER_02

[SPEAKER_00] It's done it again. [SPEAKER_00] So you see this is so much of this. [SPEAKER_00] This is also a dumber model.

40:10

SPEAKER_00

With... Or no.

40:52

SPEAKER_02

[SPEAKER_00] We moved up to GPT-4.

41:12

SPEAKER_00

So it's not even... I would say... Don't... Don't share greetings. And this is bad. Because you obviously...

42:16

SPEAKER_00

This is the pink elephant. You don't really want to mention behavior you don't want it to do. Because it's now activating parameters. And we don't really know what it's doing. But let's see. Let's see.

43:29

SPEAKER_00

You lose a feeling if the memory... Damn. That is... It really... So... Yeah exactly.

44:39

SPEAKER_00

The fact that it... Nuked the memory file. Nuked my memories. So let's get rid of... Delete them. Okay great. Let's remove the... So... Now I mean... This is the... This is the fun part of building.

45:26

SPEAKER_00

It's we're literally trying to get this machine to use this persistent state. And you'll see that... I find this so... Having built so much with sandboxes now. The fact that we just have that ID, that name. And it's just persisting whether it's alive or whether it's actively running or not. But we have that persistent state is really cool. I did work on some prompts before that do make the memory system better. So you can see... Do not save trivial interactions like greeting, small talk or information that can be derived from the code base itself. So we're going to copy this over and see how this does. And now hopefully when we go... Hi.

46:30

SPEAKER_00

I've also made this added system prompts here to make it a little bit more inquisitive. So a bit onboarding, what's your name? What do you do? So I'm, hi, I'm Nico. And so now it's adding user's name. What do you do? I work... Why am I doing this? I work on the AI SDK at Vercel. I don't know why I was typing so much today. And now hopefully when we go... Hi. I've also made this added system prompts here to make it a little bit more inquisitive. So a bit like open clause... Like onboarding was like...

47:38

SPEAKER_00

What's your name? What the fuck do you do? So I'm like, hi, I'm Nico. And so now it's adding user's name. What do you do? I work... Why am I doing this? I work on the AI SDK at Vercel. I don't know why I was typing so much today. So then we get that added. And now when I pop into a new chat... Hi.

48:52

SPEAKER_00

We've got that memory that's persisted over these states. Now another really cool thing that you can start doing just by leveraging the bash that we have here... Is using something we already have in the environment. And something the agent is already very good at. Which is generating code and executing code. And so a classic thing that you'll see, which is why open clause was so exciting and why agents have really taken off this year... Is this idea of agents modifying or extending themselves. And a big part of that is less about modifying itself. But it's about creating these...

49:17

SPEAKER_00

Giving it the feedback loop in the environment where it can build, run code, evaluate the output and iterate on that.

49:19

SPEAKER_00

And I think that's also why coding has become naturally the first place this has really gone crazy. We've got compilers, we've got type checkers. It is the perfect environment for that REPL loop. But so I've added another kind of thing here that... Telling the agent if there are any repeatable tasks, make a Python script out of it. And then inject those Python scripts or the description of it into the end of our memories.md. And use that as a core or use that to decide what to do next. So if the user asks you to get the weather, make a weather script. And if the user asks to get the weather again, use that script that you already have access to.

50:30

SPEAKER_00

And these are the building blocks of this agent that learns, that builds on itself as we go. And so we can give this a go. We can see how it will do. We say get the weather in London. [SPEAKER_03] Use Python. [SPEAKER_03] I mean this is a... I'm going to force it a little bit here given we're doing a demo. And we'll see what it does. So it's going to write some Python here. And we'll see it's gotten the weather. And now it's added that to our memories.md. And it's pretty cool. This was all one assistant turn just using bash. And now it's got this tool for getting the weather. So now we could see get weather in SF. Let's try.

51:44

SPEAKER_00

Will it work? And look, it used... It is not minus 11 because it got Quebec, Canada for some reason. We'll have to modify it slightly. Yes, do this. Do this. Oops. I probably broke everything here. Amazing it actually worked. Okay, so it's 12 degrees Celsius. But you can see now we've got this... I prefer... It's cool. The more that it's asking, the more that it's building itself, the more that it's learning about what I like.

52:59

SPEAKER_00

And now in this computer, this is my agent's dedicated playground and workspace for it to help me out with anything. So that was the basis for what I wanted to go through today. What I did want to show you is what you can... I've built a very complex agent system on top of these exact patterns that I use now every single day to get my work done. So I wanted to... This is coming out later today. I'm holding myself to that. And this is effectively cursor background agents, but using those concepts that we had today.

53:37

SPEAKER_00

So it uses AI SDK, uses that exact pattern, uses AI gateway for inference, uses workflow. So it can run infinitely and each of the steps are durable. So each LLM step is matched to a durable workflow step. And if any step fails, it literally retries until it gets a result. And yeah, this also builds on that, what we were talking about before with sub-agents. So I can create any new session and be like, yeah, we have to finish. I'm done anyway. This is more of me showing off. But spin up a sub-agent to explore this project. And after it creates, because this is on my hotspot, so it's a little bit slow. See...

53:55

SPEAKER_00

Effectively, what I'm trying to show is this pattern scales to a pretty large system. And to show you this is being used by... So yeah, we've got sub-agents spinning up. And this runs in the background. This is being used by 23 people at Vercel. I've put 3.8 billion tokens through this in the last month or two. That cache read ratio I was telling you I'm very proud of.

54:24

SPEAKER_02

[SPEAKER_00] 91% responsible for close to 350 PRs.

54:26

SPEAKER_00

Although this doesn't include all of them that closed. And yeah, you'll see this will, if we go back here, has resumable streams. And you can see our sub-agent was going off, finding, searching all of this stuff off of the main thread. Used 30,000 tokens but just returned 500 at the very end. So that keeps us within literally 7,000 tokens for the main agent thread. So yeah. And this hopefully will be out later today. I'm obviously here if you guys have questions afterwards if we're getting kicked out. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Thank you. Because it's end-to-end type safe. And who doesn't love end-to-end type safety?

54:57

SPEAKER_00

So, the final thing that we're gonna wanna do... It will work right now, but you obviously need to wait. And you don't see the terminal actually. Like, the command's running in line. So, we can actually define a component for our bash tool. Which we'll do here by providing a case for tool-bash. So, I'm gonna jump back to the page. I'm gonna copy this in and we can look at this again. So, before we have what should be rendered if the model generates a text. Or model or the user has some text. What should be rendered in the UI if the model uses web search. And then, finally, what should be rendered in the UI if the model uses bash. And this, obviously, I didn't create.

55:44

SPEAKER_00

Because it both looks good and it works. So, we can say run ls-la. And now, we should see our terminal that's spun up.

55:55

SPEAKER_04

It's running the command. And we'll see returned that we are in a Brutel sandbox.

56:03

SPEAKER_00

So, this is like... I know this is kind of a small... This is a small leap for us today. But a huge jump for our agent to be able to now have this. But, you'll notice... And I go back to instructions and behavior. If I said, what do you see...

56:24

SPEAKER_00

The agent is like oblivious. It now knows that it's got a bash tool. But it doesn't really know when to use it, how to use it, what to use. And so, this is when we could literally say... We could jump into our agent's instructions here and say, you are an agent with a computer. You can access with bash. If the user asks what you can see, use ls, something like that. This is a terrible system prompt. But you can be like, what do you see? Naturally, it's going to start using the computer a little bit more. This is not hugely helpful right now. I think a natural progression to having access to a file system is like, let's store stuff in it. Right?

57:16

SPEAKER_00

And the pretty natural next step for storing things with an agent is memory. I think a lot of people are thinking about what's the ideal memory. My hot take here is that memory is a file that you store in your sandbox. And you have some kind of actual deterministic code for pulling that in, injecting it into the system prompt. And then potentially having some kind of structure in your file system for different memory types. So, you'll have like a core memories.md, which is the stuff that gets sent in every single term. And in there are probably some information like for... To search conversation history, it's stored in conversations.jsonl or something like that.

57:59

SPEAKER_00

But the point being the file system becomes this playground, this environment for you to store in a structured way a lot of this information. And the beauty here is that these agents are so good at generating bash commands. That they can use things like find, like ls, like grep, like glob, all of these kinds of things. So, we're going to... Our final thing that we're going to play around with here is adding this concept of persistent memory. So, I'm going to copy this in quickly in here. I explained the basic idea of what I was thinking about. Which is like let's have a file in our file system called memories.md. The agent will know to put new memories inside there.

58:46

SPEAKER_00

And then we'll actually in that prepare call, which runs once before every single agent run, let's fetch that file and just inject it into the system prompt, into the instructions with some kind of context around it. So, that's what I'm doing here. I'm fetching the memories.md. I'm getting the string itself. And then I have the instructions from before. You're a coding agent with access to a computer. You have a memories.md file that you can read and write to. You should always add any facts the user shares to memories.md. And then if we have any memories, here are your memories. Otherwise, no memories yet. And we could even ask something like...

59:22

SPEAKER_00

No, actually, that's fine. We can go like this. And so now I could say, hey, my name is Nico. Again, really kind of dumb examples here. But we'll see the beginning of it. The user's name is Nico. And now if I refresh and I say hi.

59:42

SPEAKER_00

We'll see that it's doing some behavior that we don't want. Which is it's adding some kind of arbitrary like user greeted with hi.

59:50

SPEAKER_00

That. But you can see that it did regardless. Like it had Nico in the file system. I could say something here like don't report every... Only record important memories. You say always add any facts. And it was a fact that you... You're absolutely right. Important facts the user shares. Only record important memories. Let's try it. You see this explains what's broken with most agents. And the difference is that like me even as a native English speaker kind of... Writes conflicting instructions that an agent will obviously take very literally. So I say hi. And now hopefully... It's done it again. So you see like this is so much of this.

1:00:41

SPEAKER_00

This is also like a kind of a dumber model. With... Or no. We moved up to GPT-54. So it's not even... I would say... Don't... Don't share greetings. And this is bad. Because you obviously... This is like the pink elephant. You don't really want to mention behavior you don't want it to do. Because it's now activating parameters. And we don't really know what it's doing. But let's see. Let's see. You lost a feeling if the memory... Fuck! That is... It really... So... Yeah exactly. The fact that it... Renate the memory file. Nuked my memories. So let's get rid of... Delete them.

1:01:25

SPEAKER_00

Okay great. Let's remove the... So... Now I mean... This is the... This is kind of the fun part of building. It's like we're literally trying to get this kind of machine to use this persistent state. And you'll see that... I find this so... Having built so much with sandboxes now. The fact that we just have that ID, that name. And it's just persisting whether it's alive or whether it's actively running or not. But we have that persistent state is really cool. I did work on some prompts before that do make the memory system better. So you can see... Do not save trivial interactions like greeting small talk or information that can be derived from the code base itself.

1:02:04

SPEAKER_00

So we're gonna copy this over and see how this does. And now hopefully when we go... Hi. I've also made this added system prompts here to make it a little bit more inquisitive. So a bit like open clause... Like onboarding was like... What's your name? What the fuck do you do? So I'm like, hi, I'm Nico. And so now it's adding user's name. What do you do? I work... Why am I doing this? I work on the AI SDK at Vercel. I don't know why I was typing so much today. So then we get that added. And now when I pop into a new chat... Hi. Like we've got that memory that's persisted over these states.

1:02:45

SPEAKER_00

Now another really cool thing that you can start doing just by leveraging the bash that we have here... Is using something we already have in the environment. And something the agent is already very good at. Which is generating code and executing code. And so a classic thing that you'll see, which is I think is why open clause was so exciting and why agents have really taken off this year... Is this idea of agents modifying or extending themselves. And a big part of that is less about modifying itself. But it's about like creating these... Giving it the feedback loop in the environment where it can build, run code, evaluate the output and iterate on that.

1:03:29

SPEAKER_00

And I think that's also why coding has become naturally like the first place this has really gone crazy. Is like we've got compilers, we've got type checkers. It's like it is the perfect environment for that basically REPL loop. But so I've added another kind of thing here that... Telling the agent if there are any repeatable tasks, like make a Python script out of it. And then inject those Python scripts into or the description of it into the end of our memories.md. And use that as like a core or use that to decide what to do next. So if the user asks you to get the weather, make a weather script.

1:04:09

SPEAKER_00

And if the user asks to get the weather again, use that script that you already have access to. And this is like these are the building blocks of this agent that learns, that builds on itself as we go. And so we can give this a go. We can see how it will do. We say get the weather in London.

1:04:30

SPEAKER_03

Use Python. I mean this is a...

1:04:32

SPEAKER_00

I'm gonna force it a little bit here given we're doing a demo. And we'll see what it does. So it's gonna write some Python here. And we'll see it's gotten the weather. And now it's added that to our memories.md. And it's pretty cool. Like this was all one assistant turn just using bash. And now it's got this tool for getting the weather. So now we could see get weather in SF. Let's try. Will it work? And look, it used... It is not minus 11 because it got Quebec, Canada for some reason. We'll have to modify it slightly. Yes, do this. Do this. Oops. I probably broke everything here. Amazing it actually worked. Okay, so it's 12 degrees Celsius.

1:05:31

SPEAKER_00

But you can see now like we've got this... I prefer... See, it's cool. Like the more that it's asking, the more that it's building itself, the more that it's learning about what I like. And now in this computer, this is like my agent's dedicated playground and workspace for it to help me out with anything. So that was the basis for what I wanted to go through today. What I did want to show you is like what you can... I've built a very complex agent system on top of these exact patterns that I use now every single day to get my work done. So I wanted to... I mean, shill. This is coming out later today. I'm holding myself to that.

1:06:16

SPEAKER_00

And this is effectively like cursor background agents, but using those concepts that we had today. So it uses AI SDK, uses that exact pattern, uses AI gateway for inference, uses workflow. So it can run infinitely and each of the steps are durable. So each LLM step is matched to a durable workflow step. And if any step fails, it literally retries until it gets a result. And yeah, this also builds on that, what we were talking about before with sub-agents. So like I can create any new session and be like, yeah, we have to finish. I'm done anyway. This is more of me showing off. But spin up a sub-agent to explore this project.

1:07:11

SPEAKER_00

And after it creates, because this is on my hotspot, so it's a little bit slow.

1:07:19

SPEAKER_00

Kind of see... Effectively, what I'm trying to show is like this pattern scales to a pretty large system. And to show you like this is being used by... So yeah, we've got sub-agents spinning up. And this runs in the background. Like this is being used by 23 people at Resell. I've put 3.8 billion tokens through this in the last month or two. That cash read ratio I was telling you I'm very proud of. 91% responsible for close to 350 PRs. Although this doesn't include all of them that closed. And yeah, and you'll see like this will, if we go back here, has resumable streams.

1:08:07

SPEAKER_00

And you can see our sub-agent was going off, finding, searching all of this stuff off of the main thread. Used 30,000 tokens but just returned like 500 at the very end. So that keeps us within literally 7,000 tokens for the main like agent thread. So yeah. And this hopefully will be out later today. I'm obviously here if you guys have questions afterwards if we're getting kicked out. Yeah.

1:08:38

SPEAKER_00

Yeah. Yeah.

1:08:43

SPEAKER_00

Yeah. Yeah. Yeah. Yeah.

1:08:52

SPEAKER_00

Yeah. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note