Context Is the New Code — Patrick Debois, Tessl
Description
As AI coding agents become more capable, context is starting to matter as much as code. Yet while code has version control, review, testing, CI/CD, and production observability, the prompts, rules, and memory that drive agents are still often managed like ad hoc hacks. Patrick argues that context needs its own engineering discipline. He introduces the Context Development Lifecycle: Generate, Evaluate, Distribute, and Observe, along with the team practices that make context a shared, repeatable, and improvable part of software delivery. The session also explores the larger context flywheel, where better context leads to better agent output, which creates better observations, which in turn improves context again. Speaker info: - https://x.com/patrickdebois - https://www.linkedin.com/in/patrickdebois/ Timestamps: 0:00 - Introduction to the talk 1:14 - Why context is the new code 2:37 - Introducing the Context Development Lifecycle 3:50 - Generate: Creating context for agents 6:26 - Evaluate: Testing your context 13:59 - Distribute: Sharing and packaging context 17:49 - Observe: Monitoring and feedback loops 22:33 - Conclusion and the context flywheel 24:49 - Q&A session
Summary
Generated by claude-haiku-4-5-20251001Context Is the New Code — Patrick Debois, Tessl
Main Topics
- Context Development Lifecycle: A structured approach to managing AI prompts and instructions similar to software development lifecycle (SDLC)
- AI-Driven Development Paradigm Shift: Moving from traditional code-writing to context/prompt engineering
- The Four-Stage Loop: Generate → Test → Distribute → Observe (with adaptation/regeneration)
- Enterprise-Scale Context Management: Packaging, versioning, and scaling reusable context across teams and organizations
- Quality Assurance for Prompts: Testing and evaluating context effectiveness through various methods
Key Points
Context as Code
- "Context is the new code" — prompts and instructions are now the primary development artifact
- Large code blocks can be transformed into reusable "skills" with context-driven instructions
- Code generation has become prompt-driven rather than manual coding
Generation Phase
- Reusable prompts: Standardization emerging around formats like AgentMD and specification-driven development
- Context sources: Documentation, GitHub/GitLab repos, Slack, tickets, and external APIs (via MCP)
- Libraries and dependencies: Pulling in latest documentation to prevent LLM hallucination about outdated versions
Testing Phase ("Evals")
- Multiple testing levels:
- Linting: Validate context format and structure
- Grammar checking: Ensure context is verbose enough for agent comprehension
- Unit test equivalents: Verify specific rules (e.g., all API endpoints use "awesome" prefix)
- End-to-end tests: Execute code in sandboxes and validate actual runtime behavior
- Non-deterministic nature: Tests must account for AI unpredictability; success measured as percentage rate (e.g., 5/5 runs)
- Error budgets: Accept minimal failures for certain tests while maintaining strict standards for critical ones
Distribution Phase
- Skills and registries: Marketplace ecosystems (Tesla registry, Copilot marketplace) for sharing reusable context
- Packaging standards: Skills can contain context, scripts, documents, and MCP integrations
- Dependency management: Context faces same "dependency hell" challenges as traditional package management
- Security considerations:
- AI SBOM (Software Bill of Materials) tracking
- Credential scanning (Snyk)
- Supply chain security for untrusted context sources
Observation Phase
- Agent logs analysis: Identify missing context patterns at scale
- PR feedback as context feedback: Incomplete PRs indicate context deficiencies, not just code issues
- Production monitoring: Instrument generated code to catch runtime failures and create test cases
- Context filters: WAF-like filtering to prevent prompt injections and malicious patterns
- Tracing and observability: Full logging of agent execution for continuous improvement
Scaling Patterns
- Individual loop: Personal context optimization and crafting
- Team loop: Shared context libraries within teams
- Organizational flywheel: Enterprise-wide context reuse where fixes in one team benefit all teams
Notable Quotes
> "Context is the new code because it's being generated."
> "If you give the engine the wrong fuel, which is context, they're not going to perform... I can optimize my context. And that's the message."
> "You cannot do exact testing all the time. It's a different way that this works."
> "The reality is 99.9% of the skills is crap. But it's good to learn from others."
> "We're going to have dependency hell [with context], just like with code."
> "Doing this more in an engineered way than just copy and pasting things and hoping for the best."
Takeaways
For Developers
- Invest in prompt/context quality — It's now a first-class development concern, not an afterthought
- Write evals rigorously — Testing context is as important as testing code; expect to spend time on comprehensive test suites
- Use voice-to-text for better context — More elaborate prompts yield better results than typed prompts
- Leverage libraries and registries — Don't reinvent context; reuse tested, proven prompts
For Teams & Organizations
- Implement context version control — Treat prompts like code; check them into repositories
- Build organizational feedback loops — Monitor agent logs to identify missing context at scale and create shared solutions
- Establish context governance — Implement filtering, security scanning, and quality standards for enterprise context
- Create internal registries — Consider private registries for proprietary context rather than public marketplaces
For Enterprise
- Context optimization yields ROI — Better context provides organization-wide improvements in AI-assisted development
- Flywheel effect: One team's context improvements benefit all teams, creating network effects
- Security is critical — Context can be attacked via prompt injection; implement WAF-like filtering and sandboxing
General Principle
LLMs and coding agents are engines; context is the fuel. The quality of generated code depends far more on prompt engineering than on model selection. Organizations should focus on optimizing context systematically rather than chasing newer models.
Transcript
There's a few people who want to start earlier. I'm gonna take the opportunity to officially open the architect track. There's no track host so I do it myself. So thank you for coming here. I hope you already had a good conference. It's amazing that so many people showed up. Maybe before I start, who's used any AI coding agent in this room? Raise your hand. Lower it. Who hasn't? Raise your hand. Okay, my people. Perfect. All right. Context is a new code or context development lifecycle. I feel honored to be here every time I try to do a different talk at the AI engineering. So this is a little bit of thinking ahead. It's an unpolished talk. It's not like everything's there, but is there anything there in AI anyway? But so let's start. I assume you all are now coding with prompts. I barely touch any code anymore. I just tell the AI to do something different. So I would say context is the new code because it's being generated. A little bit more advanced maybe is I see myself having a tendency where I had large pieces of code that I was using maybe some helpers and some other pieces and I just turned them into a skill. We had that in our product. It was an onboarding from AI agents. People have npm, no GS, all the various things. Then they have different tools for packaging and it is impossible to actually code that like it will require a lot of coding. But if I just say a skill says please first figure out what their package manager is, then figure out what their ecosystem is, and then do these steps together with the user. You know, it solved a lot more problems that we could ever code. So that is another piece that I would say code is also transforming back into context as a skill as well as a workflow that's reusable. And leave that with you. I like to think in parallels. In 2009, I don't know if there's any DevOps people in the room. It was kind of me saying like what if ops look more like dev and then we got like, hey, collaboration, deployment, all that stuff. So kind of, you know, last year I started thinking what if context is the code. How do we deal with this in a more consistent way? And it's basically saying if we have a software development lifecycle, how does a context development lifecycle look like? Because we're basically shifting somewhere else. It's context, it's not code. How does it look like? I came up with this, you know, of course, an infinity loop with some DevOps background. But the whole idea is that we generate a lot of context, then hopefully we test the context, we distribute the context, maybe to some colleagues to some other parts of the organization, we observe whether it works. And if it doesn't work or works, we call it adapt and regenerate the context and then go from there. So that's the loop of the talk that I'll be going for with some examples. So, step by step going through. Generate. It's probably the one that you're all most familiar with. Because you're all prompting. You're the human context creation typing things, right? I was actually amazed that I just asked, tell me when my talk is at AI engineer that you'll fetch the website and would just say, here's your talk. Blew my mind. But hey, I said the context that I've given it, I'm Patrick, all that stuff, right? So, very simple context. It's what you do probably a lot in your setup. If you get a little bit more advanced, you say that prompting is tedious, I want to have reusable prompts. So, you know, depending on the flavor of your coding agents, they call it instructions. Luckily, there's a little bit of a standardization now happening where it's like an agent MD and some pieces like that. Boo Claude for still calling it Claude MD. But anyway, you get the picture. There's reusable prompts, reusable pieces of context that we're doing. We can also bring other context in. If we have documentation of libraries that we use day to day, we want to pull that in. Because the LLMs might not have the latest documentation. And so it's hallucinating. Is it version 2, version 3? We don't know. So, we give it a context and say, please download the documentation. Hopefully, then agent optimized. And then they will do a better job at generating the code for that version of the library. Another piece of getting better context and creating context from libraries. And of course, it wouldn't be complete if we would say, pull context from wherever. MCP has been instrumental. Get it from your GitLab, GitHub, Slack. All context. We're pulling in. We're creating. Even the ticket is creating context because we're pulling that in while we go there. And then maybe the new kid on the block is, okay, what if we start writing our prompts as specifications, spec-driven development, which then gets broken down by the agent into a planning mode and do step-by-step prompts that it then runs through. So, a lot of creation happening in that field. Testing. This is probably what you're closest to. But when you're typing all that context and creating all that context, you change two lines in your Claude MD. Do you know the impact? Is it YOLO? Looks good to me. Let's do it. You have to think about how do we test things. It's not just about we have a piece of code and we have a piece of context now. We need to write tests to see what is the impact. New coding agents, we don't know where the lines still work. Now, it's not new in the world of AI engineering, but it's not that common yet in the world of coding with AI that you start writing evals for which are tests for your code context. A little bit hard to read, but you know, if you think in parallels, we have different levels of testing in code and the simple one could be linting. Your IDE has these wiggly lines like, hey, this is not like, you know, there's some incorrect syntax or you could do better like that. Here's an example of a validation of a skill where we say, well, you need to have the description. It can only be so long. So, it's validating according to the spec of the format of the context. In this case, simple analogy, simple linter that you can run. And then you can do other things like, and I haven't found maybe the good coding equivalent, but think of this as a Grammarly. So, if you write context, is it actually, can the agent understand what you're writing? If you write two words, it's not verbose enough for it to actually understand the context. So, what you can do is you can say, okay, given this context, what do you think about, do you understand this? And then you can get feedback like, oh, it's not explicitly enough written or it's not complete, you're missing pieces. So, that's feedback that you can get out from tools as well. So, whenever you're writing simple analogy, simple linter that you can run. And then you can do other things like, and I haven't found maybe the good coding equivalent, but think of this as a Grammarly. So, if you write context, is it actually, can the agent understand what you're writing? If you write two words, it's not verbose enough for it to actually understand the context. So, what you can do is you can say, okay, given this context, what do you think about, do you understand this? And then you can get feedback like, oh, it's not explicitly enough written or it's not complete, you're missing pieces. So, that's feedback that you can get out from tools as well. So, whenever you're writing now your context, you get a Grammarly saying, hey, do this. That's why I like to voice code. For some reason, I'm way more elaborate voice coding than typing. I'm a bad typer, two fingers, still after so many years. But when I talk, I see the sentences come on the screen, but it helps to get good context there. All right, another kind of test. So, imagine you put in your CloudMD or AgentMD, I should say, every API point must use the prefix awesome, right? You have some convention in your company, right? Which is great. So, your prompt will be then, add me a new endpoint to save a user. And you expect actually your coding agent to just say, the code that's being generated has an awesome user. That's great. But the way we can test this is by asking an LLM, the code that was generated, does it actually start with slash awesome? Now, you can do that with regex, I know, this is just for example purposes. But you can ask it to judge your code based on your criteria and whether it did the right thing, right? So, imagine you would ask the same question without your context above. No LLM is ever going to prefix URL with awesome. So, that's where your content or your company specific, your team specific things come in and that's why you still write those tests to see if this still works. Now, maybe Gemini reacts differently than Copilot or something. Your company needs to make it more switchable of context. With this, you run the test and you can actually tell. That's the difference. And then you can make whole suites. And I would compare that almost to unit tests. I have a bunch of these tests and they tell me whether that's actually good code, the code is following the rules and everything's fine. In this case, it's even kind of infrastructure as code. It doesn't need to be code only. It could be various things. It could be config files as well. And I just have a bunch of criteria that I just run every time to do that. But if you want to test whether an endpoint has slash awesome slash user, there's a real test that we want to run, which is I want to test the endpoint. I just don't want only to check the code. I want to have it running. So, when you give the judge a tool and the judge becomes an agent and it can do things in a sandbox and execute stuff, it can actually do the curl. So, you can bind LLM as a judge with some tooling and then you can have multitude of tests. Actually, in this case, it ends up being an end-to-end test, right? Because it's not just looking at the file, it's actually running the piece with everything that it's supposed to do. And then I can do this like given a certain commit in my repo, I want to run this scenario. So, given this piece of context, did it make a difference? Yes or no? So, you're building this up while you're committing context also within your repo. And because we now have tests and it gives us feedback whether it's working yes or no or what it's missing, we can optimize context. So, that's the, we can put that in a code action or something that says like, okay, fix this context, improve this context with all the feedback the LLM has given us to improve that. So, again, coding improvements, but we start thinking more in testing that piece as well. One of the first reactions is once you have tests and optimizations, we run this in a CI CD system because that's perfect, right? That's where we run our tests and their test suites and do that. Now, there's a little bit of a real thing. If you run evals, you run it once, you run it another time, it might not give the same results. Remember, undeterministic things. So, you cannot say, well, run it once and then if it passes or not, you're going to be in for a treat because it's like, oh, I can't debug that. So, think about this, you run it five times and out of five, how many times does it succeed? And maybe in several cases, it hits 100% all the time, which is great, but in others not. And depending on how you change your context, it will influence which tests actually work or not. I find it personally helpful to think about this as error budgets. I give a set of tests an error budget that I really care about. So, it's only allowed to fail minimally and other pieces are okay. So, that's how you have to think about testing context. You cannot do exact testing all the time. It's a different way that this works. All right. So, generate, hopefully you understood what the testing could do for you and distribute. Maybe that's also something you already did. If you maybe have checked context into your repo, right, which is great, all of a sudden it becomes available, your colleague checks it out, zero friction, I can push, I can share. But we have another mechanism for doing things. Think of this like, imagine you have a reusable context that you want to reuse across multiple projects, across multiple teams. We had the concept of a library. So, what if we package pieces of context and then we are able to install pieces of context that we need for this project? Guidelines, front-end, it doesn't matter for that. And then if we take that up a notch, how to discover what packages exist, that's a registry, right? Now, in that way, it's no surprise that you'll see things like skills and the Tesla registry and the marketplace where you can find a multitude of skills. Now, the reality is 99.9 and I mean that in a very sincere way of the skills is crap. But it's good to learn from others to see what they're doing. But hardly any of them if you run any set of evals on there is actually up to a quality standard. Now, that will likely improve. But there's also a tendency that a lot of the skills and pieces people actually want to put that in their own registry. So, I'll come to that later again. But so, you start seeing the gist, a skill Skills and the Tesla registry and the marketplace where you can find a multitude of skills. Now, the reality is 99.9 and I mean that in a very sincere way of the skills is crap. But it's good to learn from others to see what they're doing. But hardly any of them if you run any set of evals on there is actually up to a quality standard. Now, that will likely improve. But there's also a tendency that a lot of the skills and pieces people actually want to put in their own registry. So I'll come to that later again. But so you start seeing the gist, a skill not only contains context, it can contain scripts, it can contain documents, contain a bunch of things. So is this the package format? Probably plugins could now also contain MCP. But you see there's a standard coming in. Skills all of a sudden, when that came out, all the coding agents said we're supporting this as almost a package format for people to distribute their context on. And then when I have one piece of context, I have dependencies. And I'm sorry, but also with context, we're going to have dependency hell. Right? I'm going to download this for frontend and maybe it's conflicting with what is in the React context package. And so you start having to deal with that as well. So you start seeing also packages that mirror your library versions, your code, your context versions, and pull that in as well. And of course, when we have packages and people are publishing things in registry, we need security. Right? OpenClaw, thank you for that. Everybody all of a sudden became aware that we need more secure things because we are able to run things on our laptop that are coming from strangers. Right? So Snyk has a way of scanning context. Right? It's doing some credential handling. It's exposing some third-party pieces. So you start seeing those scanners on the context as well. And then when you think about security, who actually built this skill? How was it built? With what model was this built? So all of this is capturing what we learned in packaging, the s-bomb, the AI s-bomb, the package of context that we're putting in. So you've seen the path, right? You generate, evaluate, distribute. Let's move into observe. When you are making libraries of scales and context for others, and I don't mean copy and paste this over Slack or something, but when you actually want to maintain this as something somebody else can use, similar to a library, when they start using that, how do you get feedback whether that still works? So a great place to get feedback is actually by looking at the agent logs. So imagine developer one coding on the project, and the agent is not doing what they want. They could put this into their context, which is great. Right? Okay. Let me do the TDD almost like I hit a problem. It's not TDD, but you get my gist. Or what if we at a team or an organization scale would look at the logs every time an agent said we're missing this piece? And we surface that and say if everybody's missing this piece, we should create context for this. And then we distribute the context to everybody, and all of a sudden the impact of improvement is for everybody. Luckily, the agent MD, there is now standards becoming for logs. So we can read from logs, and that's part of our feedback channel to see if the agent is actually using or missing some of the context. Any feedback you get on a PR that's not complete. That's feedback on your context, because that PR was created with certain pieces of context. If you say this is not correct, you can keep arguing on the PR, or you can just say let's improve the context so the next iteration actually improves and you don't hit that same problem again. What about running code in production that was generated from context? And that's not correct because yes we do our PR reviews and we say thumbs up, thumbs down and we give the feedback, but the actual feedback is also in production when it's running. So this is a tool that actually instruments your code, pushes it out, it's almost a wrapper, it pushes it out to production. When it fails, it says these pieces of code were changed and we're failing. Hey, in this case, input, output, it did something wrong. Can we create a test case for this so the next time we don't hit this again in production? Feedback loop. Now these are all pretty trivial, missing pieces of context or improvements, but if you run agents and the equivalent of scanning maybe in the CI CD is you need to make sure when it's running in production is it not doing strange things? So we need a way of looking at that. Now I've been toying myself with sandboxing agents and it is very resourceful at finding things. I like okay you know run this thing, try to figure out anything useful to get break out of the system, and okay it uses my environment variables. Okay stupid, let me remove the secret. Let me look at your memory files. So you have to really make sure that whatever it's doing you can have a way of tracing this as well. And I apologize again for the slide, but it just is we can have a sandbox where the agent runs inside. You can have a code agent, but your code agent by default without any restrictions loads your agent MD, you load your skill MD. Nothing is blocking that. So if you download this immediately it's loaded. So you can't filter that with sandboxes, you need to have another way. I call that a context filter. Think of this as a web application firewall that just filters out any patterns or prompt injections or stuff that is coming in directly in that piece. And if you take that, there's a lot of talk here as well on harness engineering. Harness engineering itself also has this full observability, looking at logs, looking at traces, looking at feedback. So it's useful for training pieces, but as much useful for running your own piece as well. Harness engineering itself. So those were the pieces for me today. I would say for a lot of people there's create context, test context, think of this as your library authoring tool loop. And then when you push this into the enterprise there's an organizational loop. Hey I made a library, somebody else is using it. I'm looking whether that's useful, whether that's still working, whether that's still working for all the other pieces. So that's the improvement, almost a Sonar CI CD model for context. And then you're currently probably doing a lot at the individual solo model. You're improving. Harness engineering itself. So, those were the pieces for me today. I would say for a lot of people, there's create context, test context, think of this as your library authoring tool loop. And then, when you push this into the enterprise, there's an organizational loop. Hey, I made a library, somebody else is using it. I'm looking whether that's useful, whether that's still working, whether that's still working for all the other pieces. So, that's the kind of improvement, almost like Sonar CI CD model for context. And then, you're currently probably doing a lot at the individual solo model. You're improving, you're honing, crafting your own markdown. What if you start doing this more with your team? Make that a reflex. If it's missing, add some context. What if you put that out to a team of teams, and you start having a flywheel? If you fix it here, the other team can reuse it. And that's scaling things out into the organization as well. And so, there's a lot of talk about LLMs and coding agents, and I love them. But the way that I see it is they're just the engine. If you give the engine the wrong fuel, which is context, they're not going to perform. So, you can't do anything on the LLMs, at least not me, right? I'm just using the coding agent. I'm using whatever they give me. But I can optimize my context. And that's the message. Doing this more in an engineered way than just copy and pasting things and hoping for the best. If you like this talk, connect on LinkedIn for the slides. Give me some feedback, good and bad. If you want to try TESOL, where we implement some of the pieces of this, have a go. And if you're also interested in another conference, I know you can never have enough conferences. Visit AI DevCon, which I curate the content for here in London, 1st and 2nd of June. And that's it. I can maybe take a few questions. Any questions? Sure. So, I'm wondering if you have any thoughts about more exotic forms of context? Like, traditional ones? So, for example, one of the things I'm working on is an automated system for scoping out architectural problems, trying to create hard definitions for them, where you can see that page and then create actual objective tests. Oh, cool. Yeah. Microphones. And one of the things I've been testing out is ability to create consistency as a form of context or as a form of eval. So, given this rough, very loose definition of what the plan is, if you put that, if you try that agent system, turn that into a really crisp definition, and you just have that done in parallel, how often do you get the same crisp definition? And if they're all over the place, then the original definition is so poor, you need to go back to basic principles or to an architect. But if they're all the same, then it's probably a pretty good definition that you can carry on with the downstream process. So, besides just code and typical evals, any other sources of context or generating context that you think is useful? I don't have a specific answer to your exotic case, but I would say that the piece that people underestimate is that once you thought you were going to save time by writing your context instead of all your code, but if you take this rigorously, you're going to spend time on writing the right evals. Right. And that's, you know, a lot of work because now you don't only have one prompt that you're trying to get right. It's all the prompts of the evals. And people do almost like the more advanced people, they almost have their own process, and they build their own process on top of building the right evals on your business case as well. So, yeah, good question. Thank you. Any other questions? If not, I'll be around. Say hi. I'm also going to be at the Tesla booth. So, thank you very much, and I'm going to make space for the next speaker. Thank you. generated. A little bit more advanced maybe is I see myself having a tendency is I had large pieces of code that I was using maybe some helpers and some other pieces and I just turned them into a skill. We had that in our into our product. It was an onboarding from, you know, AI agents. People have patent, no GS, all the various things. Then they have different tools for packaging and it is impossible to actually code that like it will require a lot of coding. But if I just say a skill says please first figure out what their package manager is, then figure out what their ecosystem is, and then do these steps together with the user. You know, it solved a lot more problems that we could ever code. So that is another piece that I would say code is also transforming back into context as a skill as well as a workflow that's reusable. And leave that with you. I like to think in parallels. In 2009, I don't know if there's any DevOps people in the room. It was kind of me saying like what if ops look more like dev and then we got like, hey, collaboration, kind of deployment, all that stuff. So kind of, you know, last year I started thinking what if context is the code. How do we deal with this in a more consistent way? And it's basically saying if we have a software development lifecycle, how does a context development lifecycle look like? Because we're basically shifting somewhere else. It's context, it's not code. How does it look like? I came up with this, you know, of course, an infinity loop with some DevOps background. But the whole idea is that we generate a lot of context, then hopefully we test the context, we distribute the context, maybe to some colleagues to some other parts of the organization, we observe whether it works. And if it doesn't work or works, we call it, you know, adapt and regenerate the context and then go from there. So that's kind of the loop of the talk that I'll be going for with some examples. So, step by step going through. Generate. It's probably the one that you're all most familiar with. Because you're all prompting. You're like the human context creation typing things, right? I was actually amazed that I just asked, tell me when my talk is at AI engineer that you'll fetch the website and would just say, here's your talk. Like, blew my mic. But hey, I said like the context that I've given it, I'm Patrick, all that stuff, right? So, very simple context. It's what you do probably a lot in your setup. If you get a little bit more advanced, you say, that prompting is tedious, I want to have reusable prompts. So, you know, depending on the flavor of your coding agents, they call it instructions. Luckily, there's a little bit of a standardization now happening where it's like an agent MD and some pieces like that. Boo Claude for still calling it Claude MD. But anyway, you get the picture. There's like reusable prompts, reusable pieces of context that we're doing. We can also bring other context in. If we have documentation of libraries that we use day to day, we want to pull that in. Because the LLMs might not have the latest documentation. And so, it's hallucinating. Is it version 2, version 3? We don't know. So, we give it a context and say, please download the documentation. Hopefully, then agent optimized. And then they will do a better job at generating the code for that version of the library. Another piece of getting better context and creating context from libraries. And of course, it wouldn't be complete if we would say, pull context from wherever. MCP has been instrumental. Get it from your GitLab, GitHub, kind of Slack. All context. We're pulling in. We're creating. Even the ticket is creating context because we're pulling that in while we go there. And then maybe the new kid on the block is, okay, what if we start like writing our prompts as specifications, spec-driven development, which then gets broken down by the agent into a planning mode and do step-by-step kind of prompts that it then kind of runs through. So, a lot of creation happening in that field. Yeah. Simple. This is probably what you're closest to. But when you're typing all that context and creating all that context, you change two lines in your CloudMD. Do you know the impact? Is it like YOLO? Looks good to me. Let's do it. You have to think about how do we test things. It's not just about we have a piece of code and we have a piece of context now. We need to write tests to see what is the impact. New coding agents, we don't know where the lines still work. Now, it's not new in the world of AI engineering, but it's not that common yet in the world of coding with AI that you start writing evals for which are tests for your kind of code context. A little bit hard to read, but you know, if you think in parallels, we have different levels of testing in code and the simple one could be linting. Your IDE is has this wiggly lines like, hey, this is not like, you know, there's some incorrect syntax or you could do better like that. Here's an example of a validation of a skill where we say, well, you need to have the description. It can only be so long. So, it's validating according to the spec of the format of the context. In this case, simple analogy, simple linter that you can run. And then you can do other things like, and I haven't found maybe the good coding equivalent, but think of this as a grammarly. So, if you write context, is it actually, can the agent understand what you're writing? If you write two words, it's not verbose enough for it to actually understand the context. So, what you can do is you can say, okay, given this context, what do you think about, do you understand this? And then you can get feedback like, oh, it's not explicitly enough written or it's not complete, like you're missing pieces. So, that's kind of feedback that you can get out from tools as well. So, whenever you're writing now your context, you get a grammarly saying, hey, do this. That's why I like to voice code. For some reason, I'm way more elaborate voice coding than typing. I'm a bad typer, two fingers, still after so many years. But when I talk, I was like, you know, I see the sentences come on the screen, but it helps to get good context there. All right, another kind of test. So, imagine you put in your CloudMD or AgentMD, I should say, every API point must use the prefix awesome, right? You have some convention in your company, right? Which is great. So, your prompt will be then, add me a new endpoint to save a user. And you expect actually your coding agent to just say, the code that's being generated has kind of an awesome user. That's great. But the way we can test this is by asking them an LLM, the code that was generated, does it actually start with slash awesome? Now, you can do that with regex, I know, this is just for example purposes. But you can ask it to kind of judge your code based on your criteria and whether it did the right thing, right? So, imagine you would ask the same question without your context above. No LLM is ever going to prefix URL with awesome. So, that's kind of where your content or your company specific, your team specific things come in and that's why you still write those tests to see if this still works. Now, maybe Gemini kind of reacts differently than Copilot or something. I mean, your company, you need to make it more, you know, switchable of context. With this, you run the test and you can actually tell. That's the difference. And then you can make like whole suites. And I would compare that almost to unit tests. I have a bunch of these tests and they tell me whether that's actually, you know, good code, the code is following the rules and everything's fine. In this case, it's even kind of infrastructure as code. It doesn't need to be code only. It could be various things. It could be config files as well. And I just have, it's hard to read, but a bunch of kind of criteria that I just run every time to do that. But, if you want to test, you know, whether an endpoint has slash awesome slash user, there's a real test that we want to run, which is I want to test the endpoint. I just don't want only to check the code. I want to have it running. So, when you give the judge a tool and the judge becomes an agent and it can do things in a sandbox and execute stuff, it can actually do the curl. So, you can bind LLM as a judge with kind of some tooling and then you can have multitude of tests. Actually, you know, in this case, it kind of ends up being an end-to-end test, right? Because it's not just looking at the file, it's actually running the piece with everything that it's supposed to do. And then I can do this like given a certain commit in my repo, I want to run this scenario. So, given this piece of context, did it make a difference? Yes or no? So, you're kind of like building this up while you're committing context also within your repo. And because we now have tests and it gives us feedback whether it's working yes or no or what it's missing, we can optimize context. So, that's kind of the, you know, we can put that in a code action or something that says like, okay, fix this context, improve this context with all the feedback the LLM has given us to improve that. So, you know, again, coding improvements, but we start thinking more in testing that piece as well. One of the first reactions is once you have tests and optimizations, we run this in a CI CD system because that's perfect, right? That's where we run our tests and their test suites and do that. Now, there's a little bit of a real thing. If you run evals, you run it once, you run it another time, it might not give the same results. Remember, undeterministic things. So, you cannot say, well, run it once and then if it passes or not, you're going to be in for a treat because it's like, oh, I can't debug that. So, think about this, like you run it five times and out of five, how many times does it succeed? And, you know, maybe in several cases, it hits 100% all the time, which is great, but in others not. And depending on how you change your context, it will influence which tests actually work or not. I find it personally helpful helpful to think about this as error budgets. I give a set of tests, an error budget that I really care about. So, it's only allowed like, you know, to fail minimally and other pieces are okay. So, that's how you have to think about testing context. You cannot do like exact testing all the time. It's a different way that this works. All right. So, generate, hopefully you understood what the testing could do for you and distribute. Maybe that's also something you already did. If you maybe have checked context into your repo, right, which is great, you know, all of a sudden it becomes available, your colleague checks it out, zero friction, I can push, I can share. But, we have another mechanism for doing things. Think of this like, imagine you have a reusable context that you want to reuse across multiple projects, across multiple teams. We had the concept of a library. So, what if we package kind of pieces of context and then we are able to install pieces of context that we need for this project? Guidelines, front-end, it doesn't matter for that. And then if we take that up a notch, how to discover what packages exist, that's a registry, right? Now, in that way, it's no surprise that you'll see things like skills and kind of the Tesla registry and the marketplace where you can find a multitude of skills. Now, the reality is 99.9 and I mean that in a very sincere way of the skills is crap. But, it's good to learn from others to see what they're doing. But, hardly of them if you run kind of any set of evals on there is actually up to a quality standard. Now, that will likely improve. But, there's also a tendency is that a lot of the skills and pieces people actually want to put that in their own registry. So, I'll come to that later again. But, so, you start seeing the gist, a skill not only contains context, it can contain scripts, it can contain documents, contain a bunch of things. So, is this kind of the package format? Probably, you know, plugins could now also contain MCP. But, you see there's like a standard coming in. Skills all of a sudden, when that came out, all the coding agents said we're supporting this as almost like a package format for people to distribute their context on. And then, when I have one piece of context, I have dependencies. And, I'm sorry, but also with context, we're going to have dependency hell. Right? I'm going to download this for frontend and maybe it's conflicting what is in the React context package. And so, you start having to deal with that as well. So, you start seeing also packages that mirror your library versions, your code, like your context versions, and kind of pull that in as well. And, of course, when we have packages and people are publishing things in registry, we need security. Right? OpenClaw, thank you for that. Like everybody all of a sudden became aware that we need more secure things because we are able to run things on our laptop that are not and coming from strangers. Right? So, Snyk has a way of scanning context. Right? It's doing some credential handling. It's exposing some third-party pieces. So, you start seeing those scanners on the context as well. And then, when you think about security, who actually built this skill? How was it built? With what model was this built? So, all kind of capturing what we learned in maybe with packaging, like the s-bomb, is kind of the AI s-bomb, like the package of context that we're putting in. So, you've seen still on the path, right? You generate, evaluate, distribute. Let's move into observe. When you are making libraries off-scales and context for others, and I don't mean copy and paste this over Slack or something, but when you actually want to maintain this as something somebody else can use, similar to a library, when they start using that, how do you get feedback whether that still works? So, a great place to get feedback is actually by looking at the agent logs. So, imagine developer one coding on the project, and the agent is not doing what they want. They could put this into their context, which is great. Right? Okay. Let me do the TDD almost, like, you know, I hit a problem. It's not TDD, but you get my gist. It's not TDD. Or, what if we, at a team or an organization scale, would look at the logs every time an agent said, we're missing this piece? And we surface that and say, if everybody's missing this piece, we should create context for this. And then we distribute the context to everybody, and all of a sudden, the impact of improvement is for everybody. Luckily, like the agent MD, there is now our standards becoming for logs. So, we can read from logs, and that's part of our feedback channel to see if the agent is actually using or missing some of the context. Any feedback you get on a PR that's not complete. That's feedback on your context, because that PR was created with certain pieces of context. If you say this is not correct, you can kind of keep arguing on the PR, or you can just say, let's improve the context, so the next iteration actually improves, and you don't hit that same problem again. What about running code in production that was generated from context? And that's not correct, because, yes, we do our PR reviews, and we say thumbs up, thumbs down, and we give the feedback, but the actual feedback is also in production when it's running. So, this is a tool that actually instruments your code, pushes it out, it's almost like a wrapper, it pushes it out to production. When it fails, it says, these pieces of code were changed, and we're failing. Hey, in this case, input, output, it did something wrong. Can we create a test case for this, so the next time we don't hit this again in production, feedback loop. Now, these are all kind of pretty trivial, like missing pieces of context or improvements, but if you run agents, and the equivalent of scanning maybe, you know, in the CICD is, you need to make sure when it's running in production, is it not doing strange things? So, we need kind of a way of looking at that. Now, I've been toying myself with, you know, sandboxing agents, and it is very resourceful at finding things. I like, okay, you know, run this thing, try to figure out, like, anything useful to get break out of the system, and okay, it uses my environment variables. Okay, stupid, let's let me remove the secret. Let me look at your memory files. So, you have to really make sure that, like, whatever it's doing, you can have a way of tracing this as well. And, I apologize again for kind of the slide, but it just is, we can have a sandbox where the agent runs inside, you can have a code agent, but your code agent, by default, without any restrictions, loads your agent MD, you load your skill MD. Like, nothing is blocking that. So, if you download this, immediately, it's loaded. So, you can't filter that with sandboxes, you need to have another way. I call that a context filter. Think of this as a web application firewall that just filters out any patterns or prompt injections or stuff that is coming in directly in that piece. And if you take that, there's a lot of talk here as well on harness engineering. Harness engineering itself also has this kind of full observability, looking at logs, looking at traces, looking at feedback. So, it's kind of, you know, useful for training pieces, but as much useful for running your own piece as well. Harness engineering itself. So, those were the pieces for me today. I would say, for a lot of people, there's like create context, test context, think of this as your library authoring tool loop. And then, when you push this into the enterprise, there's an organizational loop. Hey, I made a library, somebody else is using it. I'm looking whether that's useful, whether that's still working, whether that's still working for all the other pieces. So, that's kind of like, the kind of improvement, almost like Sonar CI CD model for context. And then, you're currently probably doing a lot at the individual solo model. You're improving, you're honing, crafting your own kind of markdown. What if you start doing this more with your team? Make that a reflex. If it's missing, add some context. What if you put that out to a team of teams, and you start having a flywheel? You know, if you fix it here, the other team can reuse it. And that's kind of like, you know, scaling things out into the organization as well. And so, there's a lot of talk about LLMs and coding agents, and I all love them. But the way that I see it, is they're just the engine. If you give the engine the wrong fuel, which is context, they're not going to perform. So, and you can't do anything on the LLMs, at least not me, right? I'm just using the coding agent. I'm using whatever they give me. But I can optimize my context. And that's, I think, the message. Doing this more in an engineered way, than just copy and pasting things and hoping for the best, you know. If you like this talk, connect on LinkedIn for the slides. Give me some feedback, good and bad. If you want to try TESOL, where we implement some of the pieces of this, have a go. And if you're also interested in another conference, I know. You can never have enough conferences. Visit AI DevCon, which I curate the content for here in London, 1st and 2nd of June. And that's it. I can maybe take a few questions. Any questions? Sure. So, I'm wondering if you have any thoughts about, like, more exotic forms of context? Like, traditional ones? So, for example, one of the things I'm working on is an automated system for scoping out architectural problems, like trying to create hard definitions for them, where you can see that page and then, you know, create actual objective tests. Oh, cool. Yeah. Microphones. And one of the things I've been testing out is, like, ability to create consistency as a form of context or as a form of eval. So, given this rough, like, very loose definition of what the plan is, if you put that, if you try that agent system, turn that into a really crisp definition, and you just have that done in parallel, how often do you get the same crisp definition? And if they're all over the place, then the original definition is so poor, you need to, like, go back to basic principles or to an architect. But if they're all the same, then it's probably a pretty good definition that you can carry on with the downstream process. So, I guess, like, besides just code and typical evals, any other sources of context or generating context that you think is useful? Um, I don't have maybe a specific answer to your, like, exotic case, but I would say that maybe the piece that people underestimate is that once you, you know, you thought you were going to save time by writing actually your context instead of all your code, but if you take this rigorously, you're going to spend time on writing the right evals. Right. And that's kind of like, you know, a lot of work to kind of, because now you don't only have one prompt that you're trying to get right. It's like all the prompts of the evals. And that, like, people do almost like, like, the more advanced people, they almost have their own process, and they build their own process on top of, like, for building the right evals on your business case as well. So, yeah, good question. Thank you. Any other questions? If not, I'll be around. Say hi. I'm also going to be at the Tesla booth. So, thank you very much, and I'm going to make space for the next speaker. Thank you.