AI Engineer

Total Recall: Agent Memory and Harness Engineering — Ignacio Martinez, Oracle

2131 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Agent reliability comes less from the frozen, interchangeable LLM than from a deliberately engineered harness for memory, retrieval, context assembly, tool access, control loops, and model routing.
  • Why it matters: It offers a practical architecture for making OpenClaw-style agent systems more repeatable, token-efficient, fault-tolerant, and capable of improving behavior over time without retraining model weights.
  • Best use: Use it as a design review checklist for an agent control plane, especially memory tiering, dynamic tool/skill retrieval, bounded execution loops, and continual promotion of successful workflows into reusable skills.

Executive Summary

Martinez frames an AI agent as an LLM plus a harness: the LLM supplies non-deterministic reasoning, while the harness supplies the engineered components that make results reliable and repeatable. His central strategic argument is that models, infrastructure, and compute are becoming commoditized and swappable; the durable implementation advantage therefore shifts to the data and harness layers.

The talk’s practical center is memory engineering. It argues against treating an ever-larger context window as a replacement for memory, because long contexts suffer “context rot”: attention becomes diluted as more tokens compete for relevance, while attention costs scale quadratically. Instead, agents should keep active context small and assemble it per loop iteration from short-term state, durable episodic/procedural memory, enterprise semantics, retrieved tools, and skills.

The proposed storage design is hybrid. Files are model-friendly and convenient for mutable, unstructured working state, but databases provide transactionality, concurrency, high availability, access control, relations, and vector/hybrid search. The speaker recommends retaining ephemeral work such as a coding-agent to-do list in files while promoting durable preferences, experiences, and reusable workflows into structured database-backed memory.

The workshop implementation demonstrates a minimal harness built incrementally around a model, including retrieval, context construction, tool and skill selection, a fault-tolerant observe-reason-act loop, and traces/context-window inspection. The technical content is interwoven with an Oracle product pitch—especially Oracle Database File System, Oracle Agent Memory Package, in-database embeddings, and OCI model inference—but the underlying patterns are portable.

Key Takeaways

  • Claim: Treat the harness—not the LLM—as the primary control surface for agent quality, because the harness is what converts probabilistic reasoning into repeatable operational behavior. | Evidence: The speaker defines an agent as a model plus a harness: reasoning comes from the LLM, while memory, tools, and perception are the customizable components. He explicitly says the objective of harness engineering is reliable, predictable outputs despite the same model input potentially producing different outputs. | Implication: Prioritize investment in agent state, tool contracts, retrieval, observability, execution controls, and routing rather than locking the system to one model vendor. | Caveat: This is an architectural framing rather than evidence that a harness alone solves model capability, grounding, or safety failures; model choice still affects tool-call accuracy and required retry limits.
  • Claim: Do not use a large context window as a universal memory system; preserve a small, salient working context and retrieve or compact information selectively. | Evidence: Martinez calls context windows a form of short-term memory and describes “context rot,” where more material dilutes attention across the prompt. He notes transformer attention relates tokens to every other token, causing quadratic scaling as the context grows. | Implication: Build explicit policies for summarization, extraction, recency, retrieval, and context budgets instead of continually appending conversation history, tool schemas, and documents. | Caveat: The talk does not quantify a universal context-size threshold; the appropriate compacting and retrieval policy is workload- and model-dependent.
  • Claim: Use tiered memory: ephemeral local state for work in progress, and durable memory for facts, preferences, prior episodes, and proven procedures. | Evidence: Examples include a coding agent’s current to-do list as short-term memory; prior conversations as episodic memory; and a successful front-end workflow captured as procedural memory so later work can reproduce its design and implementation approach. Shared memory is presented as the coordination substrate between parent agents and sub-agents or collaborating agents. | Implication: Separate operational scratchpads from governed, long-lived memory, and define promotion criteria so only useful, reusable state becomes durable organizational or user memory. | Caveat: The transcript does not specify retention, consent, privacy, or deletion policies for user preferences and conversation-derived memories.
  • Claim: A hybrid file-plus-database memory substrate is preferable to an all-files or all-database design for multi-agent systems. | Evidence: Files are easy for models to create and edit and align with OS/POSIX workflows, but lack transactional consistency, hybrid/vector search, and durable recovery. The speaker uses multiple agents concurrently updating one file as the motivating concurrency problem; worktrees are presented as a workaround, whereas database-backed storage provides ACID operations, replication/high availability, relations, security, and vector search. | Implication: Keep filesystem-native artifacts where agent tooling benefits from them, but put shared state, durable memory, permissions, concurrent writes, and searchable agent knowledge behind transactional, governed storage. | Caveat: The specific Oracle DBFS solution is vendor-specific, and the claim that consolidating data creates only a “single attack vector” oversimplifies security: consolidation reduces system sprawl but increases the blast radius of a compromised platform.
  • Claim: Retrieve tools and skills dynamically rather than placing every available capability in every prompt. | Evidence: The toolbox and skillbox patterns store tool/skill descriptions for retrieval and load only the capabilities needed for the current agent-loop iteration. For large tool sets, Martinez proposes vector retrieval using HNSW indexes; when descriptions are too similar, he suggests generating richer LLM-enhanced descriptions to improve semantic separation. | Implication: Create a capability registry with high-quality descriptions, metadata, authorization attributes, and retrieval evaluation; dynamically assemble the allowed tool set per task rather than exposing the full enterprise tool catalog. | Caveat: Semantic similarity alone is not an authorization system: tools that expose confidential data require policy-aware filtering and identity/permission checks before or alongside retrieval.
  • Claim: Agent autonomy needs bounded, fault-tolerant control loops and model routing rather than unlimited retries on one model. | Evidence: The minimal agent loop is observe, reason, act, repeated with failure resistance. In the demonstration, an agent continues after an error and eventually answers through database-backed tools. Martinez recommends a configurable “hysteresis” or patience variable; for the cited Grok 4.1 Fast Reasoning example, he found a maximum of roughly 8–12 tool calls appropriate, though one demonstration completed in 2 steps and another took 16. He also advocates routing harder tasks to frontier models and easier tasks to smaller open-weight models. | Implication: Instrument loop step counts, failure causes, tool-call success, cost, and latency; use these to establish per-model/per-task stop conditions and route requests to specialized models where they are demonstrably sufficient. | Caveat: The 8–12 range is explicitly a personal empirical setting for one model, not a general production standard; retries must be bounded by cost, latency, safety, and task-success metrics.
  • Claim: Continual learning for practical agents can happen in context and workflow space without expensive weight updates. | Evidence: The speaker separates continual learning into changes to model weights, representations such as embeddings/reranking, and context/token space. He focuses on the last because it is the least expensive and most achievable: successful multi-hour workflows can be stored, distilled into an improved version of a skill.md file, and replace a prior skill with preferences such as favored libraries, design styles, or database choices. | Implication: Implement a governed skill-promotion pipeline: capture successful traces, distill them into candidate procedures, benchmark them against the prior version, approve or gate rollout, retain provenance, and support rollback. | Caveat: Automatically promoting past outputs can institutionalize errors, stale practices, or weak preferences unless skills are versioned, evaluated, and reversible.

Detailed Brief

Harness architecture and data-plane framing

  • Claims: The speaker places agents in a five-layer stack: application surface, data, model, infrastructure/orchestration, and compute.; He positions gateway/MCP as the connective layer that lets an otherwise isolated LLM access external data and programs, using Outlook as an example of a tool-enabled external application.; The semantic layer captures the enterprise’s implicit vocabulary and institutional assumptions—the knowledge coworkers omit because they already share it.
  • Evidence: The data layer is described as containing memory, knowledge, retrieval, encoding, search, gateway/MCP, semantic context, and tools/skills.; Martinez borrows the term “Umwelt,” associated in the talk with Jakob von Uexküll, to describe the interpretive lens through which an agent processes requests.; Examples of semantic-layer content include how enterprise data is modeled, query conventions, and metadata meanings.
  • Caveats: MCP provides connectivity, not inherent correctness, least privilege, auditability, or data-loss prevention; those must be engineered into the control plane.; The presentation treats data as the area of greatest control, but production differentiation may also depend on product UX, domain-specific evaluation, and operational process.
  • Implications: Represent semantic assumptions explicitly in retrievable schemas, glossary entries, policy metadata, and tool documentation rather than relying on a model to infer tribal knowledge.; Make tool connectivity, memory, and semantic grounding first-class platform services rather than ad hoc prompt additions within each agent.

Oracle-specific implementation and workshop artifacts

  • Claims: Oracle’s product proposition is a converged database that holds relational, JSON, vector, graph, spatial, textual, and metadata-oriented workloads in one engine and development stack.; Oracle Agent Memory Package (OAMP) is presented as a managed abstraction that creates a structured context card from a conversation in one call.; The workshop provides a GitHub Codespaces notebook with 19 implementation tasks and an Appbook-style UI for inspecting agent traces, selected tools and skills, data sources, and context-token usage.
  • Evidence: The context card includes topics, a compacted summary/current intent, relevant facts and preferences, associated memories, unanswered episodic questions, and recent messages.; Oracle describes LangChain Oracle DB integration for vector-store operations and in-database embeddings intended to avoid sending embedding inputs to a third-party service.; The demo application answered a revenue-by-product-category query while displaying its tool selection, schema/context, agent steps, and an observed recovery from an error.
  • Caveats: Most implementation shortcuts described in this portion depend on Oracle products or OCI; the talk does not provide comparative performance, cost, or operational evidence against other databases and memory frameworks.; Keeping embedding inference in a database may improve data-boundary control, but enterprise data protection still requires access controls, monitoring, retention rules, and threat modeling.
  • Implications: The reusable product requirement is not Oracle-specific: expose trace-level visibility into retrieved memory, context composition, selected capabilities, loop progress, errors, and token use.; Evaluate managed memory packages by their extraction quality, promotion controls, observability, interoperability, and governance—not merely their claim to reduce code to one API call.

Notable Concepts & Terms

  • Agent harness: The engineered layer around an LLM—memory, tools, perception, context, storage, loops, and policies—that makes an agent operationally reliable.
  • Context rot: The degradation in relevance and model attention as an overloaded context window accumulates too much material; used to argue for selective retrieval and compaction.
  • Context card: A compact, structured representation of a thread containing topic, summary/current intent, facts, preferences, memories, open questions, and recent local messages.
  • Semantic layer / Umwelt: The explicit representation of the otherwise unsaid domain and organizational knowledge through which an agent interprets requests.
  • Toolbox pattern / skillbox pattern: Retrieval-based catalogs that inject only task-relevant tools and reusable skills into the agent’s context instead of exposing every capability at once.
  • HNSW: Hierarchical Navigable Small World indexing, a graph-based approximate nearest-neighbor method proposed for scalable vector retrieval of tools, skills, and memories.
  • Skill promotion: Distilling a successful execution or workflow into a versioned reusable skill, then replacing or improving the prior procedure to change behavior without retraining weights.
  • Hysteresis/patience variable: A configurable cap on how many tool calls or recovery attempts an agent loop may make before stopping or escalating to a stronger model.

Operator Notes / Why Ken Should Care

  • Create a harness scorecard for current agents: context budget, memory tiers, promotion rules, tool retrieval precision, tool authorization enforcement, retry ceilings, routing logic, and trace completeness.
  • Run an experiment comparing full tool-schema prompting against retrieval-gated tool injection; measure selection accuracy, prompt-token reduction, latency, and unauthorized-tool exposure.
  • Add a memory promotion gate that requires provenance, evaluation against a baseline skill, versioning, explicit approval thresholds, expiry/review dates, and rollback capability.
  • Set task-class-specific stop policies rather than a global retry count; log loop-step distributions and route recurring failure classes to a stronger model, a deterministic workflow, or human review.
  • Threat-model shared memory and tool registries before consolidating them: define tenant boundaries, write permissions, data classification filters, audit trails, and incident blast-radius controls.
  • Require agent observability that renders the exact retrieved memories, semantic context, tools, skills, tool calls, errors, and token composition used in each consequential run.

Source/Metadata

  • Title: Total Recall: Agent Memory and Harness Engineering — Ignacio Martinez, Oracle
  • Transcript words: 13621
  • Duration seconds: 3647
  • Timestamp note: No timestamps or chapters were present in the supplied transcript. The transcript contains substantial duplicated workshop material and live-event setup/discussion.
Full transcript 8209 words · 60 min read
0:12

Hello, hello. Perfect. Yeah. So, perfect. Yeah, this volume is perfect. It's just a website that I registered the domain on Saturday just to do the registration of the workshop. But for those of you who are going to follow along with me, just know that after you complete that process in the workshop, you will get an invitation to the repository. And just view the invitation, accept it, and then you'll be able to run the workshop. We'll be running the workshop on GitHub Codespaces, so if you want to start that up, it takes five minutes. Just for reference, I'm going to use the first 30 minutes to give you an introduction to Agent Harness and lots of other concepts. And then, in the next hour and a half, we're going to actually go through the workshop together. Sound good? Okay, perfect. And thank you, by the way, for being here, because I know it's Monday 9 a.m., so I commend you all for being here. Just two minutes before we begin, so I will shut up.

0:21

Thank you. Thank you. Thank you. All right. Well, let's begin, guys. So thank you for being on this at this time here with me. I know it's 9 a.m., so hopefully I can, and me and my team can make your time worthwhile. At the end of this session, what I would like you to leave with is some knowledge on agent memory and agent harnesses, how you can build your own agent harness, which is nowadays one of the hot topics on AI, I would say. Lots of people are talking about models constantly, but the thing is that models, language models, they are the frozen part of the reasoning, right?

1:14

We just have to accept what we are given, and during the past few weeks, if you are following the news, you will have seen that this has never been more true than now.

1:29

So the harness is what we are going to talk about.

1:51

For those of you who weren't here, you can scan this QR code or go to that website, workshopwaitingroom.com, and just register.

2:20

You will get an invitation to a GitHub repo, and on this GitHub, you will be able to create a GitHub code space where we will run the workshop, and you will have everything set up for you.

2:49

So just a little introduction on who I am.

3:22

I have been working for Oracle for seven years. I've been a developer advocate for about four of them, and you can find my talks, and I'm very active on GitHub as well, so if you are a GitHub user, just check out my GitHub profile if you like. What is the highlight of my career so far? I launched a course with Andrew Ang on agent memory, so if you are interested in the memory components of what we are going to discuss today, you can just check out that course if you would like. These are the things that we are going to talk about today.

3:41

First of all, we are going to do a little introduction on what the agent stack is, and then we are going to zoom into the data layer where the memory lives. Then we are going to explore a little bit about the shapes, the different shapes that AI applications take nowadays. What an agent is, followed by the seven different layers that make up an agent harness. So if you just follow these seven different structures, you will be able to create an agent harness, a minimal agent harness that you can connect any model to.

4:01

Then we are going to finally talk a little bit about continual learning and how us, as Oracle, we are uniquely positioned to help you achieve and develop agent harnesses and AI applications. So the agent stack. By the way, if you have any questions, just raise your hand. I am very happy to take questions as well. So the agent stack. The agent stack, every agent, every AI agent sits on these five layers. You have an application, which is the product surface, what we interact with as users. You have data, and that has lots of components. You have memory, you have knowledge, you have retrieval, encoding, search.

4:50

You have the model itself, the reasoning large language model that is behind everything. Infrastructure, which is the orchestration. What model do we serve, depending on the reasoning effort that we need, and things like that. And then we have compute, which is the cloud or GPUs and also the database engine. And the application, the model, the infrastructure and the compute, these four layers, except for the data, they are increasingly commoditized. What do I mean by that?

5:26

They are trying to take away the complexity from these layers out of our domain, right? So the thing that we have the most control over when working with AI applications is actually the data. And that's the part that we're going to focus in, because the agent harness excels and lives very closely with the data layer. So let's focus on the data layer. The data layer is where an agent harness appears. And it has many components. The first component is the gateway and MCP. It's interesting because that's the one that connects an agent harness to data and tools.

6:24

So think of a large language model as an isolated thing that wouldn't be able to work at all if it didn't have access to things like data, right? So you then have the memory layer, the semantic layer, retrieval layer, context layer, and tools and skills that are built on top of the gateway and MCP layer. Let's say, for instance, that you have an application on your computer and your large language model doesn't have access to it. For instance, Outlook, right? So if you might want to have your large language model connect to that, you can create an MCP, you specify some functions, and then the LLM all of a sudden is able to communicate with this program.

6:47

So gateway and MCP is the layer that connects a large language model to the outside world, or de facto to our computer or wherever you're working. And AI applications today take up four different shapes, right? You have LLM chatbots, which are very passive. They just respond when you ask a question. You also have RAG applications, which are semi-passive because they have to do some processing in the background, but then you also have a passive nature of it.

7:11

And then you have the more active components of AI applications, which are what we use every day, like Code Cloud, etc., which are a combination of LLM-driven workflows that provide automation and AI agents that provide autonomy. And we'll explain later what I mean by that. But first, I want you to have a very clear definition of what an AI agent is. So to me, an AI agent is essentially a model, a large language model, plus a harness. The model itself will be the reasoning and everything else will be the harness. So this definition is something like this. An autonomous entity whose cognitive functions are powered by a large language model for reasoning.

7:48

They are augmented by a database or files for memory. They are extended through tools for actions and grounded in inputs that let it perceive its environment. So an agent is a model plus the harness. So if the agent is the model plus the harness, in this diagram, right, we have reasoning. And the reasoning is the part that we don't control. It's the part that is heavily subsidized, the part that we rent. And if you're like me, you're subscribed to every imaginable subscription on earth. And that's the thing that we do not control, right? We have no control over what we are offered.

8:38

And then we have memory, tools and perception that actually we get some customizability that we can do. So the goal of harness engineering is to create reliable and predictable outputs over and over. Whereas a reasoning model is very non-deterministic. You might give it the same input and it might produce different outputs every time, right? So I said that in AI applications, the most typical ones nowadays, Code Cloud, Codex, and any other type that you can think of is a combination of automation plus autonomy. Why? Well, because automation provides reliability to a system and autonomy gives flexibility.

9:06

And this is a combination that is very convenient when we are developing ourselves. It's the part that we are, that's heavily subsidized, the part that we rent. And if you're like me, you're subscribed to every imaginable subscription on earth. And that's the thing that we do not control, right? We have no control over what we are offered. And then we have memory tools and perception that actually we get some customizability that we can do. So the goal of harness engineering is to create reliable and predictable outputs over and over. Whereas a reasoning model is very non-deterministic.

9:49

You might give it the same input and it might produce different outputs every time, right? So in AI applications, the most typical ones nowadays, cloud code, codex, and any other type that you can think of is a combination of automation plus autonomy. Why? Well, because automation provides reliability to a system and autonomy gives flexibility. And this is a combination that is very convenient when we are developing ourselves. By the way, raise of hands, who is working as an AI engineer or as an AI developer? Oh my god, okay, good. So you must all have used one of these systems, right?

10:28

So all of them, they have this commonality, which is they have autonomy and they have flexibility. So the idea is that these systems, right, they are built on top of a proprietary agent harness. And an agent harness is nothing more than everything that we've spoken about an AI agent. All the things that you need to do around that to enable it to produce reliable and repeatable outcomes, right? So the model itself, non-deterministic. Same input, different outputs every time. But the harness, what we want to do with the harness, is to turn this non-deterministic nature of a large language model and be able to produce reliable and repeatable results.

11:03

So the reasoning, which is the part that we do not control, we're not going to focus. Actually, the harnesses are built on top of models that are swappable. You just need a common interface, an OpenAI protocol or the Anthropic API specification, right? All these things make it so that the model part is swappable and the harness is what we're going to focus on today. So seven things that make up an agent harness. And as I said, we're not going to touch on the model layer because we have no control over it. But let's go a little bit more in detail into each of these, right?

11:46

You need a storage layer on your agent harness that essentially determines where the data is going to live, where the memory physically lives. And we'll see about this dilemma that has been going on about the last six months about files versus databases. And why I think a hybrid combination of both is actually the best part. And then you also have memory engineering components, which are all the encoding, the search, and the retrieval components of it. And also the semantic layer, which is the hidden things that happen or the hidden vocabulary that we assume that a large language model knows, that is proprietary to our companies or our knowledge.

12:22

What we don't say to the LLM, that's the semantic layer. We'll lightly touch on agent loops and what an agent loop is and how to implement a very, very minimalistic agent loop. And finally, go about some context engineering techniques that keep the window as salient as possible, the context window as salient as possible.

12:38

You want to minimize the context window as much as possible so that the task that you're solving stays relevant. So until here, we have done an introduction to what an AI agent is, its use cases. And now we're going to dive deeper into an agent harness and each one of the individual components. So the model layer, right? The frozen reasoning core. I say it's frozen because typically the weights of a model don't change.

13:22

And I say typically because we are actually, I'm actually in the process with Cassius sitting right there. We're going to record a new course with Andrew Ng on continued learning for AI agents. So if you're interested, just check that out in a couple weeks. But the thing is that 99.9% of the time, you will have a model and the weights of the model will never change.

13:48

Unless you have millions of dollars or a lot of time or GPUs, it's very hard to change the weights of a model. So there are other ways in which you can affect the reasoning without actually changing the weights of the model. But this is motivation for the continued learning part that we will see in the workshop. Where does the memory live, right? The files versus databases dilemma that we've been having since January. Some people are very maximalists of files and some are maximalists of the database. And both things are right. Files have convenient characteristics and databases also do. So, the files, right? They are very easy to match the model's instincts.

14:52

They are very easy to create. They are very easy to insert and append data into files, right? It has a very unstructured nature to it. And databases, on the other hand, have a very structured way, right? But that's when you're thinking about SQL. And you don't actually need to choose one or the other. You can actually use both of them. Files are attractive because the model speaks them. And they work very easily with operating systems. They follow POSIX semantics. So they are compatible on Debian, Ubuntu, any other operating system that you might want. And they also have some disadvantages. For instance, they don't have transactional consistency.

16:00

So this is one of the problems that I wanted to talk about. If you are working with eight, 16, 32 agents at a time, the problem with this is that files cannot be modified and inserted and modified at the same time. So what is the solution nowadays to not having transactional consistency and working with files? Any suggestions? Work trees, exactly. So agents, when they want to modify a file but another agent is working on this thing, they just create a different work tree, right? They will do all the progress in the work tree. And then, after the implementation is done, they will merge to master or merge to main.

16:47

So this is the way that it's a workaround against not having transactional consistency, right? And you also have other characteristics, like hybrid search. This is not available on files, but it's very easily achieved on databases. So if your operating system gets corrupted or something, you will just lose everything. And these are things that the database fixed 35, 40 years ago, and people have forgotten about. So what I want to do is to give you the best of each of the implementations, right? You can get the benefits of files and the benefits of databases in the same place, and we'll see why. But these are some of the advantages, right?

17:10

You have ACID consistency, atomic operations, consistent operations, isolated and durable. You have high availability. You might replicate the database and have it in three different places in the world with a replication factor. You also have vector search, which is very easy. In files, you just have to do regular expression matching or a derivative of that. And lots of other things. So what we will do on the actual workshop is we're going to run an actual example of trying to modify a file with three different agents that will be working on the same file and trying to update a counter on this file. And let's see who's faster.

17:37

I know the answer, of course, but you will get to know it later. But what I want to introduce to you is that we have a thing called the Oracle DBFS, or the Database File System, where you can store files inside the database on a file system. And that gives you lots of the advantages that we set. You will get files with ACID transactional consistency. You will get vector search, relations, security, high availability, etc. Right? And my suggestion is that since a hybrid system works best, we can have things like short-term memory, right, that lives in files.

18:04

And when something needs to be promoted into a long-term memory, for instance, user preferences, things like this, they can go into a more structured space like a database. And this is what we'll do in the workshop. For the encoding, the search and the retrieval, which was another of the components in the agent harness, we also have a Langchain integration that I want to mention, called Langchain Oracle DB, that makes it very easy to insert into vector stores, search in the vector stores, and retrieve from the vector stores. Also, we have another thing called in-database embeddings. Have you ever heard about in-database embeddings? Yes? Okay.

18:44

So in-database embeddings is very convenient, especially for enterprise customers, because you will get the embedding model inside the database. Right? And my suggestion is that since a hybrid system works best, we can have things like, for instance, short-term memory, right, that lives in files. And when something needs to be promoted into a long-term memory, for instance, user preferences, things like this, they can go into a more structured space like a database. And this is what we'll do in the workshop.

18:49

For the encoding, the search and the retrieval, which was another of the components in the agent harness, we also have a Langchain integration that I want to mention, called Langchain Oracle DB, that makes it very easy to insert into vector stores, search in the vector stores, and retrieve from the vector stores.

18:55

Also, we have another thing called in-database embeddings. Have you ever heard about in-database embeddings? Yes? Okay. So in-database embeddings is very convenient, especially for enterprise customers, because you will get the embedding model inside the database, so that when you're doing embeddings, you don't have to call a third-party service. That's very convenient for data retention and data security purposes.

19:03

So this is what an embedding searching and re-ranking model would look like from a chatbot interface, for instance, right? You have a lot of documents, then you use a bi-encoder. So a bi-encoder is essentially an embedding model. You will create embeddings out of this document. You will split them and create vectors, put it into a vector store, right? And then you will get a user prompt, a user query, a question. You can also create an embedding out of that, and then compare it to what you had previously on your vector store. And this is how you get the most relevant vectors in an answer.

19:08

Then you will run a cross encoder, which is a re-ranker, and take a look at the question plus the result. And that's how questions are answered in RAG applications, right? So these things we are also going to touch briefly on the workshop.

19:16

And the thing is that imagine this RAG application from the beginning. From the documents, you create, you do lots of things, right? You do tokenization, then you create the embeddings. You have to do things like duplicating the data, normalization, redacting personally identifiable information, lots of things, right?

19:22

And then on your store, you have textual data from the documents. You have metadata, which is typically stored in JSON. You have vectors, which are represented as dense embeddings of 32 bits. And then on your list, you have so many types of data that you need to work on, that typically what people have is, for instance, you might need different databases for each one of these, right? You get what I mean? So there is a very high, nowadays, very high data synchronization logic overhead for AI engineers, or even for agents, right? So lots of maintenance required, and lots of engineering effort on it. And this is something that we want to avoid.

19:38

So what I want you to do today is just to try us out as Oracle. Try our database. We have support for every type of data imaginable that you can think of. We are called the Converged Database. We are the only Converged Database in the market that supports JSON, relational, spatial, graph, JSON. Anything that you can think of, you'll be able to create with us. And you'll get one database engine, one query interface, and one single development stack. Everything will be in the same database.

19:44

So also for data security purposes, you just have to secure and save all your data in one place. So a single attack vector is what you need to worry about. You don't need to worry about updating five different databases. You can just worry on securing one database. So we can be the engine of your AI applications, not just a single step, which is what people think of when they are working with databases.

19:55

Yeah, so agent memory. Agent memory is essentially a description of all the mechanisms and the systems that allow an agent to retain, reuse, refine, and recall information. We want to reuse the data and refine it in the process. But we want to reuse the data so that the next time that an agent or us as engineers, we are working on a problem that took us three hours, the next time that we observe this problem, the problem becomes easier, either for AI agents or for us.

19:56

So we have a lot of components, right? We have short-term memory, we have long-term memory, and then we have shared memory, which is something that's relatively new. And this shared memory is what happens when a sub-agent is communicating with its parent, for instance, or two agents are trying to collaborate on solving one specific problem together.

20:01

And then what I want you to see is that depending on the type of memory or the type of thing that we want to store, it will be in one place or the other, right? So short-term memory is ephemeral, is very short-lived, and it's very useful for things that are happening right now. For instance, the to-do list on a coding agent, right? It's happening right now, but you actually don't want to save that in long-term.

20:04

But then there are things like, for instance, episodic memory, things that previous conversations that you've had, that's very useful to have, for instance. I don't know if you use Claude. Some people are using Claude here. But in Claude, you might take your previous conversations and try to refine all your workflows and your skills based on the things that you've done in the past. So this is something that makes sense to save in the long run, right?

20:06

You also have things like procedural memory, previous workflows that have worked very well for your system. For instance, you worked on this front-end and then you created a very beautiful design that you like. You might take the whole conversation and turn that into a workflow that is repeatable and reusable, so that the next time you're working on the front-end, the results will be similar to the previous one, right? So these are the things that we will see on the workshop.

20:17

And some people say, okay, why do I even need all of this? People that have animosity towards agent memory. People say, okay, let's just put 15 million context window, even though it's not possible yet, but some people really believe that this is the thing, right? But the context window is a type of short-term memory. So it's useful for some things, but not for all of them. And one of the problems that happened with working with a context is this thing called context rot, or context degradation over time. And what happens is that the more things that you put into the context window, the less attention there will be for each one of the things that are in the context.

20:33

So at the beginning of a conversation, and this is a famous problem that the context window has, is at the beginning of the conversation, it will stay on track a lot because you just started the conversation. So let's say that, for instance, in school, right? Or if I'm having a conversation with you, I might have a chat with you for 30 minutes and your attention to me is very, very high because I've just started speaking. But if the conversation goes on for eight hours, then you want to punch me, right? Because I haven't shut up and you haven't learned almost anything at the end. And the problem is that attention, us humans, is very limited. So the more things that you put in the context, the attention matrix of the neural network will also degrade and it will scale quadratically, because the attention matrix, you know, is one token. It's essentially a reference of one token for every other token in the context window. So the bigger the context window is, the matrix scales on the number of rows and on the number of columns as well, which is a problem. So you want to keep the context window as small as possible to avoid context rot. And memory engineering, the components of memory engineering, so it's designing, building and doing everything around building agent memory for AI agents. And we want to retain, recall, reuse and refine this data in some way.

20:35

So it is a discipline, right? And here we have Valentin, for instance, and we have people from Oracle, my colleagues, all over the room. So if you see them, you can say hi to them. Valentin here, he's working on the development, or he worked on the development of this agent memory package. So if you have any questions about this, you can ask him. OAMP, or Oracle agent memory package, is the managed answer that we have in Oracle to lots of the problems that you will find when working with this type of data. For instance, as an engineer, if you do not have a managed solution, you need to make a lot of decisions. You need to see, when do I do context compaction? When do I summarize? How do I write the summarizer? What do I keep? What do I not keep? What do I extract from my previous conversations and when?

20:40

So it is a discipline, right? And here we have Valentin, for instance, and we have people from Oracle, my colleagues, all over the room. So if you see them, you can say hi to them. Valentin here, he's working on the development, or he worked on the development of this agent memory package. So if you have any questions about this, you can ask him. OAMP, or Oracle agent memory package, is the managed answer that we have in Oracle to lots of the problems that you will find when working with this type of data. For instance, as an engineer, if you do not have a managed solution, you need to make a lot of decisions. You need to see, when do I do context compaction? When do I summarize? How do I write the summarizer? What do I keep? What do I not keep? What do I extract from my previous conversations and when? How many tokens do I use for this problem? And lots of these things, right? With OAMP, an Oracle agent memory package, you can actually just do all of this in one single line of code. And we want to make it easier so that we reduce the cognitive load of AI engineers and AI agents as well. And with this context card thing, you will get something like what you see on the left. And this is an explanation of what each part does on it. But essentially, the topics that you see, for instance, from any conversation, you can create a context card from the thread or from the conversation. The topics will orient the model, the summary will compact the thread, and it will state the current intent of an AI agent. And the relevant information has three different parts, which is the facts, the preferences, and the memories that are associated to this conversation. Then you have the episodic memories that explicitly track the unanswered question that is going on right now. And the recent messages give local context to the model. So whenever you're feeling unsure, what do I need to do right now with the data that I have or this conversation, you might use the Oracle agent memory package on Python, and it is all assembled by one single call. Yes.

20:48

So you're not suggesting that this goes directly into the model. This is actually something that is used by the furnace. Exactly. That's what's going into the model. Yes. Yes. That's the component that you're talking about. The Oracle uses it, knows about the structure, and uses the structure to... Exactly. Exactly. So this structure, this is an abstraction that we build on top of the model, because the model we have no control over in most cases, right? The thing is that we can build on top of that, whatever abstractions we want, to make the model perform as reliably as possible.

20:57

So the memory and the semantic layer, they work together. I know we talked only about memory components. I'm going to walk you quickly, because I don't have a lot of time, to the semantic layer, to the semantic components of it, right? And the semantic layer is the meaning of what's going on behind it, that you assume that that happens, right?

21:01

So anyone from Germany? Okay, so I apologize for my pronunciation, but I'm going to try. So the Umwelt is the ambient movement of a model, and this was coined by Jacob von Wekskuhl, and this guy said that essentially every organism in the world that is living perceives its reality through a lens, and the lens is what it has access to. For us, humans, for instance, we have our eyes, our senses, right? So everything that we perceive and everything that we live, all our experiences, are seen through this lens, right? And an agent doesn't have human-like senses, but it has also a semantic layer or a semantic lens that everything that you ask it is filtered through, and this lens is essentially what you train it with, and then what you also give context to. So everything that you talk to the agent, right? When you talk to the agent, this agent will look it through the Umwelt. And the semantic layer is essentially the agent's Umwelt.

21:07

So the unsaid is what I want to focus on, for instance, organizational knowledge or enterprise knowledge, things that when you're working with a colleague, you don't mention this, because this is already known between you and your colleague. You need to specify everything that you work on every day, but if there was someone else, a child, that wanted to start working with you, you would have to specify everything in detail. These are all the unsaids, and this is what the semantic layer captures. So the unsaid is the actual tribal knowledge that an enterprise has, right? Institutional knowledge as well, how the data is modeled, how the queries are executed, what is the metadata? All these things are in the semantic layer.

21:08

And quickly, just for you to know that we're also going to implement a very minimalistic agent loop. And an agent loop is the driver of the model, right? It lets a model be independent and autonomous. And it is what makes a model, it turns a model into an agent, right? And this is the simplest agent loop that you can find, is this observing and reasoning, and then acting part that happens all of the time. And it has to be failure-resistant so that we never exit the loop. This is the idea. That the agent, you give autonomy to the agent, well, you give autonomy to the model so that it becomes the agent.

21:13

And then on the context engineering part, which is the last part of the seven layers of an agent harness, you also have lots of things that you can do. For instance, the toolbox pattern and the skill box pattern, which we are... this is on the workshop as well. And these are ways in which you can store the available tools and the available skills of a model so that they are retrieved optimally. And what you want to do is only retrieve the tools and the skills when they are actually needed and put it on the context window only when needed. Every iteration of an agent loop, you will see if this is actually the right place, and then if it's not, you can just take them out temporarily. So...

21:19

Oh, sorry. I thought I saw a question. So we will see all of this in the workshop. I don't want to take up too much time, but the idea is that we will assemble the context at every iteration on the agent loop.

21:24

And then, continual learning is the part that we talked about before about the ability to get better over time with the things that we've done with a model, right? So a frozen model, as we saw, it doesn't get better, right? But there are ways in which we can make an agent improve in the weights, which is the part that we are going to work on, in the representation part, so on the embedding and the re-ranking part, and also in the context window. And these are the three types of continual learning techniques. We are going to focus on the workshop on the context and the token space, because it's one of the easiest ones and one of the least expensive ones as well. I assume no one is a millionaire, or not many of us are millionaires, so this is also the most achievable and the most realistic way to change the model behavior over time. So without further ado, I just want to introduce you to this part, which is skill promotion and workflow promotion. So those skills or those workflows that you've done for three or four hours, right? You've been working for the whole day on a workflow, you were able to successfully do your job, and then these things can actually be retrieved, they can be stored into the memory components that we'll see. And then we will see about skill promotion. So if we promote a skill, for instance, we can promote it through a distillation process and create a better skill.md than the original. So we will retire the old version and update from the new version. And this allows us to do some continual learning on our own skills that turn them into more customized skills for ourselves, for our tone, our way to work, our preferences, for instance, let's use this specific library because I really like the look and feel of it. Let's use this specific database engine because it has less bugs or I found it easier to work with. All these things can be promoted into reusable and improvable skills over time.

21:29

So this is what the whole harness would look like at the end of the workshop. And hopefully what you leave this, you'll leave today with a better understanding of all the specific components that make up an agent harness. So let me go to here before I forget. If you are interested in the Oracle agent memory package or are working on the agent memory package, we have a Discord server in which you can chat with us. If you're a Discord user, just feel free to join this Discord server. I'll put the link later as well. But without further ado, let's begin with the actual workshop and I will tell you how. So let me show you first what we are going to build. Right. This is an app book. Oh, sorry. You don't see this. Yes.

21:33

So this is what the whole harness would look like at the end of the workshop. And hopefully what you leave today with is a better understanding of all the specific components that make up an agent harness. So let me go to here before I forget. If you are interested in the Oracle agent memory package or are working on the agent memory package, we have a Discord server in which you can chat with us. If you're a Discord user, just feel free to join this Discord server. I'll put the link later as well. But without further ado, let's begin with the actual workshop and I will tell you how. So let me show you first what we are going to build. Right. This is an app book. Oh, sorry. You don't see this.

21:39

Yes. So the first, yes, the workshop instructions, right? So for those of you who weren't here, you can just go into this website, register with your GitHub user, and then you will get an invitation to a GitHub repo. And from here, we will create a GitHub code space. So please make sure to do that right now. And all my colleagues are around the room to answer any questions that you might have during the creation of the code space, et cetera. Yes. Yeah. The internet. Are you having issues with the internet? Okay, let's see. So do you guys have the same Wi-Fi? AI.engineer.wifi. Okay. So can someone assist people with the Wi-Fi if possible?

22:13

Yeah. Please fix the system. Whoever. Do it.

22:25

I can't. Yeah. Yeah. My colleagues will take a look at that. Right. Yeah. Yeah. It always happens. Yep. All right. Say again, sir. Yes. Yes. They will go into the AI. I believe they will go into the AI engineer. So if you go into the session, I'll make sure to go there. If not, if you either join the Discord or any other, you can just message me as well on LinkedIn. I'll gladly give you the slides if you want. All right. Video team. Can you turn me on? Sorry. Switch me on, please. Oh, perfect. Thank you.

23:34

So what I'd like to show you is that we have built also, apart from the notebook that we're going to go through, we also built this app book. And with the app book, you can actually test every individual component of an agent harness individually. Right. So just for me to show you that this is possible. And you will get this automatically deployed in GitHub code spaces as well. So you will have these already deployed. And you will say, well, how am I actually making requests? We are going to be making requests to Oracle. Oh. Okay. Not found. Demo time. The idea is that. Yeah. Okay. I know what's happening.

23:58

So I lost connection to my code space because of inactivity. Let me restart. This app book is going to allow you to create and chat and interact with the whole agent harness. And the models that we're going to use are actually deployed on a managed service that we have on Oracle called OCI, the Generative AI service. We have partnerships with Google, with Meta, and with OpenAI, and with XAI for the time being. And we can actually provide inference to their models through our managed service. So think of us as the Enterprise Open Router, if you'd like.

24:08

So let me just go so you get started. You can get started. This is the repo, right? So the Agent Harness Workshop. If you're here, and thank you for starting that, by the way, if you're here, you just have to click on Open in GitHub Codespaces. And it will take you here. And you can select as many cores as you like. If you want to create this with 8 or 16, please don't, because I'm paying for this myself. But you might also create this with more resources. But just create the Codespace. And I'm going to pay for it, as I said. So don't worry about that. And this will create a new Codespace instance. And once it finishes, which it hasn't yet, I will show you what we can do with the Appbook and the Notebook.

24:14

But the idea is to use the remainder of the time that we have 1 hour and 15 minutes to go through the Notebook. And you will actually have to let me show you on GitHub, actually. You can go here. And inside the Notebook, after you deploy the Codespace, you will get a student notebook here. And this is one part of the workshop. Right? And here we're going to implement the whole Agent Harness substrate from scratch. So we're going to start with only the model. And then we're going to keep adding layers to the Agent Harness, as we saw, the seven layers. Right?

24:21

And we're going to be here to assist you. You will have to do some to-dos. So let me show you. There are a couple of things to do for you. So, for instance, the first thing that you need to do, you need to create a question. Right? The simplest thing of everything. You just have to communicate with a model with no Agent Harness implemented. Right? So the first thing you'll need is to ask any question that you like. This will go through the OpenAI Completions API, and it will return you a response. This is the simplest of all. And then we will start adding search, retrieval, encoding, and all the other components that we have seen. There are a total of 19 things that you need to do. If you finish first, raise your hand, and I will give you a hug because I don't have anything else.

24:29

And yeah. So anyone already deployed the Codespace? Okay. One person. Okay. Good job. So any. Yeah. If you have any questions or any problems, let me know. But this is what it looks like when you have it deployed. Okay. So let me go through this quickly. So you will get an app, right? And the app will already have everything that you need. If you want to deploy this app yourself, you might change this total recall port here.

24:46

Let me show you how I did it again. I go into ports. I clicked on the visibility of the port, and I changed this to public. And then this is now using a public gateway so that if I open the browser, I can actually get access to my individual total recall instance. So for instance, if I ask a question, I can show the total revenue by product category. And of course, this is mission control. So this is. This has all of the components that we've spoken about implemented already. You will get also a context window visualization of the things that are going on in the background. For instance, these are the tools that were selected by the agent harness to be loaded into the context to answer this question. This is the schema that's happening. And then we can also take a look at the individual agent traces that are going on. For instance, which skills are being loaded? What sources of data are we taking? And what are the tool calls being used? For instance, running a skill command, et cetera, to answer your question.

24:54

So the question is still being built. It's taking 16 steps. And it's going to. For instance, here, it detected an error, right? But because our agent harness is fault tolerant, it will keep trying because it has an agent loop implemented, et cetera, right? So all these things will actually yield you this result from the data in the database, right? And you can actually go into the context window, see how many tokens we're using. And if you're particularly interested in some of these parts, for instance, the Oracle agent memory package, for instance, you can interact also with only the context card, how the context card is being created, et cetera. So you will all get this deployed in your code base. Yeah?

25:00

Yeah. Yeah.

25:13

So his question for those of you who didn't listen, what happens if you have thousands of tools in an organization, right? Well, we introduced this concept called the toolbox pattern in this course with Andrew Ng. And the thing is that you can optimize so that the retrieval of these tools is negligible. So you will use hierarchical navigable small world indexes that use a graph structure, and then each node in the graph is a vector index or a vector store. And then you can actually, like HNSW indexes, be created for these types of problems only in the data, not in files. So great question. It doesn't have to worry you until you reach millions and millions of users. And tool calls, like different specific tool calls, you might not get five million. It's more like reading a file, writing a file, grepping, and all these kinds of tool calls that we do every day. They typically don't exceed 100 or 1000. But by being on a vector store, you abstract the complexity and the amount of it. You can just make a query 2000, just as simply as you would 10,000, because of the storage component that we choose, which is an HNSW index.

25:18

So you will use hierarchical navigable small world indexes that use a graph structure, and then each node in the graph is a vector index or a vector store. And then you can actually, like HNSW indexes, you can be created for these types of problems only in the data, not in files. So great question. It doesn't have to worry you until you reach millions and millions of users. And tool calls, like different specific tool calls, you might not get five million. It's more like reading a file, writing a file, grepping, and all these kinds of tool calls that we do every day. They typically don't exceed 100 or 1000. But by being on a vector store, you abstract the complexity and the amount of it. You can just make a query 2000, just as simply as you would 10,000, because of the storage component that we choose it, which is an HNSW index.

25:28

Yeah, some of them they have access to confidential data, for instance, some of them don't. So what do you think is that the tool descriptions are similar? Yes. So what do you want to start with the tool tools and figure out which one is the number?

25:52

Great question. So his question was, what happens if the tool descriptions that two different companies have are very similar, right? And one of the things we can do on the toolbox pattern is actually generate with LLM enhanced toolbox descriptions for specific tools to increase the separability of the tools. So if you think that the current descriptions of a tool or of a skill as well are not enough, you can actually enhance them with LLM retrieval like you would instead of running, for instance, named entity recognition, which is very caveman style. You can also do something more sophisticated, which is enhancing the kind of docstring enhanced representations of a tool, so that you increase the separability when you're doing vector search. Does that answer the question?

25:57

The agent loop? To come back, do you have a list, how do you find the sweet spot?

26:14

Yes, so there is a limit, of course, because we don't have infinite money, so we can't just keep trying and trying over and over if the generations are just hallucinations, right? There is a cutoff point that I set, depending on the frontier LLM that I'm using. For instance, for Grok 4.1 fast reasoning, which is the one that we're using here, I found that a value of 8 to 12, like maximum number of tool calls before giving up, is correct. Depends also on the accuracy and the correctness of the model, like for instance, in this case it was just able to show it in 2, before it was able to find it in 16. So sometimes it will have a faster retrieval, sometimes you will need to be a little bit more patient. But what I like to define is a variable, like a hysteresis variable, that holds the amount of patience that the harness will get with the model. Then you can do some other things, like for instance, if the model is garbage, you can just use, or use like a router for more different, like for difficult types of problems, you will route this problem to a frontier LLM. And then for the easier types of problems, you can just attach an open weights SLM, for instance, which will be more interesting.

26:23

Do you recommend multiple models for our experience?

26:32

Yes, yes. I think that's my personal opinion is that the future is a mixture of small experts for each type of problem. Some companies that they have developed like 100 million parameter models that work exceptionally well for one type of problem. And if you just have an aggregator or an orchestrator that routes the correct model to that, like the correct query to that model, then you will have a very token efficient type of agent harness. So you can actually do model routing inside the agent harness. Some companies are actually essentially only doing that. And they will charge you like, let's do, I'm going to charge you 10% of the tokens that I'm going to save you from the original amount of money that you were going to spend. Right? So let me show you the student notebook, right? So once you are inside the student notebook, for those of you who are not familiar with Visual Studio Code, you might need to select a kernel here, so that you run the notebook. So you might select Python 3.12 here, and then you can just start reading. If you stumble into a to-do that you need to do, you have a docs folder with all the explanations, the individual explanations that you need to solve this specific problem. For instance, the first to-do, which is just talking to the reasoning core, to the model layer, without doing anything else, it will just explain what you need to implement on that cell so that it works and you can proceed to the next one. And also have a solution. But if you're not lazy, you will try. And I hope that you try. And we will be here answering questions around the room. I'm going to turn off my microphone, just come down with my colleagues, and then let's chat about it for the remainder of the session. And if you have any questions or you like to talk more to us, please come by and swing by the booth, the Oracle booth. We'll be there every day, all the time. And you know, it makes it feel good. Like we are wanted and we have friends. So if you want to come up to us, just chat with us a little bit. It will be nice.

26:38

Internet? How's the internet? Internet? I feel like Caesar. All right, so I'm gonna leave this here. I'm gonna keep this here and I'm gonna come down. I'm gonna keep the internet. Thank you. and it's very useful for things that are happening right now. For instance, the to-do list on a coding agent, right? It's happening right now, but you actually don't want to save that, you know, in long-term. But then there are things like, for instance, episodic memory, things that previous conversations that you've had, that's very useful to have, for instance. I don't know if you use Claude. Some people are using Claude here.

27:40

But in Claude, you might take your previous conversations and try to refine all your workflows and your skills based on the things that you've done in the past. So this is something that makes sense to save in the long run, right? You also have things like procedural memory, previous workflows that have worked very well for your system. For instance, you worked on this front-end and then you created a very beautiful design that you like. You might take the whole conversation and turn that into a workflow that is repeatable and reusable, so that the next time you're working on the front-end, the results will be similar to the previous one, right?

28:20

So these are the things that we will see on the workshop. And some people say, okay, why do I even need all of this? Like people that are very... have animosity towards agent memory. People say, okay, let's just put like 15 million context window, even though it's not possible yet, but some people really believe that this is the thing, right? But the context window is a type of short-term memory. So it's useful for some things, but not for all of them. And one of the problems that that happened with working with a context is this thing called context rot, or context degradation over time. And what happens is that the more things that you put into the context window,

29:07

the less attention there will be for each one of the things that are in the context. So at the beginning of a conversation, and this is a famous problem that the context window has, is at the beginning of the conversation, it will stay on track a lot because you just started the conversation. So let's say that, for instance, like in school, right? Or if I'm having a conversation with you, I might have a chat with you for 30 minutes and your attention to me is very, very high because I've just started speaking. But if the conversation goes on for eight hours, then you want to punch me,

29:45

right? Because I haven't shut up and you haven't learned almost anything at the end. And the problem is that attention, like us humans, is very limited. So the more things that you put in the context, the attention matrix of the neural network will also degrade and it will like scale quadratically, because the attention matrix, you know, is one token. It's essentially a reference of one token for every other token in the context window. So the bigger the context window is, the matrix scales on the number of rows and on the number of columns as well, which is a problem. So you want to keep the context

30:24

window as small as possible to avoid context rot. And memory engineering, the components of memory engineering, so it's like designing, building and doing everything around building agent memory for AI agents. And we want to retain, recall, reuse and refine this data in some way. So it is a discipline, right? And here we have Valentin, for instance, and we have people from Oracle, my colleagues, all over the room. So if you see them, you can say hi to them. Valentin here, he's working on the development, or he worked on the development of this agent memory package. So if you have any questions about this, you can ask him. OAMP, or Oracle agent memory package, is the

31:18

managed answer that we have in Oracle to lots of the problems that you will find when working with this type of data. For instance, as an engineer, if you do not have a managed solution, you need to make a lot of decisions. You need to see, when do I do context compaction? When do I summarize? How do I write the summarizer? What do I keep? What do I not keep? What do I extract from my previous conversations and when? How many tokens do I use for this problem? And lots of these things, right? With OAMP, an Oracle agent memory package, you can actually just do all of this in one single line of code.

32:03

And we want to make it easier so that we reduce the cognitive load of AI engineers and AI agents as well. And with this context card thing, you will get something like what you see on the left. And this is kind of an explanation of what each part does on it. But essentially, the topics that you see, for instance, from any conversation, you can create a context card from the thread or from the conversation. The topics will orient the model, the summary will compact the thread, and it will state the current intent of an AI agent. And the relevant information has three different parts, which is the facts, the preferences,

32:50

and the memories that are associated to this conversation. Then you have the episodic memories that explicitly track the unanswered question that is going on right now. And the recent messages give like local context to the model. So whenever you're feeling like unsure, what do I need to do right now with the data that I have or this conversation, you might use the Oracle agent memory package on Python, and it is all assembled by one single call. Yes. So you're not suggesting that this goes directly into the model. This is actually something that is used by the furnace. Exactly. That's what's going into the model. Yes. Yes.

33:35

That's the component that you're talking about. The Oracle uses it knows about the structure, and it uses the structure to... Exactly. Exactly. So this structure, this is an abstraction that we build on top of the model, because the model we have no control over in most cases, right? The thing is that we can build on top of that, whatever abstractions we want, to make the model perform as reliably as possible.

34:03

So the memory and the semantic layer, they kind of work together. I know we talked only about memory components. I'm going to walk you quickly, because I don't have a lot of time, to the semantic layer, to the semantic components of it, right? And the semantic layer is the meaning of what's going on behind it, that you kind of assume that that happens, right? So anyone from Germany? Okay, so I apologize for my pronunciation, but I'm going to try. So the Umwelt is like the ambient movement of a model, and this was coined by Jacob von Wekskuhl, and this guy said that essentially every organism in the world

34:50

that is living perceives its reality through a lens, and the lens is what it has access to. For us, humans, for instance, we have our eyes, our senses, right? So everything that we perceive and everything that we live, all our experiences, are seen through this lens, right? And an agent doesn't have human-like senses, but it has also a kind of semantic layer or a semantic lens that everything that you ask it is filtered through, and this lens is essentially what you train it with, and then what you also give context to. So everything that you talk to the agent, right? When you talk to the agent, this agent will look it through the Umwelt.

35:39

And the semantic layer is essentially the agent's Umwelt. So the unsaid is what I want to focus on, like for instance, organizational knowledge or enterprise knowledge, things that when you're working with a colleague, you don't mention this, because this is already, you know, known between you and your colleague. You need to specify everything that you work on every day, but if there was someone else, like a child, that wanted to start working with you, you would have to specify everything very, very in detail. These are all the unsets, and this is what the semantic layer captures. So the unset is the actual tribal knowledge that an enterprise has, right?

36:27

Institutional knowledge as well, like how the data is modeled, how the queries are executed, what is the metadata? All these things are in the semantic layer. And quickly, just for you to know that we're also going to implement a very minimalistic agent loop. And an agent loop is like the driver of the model, right? It lets a model be kind of independent and autonomous. And it is what makes a model, it turns a model into an agent, right? And this is like the simplest agent loop that you can find, is kind of this observing and reasoning, and then acting part that happens all of the time, all of the time. And, you know, it has to be failure,

37:19

failure-resistant so that we never exit the loop. This is the idea. That the agent, you give autonomy to the agent, well, you give autonomy to the model so that it becomes the agent. And then on the context engineering part, which is the last part of the seven layers of an agent harness, you also have lots of things that you can do. For instance, the toolbox pattern and the skill box pattern, which we are... this is on the workshop as well. And these are ways in which you can store the available tools and the available skills of a model so that they are retrieved optimally. And what

37:57

you want to do is only retrieve the tools and the skills when they are actually needed and put it on the context window only when needed. Every iteration of an agent loop, you will see if this is actually the right place, and then if it's not, you can just take them out temporarily. So... Oh, sorry. I thought I saw a question. So we will see all of this in the workshop. I don't want to take up too much time, but the idea is that we will assemble the context at every iteration on the agent loop. And then, continual learning is the part that we talked before about the ability to get better over

38:41

time with the things that we've done with a model, right? So a frozen model, as we saw, it doesn't get better, right? But there are ways in which we can make an agent improve in the weights, which is the part that we are going to work on, in the representation part, so on the embedding and the re-ranking part, and also in the context window. And these are the three types of continual learning techniques. We are going to focus on the workshop on the context and the token space, because it's one of the easiest ones and one of the least expensive ones as well. I assume no one is a millionaire, or not many of us

39:25

are millionaires, so this is also the most achievable and the most realistic way to change the model behavior over time. So without further ado, I just want to introduce you to this part, which is skill promotion and workflow promotion. So those skills or those workflows that you've done for three or four hours, right? You've been working for the whole day on a workflow, you were able to successfully do your job, and then these things can actually be retrieved, they can be stored into the memory components that we'll see. And then we will see about skill promotion. So if we promote a skill, for instance, we can promote it

40:14

through a distillation process and create a better skill.md than the original. So we will retire the old version and update from the new version. And this allows us to do some kind of continual learning on our own skills that turn them into more customized skills for ourselves, for our tone, our way to work, our preferred, like our preferences, like for instance, let's use this specific library because I really like the look and feel of it. Let's use this specific database engine because it has less bugs or I found it easier to work with. All these things can be promoted into reusable and improvable skills over time.

41:01

So this is what the whole harness would look like at the end of the of the workflow, sorry, at the end of the workshop. And hopefully what you leave this, you'll leave today with a better understanding of all the specific components that make up an agent harness. So let me go to here before I forget. If you are interested in the Oracle agent memory package or are working on the agent memory package, we have a discord server in which you can just chat with us as stuff. If you're a discord user, just feel free to to join this discord server. I'll put the link later as well. But without further ado, let's begin with the

41:50

actual workshop and I will tell you how. So let me show you first what we are going to build. Right. This is an app book. Oh, sorry. You don't see this. Yeah.

42:11

Yes. So the first, yes, the workshop instructions, right? So for those of you who weren't here, you can just go into this website, register with your GitHub user, and then you will get an invitation like this to a GitHub repo. And from here, we will create a GitHub code space. So please make sure to do that right now. And all my colleagues are around the room to answer any questions that you might have during the creation of the code space, et cetera.

42:49

Yes. Yeah.

42:55

The internet. Are you having issues with the internet? Okay, let's see. So do you guys have the same Wi-Fi? AI.engineer.wifi. Okay. So can someone assist people with the Wi-Fi if possible?

43:16

Yeah. Please fix the system. Whoever. Do it. I can't. I can't.

43:28

Yeah. Yeah. My colleagues will take a look at that.

43:36

Right.

43:43

Yeah. Yeah. It always happens, you know. Yep.

43:59

All right.

44:04

Say again, sir. Yes. Yes. They will go into the AI. I believe they will go into the AI engineer. So if you go into the session, I'll make sure to go there. If not, if you either join the Discord or any other, you know, you can just message me as well on LinkedIn. I'll gladly give you the slides if you want. All right.

44:34

Video team. Can you turn me on? Sorry. Switch me on, please.

44:43

Oh, perfect. Thank you. So what I'd like you to show you is that we have built also, apart from the notebook that we're going to go through, We also built this app book. And with the app book, you can actually test every of the individual components of an agent harness individually. Right. So just for me to show you that this is possible. And you will get this automatically deployed in GitHub code spaces as well. So you will have these already deployed. And you will say, well, how am I actually making requests? We are going to be making requests to Oracle... Oh. Okay. Not found. Demo time. The idea is that... Yeah. Okay. I know what's happening.

45:35

So I lost connection to my code space because of inactivity. Let me restart. This app book is going to allow you to create and chat and interact with the whole agent harness. And the models that we're going to use are actually deployed on a managed service that we have on Oracle called OCI, the Generative AI service. We have partnerships with Google, with Meta, and with OpenAI, and with XAI for the time being. And we can actually provide inference to their models through our managed service. So think of us as the Enterprise Open Router, if you'd like. So let me just go so you get started. You can get started. This is the repo, right?

46:22

So the Agent Harness Workshop. If you're here... And thank you for starting that, by the way. If you're here, you just have to click on Open in GitHub Codespaces. And it will take you here. And you can select as many cores as you like. If you want to create this with 8 or 16, please don't, because I'm paying for this myself. But you might also create this with more resources. But just create the Codespace. And I'm going to pay for it, as I said. So don't worry about that. And this will create a new Codespace instance. And once it finishes, which it hasn't yet, I will show you what we can do with the Appbook and the Notebook.

47:13

But the idea is to use the remainder of the time that we have 1 hour and 15 minutes to go through the Notebook. And you will actually have to... Let me show you on GitHub, actually. You can go here. And inside the Notebook, after you deploy the Codespace, you will get a student notebook here. And this is one part of the workshop. Right? And here we're going to implement the whole Agent Harness substrate from scratch. So we're going to start with only the model. And then we're going to keep adding layers to the Agent Harness, as we saw, the seven layers. Right? And we're going to be here to assist you. You will have to do some to-dos. So let me show you.

48:05

There are a couple of things to do for you. So, for instance, the first thing that you need to do, you need to create a question. Right? The simplest thing of everything. You just have to communicate with a model with no Agent Harness as implemented. Right? So the first thing you'll need is to ask any question that you like. This will go through the OpenAI Completions API, and it will return you a response. This is the simplest of all. And then we will start adding search, retrieval, encoding, and all the other components that we have seen. There are a total of 19 things that you need to do. If you finish first,

48:48

raise your hand, and I will give you a hug because I don't have anything else. And yeah. So anyone already deployed the Codespace? Okay. One person. Okay. Good job. So any... Yeah. If you have any questions or any problems, let me know. But this is what it looks like when you have it deployed. Okay. So let me go through this quickly. So you will get an app, right? And the app will already have everything that you need. If you want to deploy this app yourself, you might change this total recall port here. Let me show you how I did it again. I go into ports. I clicked on the visibility of the port,

49:46

and I changed this to public. And then this is now using a public gateway so that if I open the browser, I can actually get access to my individual total recall instance. So for instance, if I ask a question, I can show the total revenue by product category. And of course, this is mission control. So this is... This has all of the components that we've spoken about implemented already. You will get also a context window visualization of the things that are going on on the background. For instance, these are the tools that were selected by the agent harness to be loaded into the context to answer this question.

50:29

This is the schema that's happening. And then we can also take a look at the individual agent traces that are going on. For instance, which skills are being loaded? What sources of data are we taking? And what are the tool calls being used? Like for instance, running a skill command, etc., to answer your question. So the question is still being built. It's taking 16 steps. And, you know, it's gonna... For instance, here, it detected an error, right? But because our agent harness is... fault tolerant, it will keep trying because it has an agent loop implemented, etc., right? So all these

51:11

things will actually yield you this result from the data in the database, right? And you can actually go into the context window, see how many tokens we're using. And if you're particularly interested in some of these parts, for instance, the Oracle agent memory package, for instance, you can interact also with only the context card, how to... how the context card is being created, etc., etc. So you will all get this deployed in your code base. Yeah?

51:50

Yeah. Yeah.

51:59

So his question for those of you who didn't listen, where what happens if you have thousands of tools in an organization, right? Well, we introduced this concept called the toolbox pattern in this course with Andrew Ang. And the thing is that you can optimize so that the retrieval of these tools is negligible. So you will use hierarchical navigable small world indexes that use a graph structure, and then each node in the graph is a vector index or a vector store. And then you can actually, like HNSW indexes, you can be created for these types of problems only in the data, not in files. So great question.

52:48

It doesn't have to worry you until you reach millions and millions of users. And tool calls, like different specific tool calls, you might not get five million. It's more like reading a file, writing a file, grepping, and all these kinds of tool calls that we do every day. They typically don't exceed 100 or 1000. But by being on a vector store, you abstract the complexity and the amount of it. You can just make a query 2000, just as simply as you would 10,000, because of the storage component that we choose it, which is an HNSW index.

53:44

Yeah, some of them they have access to confidential data, for instance, some of them don't. So what do you think is that the tool descriptions are similar? Yes. So what do you want to start with the tool tools and figure out which one is the number? Great question. So his question was, what happens if the tool descriptions that two different companies have are very similar, right? And one of the things we can do on the toolbox pattern is actually generate with LLM enhanced toolbox descriptions for specific tools to increase the separability of the tools. So if you think that the current descriptions of a tool or of a skill as well are not enough,

54:34

you can actually enhance them with LLM retrieval like you would instead of running, for instance, named entity recognition, which is very caveman style. You can also do something more sophisticated, which is enhancing the kind of like docstring enhanced representations of a tool, so that you increase the separability when you're doing vector search. Does that answer the question?

55:04

The agent loop? I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean, I mean to come back, do you have a list, how do you find the sweet spot?

55:45

Yes, so there is a limit, of course, because we don't have infinite money, so we can't just keep trying and trying over and over if the generations are just hallucinations, right? There is a cutoff point that I set, depending on the frontier LLM that I'm using. For instance, for grog 4.1 fast reasoning, which is the one that we're using here, I found that a value of 8 to 12, like maximum number of tool calls before giving up, is correct. Depends also on the accuracy and the correctness of the model, like for instance, in this case it was just able to show it in 2, before it was able to find it in 16. So sometimes it

56:34

will have a faster retrieval, sometimes you will need to be a little bit more patient. But what I like to define is a variable, like a hysteresis variable, that holds the amount of patience that the hardness will get with the model. Then you can do some other things, like for instance, if the model is garbage, you can just use, or use like a router for more different, like for difficult types of problems, you will route this problem to a frontier LLM. And then for the easier types of problems, you can just attach an open weights SLM, for instance, which will be more interesting. Do you recommend multiple models for our experience?

57:16

Yes, yes. I think that's, like my personal opinion is that the future is a mixture of small experts for each type of problem. Some companies that they have developed like 100 million parameter models that work exceptionally well for one type of problem. And if you just have an aggregator or an orchestrator that routes the correct model to that, like the correct query to that model, then you will have a very token efficient type of agent harness. So you can actually do model routing inside the agent harness. Some companies are actually essentially only doing that. And they will charge you like, let's do,

58:02

I'm going to charge you 10% of the tokens that I'm going to save you from the original amount of money that you were going to spend. Right? So let's, let me show you the student notebook, right? So once you are inside the student notebook, for those of you who are not familiar with Visual Studio Code, you might need to select a kernel here, so that you run the notebook. So you might select Python 3.12 here, and then you can just start reading. If you stumble into a to-do that you need to do, you have a docs folder with all the explanations, the individual explanations that you need to solve this specific problem. For instance, the first to-do,

58:51

which is just talking to the reasoning core, to the model layer, without doing anything else, it will just explain what you need to implement on that cell so that it works and you can proceed to the next one. And also have a solution. But if you, if you're not lazy, you will try. And I hope that you, that you try. And we will be here answering questions around the room. I'm going to turn off my microphone, just come down with my colleagues, and then let's chat about it for the remainder of the session. And if you have any questions or you like to talk more to us, please come by and swing by the booth, the Oracle booth. We'll be there every day,

59:38

all the time. And you know, it makes, it makes us feel good. Like we are wanted and we have friends. So if you want to come up to us, just chat with us a little bit. It will be nice.

59:55

Internet? How's the internet?

1:00:06

Internet? I feel like Caesar.

1:00:20

All right, so I'm gonna leave this here. I'm gonna keep this here and I'm gonna come down.

1:00:43

I'm gonna keep the internet. I'm gonna keep the internet. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note