AI Engineer

Context Graphs for Explainable, Decision-Aware AI Agents — Andreas Kollegger & Zaid Zaim, Neo4j

1164 summary words 5 min summary Watch video

Start with the signal

5 min read

Summary

30-second take

Neo4j proposes "context graphs" as a shift from giving AI agents mere knowledge to encoding rules and policies that enable autonomous decision-making—the "why" behind actions, not just the "what." Andreas Kollegger presents a structured decision-making workflow (framing, global context, risk/value analysis, authority checks, and self-learning) designed to prevent agents from making dumb moves when they lack explicit instructions—like overspending on Red Bull when rent is due or prescribing the wrong drug to a 1% edge-case patient. The framework is domain-specific and non-trivial to generalize, but Neo4j is positioning graphs as the memory and policy layer for multi-agent systems that need auditability, precedent, and safe escalation paths.

Key takes

  • Context graphs encode decision logic, not just data: Unlike standard knowledge graphs, they store policies, rules, and precedents so agents understand why they should act, enabling them to operate autonomously in environments where every edge case wasn't pre-programmed. This matters when agents control real resources (credit cards, medical decisions) and can't just hallucinate their way through ambiguity.
  • Decision-making framework has seven stages: Framing (local context, causality, objective), global context (precedents and business rules), risk analysis (reference class validation, reversibility, cost of being wrong), value analysis (what you're maximizing), proposal generation (options with pros/cons), authority check (act or escalate), and logging (store reasoning for future agents). This structures the "decide whether to decide" step that multi-agent systems lack.
  • Reference class validation prevents statistical averages from killing outliers: Before applying "99% of the time, prescribe drug X" logic, an agent must first determine if the patient is in the 1% for whom that drug is fatal. Kollegger emphasizes making implicit human judgment (which matters, which doesn't) explicit for agents, especially in high-stakes domains like healthcare or finance.
  • Agents should propose, not decide alone: The framework separates analysis (generating options) from action (choosing and executing). An agent lacking certainty or authority escalates to another agent or human rather than guessing. This reduces blast radius and enforces accountability in multi-agent orchestration.
  • Memory graph = short-term + long-term + reasoning: Short-term captures conversation state, long-term holds entities (people, orgs), and reasoning layer embeds policies/rules. This tri-part structure feeds the decision framework with both historical behavior and current constraints.
  • Explainability comes from logging the entire reasoning chain: Every decision—what was considered, what wasn't, why options were ranked—gets stored in the graph as precedent. Future agents learn from this, and humans can audit when things go sideways.
  • Domain specificity is the caveat: Kollegger admits the framework is "very hard to generalize"—each step requires custom implementation per vertical (finance, healthcare, e-commerce). Neo4j is building a catalog of domain examples but doesn't have a one-size-fits-all playbook yet.

Useful details

  • Red Bull example: Agent tasked with "keep fridge stocked with Red Bull at all times" might order Red Bull when rent is due and money is low, unless decision rules specify budget priorities. Illustrates why agents need policy layers, not just task instructions.
  • Medical edge case: Prescribing a drug that's correct 99% of the time but fatal for 1% of patients. Agent must validate which reference class (99% or 1%) the patient belongs to before acting—statistical defaults fail in high-stakes scenarios.
  • Memory graph structure (visual): Green = short-term (conversations, state history), generalized = long-term (people, orgs), reasoning = policies/rules for why an action should happen.
  • Agentic GraphRAG flow: User query → agent checks internal knowledge → if missing, query graph DB via text-to-Cypher → traverse graph → return enriched, reliable content. Decision framework sits on top of this retrieval layer.
  • Multi-agent compartmentalization: Kollegger prefers "highly focused agents working together" over monolithic agents. One agent analyzes and proposes; another checks authority and acts. Reduces single-point-of-failure risk.
  • Implementation options: Framework can be built in LangGraph, ADK, or custom skills. Neo4j offers free online courses at Graph Academy to learn context graphs and decision workflows.
  • Reversibility matters: If a decision is easily reversible ("oh sorry, bad idea, let's undo"), risk tolerance changes. Agents should factor this into their analysis before committing.

Caveats / counterpoints

  • Generalization is hard: Kollegger explicitly states "it's very hard to generalize"—the framework is a template, but each step is "very domain specific." You can't drop this into any agent system and expect it to work without heavy customization.
  • No worked examples shown: The talk is conceptual. No live demo or code walkthrough of an agent using this framework in action. It's a blueprint, not a proven product.
  • Domain catalog is incomplete: Neo4j is "working on lots of different examples in different domains," but Kollegger invites attendees to propose domains they haven't tackled yet. This implies limited real-world validation outside their initial verticals.
  • Assumes access to structured policies: The framework depends on having "hard and soft rules" encoded—business process language, Slack docs, etc. If your org's policies are tribal knowledge or contradictory, the graph won't magically fix that.
  • Human-in-the-loop overhead: Escalation to humans or higher-privilege agents is good for safety, but if agents escalate constantly due to low certainty thresholds, you've just built an expensive approval queue, not an autonomous system.
  • No discussion of cost/latency: Running a seven-stage decision process with graph lookups, risk analysis, and precedent checks on every uncertain action could be slow and expensive. Transcript gives no benchmarks.

Ken relevance

High relevance for Ken's agent orchestration and ops work. If Ken is building multi-agent systems that need to handle edge cases, budget constraints, or high-stakes decisions (e.g., customer onboarding, content moderation, financial workflows), this framework offers a template for when and how agents should decide vs. escalate. The emphasis on explainability (logging reasoning chains) aligns with Ken's interest in auditability and debugging agent behavior. The tri-part memory model (short-term, long-term, reasoning) could inform how Ken structures agent state and retrieval in Neo4j or other graph DBs. Caveats: The lack of generalization and domain specificity means Ken would need to invest in custom implementation per use case—no plug-and-play solution. If Ken's agents operate in low-stakes environments (content tagging, summarization), the overhead of this framework may not justify the complexity. But for scenarios where "the agent might spend all your money" or "prescribe the wrong thing," this is directly actionable. Also useful for investing thesis: companies solving decision-aware agent orchestration (especially with graph-based memory and policy layers) are addressing a real gap as autonomous agents scale.

Watch verdict

Skim. The decision framework is conceptually solid and Ken will want the seven-stage breakdown for reference, but there's no code, no demo, and no proof it works at scale. The Red Bull and medical examples clarify the problem space well, but the talk is a pitch for an approach, not a case study. Ken can extract the framework diagram and key principles in 10 minutes without watching the full 20-minute talk, especially since the speakers admit implementation is domain-specific and non-trivial. Useful if Ken is designing agent decision logic; skippable if he's looking for turnkey tools or empirical results.

Full transcript 2800 words · 17 min read
0:00

[SPEAKER_00] Hello, good evening developers.

0:14

SPEAKER_00

How is everyone doing? Welcome to AI Engineer and welcome to the Neo4j context graph session number two. Excited to be here together with my peer at ABK to share with you more about context graphs, how they can help AI agents to become more decision aware. So one of the big challenges that we are trying to solve these days in the AI era is using knowledge graphs to unlock AI where we want to use graphs to fill the gap of knowledge, give AI agents the right tools and enrich them with the right content to solve certain challenges.

0:25

SPEAKER_00

So we all know that AI agents are really good in language, reasoning and creativity, but we want to fill the missing puzzle piece of knowledge for the agents. A big topic we have been discussing, we also saw before at Steve's talk, is memory. We talk about different types of memories: short term, long term and reasoning memory. In short term memory, we try to understand and capture the conversations that a user had with an AI agent. In long term memory, we capture more contextual knowledge, in this case, things around organizations, people and things. That gives us more of a generalized overview of the context.

0:56

SPEAKER_00

And of course we need reasoning here, helping the AI to make the right decisions to achieve a certain task. We saw also that for the background of this, we need a graph. A graph is made out of nodes and relationships to help you understand the deep and complex relations and connections between different types of data. So the big question is now why context graphs? We also have been talking lately a lot about context engineering.

1:29

SPEAKER_00

Now the big ask again is why context graphs? We are still within the sphere of context engineering when we talk about context graphs, but context graphs really represent a shift of having AI agents that already are very good at providing knowledge to their users, having the right tools and content and context. But with context graphs, we would really like to provide additionally to the knowledge the right rules and policies to the agents to help them become more capable to drive decisions. So not only knowledge, but also decisions.

1:45

SPEAKER_00

That's why we're moving not only to how an agent or what an agent can do, but also now we want to really capture within a context graph the missing why. So why an agent needs to do something. And this is something we capture with data that focuses on policies and rules. When we talk about Neo4j, we say also that graphs are everywhere. This is a basic replica of an organization. Each organization is basically also reflected here from the finance department to product to suppliers and so on and so forth. We saw also a good example by Steve before on financial services. So, is someone eligible to get a certain amount of money here.

2:35

SPEAKER_00

Yes or no. As part of this chapter, we are really going deeper with our customers and verticals to understand their needs and basically how we can help them solve the why question and help them drive better agents that can be more decision aware. This is also how a memory graph would look like. If we want an agent to also drive the right decisions, we need to give them memory capabilities. This is a quick overview of how the memory graph would look like. In green, we see the short term memory. We are capturing the conversations and the state history. In long term memory, we see more of a generalized overview of organizations and people.

3:12

SPEAKER_00

And most importantly, we also have the reasoning. So why an AI agent should do a certain task based on predefined policies and rules. Behind the story is always our foundation. If we would like to build Agentic GraphRag applications, we always have the AI agent where the user sends a query first. The agent looks for this knowledge or for this topic if it's available in its knowledge source. If not, we jump to the graph database. We have a certain set of tools like text to cipher that translates human text to our query language and then we traverse the graph for the right content and hopefully come back to the user with more reliable and qualitative content.

3:51

SPEAKER_00

So, Abike, you have been building a decision framework in the last couple of days. Let's have a look at that. [SPEAKER_01] Sure, thank you Zaid. Thank you Abike.

4:13

SPEAKER_00

[SPEAKER_01] So we've been hearing a lot about what the data model looks like for how you store memories. [SPEAKER_01] We're talking about context graphs and it's all about decision making. [SPEAKER_01] Why do we really care about decision making? [SPEAKER_01] And maybe you noticed the awesome animation. [SPEAKER_01] How many people have a claw? [SPEAKER_01] Claw owners, some of you. [SPEAKER_01] Some of us have claws these days, right? [SPEAKER_01] At least you have heard about claws. [SPEAKER_01] What's awesome about open claw, all the different claws, is they really exaggerate the need for good decision making.

4:59

SPEAKER_00

[SPEAKER_01] Where decision making is basically when you run into a point in the workflow or the life of the agent where it has to do something that it doesn't have in its setup. [SPEAKER_01] You haven't given it instructions about what to do in a certain circumstance. [SPEAKER_01] Autonomous agents run into this all the time. [SPEAKER_01] If you just let them go and do stuff, give them a credit card, give them access to your Amazon account, and say, hey, keep my stock completely, my fridge stocked with Red Bull at all times. [SPEAKER_01] Okay, if they notice that you need some Red Bull, they're going to go order some Red Bull.

5:29

SPEAKER_00

[SPEAKER_01] They may not know, oh yeah, I should order Red Bull unless the rent is coming up and I don't have enough money for rent. [SPEAKER_01] You may not have anticipated that. [SPEAKER_01] You haven't given it instructions about what to do in a certain circumstance. [SPEAKER_01] Autonomous agents run into this all the time.

5:56

SPEAKER_00

[SPEAKER_01] If you just let them off and go and do stuff for free, give them a credit card, give them access to your Amazon account, and say, hey, keep my stock completely, my fridge stocked with Red Bull at all times. [SPEAKER_01] Okay, if they notice that you need some Red Bull, they're going to go order some Red Bull. [SPEAKER_01] They may not know, oh yeah, I should order Red Bull unless the rent is coming up and I don't have enough money for rent.

6:15

SPEAKER_00

[SPEAKER_01] You may not have anticipated that. [SPEAKER_01] You can go and fix that with some prompt engineering and keep improving the instructions for the agent. [SPEAKER_01] But really there's a meta problem, which is how do you make good decisions? [SPEAKER_01] I'm going to talk through, given if you have some memory, if you have a graph around, you've got an agent running, what does an agentic workflow look like for actually making good decisions? [SPEAKER_01] Because the crab wants help. [SPEAKER_01] Help the poor crab. [SPEAKER_01] Okay, and actually it's even worse, of course, if you've gone down this path at all.

6:42

SPEAKER_00

[SPEAKER_01] You start with maybe one agent that's running. [SPEAKER_01] That becomes multiple agents that have specialized tasks, they're all off doing things. [SPEAKER_01] They all have to collaborate together. [SPEAKER_01] They all have to be self-aware of each other and themselves.

6:55

SPEAKER_01

The problem gets worse and worse at scale.

6:56

SPEAKER_00

[SPEAKER_01] So you need a thoughtful framework for how to make decisions.

6:57

SPEAKER_01

If you have friends who are in MBAs that went through some business school, perhaps here in the Judge Business School in the UK, they probably have entire courses on how to make good decisions. I'm going to talk through a framework that you can give to an agent. You can implement yourself in your choice of framework, whether it's LangGraph or if you just want to write some skills, you can give that a try as well. But this overall workflow is just how you make decisions generally as a human.

7:04

SPEAKER_01

And as we know for much of the engineering we do these days, transferring our implicit understanding about how to do things in an explicit way to the agents is what helps agents do their job really well. So given a context graph, given some memory, here's a workflow for actually making good decisions. Okay. Starts at the top with just framing the problem. It's the local context that matters. And for the local context, there's how did you get here in the first place? That's the causality of the second one actually. I would probably lead with that.

7:21

SPEAKER_01

You went through some reasoning chain, you went through some actions, and suddenly you're at a point where you have some amount of uncertainty. And rather than just making a choice about what to do next, you're, hang on, let me go into this sub-process, which is decision making, and figure out what to do. You've got an objective in mind, something you want to get out of that process. And the example we had before for finance was actually instigated by somebody who wants to increase their loan or something like that. Fine. Could be that, could be ordering Red Bull, whatever the decision might be.

7:38

SPEAKER_01

An objective, the causality that led to having that objective be a thing you want to resolve. And then there's an environment within which that's actually operating, that's purchasing, that's actually perhaps medical decision making guidance in that realm. So it goes from, hey, can I order Red Bull to there's a life at stake? Those are very different environments, but those environments matter for what kinds of decisions you end up making. That's the kind of thing that we understand because we experience these things. You have to tell the agent these things so that they understand it as well.

7:54

SPEAKER_01

Once you've got the framing, that feeds into a larger context, which is the global context around what did you do before in this situation? That's always a good idea. If you can stay consistent, that's really good. But then also there are global rules, both the hard and soft rules of a business. Maybe they're formally described in a business process language or something like that. Or they're just informal in some kind of guidance in Slack channels and Google Docs, whatever. Hard and soft rules come to place here for having alignment over the decisions you're making. And it's important to keep both of those things in balance. What did you do before?

8:18

SPEAKER_01

Could be the right thing to do again. But also what are the global rules that are applying now that might have been different before? So there might be a reason to overrule what prior decisions were. These are all part of the framework. You've got the inputs. You've got a framework for the decision making that is global. And now you've actually got to go through and do the analysis. And this is of course just the classic risk value analysis. And you're going to figure out for the risk side, that's probably where an agent should spend most of their time. If it's a serious thing, then you spend even more time.

8:42

SPEAKER_01

And by serious, I would mean things like this first point is called a reference class validation, where a reference class is rather than assuming what is good and what matters, you have to try to decide for the players involved, what is the most important thing for them? And so what is your reference point for what's important to them? The example I love here is if you have, in medical care, this is a very dangerous area to be doing, letting open call loose, right? If you're prescribing drugs for somebody, and 99% of the time, you might prescribe drug X for symptom Y. And that's the right thing to do 99% of the time.

8:59

SPEAKER_01

But for the 1% of the time, if you're in the small 1% of the population, giving that same drug might be fatal. So before deciding what you should do, it's very important to know if you're part of the 99% or the 1%. Statistical behavior does not really help you there. That's an extreme example, but there's flavors of that in many decision-making circumstances that right now are kind of implicit for us. A lot of our practice as AI engineers is being explicit about the implicit knowledge that we carry with us. So keep that in mind. Actually classifying what matters, what doesn't matter, what is the actual risk involved. Is the decision reversible? Can you take it back?

9:27

SPEAKER_01

drug might be fatal. So before deciding what you should do, it's very important to know if you're part of the 99% or the 1%. Statistical behavior does not really help you there. That's an extreme example, but there are flavors of that in many decision-making circumstances that right now are implicit for us. A lot of our practice as AI engineers is being explicit about the implicit knowledge that we carry with us. So keep that in mind. Actually classifying what matters, what doesn't matter, what is the actual risk involved. Is the decision reversible? Can you take it back? Be able to say, oh sorry, that was obviously a bad idea. Let's just back it off and do it again.

9:57

SPEAKER_01

That changes the dynamics of how you assess things. And then also, what's the cost of being wrong generally? Again, is a life at stake? Or is it, okay, I'm going to not have Red Bull in my refrigerator? Not great, but not the end of the world, so it's okay if that's a bad decision. Then also, of course, on the value side, it isn't obvious what you're maximizing. Are you trying to maximize your savings, your budget? So actually, don't spend all your money on Red Bull. Save some aside because you're saving for a vacation or for buying a house or whatever it might be. What is it, are you maximizing some value or are you trying to minimize some cost?

10:31

SPEAKER_01

That has to be part of it as well and it has to be explicit. Without that, an agent is going to take its general knowledge of conditions and be like, okay, well, most of the time, this is the right thing to do. But the particulars matter. The particulars always are important. And then the final thing is that the output of this stage is actually not making a decision. I like to think in terms of multi-agent systems. For me, I love compartmentalized, highly focused agents working together. This agent's job, the only thing would be to actually come up with a proposal of some alternatives and to give some pros and cons for those alternatives, not to actually make the choice.

10:57

SPEAKER_01

The choice then would be handed off to the next agent, which would be actually deciding whether or not it can act. Does it have the authority? And if it does, it can go ahead and rank the options that it has available to it and take those actions and have the impact actually be felt. Or it can decide, actually, I don't have certainty here or I don't have authority here. So let me escalate this to either another agent or another actor that could be a human that does have the authority that can make the decision. So this is incredibly important here. Either act or don't act. Let somebody else make the call.

11:22

SPEAKER_01

And the oversight here is basically the sub-process that would be kicked off for a human in the loop to step in or an agent with higher privileges. The final part of this is really an important self-learning part. Once you're done, once the decision has been made, either a decision has been made or there was not enough information or certainty to make the decision, you might defer actually doing anything. That's a possible state as well. Or if you make the decision, record the entire reasoning process, what you considered, what you didn't consider. This is part of the tracing, the accountability that we talked about in all of the other moments around context graphs.

11:39

SPEAKER_01

That all gets saved into the graph along with the decision itself and the actions that were taken that lets future agents do better at their job because they have this now as precedent that they can refer to. Okay, that was my speed run. This is the full diagram of the workflow. You can implement this again if you want to in LangGraph, in ADK, whatever you've got on hand. Write up some skills doing this. That works well. I will say it's very hard to generalize. Each of these things, this is a general framework. The particulars of every step end up being very domain specific. So a bit of a caveat there.

12:17

SPEAKER_01

We're obviously working on lots of different examples in different domains for this stuff. So if you've got a domain that we haven't really talked about yet, I'd love to talk about it so we can add it to our catalog of things. [SPEAKER_00] Okay. Very quickly, that was me. Call to action is the same as Steve had before. If you want to learn more about this stuff, we have free online courses at Graph Academy. You can learn more about context graphs, learn more about graphs. Come talk to us at the booths. Come talk to us after this talk. We love to talk about this all day long. Thank you.

12:49

SPEAKER_01

Actually classifying what matters, what doesn't matter, what is the actual risk involved. Is the decision reversible? Can you take it back? Be like, oh sorry, that was obviously a bad idea. Let's just back it off and do it again. That changes the dynamics of how you assess things. And then also, what's the cost of being wrong generally? Again, like is the life at stake? Or is it the, okay, I'm going to not have Red Bull in my refrigerator? Not great, but not the end of the world, so it's okay if that's a bad decision. Then also of course, on the value side, it isn't obvious what you're maximizing. Are you trying to maximize your saving your budget?

13:24

SPEAKER_01

So actually, don't spend all your money on Red Bull. Save some side because you're saving for a vacation or for buying a house or whatever it might be. What is it, are you maximizing some value or are you trying to minimize some cost? That has to be part of it as well and it has to be explicit. Without that, an agent is going to take its general knowledge of conditions and be like, okay, well, most of the time, this is the right thing to do. But the particulars matter. The particulars always really are important. And then the final thing is that the output of this stage is actually not making a decision. I like to think in terms of multi-agent systems.

13:55

SPEAKER_01

For me, I love compartmentalized, highly focused agents working together. This agent's job, only thing would be to actually come up with a proposal of some alternatives and to give some pro cons for those alternatives, not to actually make the choice. The choice then would be handed off to the next agent, which would be actually deciding whether or not it can act. Does it have the authority? And if it does, it can go ahead and rank the options that it has available to it and take those actions and have the impact actually be felt. Or it can decide, actually, I don't have certainty here or I don't have authority here.

14:29

SPEAKER_01

So let me escalate this to either another agent or another actor that could be a human that does have the authority that can make the decision. So this is incredibly important here. Either act or don't act. Let somebody else make the call. And the oversight here is basically the sub-process that would be kicked off for a human in the loop basically to step in or an agent with higher privileges. The final part of this is really an important self-learning part. Once you're done, once the decision has been made, either a decision has been made or there was not enough information or certainty to make the decision, you might defer actually doing anything.

15:05

SPEAKER_01

That's a possible state as well. Or you make the decision if you had, record the entire reasoning process, what you considered, what you didn't consider. This is part of the tracing, the accountability that we talked about in all of the other moments around context graphs. That all gets saved into the graph along with the decision itself and the actions that were taken that lets future agents do better at their job because they have this now as precedent that they can refer to. Okay, that was my speed run. This is the full diagram of the workflow. You can implement this again if you want to in LangGraph, in ADK, whatever you've got on hand.

15:38

SPEAKER_01

Write up some skills doing this. That works well. I will say it's very hard to generalize. Each of these things, this is a general framework. The particulars of every step end up being very domain specific. So a bit of a caveat there. We're obviously working on lots of different examples in different domains for this stuff. So if you've got a domain that we haven't really talked about yet, I'd love to talk about it so we can add it to our kind of catalog of things.

16:00

SPEAKER_00

Okay.

16:02

SPEAKER_01

Very quickly, that was me. Call to action is the same as Steve had as well before. If you want to learn more about this stuff, we have free online courses at Graph Academy. You can learn more about context graphs, learn more about graphs. Come talk to us at the booths. Come talk to us after this talk. We love to talk about this all day long. Thank you. .

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note