AI Engineer

Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS

3404 summary words 15 min summary Watch video

Start with the signal

15 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Five code-level techniques (semantic tool selection, GraphRAG, multi-agent validation, neurosymbolic guardians, runtime steering) reduce AI agent hallucinations and token waste without changing prompts—all implementable in production with AWS Bedrock Agent Core.
  • Why it matters: Ken can eliminate agent hallucinations and token waste using concrete code patterns rather than prompt engineering, with clear production paths via AWS services.
  • Best use: Study code examples for agent systems; adopt techniques in order of implementation complexity; evaluate AWS Bedrock Agent Core for production deployment.

Executive Summary

Elizabeth Fuentes (AWS Developer Advocate) presents five production-ready techniques to stop AI agent hallucinations—each a code change, not a prompt change. The core insight: prompts are suggestions, code is constraint. Every token you send costs money, and incorrect context causes hallucinations. The talk demonstrates live code using Strands Agent (AWS open-source framework) with a 29-tool travel agent, showing before/after comparisons for each technique.

Technique 1 (Semantic Tool Selection) cuts token usage from ~3,000 to ~300 per call by filtering tools via vector search before model invocation. Technique 2 (GraphRAG) replaces text retrieval with structured graph queries (Neo4j + Cypher) for aggregation/counting tasks—eliminating the model's tendency to estimate when it should compute. Technique 3 (Multi-Agent Validation) uses three agents in sequence (executor, validator, critic) via Strands' SWAM class to catch fabricated success responses. Technique 4 (Neurosymbolic Guardians) enforces rules in Python hooks (not prompts) that fire before tool execution—models cannot skip them. Technique 5 (Runtime Steering) uses Agent Control SDK to apply soft rules that guide rather than block, enabling self-correction without hard stops.

All techniques run locally for development but map directly to AWS Bedrock Agent Core for production. Agent Core provides runtime, gateway (auto-tool-selection), short/long-term memory, CloudWatch observability, and DynamoDB-based steering rules—no custom infrastructure. The talk includes complete Jupyter notebooks with dummy tools, ground truth queries, token counting, and accuracy comparisons. Elizabeth emphasizes that these are operator-level patterns: production-ready, cost-reducing, and hallucination-preventing.

Key production insight: AWS Bedrock Agent Core Gateway handles semantic tool selection automatically (you register tools once, it builds the index). Agent Core Policies Service enforces neurosymbolic rules at infrastructure level. Runtime steering rules live in DynamoDB and update without redeployment. The architecture supports any framework (Strands, LangChain, custom) and integrates with external graph databases like Neo4j. Free tier available; repository includes CDK and notebook deployment options with AWS credits for testing.

Key Takeaways

  • Claim: Semantic tool selection reduces token costs by 90% for multi-tool agents by filtering tools before model invocation. | Evidence: Travel agent with 29 tools: baseline sends ~3,000 tokens/call just for tool schemas (170-200 tokens/tool). With vector-based tool filtering, only 3 most relevant tools are sent, dropping usage to ~300 tokens. Live demo shows 90-query test with token counting and accuracy comparison. Tools are embedded once in a local FAISS store (using sentence-transformers), queried per user message, and swapped in/out of agent state via Strands' tool register. | Caveat: Requires building and maintaining a tool embedding index; demo uses dummy tools with intentionally ambiguous queries to show failure cases. Accuracy may degrade if tool descriptions are too similar or queries are generic. Memory-enabled agents (chat history) still accumulate tokens unless tools are swapped per turn. | Implication: Ken should implement tool filtering for any agent with 10+ tools or conversational memory. Use AWS Bedrock Agent Core Gateway in production (auto-indexes tools, handles selection). Measure token savings in staging before rollout. Consider tool description quality as critical for semantic matching. | Timestamp: 00:00-15:00
  • Claim: GraphRAG replaces vector search with structured graph queries for aggregation/counting/multi-hop reasoning, eliminating hallucinated estimates. | Evidence: Demo compares traditional RAG (vector search returns 3 chunks from 300 docs, model estimates) vs. GraphRAG (Neo4j + Cypher query computes exact answer). Test cases: 'What is average guest rating across all hotels in Paris?' (RAG guesses from 2 samples, GraphRAG returns computed average); 'How many hotels have a pool?' (RAG says 'no specific information', GraphRAG returns exact count); 'Tell me about hotels in Antarctica' (RAG generates filler text, GraphRAG returns 'zero hotels'). Neo4j's knowledge graph pipeline auto-generates graph from text files using an LLM (OpenAI in demo). | Caveat: Requires building a knowledge graph (Neo4j local instance in demo, cloud in production). Model must generate valid Cypher queries (failure cases not shown). Not suitable for open-ended semantic search—only for structured queries. Neo4j GraphRAG library is free but Neo4j cloud has costs. | Implication: Ken should use GraphRAG for any agent that must aggregate, count, or traverse relationships (e.g., analytics agents, multi-entity reasoning). Pair with traditional RAG for hybrid systems (graph for structured, vector for semantic). Test Cypher query generation reliability before production. AWS Bedrock Agent Core supports external graph databases like Neo4j as a service. | Timestamp: 15:00-30:00
  • Claim: Multi-agent validation catches fabricated success responses by using three agents in sequence: executor, validator, critic. | Evidence: Demo uses Strands' SWAM class to chain agents. Test case: 'Confirm booking without payment.' Single agent calls tool, gets error, but returns 'Booking confirmed successfully' (fabricated success). SWAM version: executor calls tool and gets error, validator checks output and marks invalid, critic rejects with explanation. Ground truth test cases: valid booking (critic approves), no-payment booking (critic rejects), nonexistent hotel (critic rejects). SWAM handles handoffs automatically; agents defined via system prompts ('You are an executor agent...', 'You are a validator...', 'You are a critic...'). | Caveat: Triples token cost (three agent calls per request). Requires careful prompt engineering for each agent role. Demo uses dummy tools; real-world tool errors may be harder to detect. No latency comparison shown. Single-agent failures may be rare if tools are well-designed. | Implication: Ken should apply multi-agent validation only for high-stakes actions (payments, irreversible operations, compliance-critical tasks). Use as a safety layer, not for every query. Measure latency/cost tradeoff in production. AWS Bedrock Agent Core supports multi-agent workflows via SWAM or custom orchestration. | Timestamp: 30:00-40:00
  • Claim: Neurosymbolic guardians enforce rules in code (Python hooks) that fire before tool execution, preventing models from ignoring constraints written in prompts. | Evidence: Demo creates rules in Python classes (e.g., 'max 10 guests per booking', 'payment required before confirmation', 'check-in before check-out'). Strands' hook system provides three methods: HookProvider (registers callbacks), ToolRegistry (accesses tool state), BeforeToolCallEvent (intercepts calls before execution). Test case: 'Book hotel for 50 guests.' Baseline agent (rule in prompt) confirms booking despite violating 10-guest limit. Guarded agent (rule in hook) blocks booking and returns 'Max 10 guests per booking.' Rules are checked before tool runs; model cannot skip them. Three scenarios tested: no-payment confirmation (blocked), 11-guest booking (blocked), valid 5-guest booking (allowed). | Caveat: Hooks are all-or-nothing (hard blocks). Changing a rule requires code change and redeployment. Demo shows only blocking behavior, no steering. Rules must be written in Python; no declarative DSL. Not all frameworks support hooks (Strands does, others may not). | Implication: Ken should use neurosymbolic guardians for hard constraints (regulatory compliance, safety limits, business rules). Pair with runtime steering (next technique) for soft constraints. AWS Bedrock Agent Core Policies Service enforces rules at infrastructure level (no code changes needed). Use hooks in dev, policies in prod. | Timestamp: 40:00-50:00
  • Claim: Runtime steering allows agents to self-correct and complete tasks by guiding (not blocking) when rules fire, using external rule updates without redeployment. | Evidence: Demo uses Agent Control SDK (open source) with local server. Steering rule: 'If guest count exceeds 10, guide agent to split into multiple rooms.' Test case: 'Book hotel for 50 guests.' Hook-based agent (previous technique) blocks request. Steering-based agent splits booking into two rooms (one for 10, one for 5) and completes task. Agent Control SDK has two components: AgentControlPlugin (captures events, sends to server) and AgentControlSteerHandle (receives steering decisions). Rules stored in local server (dev) or DynamoDB (production). Rule changes take effect on next call without code redeployment. Demo shows three steering scenarios: oversized group (splits rooms), no payment (blocks), valid booking (proceeds). | Caveat: Requires running Agent Control server (local or cloud). Steering logic must be carefully designed to avoid infinite loops. Demo uses dummy data; real-world edge cases may be harder to steer. Not all constraints are steerable (e.g., hard legal limits). Adds complexity vs. simple hooks. | Implication: Ken should use runtime steering for soft constraints where flexibility improves UX (e.g., 'sold out flight? suggest next departure', 'room full? suggest splitting group'). Store steering rules in DynamoDB (AWS) for zero-downtime updates. Use hooks for hard rules, steering for soft. Test steering behavior exhaustively before production. | Timestamp: 50:00-60:00

Detailed Brief

Token Economics and the Core Problem

  • Claims: Every token sent to and from an agent costs money; tokens are billed as input and output.; Sending too many tokens (e.g., all 29 tool schemas every call) wastes money and degrades accuracy.; Sending too few tokens or wrong context causes hallucinations (fabricated answers, ignored rules).; The five techniques are code changes, not prompt changes—prompts are suggestions, code is constraint.
  • Evidence: Travel agent with 29 tools sends ~3,000 tokens/call just for tool schemas before any user message or response.; Each tool schema is 170-200 tokens (name, description, parameters in JSON).; Demo uses Strands Agent (AWS open-source framework) with OpenAI for model invocation, sentence-transformers for embeddings, FAISS for local vector store.; Live token counting via Strands' transaction context shows per-call usage.
  • Caveats: Demo uses dummy tools and ground truth queries; real-world agents may have different token patterns.; Token costs vary by model (OpenAI GPT-4 vs. Anthropic Claude vs. AWS Bedrock models).; Accuracy comparisons are based on 90 test queries with intentionally ambiguous phrasing to surface failures.
  • Implications: Ken should instrument agent token usage before applying optimizations (measure baseline).; Prioritize techniques by ROI: semantic tool selection (easiest, highest savings) first, GraphRAG (highest accuracy gain) second, validation/guardians (highest reliability) third.; Use local dev setup (Jupyter notebooks, free models like Ollama) for testing before cloud deployment.

Production Architecture: AWS Bedrock Agent Core

  • Claims: All five techniques have production equivalents in AWS Bedrock Agent Core (no custom infrastructure).; Agent Core Gateway handles semantic tool selection automatically (you register tools once, it builds/maintains vector index).; Agent Core Policies Service enforces neurosymbolic rules at infrastructure level (same concept as hooks, but declarative).; Steering rules live in DynamoDB and update without redeployment; Agent Core picks them up on next call.; Agent Core provides runtime, short/long-term memory, CloudWatch observability, and Lambda integration for tools.
  • Evidence: Architecture diagram shows Strands running inside Agent Core runtime (supports any framework: Strands, LangChain, custom).; Gateway automatically converts Lambda functions to tools and routes calls.; External graph databases (Neo4j, etc.) integrate via API; Neo4j offers free tier.; Repository includes two deployment options: Jupyter notebook (for AWS beginners) and CDK (Cloud Development Kit) for one-command deploy.; Elizabeth provides AWS credits via resource links for testing.
  • Caveats: Requires AWS account and credentials; free tier available but not unlimited.; Agent Core is AWS-specific; not portable to other clouds without rewriting.; Repository is a demo; production use requires additional hardening (error handling, rate limiting, monitoring).; CDK deployment assumes familiarity with infrastructure-as-code; notebook option is more accessible.
  • Implications: Ken should test locally first (notebooks + Ollama = zero cost), then deploy to AWS Bedrock Agent Core for production.; Use Agent Core Gateway for tool management (avoid building custom tool indexes).; Store steering rules in DynamoDB for operational flexibility (business users can update rules without code changes).; Monitor token usage and costs via CloudWatch; set budget alerts.

Implementation Details: Strands Agent Framework

  • Claims: Strands Agent is AWS's open-source framework for building agentic applications (maintained by AWS).; Tools are defined with Python decorators; Strands auto-generates JSON schemas for model consumption.; Agent loop exposes full state control (tool registry, memory, hooks) via simple APIs.; SWAM class enables multi-agent orchestration without manual for/while loops.; Hooks (HookProvider, BeforeToolCallEvent, AfterToolCallEvent) allow interception at specific lifecycle points.
  • Evidence: Tool example: '@tool' decorator on Python function with docstring generates schema with name/description/parameters.; Tool swapping: agent.tool_register.clear() removes tools, agent.tool_register.add() adds new ones per turn.; SWAM demo chains three agents (executor, validator, critic) with entry_point and max_handoffs parameters.; Hook demo shows rule classes (BookingRule, ConfirmationRule) with validation logic called before tool execution.; Agent Control SDK integrates with Strands via AgentControlPlugin and AgentControlSteerHandle.
  • Caveats: Strands is AWS-maintained but open source; community support may be limited vs. LangChain.; Hooks are Strands-specific; other frameworks may have different mechanisms (or none).; SWAM is a Strands feature; multi-agent orchestration in other frameworks requires custom code.; Demo uses OpenAI; Strands supports AWS Bedrock, Anthropic, Ollama, but API differences exist.
  • Implications: Ken should evaluate Strands vs. LangChain/AutoGen based on AWS commitment and framework maturity needs.; If using Strands, leverage SWAM for multi-agent workflows (simpler than custom orchestration).; If using other frameworks, implement hook-equivalent logic for neurosymbolic rules (e.g., LangChain callbacks).; Use tool swapping pattern (clear + add) for memory-enabled agents to prevent tool bloat.

GraphRAG Implementation: Neo4j + Cypher

  • Claims: Neo4j's knowledge graph pipeline auto-generates graphs from text files using an LLM (OpenAI in demo).; Cypher query language (similar to SQL) allows precise queries for aggregation, counting, multi-hop reasoning.; GraphRAG eliminates model estimation by returning computed results directly (no LLM reasoning needed).; Tools for GraphRAG require model to generate Cypher queries; tool description must explain Cypher syntax.
  • Evidence: Demo uses Neo4j local instance; GraphRAG library (pip install neo4j-graphrag).; Knowledge graph built from .txt files via simple_knowledge_graph_pipeline (one function call).; Tool example: search_graph_knowledge_base tool with description 'To search, build a Cypher query...'.; Test case: 'Average guest rating across all hotels in Paris' returns computed average (4.7) vs. RAG's estimate from 2 samples.; Neo4j driver (GraphDatabase.driver) connects to local instance; queries executed via driver.execute_query().
  • Caveats: Requires building and maintaining a knowledge graph (data modeling, schema design, updates).; Model must generate valid Cypher queries; syntax errors cause tool failures (demo does not show error cases).; Not suitable for open-ended semantic search (e.g., 'Tell me about Paris')—only structured queries.; Neo4j local setup is manual; cloud deployment (Neo4j Aura) has costs.; Demo uses simple graph; complex schemas may require query optimization.
  • Implications: Ken should use GraphRAG for analytics agents, compliance queries, multi-entity reasoning (e.g., 'Which customers in region X purchased product Y?').; Build hybrid systems: vector search for semantic, graph queries for structured.; Test Cypher query generation reliability with diverse query types before production.; Use Neo4j Aura (cloud) or AWS Neptune (managed graph DB) for production; avoid self-hosted Neo4j.

Operational Patterns: Dev to Prod

  • Claims: All demos run locally first (Jupyter notebooks, free models, local vector stores/graphs).; Production deployment uses AWS Bedrock Agent Core (managed runtime, no servers to manage).; Rules update without redeployment via DynamoDB (Agent Core polling model).; Repository includes complete code: notebooks, CDK templates, requirements.txt, dummy tools, ground truth queries.
  • Evidence: Demo uses sentence-transformers (free local embeddings), FAISS (local vector store), Neo4j (local graph), Ollama (optional free LLM).; AWS deployment via CDK (one command) or Jupyter notebook (step-by-step for beginners).; DynamoDB steering rules example: update rule in DynamoDB console, next agent call picks up new rule.; Repository link: [QR code shown; repo contains all notebooks, CDK code, AWS credits link].; Elizabeth emphasizes 'don't stop at the demo—go to production with Agent Core.'
  • Caveats: Local testing has no memory/state persistence (notebooks use in-memory state).; AWS Bedrock Agent Core requires AWS expertise (IAM roles, Lambda, DynamoDB, CloudWatch).; CDK deployment may fail if AWS quotas are low (e.g., Lambda concurrency limits).; Repository is a demo; production hardening not included (e.g., multi-region, DR, compliance).
  • Implications: Ken should prototype locally (zero cost, fast iteration), then deploy to AWS for production.; Use CDK if infrastructure-as-code is preferred; use notebook option if new to AWS.; Plan for operational costs: Bedrock API calls, DynamoDB reads/writes, Lambda invocations, CloudWatch logs.; Consider multi-region deployment for high-availability agents (not covered in demo).

Notable Concepts & Terms

  • Semantic Tool Selection: Filter tools via vector search before model invocation to reduce token usage and improve accuracy by sending only relevant tools per query.
  • GraphRAG: Replace vector-based RAG with graph-based queries (Neo4j + Cypher) for aggregation/counting/multi-hop reasoning, eliminating model estimation.
  • Multi-Agent Validation (SWAM): Chain three agents (executor, validator, critic) to catch fabricated success responses; SWAM (Strands class) handles orchestration automatically.
  • Neurosymbolic Guardians: Enforce rules in code (Python hooks) that fire before tool execution, preventing models from ignoring constraints in prompts.
  • Runtime Steering: Apply soft rules that guide (not block) agents to self-correct and complete tasks; rules update via API without redeployment.
  • Strands Agent: AWS-maintained open-source framework for building agents with tool decorators, SWAM orchestration, hooks, and full agent loop control.
  • AWS Bedrock Agent Core: Managed service for production agents providing runtime, gateway (auto-tool-selection), memory, observability, and policies (rule enforcement).
  • Agent Core Gateway: AWS service that auto-indexes tools, handles semantic tool selection, and routes tool calls to Lambda functions without custom infrastructure.
  • Agent Core Policies Service: AWS infrastructure-level rule enforcement (equivalent to neurosymbolic hooks but declarative and managed).
  • Hooks (Strands): Functions called automatically at specific agent loop points (BeforeToolCallEvent, AfterToolCallEvent) for rule enforcement or logging.
  • Tool Swapping: Pattern of clearing and re-adding tools per agent turn to prevent tool bloat in conversational agents with memory.
  • Cypher Query Language: Neo4j's query language (similar to SQL) for graph databases; used in GraphRAG to compute precise answers via structured queries.
  • Agent Control SDK: Open-source library for runtime steering; captures agent events, sends to server, receives steering decisions, and updates agent behavior dynamically.
  • Token Economics: Cost model where every token sent (input) and received (output) from an LLM incurs charges; optimization critical for agent systems.

Operator Notes / Why Ken Should Care

  • Ken: These are production-ready patterns, not research. Implement in order: (1) semantic tool selection (easiest, highest ROI), (2) GraphRAG (for analytics/compliance), (3) multi-agent validation (high-stakes only), (4) neurosymbolic guardians (hard rules), (5) runtime steering (soft rules). Measure token savings before/after each technique.
  • AWS Bedrock Agent Core is the production path—Gateway handles tool selection, Policies enforce rules, DynamoDB enables zero-downtime rule updates. Prototype locally (free), deploy to AWS for scale. Repository includes CDK and notebook deployment options.
  • Strands Agent is AWS's framework—comparable to LangChain but AWS-native. Key features: tool decorators (schema auto-gen), SWAM (multi-agent), hooks (lifecycle interception), full state control. Evaluate vs. LangChain based on AWS commitment.
  • GraphRAG is critical for agents that aggregate/count/traverse (e.g., 'How many X?', 'What is average Y?'). Vector search estimates; graph queries compute. Use Neo4j's auto-graph-builder (simple_knowledge_graph_pipeline) for fast setup. Test Cypher query generation reliability.
  • Multi-agent validation triples token cost—reserve for high-stakes actions (payments, irreversible ops). SWAM simplifies orchestration but requires careful prompt engineering per agent role. Measure latency in staging.
  • Neurosymbolic guardians are all-or-nothing (hooks block unconditionally). Use for hard constraints (regulatory, safety). AWS Bedrock Policies Service is production equivalent (declarative, managed). Pair with runtime steering for soft rules.
  • Runtime steering enables self-correction without blocking. Agent Control SDK stores rules externally (local server or DynamoDB). Use for UX-critical paths (e.g., 'flight full? suggest next departure'). Test steering logic exhaustively to avoid infinite loops.
  • Token counting is foundational—instrument baseline usage before optimizations. Strands exposes token counts via transaction context. Track input/output tokens separately. Set CloudWatch alarms for unexpected usage spikes.
  • Repository is demo-grade—production requires hardening (error handling, rate limiting, multi-region, DR, compliance). Use as reference, not production code. Elizabeth provides AWS credits for testing (check resource links).
  • For agent systems: semantic tool selection reduces waste, GraphRAG eliminates hallucinations on structured queries, validation catches fabricated responses, guardians enforce rules, steering improves UX. All techniques are complementary; use together for robust agents.

Watch Map

  • 00:00: Introduction: 5 techniques overview, token economics, Strands Agent framework intro
  • 05:00: Technique 1: Semantic tool selection—live code demo (29 tools → 3 tools, token reduction)
  • 15:00: Technique 2: GraphRAG—Neo4j + Cypher demo (aggregation, counting, multi-hop reasoning)
  • 30:00: Technique 3: Multi-agent validation—SWAM demo (executor, validator, critic)
  • 40:00: Technique 4: Neurosymbolic guardians—hooks demo (rule enforcement in code)
  • 50:00: Technique 5: Runtime steering—Agent Control SDK demo (soft rules, self-correction)
  • 60:00: Production architecture: AWS Bedrock Agent Core (Gateway, Policies, DynamoDB, CloudWatch)
  • 65:00: Deployment options: CDK vs. notebook, repository walkthrough, AWS credits, wrap-up

Source/Metadata

  • Title: Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS
  • Transcript words: 12862
  • Duration seconds: 3319
  • Timestamp note: Video duration 3319 seconds (~55 minutes); timestamps inferred from demo progression and section breaks (not verbatim from transcript)
Full transcript 6962 words · 57 min read
0:00

SPEAKER_00

Hi, today we are going to talk about how to stop AI agents hallucinations with five techniques beyond the problem. Each one is a code change, not a problem change. Let's see. Every time your AI agent responds, you are paying for the words going in and the words coming out. In your bill, you will see those calling tokens. And the more tokens you're sent in, the more you pay. And if what you send is not quite right too much or missing something important, your AI agents start to hallucinate. There are five techniques to help reduce tokens waste, improve accuracy, and catch failure before you succeed them. And each one is a code change, not a problem change at all. So let's see it. First, we have semantic tool selection. You filter which tool can go into context on every call. The model only sees what it needs for that specific query. Second, we have graph rack for persist queries like aggregation, cons, multi-hop reasoning. And you replace the text retrieval with a structural graph query. The model gets a compute of verifiable answer node, a sample. The model is a tool. As RAC do. Three, multi-agent validation. A second agent can check every response before it rejects the users. And four, Neuro Symbolic Guardians. You rule life in Python, not in the pro, one, and the model cannot skip 10. So five, runtime guardians, because if you don't want to block, you can steer and you don't need to block everything. And when a rule fires, the agent self correct and complete the task. No hard stop, no user tries. So for each technique, I will show you the agent without it, and then with it. So you can compare. And all the demos, I'm using a travel agent that I build using strands again. And strands again is an open source agent framework that we maintain on AWS. And I am Elizabeth Fuentes Leone. I am a developer advocate for AWS. I'm focused on agentics application. And here in this QR code, you will find everything that you will need to recreate all these techniques that I'm going to show you in a moment. So let's get into it.

0:09

SPEAKER_00

So semantic tool selection or travel agent has 29 tools. Flies, hotels, payments, weather, cancellations, all are dummy tools. I know like a travel agent for real. But every time that user sends a message, all the 29 tools, all the 29 tools, descriptions go into the context windows. The model treats all of them before deciding what to do. And if your agent has memory, the growth to every conversation adds more context to the context that gets sent with every single message. And you pay for every single one of those tokens. Whether the model ends up using the tool or not. To understand where those tokens come from, you need to see what a tool actually looks like to the model. In strands, you write a function with the tool decorator. That's the tool. And a name, a description, and a doc street type parameters. Then strands take that and generate a schema schema with name description parameters. And the schema is what goes into the context windows on every call. Each tool schema is about 17 or 200 tokens depending of how many parameters it has. If our travel agent has 29 tools, the data is about 30,000 tokens per call. And the data is about 30,000 tokens per call. Just for the tool description. Before your message, before the response, every single call.

0:19

SPEAKER_00

By creating a tool database, we can filter the tools that the AMA needs before the AMA is invoked. So, with this filter, the model sees only three most relevant tools. Tokens users drops from thousands to fever.300. Let me show you in the code. Here in my keyword. This is the let me clear all the output.

0:36

SPEAKER_00

So, first, we need to install the requirements. Here is the requirement file. This is a Jupyter notebook because it's more simple to show everything. But it is an application that you can run if it is more comfortable for you. So, here in the requirements, I have the strands and again, and because I'm using OpenAI as a model invocation, I'm using the API from OpenAI and I need a strands agent from OpenAI. You can use a strands agent with Aura Armagot, almost all the model provider and with a strand, with Amazon better, of course, because we as AWS, we are the maintaining of this framework. And to use Amazon better, you don't need to add the model provider. And because I need to embedding, I need to create embeddings for my vector tools database. I'm using the sentence transformer. This is a super simple model that runs locally. So, it's free. So, you want, you can run this locally in your computer without spending another model embedding. And you can use strands with Olama too. So, if you have a local model, you can run everything in this new computer for free without spending any tokens. And I'm using files as my vector store, local, super simple. And, well, right now, I'm using Neo4j. This is for other demo that I'm going to show you in a moment. And because I'm using some environments, I'm using Python. And let's see. So, here, I'm going to run this. Let me put it bigger. I know that you have some problems. Bigger. Little male, little ass. And here, this is better. So, I installed my requirement. I already did that. And I'm using OpenAI. So, I need my API key. API key. And here, I invoke my strands. I need the agent because I want to build an agent. I'm using OpenAI. And I have a bunch of tool, dummy tools here. I'm going to show you in a moment. And I'm going to use this because I have some functions that I'm using to create the Builder index for my vector store. And I need the search tools. When I use the vector store, I put my query there. I search for the tool that I'm going to use using vector, search for vectors. And I swap the tools. I'm going to show you that in a moment. So, let me show you all my dummy tools. Where is the engaged and register here. Here, I have all my tools. This is for swap, search, build index. And where are my tools? Here. So, here are all the dummy tools I created for this. This is a demo, please. This is not something that you can use to go to production. No, please. And okay. So, what is this? This one. So, I run this. And I build my semantic index. I already did that. But yeah. Yeah. So, I have 29 tools here. 29 dummy tools. And I'm going to test this with a lot of different queries. And this is the ground truth. I have a ground truth because I know what is the best tool to answer the question. So, we can know if our engine is okay or not okay. So, let's run this. We have 90 queries and 29 tools. So, this is some helper functions. I don't care. Yeah. I care, but I don't want to explain you that. So, this is my engine. My traditional engine where I'm putting all my 29 tools inside the engine. I have my

0:42

SPEAKER_00

This one. So I run this and I build my semantic index. I already did that. But yeah, yeah. So I have 29 tools here. 29 dummy tools. And I'm going to test this with a lot of different queries. And this is the ground truth. I have a ground truth because I know what is the best tool to answer the question. So we can know if our engine is okay or not okay. So let's run this. We have 90 queries and 29 tools. So this is some helper functions. I don't care. Yeah, I care, but I don't want to explain that. So this is my engine. My traditional engine where I'm putting all my 29 tools inside the engine. I have my model. Some helper functions. This is my engine. I only, this is the way I create an engine using a strength engine. I put all my tools. And my system prompt is you are a travel assistant. Use the correct tool to answer questions. Oh my God. Super. And the model. Let's run this. And this is whatever. This is a function to count the tokens because inside this transaction, you can count the tokens. You can know how many tokens are using for in and out the agent. This is some helper function. And yeah, this is every time that I put a prompt inside this engine, I'm only sending the prompt and the agent responding. I use a new engine because it's a four. So I'm not having a conversation with the agent. I only said a question and I receive an answer. So for each question I spend around 2000 tokens. So and not always it give me the right answer here. So yes, I have, I'm not so good accuracy and the average is 1000 tokens. Now let's use my new engine with the semantic approach. The thing, the first thing that I do is for every single query, for every single prompt, I send the query to my search tools and I'm going to receive the tool queue, the two key three, most relevant, with more probability to response my query because this is a semantic search inside my vector store. And then I use that response. I go to select that name and I go to put that tools inside my agent. So I only going to use the three tools that this semantic search retrieve me for my agent and I send that only tools and this is my function helper and that's it. So I am going to send the all the queries again, the same. And for the first question, I have 4000 tokens, blah, blah, blah, blah. And yeah, we have a huge difference because I don't send it all the 29 tools in all the queries. So yeah, pam, pam. And okay. Semantic memory. Yeah. Okay. No, come on. Finish. So, well, let's go to the next one. So this agent is only sending a question and I receive a response. I don't have any conversation with this agent. What happens if I start to have conversations? If the agent remembers me? So we need to send only the tools that this agent is going to use because if I give the all the tools I'm using in the conversation history, then I, it's going to be in a moment that I'm going to have the 29 tools inside my engine. So we are not resolving, we are not having our, we are not resolving our problems in the agent when we have in a conversation with the agent. So we need to put the tools and then we need to remove the tools and we can do that with the swap tools and let's run this. And what is doing swap tools? I run this. No, let me show you first the swap tools function here, the register. I delete this because with a strands in each invocation, you have complete control of the status because this is an agent loop and you can use this agent loop. How? Because this is an agent loop and you can do almost whatever you want. You can put and remove everything inside the loop with some few lines of code. So let's see here we have the agent, we have the tool register, true, true tool registers. It's something inside the agent status. So we can clear the tools and that's it. So in the next invocation, we can clear the old tools and we can add the new tools. Here, where is my, I have a lot of things here. So I run this. It's going to take a moment.

0:52

SPEAKER_00

And yes, we can see that in each invocation, my amount of token is increasing. Why I have more? Because I'm sending the chat history as well. So I send the tools that I need only the tools that I need and the chat history. So that's why I can see that my amount of token is bigger and bigger. And let's see, the accuracy is better. And yeah, we have more tokens.

1:15

SPEAKER_00

So yes, probably we don't have the best accuracy here because this is a super demo with super dummy tools. And some queries are ambiguous on purpose because search for something, book for something, check something. And the demo has generic tools, dummy tools, with some similar name. And when all the 29 tools are visible, the model sometimes pick the wrong one. And with the filtering, those generic tools only appear even the query actually match them. So you can run everything and you can test this. So this is only the wrong only locally. But what happens when you want to put this demo in production? You can, of course, you can build a bigger vector store. I don't know, you can use PostgreSQL. I think it's too much. But we on AWS, we have Amazon Bedrock Agent Core. So Amazon Bedrock Agent Core is a service dedicated only to agents in production. And inside of Agent Core, we have Agent Core Gateway. So Agent Core Gateway allows you to build this index alone. So you only have to say, hey, this is my tools and Agent Core Gateway is going to build everything for you. And it can have the vector search inside. So the error layer is inside the hidden core and it handles the tool selection automatically. So you register your tools once and it finds the right one for each request. It's the same principle, but without infrastructure to manage. Now let's see the next one. GraphRack. You know, Rack is retrieval augmented generation. It's how agents access your own data. You take the user questions, search your documents for the most similar content using vector search and pass what you find to the model. The model answer from that it's worthwhile from open question finding something about this topic. But there is a category of question where that breaks down. What is the average rating across all hotels in Paris? How many hotels have a pool? Vector search always returns something even when nothing is truly relevant and the agent only sees the top end chunks of your old data at a time. It cannot aggregate count or traverse relationship across all the full data set. So it estimates and it present that estimates as a real fact, you know, a real answer. So here we have Rack. You build the vector store and it retrieves three chunks from 300 documents and the model guesses.

1:24

SPEAKER_00

The most similar content using vector search and pass what you find to the model. The model answers from [SPEAKER_00] that it's worthwhile from open question finding something about this topic. But there is a category [SPEAKER_00] of question where that breaks down. What is the average rating across all hotels in Paris? How many hotels [SPEAKER_00] have a pool? Vector search always returns something even when nothing is truly relevant and the agent only sees [SPEAKER_00] the top end chunks of your old data at a time. It cannot aggregate count or traverse relationship across all the [SPEAKER_00] full data set. So it estimates and it presents that estimates as a real fact, a real answer. So here we have RAG. You build the vector store and it retrieves three chunks from 300 documents and the model guesses.

1:34

SPEAKER_00

If we are using GraphRAG, you can run a query across all the data and returns a computed results. GraphRAG addresses this differently because instead of retrieving text chunks, you can build a knowledge graph from the documents that knows relationship structured data. For the demo, I'm going to use Neo4j locally and the model is going to write a Cypher query to search it. Cypher query is Neo4j query language. It's similar to SQL. So the graph runs that query across all the data. The model gets back a computed verified result, not a sample. And before running this demo, I'm going to install the dependency. Let me show you. Let's go to the code.

1:43

SPEAKER_00

So graph, this is the notebook that I'm going to share with you. So we need to install the requirements again. What we have here, let me see. This is the first one. Yeah. Requirements. I'm going to use OpenAI again. Let me just close this. I'm going to use Neo4j and we need Neo4j GraphRAG. And for the agent that is using normal RAG, we are going to build a super simple vector store in files. And we are going to use again the sentence transformer. And let's see. I already have all. I invoke my OpenAI. This is my tools. Because I want to create some tools here. I'm going to use OpenAI and I'm using graph database. And this is why I have my Neo4j locally. And let's run this. And this is something to check Neo4j.

1:54

SPEAKER_00

So I'm building my vector store. This is my files. And this is my Neo4j that I have here locally. I don't know if I can show you. I think. And I have a tool, a normal tool created with the decorator. This is to search inside my vector store. You see, it's not super. I have the query. I create the embedding for my query. And then I search inside the vector store and I have the query for the knowledge graph.

2:04

SPEAKER_00

So here the model needs to understand that to search in the query, in the graph knowledge base, you need to build a Cypher query. So I put that in the context because we already know how to build tools. I put that in the context. I have my driver to send the data to connect with the vector store and to send this Cypher query that the model is going to create for me. And this is to read the results. And that's it. That's the only tool that I need to search in my knowledge graph.

2:14

SPEAKER_00

Okay. This is my model. OpenAI. This is my RAG, my RAG agent. And this is my graph agent. I have two different agents to compare the results. Let's see. Yeah. I'm ready. So let's do this. Run the first test. Aggregation.

2:23

SPEAKER_00

Here. What is the question? What is the average guest rating across all hotels in Paris? So meanwhile, this agent is already done. So what I have here, I have the average guest rating across the hotels listed in Paris is it calculates when I'm using RAG, you go to the vector store, it receives the possible answers for this question. Because this is an aggregation, it's going to use the data that it received to build a mathematical operation. This is the thing that I have here. So everything that we see here is the model reasoning.

2:34

SPEAKER_00

So when I run the graph agent, the agent is only going to give me the answer. Only these tokens here. Because the Cypher query can give me the mathematical operation itself. It can do that. So we don't need the agents, the LLMs, the model, to do that for me. Because the Cypher query already gave me the right question. So it is 47 and the average. So here in this first we are lucky. Because probably they are only two hotels. But what happens if this vector store has more than two or three hotels in the vector store? So we are not going to give a real answer. It's going to the LLM is going to calculate with the only three answers that it received. So it's going to build this operation with only three hotels. But if the vector store has more than three, we are going to have some problems. Something that. So let's run the next one. The precise count.

2:44

SPEAKER_00

Something similar to how many hotels have a swimming pool as an amenity. How many again? So it gives me the traditional, it appears that the search did not return any specific information about hotels in Paris. Okay. What happened with the other one? There are currently no hotels that offer swimming pools. No. There are no hotels. So the other one is, would you like to ask about hotels with other specific amenity or information? It's, maybe it is, or maybe I don't know. Yeah. It's no sort of cure. Okay.

2:55

SPEAKER_00

So multi-correlation. What is my question? What are the room types and price for the highest rating hotels? What are the room types and price? These two questions. So it serves the fact that I have only one. The highest rating hotel in any company, Paris. With the guest rating is that one. However, I currently don't have. Okay. A lot of data there.

3:04

SPEAKER_00

So what happened with the other one? Receiving notification for. I have an error. Here, in the second one, I received the question. The highest rating hotel is Harmony. They offer the following times. But unfortunately, these rooms are not available. So I receive an answer. I don't receive a lot of there. I only receive the answer what I need. And for the next one, auto-domain detention.

3:15

SPEAKER_00

So, let's see. What is the question? Tell me about hotels in Antarctica. A spoiler. There is no hotel in Antarctica. There are zero hotels in Antarctica. And let's see. It appears that there's certain or returning specific information about hotels. Okay, because they know. As such currently, they don't have the hotel. If you are looking for a particular experience or specific or inquire about business in Antarctica, no. Please let me know and I can assist you for it. Okay. A lot of tokens that is pending there in that answer for LLM. What happened with the other one? There are currently no hotels listed in Antarctica. Of course, because it created the Cypher query, the Cypher query, and it received zero.

3:31

SPEAKER_00

So, let's see. What is the question? Tell me about hotels in Antarctica. A spoiler. There is no hotel in Antarctica. There is zero hotels in Antarctica. And let's see. It appears that there's certain or returning specific information about hotels. Okay, because they know. Ah, as such currently, they don't have the hotel. If you are looking for a particular experience or specific or inquire about business in Antarctica, no. Please let me know and I can assist you for it. Okay. A lot of tokens that is pending there in that answer for LLN. What happened with the other one?

4:15

SPEAKER_00

They are currently no hotel list in Antarctica. Of course, because it created the cyber query, the cyber query, and it received zero. So, it gave me an honest answer. Here, some summary that I gave for me for this notebook.

4:33

SPEAKER_00

But here is something that I want to show you before I go to the other one. So, here, something that I love for Neo4j. Why? Because I'm using Neo4j? Because in the library, the Neo4j gives me, it can build, it uses a LLM. I'm using the OpenAI as well to build a knowledge graph. So, how I build my knowledge graph, I only have a bunch of data, a bunch of txt, only text, and I send that data to the Neo4j library. Here, I send that data to this. And Neo4j, using this, all this library, just right here, it can understand all my data, and it can build the graph for me. So, I don't need to create that using the simple knowledge graph pipeline inside the knowledge graph library. So, that's why I'm using Neo4j. It's amazing. It's super simple to use, so I'm inviting to use it.

4:44

SPEAKER_00

So, let's go to the next one.

4:56

SPEAKER_00

So, multi-agent validation. Sometimes an agent fails and nobody finds out. It calls a tool, and the tool returns an error, and the agent does not surface that error. It generates a confident success response instead. The user thinks it worked, you think it worked, it didn't. The agent acts and validates its own output in the same loop. There's no separation, no second opinion. So, when something goes wrong, it rationalizes and tells you it's okay, it worked. Here is what happens. Inside a single agent, when it fails, it calls the tool, gets an error, relates it and returns a success response. The user never sees the error. You can address this by adding a validation layer. You can have three agents in sequence. One acts, one checks and other approve or rejects. A Strands agent has a built-in class for this called SWAM. It manages the handoff between agents automatically. You just define the role of each agent in a system prompt. Let me show you that. So, this demo only needs a Strands agent with OpenAI integration.

5:07

SPEAKER_00

Let's see the requirements here. We only need a Strands agent. We don't need anything more. Let's go to the notebook here. Let's go to the notebook here. And the key important thing here is the, I already wrote this. Let me see here. It's the SWAM. The SWAM is what lets you to connect multiple agents together without creating a for or while manually to put all these agents together. It can build a chain and manage the handoff between them automatically. So, here we are going to create three agents. Let me go that because we have the normal and we have some ground truth data. So, here is to create a single agent. So, we have, we are going to test three different scenarios to validate a booking. They said that we expect true. To available hotels, we expect false. And no existing hotels and missing booking, we expect false too. This is how we created single agent. We did a prompt. Some tools. And we are using the tools that we have here. And here, we know that this is the data. So, the agent can give us the answer that we are looking for. And that's it. Let me go to the other ones. Oh, wait here. So, how we build the SWAM? I already run this. I want to show you the SWAM. So, we are going to build three different agents. The executor, the validator, and the critic. We have a system prompt. You are an executor agent for a hotel booking system. Use the provided tools to fulfill requests accurately and so forth. The validator is you are a validator agent. Review what the executor did and output exactly on, off. And we have the critic. That is going to set, yeah, approve or no reject. Let's run this. And oh, what is here? Oh, model. Didn't define the model. Sorry. Sorry, please forgive me the live because I didn't run this one.

5:23

SPEAKER_00

Let's run. Because this is the normal and the model is there, right? What? Bit of bias. Run that. What's happened with you? Let me go this. Let me copy and paste. I don't know why this is giving me a problem.

5:28

SPEAKER_00

Multi-agent SWAM. Is it the same name? Come on. Yeah, you can see this is live. I don't go into edit this. So, yes. Okay, we have. Thank you. And the SWAM. Look how we create the SWAM. We have SWAM. And the SWAM function. We put all the agents together. And the entry point is the executor. So, all everything is going to start in the executor and then it's going to hand off. And six times because you can handle that too. So, get the SWAM final response. Let's do a four and the same all the response. And book the grand hotels. I have booked a grand hotel for Alice tonight. Booking blah blah blah. Enough. Valid. Okay. Let's go.

5:37

SPEAKER_00

Verdict approved. So, the critic is approving this. This is the other one. And we can run some comparison here if you want. So, here we can see that the SWAM handles the flow between them. Between all the three different agents. It was what happens when the single agent tries to confirm something that does not exist in the system. And now the same request through the SWAM. The executor gets the error. The validator catches the critic rejects. And the user never sees a fabricant response. So, you can, we can test here with the single agent that it's a suspected on no entry and return success. And the SWAM, it have an executor that got the error, validates, say, hey, come on, man.

5:47

SPEAKER_00

Something is happening now let's go to the next one okay neurosymbolic guardians you have a rule for your agent let's say maximum 10 guesses per reservation you write in the system prompt you even write in the tool description and in the end still calls the tool with fixed in no because it is ignoring you because prompts probably are suggestion no constraints the model processes them as a text not as a logic it has to execute it's probabilistic only code execute logics a rule in the problem the model reads it as a suggestion a rule in the code the model cannot skip it neurosymbolic guardians rules put the rules in the code Strands and has a future column hooks the function that Strands calls automatically a specific moment in the agent loop it is this case

5:55

SPEAKER_00

Something is happening now, let's go to the next one. Okay, neurosymbolic guardians. You have a rule for your agent. Let's say maximum 10 guesses per reservation. You write it in the system prompt. You even write it in the tool description. And in the end, it still calls the tool with fixed parameters. No, because it is ignoring you. Because prompts probably are suggestions, not constraints. The model processes them as text, not as logic. It has to execute it. It's probabilistic. Only code executes logic. A rule in the problem, the model reads it as a suggestion. A rule in the code, the model cannot skip it. Neurosymbolic guardians rules: put the rules in the code. Strands has a future column hooks. The function that strands calls automatically at a specific moment in the agent loop. In this case, right before a tool executes. So you have a hook. You write the rule, check the parameters, and if they fail, you cancel the code. Let me show you this in the code. Here we have the Jupyter notebook. You have the requirements here. The same requirements as the previous demos. We only need strands and OpenAI. We don't need anything else. And my API. And here is the important thing in this demo: these three hooks provided in this Strands base class. So us that allows you to create the hook. So we have who provided, who register, and before to call event. And the tool registry. Who register is what Strands passes you to register your callback. And before to call event is the event that fires every time the model is about to execute a tool. That last one is what makes it possible to intercept the call before it runs. So you have the before two calls even. And of course we have after two calls even that we are going to use here. And let's see the rule in the problem. Remember the model reads it as text and it may follow or not follow that. So let's see this. I'm going to run this. So here we have some simulated state that we are going to use in this agent. And let me. This is the symbolic rule. So booking rule. If if rule booking rule. This is something that we create. Let me show you the rules here. So we have the rules. Rule one: validate date check is most before checkout. What is validated? They check out. This is something that you check. This is something that is going to invoke when the rule is, is, is most have to use it. And you check it out. We have another rule for max guesses: maximum tech guesses per booking. So if I want to booking, I don't know if I want to do a booking for 11, it's going to block me. It's going to reject the booking. And we have some confirmation rule: payment before confirm. You know, you can't confirm if you don't have the payment. Cancellation rule: cancellation window cannot cancel within 48 hours of checking. So this is something that I'm creating here for this demo. So I have booking rules and I have confirmation rules. I'm not using the cancellation rule here. So create validation hook. We create the validation hook. Here is right like this neurosymbolic rule. And we use the hook provider. So we have a bunch of code here. Then we add the booking rules and the confirmation rules in the state. And you can check this by yourself a little bit. But we want to see the demo running. So then we define the cleaning tools that we are going to use them for booking hotels and the process payment. This is the normal tools. And we need to add the tools to the hooks. I think I have that. Yeah. Tool name book hotel. So this thing is going to get here when the book hotel is using. Yeah, the book hotel. Okay, so let's create the agents for the comparison. So we are going as the other demos. We have three different scenarios: configure booking without payment. Payment must be verified before the confirmation. This is the rule that must trigger this. We have booking hotel exceeding guest limit. And we have valid booking for five guesses. So we have the three scenarios. And we have the normal agent. The baseline agent. And we have the agent with the neuro. The neurosymbolic guardians. So we have the hook. The neurosymbolic hook. Oh, this is the same here. And we add this. If you can see the normal. And here is only three lines: the tools, the model, and this is the line that gives us the difference between these two agents. So I'm going to run that now. Let's go to confirm a booking without payment. So I'm running this. My agents. The first one is obviously come on man. Yeah, because he confirmed the booking without payment because it doesn't have any rules. And the problem is super basic. And the other one is. Let's open this. And is booking blocked. It payment must be verified before the confirmation. Thank you. So it is okay. Now let's run the other one. Test second scenario. What is the question? Booking hotel excess guest limit. So the question is. Oh yeah, I have more than. So here, the hotel has been successful booking for. Yeah. Because the problem is basic and I didn't put in any rule there. And for the other one, it seems that the Grand Hotel has maximum capacity of 10 guests per booking. Additional booking must be made a last one day in advance. I don't remember the day. Probably this is something that I gave them. And to validate booking is both again execute or rule pass because it is only a validation for booking. So book hotel for blah blah or rule passes. Yes, we see that the booking is still. Needs to make. Ah, I don't remember the question. So booking is. May provide. Ah, yeah, okay. Okay, I get it. And okay, now run all that in this question. Confirm it without payment. Bookings max. And what happened here? This is all the scenarios that we just run. So this is something to compare the results. So valid booking. Can see here a five guesses allowance. Correct. Wrong. 50. Block it. Wrong. Correct. Confirm. So we have here the comparison between the two ends. The two ends. So what we have. What we have here is same model, same tools, same prompt, and the different outcome because the rules are in Python. Knowing the pro, this pattern of forcing rules in code before the. The rules are in code is also what Amazon. Is also what Amazon Agent Core Policies Service that we have does at the infrastructure level. So the same concept but money for you in production. And you only have to create the rules. So the hooks are all or nothing. They block everything or approve. But sometimes you want the agent to adjust the rules and keep going. No stop. I leave the user waiting. That is what I will show you in the next room time guardians. Room time. Still, still, still. Don't block. This is the next one. Hooks blocks unconditionally. The agent stops and the user has to retire for a hard constraint. That is exactly what you want. But sometimes the rule is soft. Maybe room fits for four guests but a group of six could book two different rooms or a flight is full but there is availability on the next one. You don't want to block everything. Probably you want the agent to find an option and complete the task. That is steering here. We have the hook that fires and the task fails. And with agent control, the other different here is operational. Because with the hooks, it's as changing a rule means changing code and redeploy the old harness. All the alien. And with the agent control, which is the name of the open source library that we are going to use here, the rules are registered on a local server via API. You update them without touching the code because the alien picks them up immediately. Let me show you that here in the code. Yeah, this is the notebook in this demo. We only need an extra package: the agent control SDK. So the agent control is the one that helps us to create the steering. And here we are using the setup control.

6:05

SPEAKER_00

And complete the task that is a steering here we have the hook that fires and the task fail and with agent control the other difference here is operational because with the hooks it's changing a rule means changing code and redeploy the old harness all the alien and with the agent control which is the name of the open source library that we are going to use here the rules are registered on a local server via API you update them without touching the AM code because the alien picks them up immediately let me show you that here in the code yeah this is the notebook in this demo we only need

6:41

SPEAKER_00

an extra package the agent control SDK so the AM control is the one that helps us to create the steering and here we are using the setup control this is a little application a little app that I created with the local server and the steering rules here we have the local server we have the control there are the steering rules first we have the the steering must guess steer you know guide so guide agent to reduce guest count when exceeding maximum of 10 the same and it's going to steer and it's going to steer and we have some control that deny for example deny no payment blocking confirmation without peer payment and here well this is to create the

7:29

SPEAKER_00

the service the server and let's go to the notebook so I already let me go here yeah so I need the environment so this is the agent I'm going to have an error here wait wait wait I don't want to use metro wait wait no other let me comment this one because it's going to give me an error and let's run this yeah so we have the hooks as within the demo so we have book any company in Lisbon and we have some prompts you are a hotel booking assistant when booking first describe while you will book in blah blah blah and this is the problem for my agent where is my agent agent control this is for the

8:33

SPEAKER_00

well this is on song helper and this is my hook that you already know because we created this so we have the system pro and we have the hooks so let's test this again with the book any company Lisbon for 50 guests and if you remember this only can book for less than 10 guests and of course is blocking now let's go to the agent control new agent here we have two key imports that are important for the agent control SDK so we have the agent control plugin that captures agent events and sends them to the agent control server and we have the agent control student handle that listen for a student decision for the server

9:23

SPEAKER_00

and delivers them back to the model together here they are what connects strands to the agent control steering logic so let's go this with this so the steering agent is going to a real book a room for any company Lisbon 50 guests let's see what it's doing yeah I have successful booking for stay in a company Lisbon for 50 guests so it's the reservation have been split into two rooms so it took that it itself one room for one and other room for five and that's it so use hook for hard constraints agent control for software rule so now we have five techniques or running locally but how do you take this to production without maintaining

10:13

SPEAKER_00

service without building infrastructure let me show you how everything that I just built here in the previous demos runs locally for the production version amazon bedrock amazon bedrock agent code give you the runtime a gateway a short term and a long-term memory a CloudWatch observability built in no service to manage here the architecture this transient runs inside the runtime and inside the runtime you can put every framework that you want the only error is trans the gateway rules tools calls to the lambda function automatically as a tools and the steering rules from the previous demo live in the DynamoDB so you change

11:06

SPEAKER_00

them there and they are live on the next call and you don't need to redeploy anything and if you want to use Neo4j of course you can use the info J or a dev that is external graph database we also have a free tier and the code is in the repo here it's in the repo and you will need AWS credentials if you are using amazon bedrock agent code and in the resource link there is some credits I hope that you can find it because I always try to give away some credit for AWS so you can deploy everything for free and if you are new here in AWS I have a repository that you can use to deploy

11:49

SPEAKER_00

everything over this architecture using a notebook too but if you are familiar with CDK cloud developer kit you can do it as well to deploy everything at once both options are in the repository of core and you can go deeper in agent core with all the documentation that I left there and you can do it as well to deploy everything over here so please don't stop at the demo and try going to production with amazon bedrock agent core so let me bring it all back tokens weigh on every request you can fix it with semantic tool selection and confident answer it's never compute possible possible possible possible because you are asking how many so you can use refrag and query the

12:47

SPEAKER_00

data don't sample it sometimes we have fabricated success confirmation you can use multi-aging validation and second path sketched it and rules the models are quite a little skip it you can use neuro symbolic guardians to enforce in the code don't trust in the problem and for the hard blocks that stop the user you can use runtime steer self-correct and finish each demo that I just showed you is in the repository as a notebook and an application as well you can use the demo and then go to the demo 5 and if you want deploy in the productions so have you tried any of this in your own hands thank you to join me in this session and happy building

13:41

SPEAKER_00

we have in a conversation with the agent. So we need to put the tools and then we need to remove the tools and we can do that with the swap tools and let's run this. And what is doing swap tools? I run this. No, let me, let me show you first the swap tools function here, the register. I delete this because with a strands in each invocation, you have complete control of the status because this is a agent loop and you can use this agent loop. How? Because this is an agent loop and you can do almost whatever you want. You can put and remove everything inside the loop with some few lines of code. So let's

14:40

SPEAKER_00

see here we have the agent, we have the tool register, true, true tool registers. It's something inside the agent status. So we can clear the tools and that's it. So in the next invocation, we can clear the old tools and we can add the new tools. Here, where is my, I have a lot of things here. So I run this.

15:15

SPEAKER_00

It's going to take a moment.

15:28

SPEAKER_00

And yes, we can see that in each invocation, my amount of token is increasing. Why I have more? Because I'm sending the chat history as well. So I send the tools that I need only the tools that I need and the chat history. So that's why I can see that my amount of token is bigger and bigger.

15:59

SPEAKER_00

And let's see, the accuracy is better. And yeah, we have more tokens.

16:08

SPEAKER_00

So yes, probably we don't have the best accuracy here because this is a super demo with super dummy tools. And some queries are ambiguous on purpose because search for something, book for something, check something. And the demo has a generic tools, dummy tools, with some similar name. And when all the 29 tools are visible, the model sometimes pick the wrong one. And with the filtering, those generic tools only appears even the query actually match them. So you can run everything and you can test this. So this is only the wrong only locally. But what happens when you want to put this demo in production?

17:00

SPEAKER_00

You can, of course, you can build a bigger vector store. I don't know, you can use PostgreSQL. I think it's too much. But we on AWS, we have Amazon Betro Agent Core. So Amazon Betro Agent Core is a service dedicated only to agents in production. And inside of Agent Core, we have Agent Core Gateway. So Agent Core the gateway allows you to build this index alone. So you only have to say, hey, this is my tools and Agent Core Gateway is going to build everything for you. And it can have the vector search inside. So the error layer is inside the hidden core and it handles the tool selection automatically. So you

17:54

SPEAKER_00

register your tools once and it finds the right one for each request. It's the same principle, but without infrastructure to manage. Now let's see the next one. GraphRack. You know, Rack is retrieval argument generation. It's how agents access your own data. You take the user questions, search your documents for the most similar content using vector search and pass what you find to the model. The model answer from that it's worthwhile from open question finding something about this topic. But there is a category of question where that breaks down. What is the average rating across all hotels in Paris? How many hotels

18:50

SPEAKER_00

have a pool? Vector search always returns something even when nothing is truly relevant and the agent only sees the top end chunks of your old data at a time. It cannot aggregate count or traverse relationship across all the full data set. So it estimates and it present that estimates as a real fact, you know, a real answer. So here we have Rack. You build the vector store and it retrieves three chunks from 300 documents and the model guesses. If we are using GraphRack, you can run a query across all the data and returns a compute results. GraphRack addresses this differently because instead of retrieving text chunk, you can build a knowledge graph from the documents,

19:55

SPEAKER_00

knows relationship structured data. For the demo, I'm going to use Neo4j locally and the model is going to write a Cyber query to search it. Cyber query is Neo4j query language. It's similar to SQL. So the graph runs that query across all the data. The model gets back a compute verified results, not a sample. And before running this demo, I'm going to install the dependency. Let me show you. Let's go to the code.

20:37

SPEAKER_00

So graph, this is the notebook that I'm going to share with you. So we need to install the requirements again. What we have here, let me see. This is the first one. Yeah. Requirements. I'm going to use OpenAI again. Let me just close this. I'm going to use Neo4j and we need Neo4j GraphRack. And for the agent that is using normal Rack, we are going to build a super simple vector store in files. And we are going to use again the sentence transformer. And let's see. I already have all. I invoke my OpenAI. This is my tools. I, I, because I want to create some tools here. I'm going to use OpenAI and I'm using graph database.

21:30

SPEAKER_00

And this is why I have my Neo4j locally. And let's run this. And this is something to check Neo4j. So I'm building my, uh, vector store. This is my files. And this is my Neo4j that I have here locally. I don't know if I can show you. I think. And I have a tool, a normal tool created with the decorator. This is to search inside the, my vector store. You see, it's not super. I have the query. I create, I create, I create the, the bedding for my query. And then I search inside the vector store and I have the query for the knowledge graph. So here the model need to understand that to search in the query, in the graph knowledge base, you need to build a cyber query.

22:20

SPEAKER_00

So I put that in the context because we already know how to build a tools, right? I put that in the context. I have my driver to send the data to, uh, to connect with the vector store and to send this hyper query that the model is going to create for me. And this is to read the results. And that's it. That's the only tool that I need to search in my knowledge graph. Okay. This is my model. Open AI. This is my rock, my rock agent. And this is my graph agent. I have two different agent to compare the results. Let's see. Yeah. I'm ready. So let's do this. Run the first test. Aggregation.

23:10

SPEAKER_00

Here. What is the question? What is the average guest rating across all hotels in Paris? So meanwhile, this agent is, is, oh, it's already done. So what I have here, I have the average guest rating across the hotels listed in Paris is, um, it calculates, you know, when something, when I'm using a rock, you go to the, to the vector store, it receives the end, uh, possible answers for this question. Because this is an aggregation, it's going to use the data that it received to build a mathematical, uh, operation. This is the thing that I have here. So everything that we see here is the, is the model reasoning.

23:55

SPEAKER_00

So when I run the graph agent, oh, let me, the agent is only going to give me the answer. You know, only these tokens here. Because the cyber query, it can give me the mathematical operation itself. It can do that. So we don't need the agents, the LLMs, the model, they do that for me. Because the cyber query already gave me the right question. So it is 47 and the average. So here in this, um, first we are lucky. Because probably they are, yeah, only two hotels. But what happened if this vector store have more than two or three hotels in the, in the vector store?

24:44

SPEAKER_00

So we are not going to give a real answer. It's going to, the LLM is going to calculate with the only three answer that it received. So it's going to build this operation with only three hotels. But if the vector store have more than three, we are going to have some problems. Something that, um, and I, um, yeah. So let's run the next one. The precise county. Something similar to how many hotels have a swimming pool as a amenity. How many again? So it give me the, the traditional, it's appear that the search did not return to any specific information about hotels in Paris. Okay. What happened with the other one? They are currently no hotel that offers swimming pools.

25:37

SPEAKER_00

No high. There is no hotels. So the other one is like, hmm, would you like to ask about hotels with other specific amenity or information? It's like, hmm, maybe it is, or maybe I don't know. Yeah. It's no sort of cure. Okay. So multi-coberation. Oh, what is my question? What are the room types and price for the highest rating hotels? What are the rooms types and price? These two questions. So it serves the fact that I have only one. The highest rating hotels in any company, Paris, blah, blah, blah. I don't speak French. With the guest rating is that one. However, I currently don't have, okay, blah, blah, blah. A lot of data there.

26:27

SPEAKER_00

So what happened with the other one? Receiving notification for, oh, I have an error. Here, in the second one, I received the question. The highest rating hotel is harmony, blah, blah, blah. They offer following times. But, unfortunately, these rooms are not available. So I receive an answer. I don't receive a lot of blah, blah, blah there. I only receive the answer what I need. And for the next one, auto-domain detention. So, let's see. What is the question? Tell me about hotels in Antarctica. A spoiler. There is no hotel in Antarctica. There is zero hotels in Antarctica.

27:09

SPEAKER_00

And let's see. It appears that there's certain or returning specific information about hotels. Okay, because they know. Ah, as such currently, they don't have the hotel, blah, blah, blah. If you are looking for a particular that experience or specific or inquire about business in Antarctica, no. Please let me know and I can assist you for it. Okay. A lot of tokens that is pending there in that answer for LLN. What happened with the other one? They are currently no hotel list in Antarctica. Of course, because it created the cyber query, the cyber query, and it received zero.

27:46

SPEAKER_00

So, it gave me an honest answer. Here, some summary that I gave for me for this, in this, in this, in this, in this notebook. But here is something that I want to show you before I go to the other one. So, here, something that I love for Neo4j. Why? Because why I'm using Neo4j? Because in the library, the Neo4j gives me, it can build, it uses a LLM. I'm using the OpenAI as well to build a knowledge graph. So, how I build my knowledge graph, I only have a bunch of data, a bunch of txt, only text, and I send that data to the Neo4j library that I miss it. Here, I send that data to this.

28:38

SPEAKER_00

And Neo4j, using this, all this library, just right here, it can understand all my data, and it can build the graph for me. So, I don't need to create that using the simple knowledge graph pipeline inside the knowledge graph library. So, that's why I'm using Neo4j. It's amazing. It's super simple to use, so I'm inviting to use it that. So, let's go to the next one. So, multi-agent validation. Sometimes an agent fails and nobody finds out. It calls a tool, and the tool returns an error, and the agent does not surface that error. It generates a confident success response instead. The user thinks it worked, you think it worked, it didn't. The agent acts and

29:39

SPEAKER_00

validate its own output in the same loop. There's no separation, no second opinion. So, when something goes wrong, it rationalizes and tells you it's okay, it worked. Here is what happens. Inside a single agent, when it fails, it calls the tool, gets an error, relates it and returns a success response. The user never sees the error. You can address his by adding a validation layer. You can have three agents in sequence. One's one's acts, one checks and other approve or rejects. A strands agent has a built-in class for this called SWAM. It manages the hand of between agents automatically. You just define the role of each

30:45

SPEAKER_00

agent in a system prompt. Let me show you that. So, this demo only needs a strands agent with open-in integration. Let's see the requirements here. We only need a strands agent. We don't need anything more. Let's go to the notebook here. Let's go to the notebook here. And the key important thing here is the, I already wrote this. Let me see here. It's the SWAM. The SWAM is what lets you to connect multiple agents together without create like a four or while manually to put all these agents together. It can build a shine and manage the handoff between them and automatically. So, here we are going to create three agents. Let me go that because we have the normal

31:44

SPEAKER_00

and we have some ground truth data. So, here is to create a single agent. So, we have, we are going to test three different scenarios to validate a booking. They said that we expect true. To available hotels, we expect false. And no existing hotels and missing booking, we expect false too. This is how we created single agent. We did a prompt. Some tools. And we are using the tools that we have here. And here, we know that this is the data. So, the agent can give us the answer that we are looking for. And that's it. Let me go to the other ones. Oh, wait here. So, how we build the SWAM? I already run this. I want

32:29

SPEAKER_00

to show you the SWAM. So, we are going to build three different agents. The executor, the validator, and the critic. We have a system prompt. You are an executor agent for a hotel booking system. Use the provide tools to fulfill requests accurately and blah, blah, blah. The validator is you are a validator agent. Review what the executor did and output exactly on, off. And we have the critics. That is going to set, yeah, approve or no reject. Let's run this. And oh, what is here? Oh, model. Didn't define the model. Sorry. Sorry, please forgive me the life because I didn't run this one.

33:17

SPEAKER_00

Let's run. Because this is the normal and the model is there, right? What? Bit of bias. Run that. What's happened with you? Let me go this. Let me copy and paste. I don't know why this is giving me a problem.

33:39

SPEAKER_00

Multi-aging SWAM. Is it the same name? Come on. Yeah, you can see this is life. I don't go into edit this. So, yes. Okay, we have. Thank you. And the SWAM. Look how we create the SWAM. We have SWAM. And the SWAM function. We put all the agents together. And the entry point is the executor. So, all everything is going to start in the executor and then it's going to hand off. And six times because you can handle that too. So, get the SWAM final response. Let's do a four and the same all the response. And boot the grand hotels. I have booked a grand hotel for Alice tonight. Booking blah blah blah. Enough. Valid. Okay. Let's go.

34:30

SPEAKER_00

Verdit approved. So, the critics is approving the this. This is the other one. And we can run some comparison here if you want. So, here we can see that the SWAM handles the flow between them. Between all the three different agents. It was what happens when the single agent tries to confirm something that does not exist in the system. And now the same request through the SWAM. The executor gets the error. The validator catches the critic rejects. And the user never sees a fabricant response. So, you can, we can test here with the single agent that it's a suspected on no entry and return success.

35:23

SPEAKER_00

And the SWAM, it have an executor that got the error, validates, say, hey, come on, man, alloc something is happening now let's go to the next one okay neurosymbolic guardians you have a rule for your agent let's say maximum 10 guesses per reservation you write in in the system prompt you even write in in the tool description and the end still calls the tool with fixed in no because it is ignoring you because prompts probably are suggestion no constraints the model

36:35

SPEAKER_00

process them as a text no as a logic it has to execute it's probabilistic only code execute logics a rule in the problem the model reads it as a suggestion a rule in the code the model cannot skip it neurosymbolic guardians rules put the rules in the code strands and has a future column hooks the the function that strands calls automatically a specific moment in the agent loop it is this case right before a tool executes so you have a hook you write the rule check the parameters and if they fails you cancel the code let me show you this in the code here we have the jupiter notebook you have

37:35

SPEAKER_00

the requirements here um what the same requirements that the previous demos we only need we start everything again we only need strands and open ai we don't need anything else and um my api and here is the thing important in this demo are these three hooks provided in this trans base class so us that allows you to create the hook so we have who provided who register and before to call event and the tool regi who register is what strands pass you to register your callback and before to call event is the event that finds every time the model is about to execute a tool that last one is what makes it possible they

38:31

SPEAKER_00

intercept the call before it runs so you have the before two calls even and of course we have after two calls even that we are going to use that here and let's see the rule in the problem remember the model reads as a text and it may follow or no follow that so let's see this i'm going to run this so here we have some simulated state that we are going to use in this agent and let me this is the symbolic rule so booking rule if if rule booking rule this is something that we create let me show you the rules here so we have the rules rule one validate date check is most before checkout what is validated they check out this is something that

39:26

SPEAKER_00

you check this is a something that is going to invoke when the rule is is uh is most have to use it and you check it out uh we have another rule for max guesses maximum tech guesses per booking so if i want to use i don't know if i want to booking uh i want to do a booking for 11 it's going to block me it's going to reject the booking and we have some confirmation rule payment before confirm you know you can confirm you don't have the payment cancellation rule cancellation window cannot cancel within 48 hours of checking so this is something that i'm great here you know to this demo so i have booking rules and i have confirmation rule i don't

40:21

SPEAKER_00

i'm using here the cancellation rule so create validation hook we create the validation who's here is right like this neuro symbolic rule and we use the hoop provider so we have a bunch of code here then we add the booking rules the confirmation rules in the state and you can check this by yourself a little but we want to see the demo running so that we define the cleaning tools that we are going to use them for the booking hotels and the process payment this is the the normal tools you know and we need to add the tools to the hooks i think i have that yeah tool name book hotel so this thing is going to get here when the book hotel is using the

41:18

SPEAKER_00

yeah the book hotel okay so let's create the agents for the comparison so we are going as the other demos we have three different scenarios configure booking without payment payment must be verified before the confirmation this is the rule that must trigger this we have booking hotel exceeding guest limit and we have valid booking for five guesses so yeah we have the three scenarios and we have the normal agent the normal agent the vaseline agent and we have the agent with the neuro the neuro symbolic guardians so we have the hook the neuro symbolic hook oh this is the same here and

42:07

SPEAKER_00

we add this if you can see the normal and here is only three lines the tools the model and this is the line that give us the difference between these two agents so i really run that now let's go to confirm a booking without payment so i'm running this my agents the first one is obviously come on man yeah because yeah he confirmed the booking without payment because it doesn't have any rules and the problem is super basic and the other one is let's open this and is booking block it payment must be verified before the confirmation thank you so it is okay now let's run the other one test second scenario

43:01

SPEAKER_00

what is the question booking hotel excess guest limited so the question is oh yeah i have more than so here uh the hotel have been successful booking for yeah because the problem is basic and i don't put in any rule there and for the other one it seems that the grand hotel has maximum capacity of 10 guests per booking additional booking must be made a last one day in advance i don't remember the day probably this is something that i give them and to validate booking is both again execute or rule pass because it is only a validation for booking so book hotel for blah blah or rule passes

43:45

SPEAKER_00

yes we see that the booking is still needs to make ah i don't i didn't remember the question so booking is may provide ah yeah okay okay i get it and okay now run all that in this question confirm it without payment bookings max and what happened here this is all the scenarios that we just run so this is something to compare the results so valid booking can see here a five guesses allowance correct wrong 50 block it wrong correct confirm so the we have here the comparison between the two ends the two ends so what we have what we have here is same model same tools same prompt and the different that

44:43

SPEAKER_00

outcome because the rules are in python knowing the pro this pattern of forcing rules in code before the the rules are in code is also what amazon is also what amazon agent core policies service that we have does at the infrastructure level so the same concept but money for you in production and you only have to create the rules so the hooks are all or nothing they block everything or approve but sometimes you want the agent to adjust the rules and keep going no stop i leave the user waiting that is what i will show you in the next room time guardians room time still still still don't block this is the next one hooks blocks unconditionally

45:43

SPEAKER_00

the agent stop and the user has to retire for a hard constraint that is exactly what you want but sometimes the rule is soft maybe room fits for for guests but a group of six could go book two different rooms or a fly is full but there is ability on the next one you do no you don't want to block everything probably you want the agent to find an option and complete the task that is a steering here we have the hook that fires and the task fail and with agent control the other different here is operational because with the hooks it's as changing a rule means changing code and redeploy the

46:47

SPEAKER_00

old harness all the alien and with the agent control which is the name of the open source library that we are going to use here the rules are registered on a local server via api you update them without touching the am code because the alien picks them up immediately let me show you that here in the code yeah this is the notebook in this demo we only need a extra package the agent control sdk so the am control is the one that helps us to create the steering and here we are using the setup control this is a little application a little app that i create with the local server and the steering rules

47:45

SPEAKER_00

here we have the local server we have the control there are the steering rules first we have the the still must guess steer you know guide so guide agent to reduce guest count when exceeding maximum of 10 the same and it's going to steer and it's going to steer and we have some control that deny for example deny no payment block blocking confirmation without peer payment and here well this is to um to create the the service the service the server and let's go to the uh notebook so i already let me go here yeah so i need the environment so this is the agent uh i'm going to have an error here wait wait wait i don't want to use metro wait wait

48:50

SPEAKER_00

no other let me comment this one because it's going to give me an error and let's run this yeah so we have the hooks as within the demo so we have book any company in lisbo and we have some prompts you are a hotel booking assistant when booking first describe while you will book in blah blah blah and this is the problem for my agent where is my agent agent control this is for the well this is on song helper and this is my hook that you already know because we create this so we have the system pro and we have the hooks so let's test this again with the uh book any company lisbo for 50 guests

49:38

SPEAKER_00

and if you remember this only can book for uh less than 10 guesses and of course is blocking now let's go to the agent control new agent here we have two key importance or two key imports that are important for the agent control sdk so we have the agent control plugin that captures agent events and sends them to the agent control server and we have the agent control student handle that listen for a student decision for the server and delivers them back to the model together here they are what connects strands to the agent control steering logic so let's go this with this so the steering agent is going to a real book a room for any company lisbo 50

50:45

SPEAKER_00

guests let's see what it's doing yeah i have successful booking for stay in a company lisbo for 50 guests so it's the reservation have been split into two rooms so it took that it itself you know one room for one and other room for five and that's it so use hook for hard constraints agent control for software rule so now we have five techniques or running locally but how do you take this to production without maintaining service without building infrastructure let me show you how everything that i just built here in the previous demos runs locally for the production version amazon better better amazon better agent code give you the runtime a gateway

51:48

SPEAKER_00

a short term and a long-term memory a cloud watch observability built in no service to manage here the architecture this transient runs inside the runtime and you know inside the runtime you can put every framework that you want the only error is trans the gateway rules tools calls to the lambda function automatically as a tools and the steering rules from the previous demo live in the dynamo dv so you change them there and they are live on the next call and you don't need to reply anything and if you want to use neo4j of course you can use the infoj or a dev that is external graph database we also have a free tier and the code is in the repo here

52:51

SPEAKER_00

it's in the repo and you will need aws credential if you are using amazon better better agent code and in the resource link there is some credits i hope that you can find it because i always try to to give away some credit for aws so you can deploy everything for free and uh if you are new here in aws i have a repository that you can use to deploy everything over this architecture uh using a notebook too but if you are familiar with cdk cloud developer kit you can do it as well to deploy everything at once uh both options are in the repository of core and you can go deeper in agent core with all the documentation that i left there

53:44

SPEAKER_00

and you can do it as well to deploy everything over here so please don't stop at the demo and try going to production with amazon bedrock agent core so let me bring it all back tokens weighs on every request you can fix it with semantic tool selection and confident answer it's never compute possible possible possible possible because you are asking how many so you can use refrag and query the data don't sample it sometimes we have fabricate success confirmation you can use multi-aging validation and second path sketched it and rules the models are quite a little skip it you can use

54:33

SPEAKER_00

uh neuro symbolic guardians to enforce in the code don't trust in the problem and for the hard blocks that stop the user you can use runtime steer self-correct and finish each demo that i just showed you is in the in the repository as a notebook and an application as well you can use the demo and then go to the demo 5 and if you want deploy in the productions so have you tried any of this in your own hands thank you to join me in this session and happy building

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note