AI Engineer

The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents

3325 summary words 15 min summary Watch video

Start with the signal

15 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Domain-specific agents—small, specialized agents with isolated contexts that communicate via natural language—will replace the current inheritance pattern of bloating general-purpose agents with MCP/skills/tools, driven by token cost increases, efficiency gains (80%+ token savings), and enterprise integration needs.
  • Why it matters: Token costs reversed trend in 2026 (up 29% IQ-adjusted, 76% raw), making monolithic general-purpose agents economically unsustainable for most use cases; domain-specific agents enable 137x cost reduction via small models, parallelization, stricter security, and composability.
  • Best use: Watch fully for the architectural blueprint and economic argument; then use the detailed_brief sections to extract specific implementation patterns (hooks, sandboxed filesystems, recursive sub-agents) and market timing predictions (2026-2027 inflection points).

Executive Summary

Justin Schroeder from StandardAgents (stealth) argues that the current agent integration strategy—loading general-purpose agents like Claude with MCP servers, skills, and tools—is a classic inheritance anti-pattern that will break down as enterprises scale. He defines agents as 'deterministic software that harness non-deterministic model outputs in pursuit of objectives,' and positions the agentic era as analogous to the Industrial Revolution (energy → machines, intelligence → agents). Current approaches suffer from portability, composability, observability, and security failures, forcing IT departments to choose between bypassing permissions or limiting agent capabilities.

His proposed alternative is composition over inheritance: small domain-specific agents (e.g., Figma, Gmail, Salesforce) each with isolated system prompts, minimal tool sets, sandboxed filesystems, code execution environments, and message histories. These agents communicate in natural language via a coordinator agent, mirroring how Apollo 11 mission control used teams of experts (biomimicry). StandardAgents has built this architecture internally and reports >80% token efficiency, 137x cost reduction vs. Claude when using DeepSeek V4 Flash for specialized tasks, better security (explicit approval per domain), and excellent scaling (parallelizable, no VPC bloat).

Schroeder makes a public prediction: by end-2026, domain-specific agents will see 'dramatic uptick' in adoption and frameworks (citing Vercel's Eve framework released days before this talk), with 2027 as 'the year of multi-agent orchestration.' His rationale: token costs reversed their downward trend in 2026 (despite conventional wisdom), enterprises cannot put expensive models like Claude in front of low-LTV customers, and the current agent stack (model → system prompt → tools → skills → MCP → messages) concentrates all innovation in context/model layers while ignoring structural composition. He sketches an ideal agent architecture with function/prompt/sub-agent tools, hooks for side effects (e.g., injecting timestamps), agent-specific rules, sandboxed filesystems, and recursive sub-agent trees (e.g., Salesforce agent with asset-generation sub-agent with GDPR/OSHA compliance sub-sub-agents).

Critical caveats: domain-specific agents 'don't exist' publicly yet beyond StandardAgents' internal work; task definitions must be more explicit upfront; and while small models are cheaper, they fail more often unless scoped tightly. The talk is a mix of architectural vision, economic argument, and soft product pitch for StandardAgents' early access.

Key Takeaways

  • Claim: Token costs reversed their multi-year downward trend in 2026, up 29% IQ-adjusted and 76% raw by mid-year. | Evidence: Schroeder cites tracking data from an unnamed website showing cost increases despite conventional wisdom that intelligence costs are falling; attributes this to memory crunch and other factors, though acknowledges 10-year trend may still be downward. | Caveat: Source of token cost tracking is not named; does not specify which models/providers drive the increase or whether this is regional/model-specific; long-term trend over a decade may still favor cost reductions. | Implication: Economic pressure to reduce token consumption via architectural changes (domain-specific agents) rather than waiting for model cost deflation; this contradicts widespread assumption that scaling laws will make efficiency optimizations unnecessary. | Timestamp: 18:30
  • Claim: Domain-specific agents achieve >80% token efficiency vs. general-purpose agents for defined tasks, enabling 137x cost reduction when using DeepSeek V4 Flash vs. Claude 3.5. | Evidence: StandardAgents' internal testing shows these gains because each agent operates with minimal context (system message + tools + single incoming message) rather than full conversation history and all MCP/skill layers; example: coordinator asks Gmail agent 'get last email from Debbie' rather than passing entire conversation context. | Caveat: Requires defining tasks 'a little bit more ahead of time'; if DeepSeek V4 Flash 'fails over and over again,' cost savings evaporate and user experience degrades; gains only apply when tasks are appropriately scoped to small model capabilities. | Implication: Enterprises can dramatically cut inference costs and use customer-facing AI economically by decomposing workflows into domain-scoped tasks; but architectural complexity shifts from runtime to design-time (defining agent boundaries and orchestration logic). | Timestamp: 15:45
  • Claim: Current agent integration strategies (MCP, skills) are inheritance anti-patterns that will break down at scale, similar to how inheritance breaks down in software engineering. | Evidence: Schroeder diagrams how MCP/skills/tools all inflate the context layer of a single agent, comparing this to object-oriented inheritance (adding attributes to one object). He notes research shows 'if you use very many [skills], it actually makes your agent substantially worse' and asks rhetorically whether 100 or 1000 skills would still work effectively. | Caveat: Does not cite specific research on skill degradation; does not provide empirical threshold for when inheritance pattern fails (e.g., X tools/skills/MCP servers); some users may never hit limits if use cases remain simple. | Implication: Teams investing heavily in MCP server development may be building on a fragile foundation; better to adopt composition patterns (multi-agent orchestration) now rather than after hitting scale/complexity walls. | Timestamp: 09:15
  • Claim: MCP has become 'a de facto tool distribution mechanism' and has not proven value beyond tools yet, despite supporting prompts/resources/sampling in spec. | Evidence: Shows screenshot from MCP website where only the 'tools' column is filled out across all MCP clients; states 'It has not proven to be great at providing other value yet.' | Caveat: MCP is relatively new (released late 2024); lack of adoption of non-tool features may reflect early-stage ecosystem rather than fundamental limits; Anthropic/others may be building non-tool use cases not yet public. | Implication: If evaluating MCP for enterprise integration, focus on tool distribution use case; do not expect prompts/resources to solve deeper compositional/orchestration problems; consider whether tool-only integration is sufficient or whether domain-specific agent orchestration is needed. | Timestamp: 08:00
  • Claim: Domain-specific agents enable stricter security boundaries because each agent can only perform pre-approved actions within its domain, unlike monolithic agents that 'can do anything.' | Evidence: References widespread practice of bypassing permissions to let coding agents operate freely (slide shows permissions screen); contrasts with domain-specific model where Gmail agent cannot access Salesforce, Figma agent cannot read email, etc.; claims 'Doug in IT' will be reassured by explicit domain scoping. | Caveat: Does not address how coordinator agent's permissions are scoped or how to prevent privilege escalation via natural-language requests to sub-agents; security boundary enforcement mechanisms not detailed; social engineering attacks via coordinator may still succeed. | Implication: Enterprises concerned about AI security can use domain-specific architecture to implement defense-in-depth (each agent is a security boundary); but must carefully design coordinator's capabilities and sub-agent communication protocol to prevent bypass. | Timestamp: 16:30
  • Claim: Vercel's Eve framework (released days before this talk) is the first major framework to explicitly support building domain-specific agents, signaling market inflection. | Evidence: Schroeder shows Eve homepage tagline: 'build a company brain, personal assistant, or domain-specific agent'; says this is the first time he saw his own terminology 'come back and hit me in my own face' after promoting it for months. | Caveat: Eve is brand-new (announced ~May 2025 based on context); actual adoption, production readiness, and whether it delivers on domain-specific agent vision remain to be seen; Schroeder may be over-interpreting one framework's marketing copy. | Implication: If building agent infrastructure now, evaluate Eve as potentially the first standardized way to build domain-specific agents; but expect rapid iteration and possible competitors from other framework vendors (LangChain, CrewAI, etc.) as market matures. | Timestamp: 18:00
  • Claim: Every domain-specific agent should have a sandboxed filesystem and code execution environment as primitives, not add-ons. | Evidence: Notes that ChatGPT, Claude, and Codex already provide per-agent filesystems when users ask them to create files (e.g., 'make a PDF for my son's birthday party'); argues this proves big labs recognize filesystems are necessary for effective agents, so domain-specific agents should bake this in rather than rely on external tooling. | Caveat: Does not specify sandboxing technology (containers, WASM, E2B, etc.) or how to prevent exfiltration/OS-level attacks; does not address multi-tenancy, resource limits, or cost of per-agent sandboxes at scale. | Implication: Agent runtime platforms should provide filesystem and code execution as first-class primitives, not bolt-ons; this may favor platforms like E2B, Modal, or custom container orchestration over lightweight SDKs. | Timestamp: 21:30

Detailed Brief

Why everyone is building custom agents (and why it's failing)

  • Claims: Real estate agencies, insurance brokers, and Fortune 500 companies are all trying to build custom agents despite lack of public awareness of named agents beyond Claude/Codex; Root cause is integration: businesses want their proprietary data integrated into AI to achieve 'dramatic gains'; Building robust agents is 'an absolute nightmare' for IT departments: no defined build pattern, poor portability (environment vars/runtimes break cross-machine), no composability (can't reuse university chatbot for other purposes), telemetry/observability extremely hard at scale; After failed attempts, companies retreat to MCP to shove corporate data into large general-purpose agents like Claude, which 'works okay' but is limited
  • Evidence: Anecdotal examples: 'real estate agency down the street,' 'independent private insurance brokers,' Fortune 500 companies; Notes that public cannot name agents beyond Claude/Codex, and even Claude may not be recognized as an agent; Mentions Vercel AI SDK as good tooling for provider abstractions, but still insufficient for durable execution/orchestration at scale; MCP website screenshot shows only 'tools' column fully supported across clients, not prompts/resources/sampling
  • Caveats: No quantitative data on how many companies are building agents or failure rates; MCP may mature beyond tool distribution as ecosystem evolves; Some companies may succeed at building custom agents if they have sufficient engineering resources
  • Implications: Enterprises currently caught between two bad options: build fragile custom agents or surrender control to general-purpose agents via MCP; Market opportunity for agent frameworks/platforms that solve portability, composability, observability, and security problems; IT risk and compliance concerns may accelerate search for architectures with better security boundaries than monolithic agents

Composition over inheritance: the domain-specific agent architecture

  • Claims: Current agent stack is model → system prompt → tools → skills → MCP → messages, almost all of which is context; This is inheritance pattern: taking one object (agent) and adding attributes (tools/skills/MCP) to expand capabilities; Composition alternative: many small agents (Figma, Gmail, Salesforce, travel, legal, GDPR, OSHA, etc.) each with isolated system prompt, tools, message history, and agentic loop; Agents communicate in natural language (English) via coordinator agent, mirroring Apollo 11 mission control (teams of experts, each with specialized tools/knowledge); Can be recursive: Salesforce agent has asset-generation sub-agent, legal team agent has GDPR and OSHA sub-agents, avoiding '45 megabytes of context just on GDPR'
  • Evidence: Diagrams showing single agent with bloated context vs. multiple small agents with minimal contexts; Apollo 11 control room photo with annotation identifying agent (engineer), LLM (brain), tools (dashboard), messages (mouth); Specific example: coordinator asks Gmail agent for 'last email from Debbie' → Gmail agent operates with only system message + tools + that one message, not full conversation history
  • Caveats: Requires 'defining tasks a little bit more ahead of time' compared to general-purpose agents; Orchestration logic (how coordinator routes to sub-agents, how sub-agents know when to escalate) not detailed; Failure modes of natural-language inter-agent communication not discussed (misunderstandings, hallucinated sub-agent responses, etc.)
  • Implications: Shift complexity from runtime (managing huge context) to design-time (decomposing workflows into domain-scoped agents and designing orchestration); Agent frameworks should focus on orchestration primitives, not just model wrappers; Enterprises can build reusable agent libraries (e.g., standardized GDPR agent) rather than re-implementing compliance logic in every project

Ideal agent architecture: functions, prompts, hooks, rules, sandboxes

  • Claims: Tools layer should support: (1) functions (e.g., write file to filesystem), (2) prompts (inject smaller prompts, or call different LLM like Nano Banana for image generation while using GPT-4.5.2 as primary), (3) sub-agents (full domain-specific agents as tools); Hooks: mechanisms to mutate/inject side effects, e.g., inject artificial message/tool call to tell LLM current time ('6:45 PM Pacific'), or fire side effects; Agent rules: how many turns/steps allowed per side, whether tool calls require validation, other agent-specific constraints; Every agent should have sandboxed filesystem (like ChatGPT/Claude already provide) and sandboxed code execution (safe from exfiltration/OS interaction); All of the above bundled = complete agent primitive
  • Evidence: Examples: using Nano Banana sub-agent for image generation within GPT-4.5.2 primary agent; injecting timestamp via hook as artificial message; Notes big labs (ChatGPT, Claude, Codex) already provide per-agent filesystems when users request file creation; Diagram showing agent stack with functions/prompts/sub-agents at tool layer, hooks layer above, agent rules layer above that
  • Caveats: Sandboxing technology not specified (containers, WASM, etc.); Multi-tenancy, resource limits, cost implications of per-agent sandboxes not addressed; How to prevent side effects from hooks becoming attack vector not discussed
  • Implications: Agent runtimes should be more opinionated and feature-rich than current SDK wrappers (need built-in filesystem, code execution, hooks, orchestration); Likely favors platforms like E2B, Modal, or custom container orchestration over lightweight libraries; Agent developers can use prompt-as-tool pattern to mix models within one workflow (e.g., cheap model for routing, expensive model for complex reasoning, image model for assets)

Market timing predictions and economic drivers

  • Claims: By end of 2026, 'dramatic uptick' in domain-specific agent adoption, frameworks, and discussion; 2027 is 'year of multi-agent orchestration'; Token costs reversed downward trend in 2026: up 29% IQ-adjusted, 76% raw, driven by memory crunch and other factors; Long-term (10-year) trend may still favor falling costs, but enterprises cannot assume near-term deflation; Cannot put expensive models like Claude in front of customers unless customer has 'massive lifetime value'; need efficient alternatives for customer-facing AI; Vercel Eve framework (released days before this talk) is first major framework to explicitly support domain-specific agents
  • Evidence: Unnamed website tracking token cost increases; 137x cost difference between DeepSeek V4 Flash and Claude 3.5 per task; Eve framework homepage screenshot with 'domain-specific agent' in tagline; Schroeder notes 'we're about halfway through 2026' as timing anchor
  • Caveats: Token cost data source not cited, may be specific to certain models/providers; Predictions are speculative; past pattern (MCP hype, then skills hype) shows rapid shifts in agent ecosystem; Eve is brand-new; actual production use and competitor responses unclear
  • Implications: Teams building agent infrastructure now should design for multi-agent orchestration patterns rather than monolithic agents; Economic tailwind for domain-specific agents if token costs continue rising; headwind if costs resume decline; Watch for competing frameworks from LangChain, CrewAI, Vercel, Anthropic, OpenAI that address orchestration/composition; Enterprises evaluating agent ROI should model token costs increasing, not decreasing, over next 12-18 months

Notable Concepts & Terms

  • Domain-specific agents: Small agents scoped to a single domain (e.g., Figma, Gmail, Salesforce) with isolated system prompts, minimal tool sets, and message histories; communicate via natural language to coordinator agent. Core architectural alternative to inheritance pattern.
  • Composition over inheritance: Software engineering principle applied to agents: instead of expanding one agent with more context (inheritance), use many small agents that compose (communicate). Schroeder's central argument for why domain-specific agents will win.
  • StandardAgents: Schroeder's stealth-mode company building domain-specific agent platform; claims >80% token efficiency and 137x cost reduction vs. Claude. Early access available at standardagents.ai; primary real-world validation of talk's claims.
  • Vercel Eve: Framework released days before talk, first major framework to explicitly support domain-specific agents in marketing. Schroeder cites as signal of market inflection; may become dominant framework or inspire competitors.
  • Agent hooks: Mechanisms to inject side effects or mutate agent state (e.g., insert artificial timestamp message, fire external events). Part of Schroeder's ideal agent architecture; currently not standard in most frameworks.
  • Multi-agent orchestration: Pattern of coordinator agent routing tasks to specialized sub-agents; Schroeder predicts this becomes dominant paradigm in 2027. Requires natural-language inter-agent communication and orchestration primitives in frameworks.
  • Token cost reversal: Schroeder's claim that token costs stopped falling and started rising in 2026 (29% IQ-adjusted, 76% raw), contradicting conventional wisdom. Key economic driver for domain-specific agents if trend continues.
  • MCP as tool distribution: Observation that Model Context Protocol has only succeeded at distributing tools to agents (not prompts/resources/sampling), limiting its value for deeper integration. Contrasts with domain-specific agents which are full compositional units.

Operator Notes / Why Ken Should Care

  • If building agent systems: design for composition (multiple small agents) rather than inheritance (one big agent with many tools/skills/MCP servers). Schroeder's 80%+ token efficiency and 137x cost reduction claims, if validated, are massive operational levers.
  • Watch token cost trends closely: if 2026 reversal continues, economic case for domain-specific agents strengthens dramatically. If costs resume decline, less urgency but composition pattern may still win on security/composability/observability.
  • Evaluate Vercel Eve as potentially first standardized framework for domain-specific agents, but expect rapid iteration and competitors. LangChain, CrewAI, Anthropic, OpenAI likely to release orchestration-focused tools in next 6-12 months.
  • For customer-facing AI, domain-specific agents may be only economically viable path if LTV is low. Calculate whether you can afford Claude-class models for every customer interaction; if not, start designing domain-scoped tasks for small models.
  • Security/compliance angle: domain-specific agents with explicit permission boundaries may be easier to audit/approve than monolithic agents with broad permissions. Frame pitch to IT/security as defense-in-depth, not just cost savings.
  • If StandardAgents' stealth product delivers on promises, could be interesting early investment/partnership target for enterprises needing custom agent orchestration. Justin Schroeder contactable at info@standardagents.ai.
  • Architectural implication: agent runtimes need to provide sandboxed filesystems and code execution as primitives, not add-ons. This may require heavier infrastructure (containers, not just SDK wrappers), favoring platforms like E2B, Modal, or custom orchestration.
  • Timing: Schroeder predicts late 2026 for domain-specific agent adoption acceleration, 2027 for multi-agent orchestration becoming mainstream. If building now, design for orchestration patterns to avoid rewrite in 12-18 months.

Watch Map

  • 00:00: Intro: StandardAgents (stealth), open-source projects (DMUX, Arrow.js)
  • 01:30: Thesis: agentic era = Industrial Revolution (energy → machines, intelligence → agents)
  • 03:00: Definition of agents: deterministic software harnessing non-deterministic model outputs
  • 04:15: Problem: everyone building custom agents (real estate, insurance, F500) but failing
  • 06:00: Why custom agents fail: hard to build robust, not portable, not composable, poor observability
  • 08:00: MCP retreat: people give up on custom agents, use MCP to shove data into Claude/ChatGPT; but MCP is just tool distribution, not enough
  • 09:15: Inheritance anti-pattern: current stack inflates context with tools/skills/MCP; breaks down at scale (composition over inheritance)
  • 11:30: Domain-specific agent architecture: small isolated agents (Figma, Gmail, Salesforce) with coordinator, communicate in natural language
  • 13:00: Apollo 11 analogy: teams of experts = domain-specific agents (biomimicry)
  • 14:30: StandardAgents results: >80% token efficiency, 137x cost reduction (DeepSeek V4 Flash vs. Claude 3.5), better security, excellent scaling
  • 16:30: Security argument: domain-specific agents enforce strict permission boundaries vs. monolithic agents that bypass permissions
  • 18:00: Market prediction: end-2026 dramatic uptick in domain-specific agents, 2027 = year of multi-agent orchestration; Vercel Eve as signal
  • 18:30: Token cost reversal: up 29% IQ-adjusted, 76% raw in 2026, contradicting conventional wisdom; economic driver for efficiency
  • 20:00: Customer-facing AI: cannot afford Claude for low-LTV customers; domain-specific agents enable economical efficacy
  • 21:00: Ideal agent architecture: functions/prompts/sub-agents as tools, hooks for side effects, agent rules, sandboxed filesystem and code execution as primitives
  • 24:00: Recursive sub-agents: coordinator → Salesforce → asset generation → GDPR/OSHA; each maintains minimal context
  • 26:30: Wrap: StandardAgents early access at standardagents.ai, contact info@standardagents.ai for ambitious businesses

Source/Metadata

  • Title: The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents
  • Transcript words: 10339
  • Duration seconds: 1838
  • Timestamp note: Timestamps manually estimated from 1838-second (30:38) duration; transcript does not include native timestamps, so MM:SS entries are approximate based on content flow.
Full transcript 4935 words · 46 min read
0:02

SPEAKER_00

Okay, so I'm going to be talking about domain-specific agents and why I really think that they are going to play an unbelievably important role in the future of AI and in the future of how we build agents. To get started real quick, my name is Justin Schrader. You can find me on X at JPSchrader. And I work at a small company called Standard Agents, which nobody's heard of right now because we're still in stealth mode. After this talk, if you're interested, feel free to reach out to me and I can let you know a little bit more. Mostly, I'm known for doing a lot of different open-source projects. DMUX, which is a great multiplexer for all of your coding agents. Arrow.js, which is a UI framework, similar to React for the agentic era. A bunch more that I won't get into, but maybe check them out if you're interested.

0:09

SPEAKER_00

Okay, I think we can all agree that the moment in time that we are in is very similar to the Industrial Revolution. In fact, it might be an accelerated Industrial Revolution. Maybe it's a bigger deal, but it's certainly not smaller. I probably don't need to convince you of that if you're listening to one of these talks. But that is the moment we find ourselves in. So I actually think it's helpful to go back and look at what was the key catalyst of the Industrial Revolution. And ultimately, it was that we learned how to harness energy with machines. We learned how to harness energy with machines. And what's interesting is that in this next era, we are essentially learning to harness intelligence with agents. And agents, I think, can be thought of as the machine of yesteryear. It's the thing that is going to use the intelligence, not so much us, but the agents. What's interesting about that is, I bet if I was in an actual room with you guys and we all put up our hands, I bet a lot of you when I say what is an agent instantly have examples that pop into mind, but also probably can't pull out a definition immediately. Some of you are not going to be able to do that. But I think that's interesting.

0:15

SPEAKER_00

But the reality is that we haven't even coalesced on a definition of what an agent is, even though we're well into the agentic era at this point. And I think that's interesting. Here's my definition. You can feel free to agree with it or not. But agents are deterministic software that harness the non-deterministic results produced by models in pursuit of some desired objective. Now, deterministic software might make you think more like a harness. And I actually think the distinction between an agent and a harness is really pedantic, not very helpful. And for the most part, in most cases, you can just conflate the two. A harness is an agent and an agent is a harness. Okay. And for the purposes of this talk, we're going to move forward with that. I think you could probably make some good arguments for why one is the other and vice versa, but really not important right now. Now, if you did have some examples pop to mind, they might have been Claude or Codex, OpenClaw, Hermes. But you know what's interesting is I bet if you went out onto the streets of corporate America in any city, maybe not San Francisco, but any city in America, and you asked somebody in an office building, could you name an agent by name? I think some people are going to get Claude. Some people might get Codex and that's about it. I don't think hardly anybody's going to be getting OpenClaw or Hermes. And really even Claude, I don't know that people would even know that that's an agent. These things are not well understood. And yet what's so crazy is everybody is building agents. I have a real estate agency down the street that's building agents. I know independent private insurance brokers building their own agents. I know Fortune 500 companies, lots of them building their own custom agents. Everybody is trying to build their own custom agents. And I know people don't believe me, but go talk to them. Just go talk to people. They are trying to build custom agents. And I can't help but wonder why. Nobody seems to be asking this question why. There's already AI everywhere. You can get on ChatGPT, all the way down to some open source model from China on some rickety website. There's everything in between, but still people want to build custom agents. And ultimately it comes down to integration. Businesses want their data properly integrated into AI. They believe and are probably right that if they appropriately leverage AI, they're going to have dramatic gains in their business and so on and so forth. So they need to figure out how to get integrated and build their own custom agents is obviously a way to do that. And it's one of the first ways that they discover as a mechanism for doing it. The problem though is that agents are really hard. You have to take very careful care of the agentic loop and make sure that it's properly orchestrated. There are tons of different provider abstractions you need to think about. Fortunately, there's some good tools coming out around that, like the Vercel AI SDK is great. Durable execution, you need to make sure if there's faults, we can pick back up. These are relatively hard problems, especially if you're thinking about it at scale. And the reality is there's tons more. There's all kinds of validations and stop conditions and so on and so forth. And so what often happens is people do try to build their own custom agents and they work as a demo, but not much more than that. And really, it turns out that it's an absolute nightmare for people. Building robust agents is just hard. And if you go talk to anybody in an IT department, they are pulling their hair out because there are so many different concerns. There's no defined way to build an agent right now, actually no defined way. The closest thing maybe is Eve that just came out from Vercel, maybe the closest thing. But in reality, everybody's coming up with their own way to do it. Telemetry and observability on these agents is unbelievably hard, especially at scale. If you want to know exactly what is getting transmitted on every single step of every single turn of your agent, so that way you can diagnose it and fine tune it and make sure things aren't going off the rails. That is very hard to do. Agents are also not portable. So if I do get a good agent working, if I've managed to climb to the top of this mountain and I've got a good agent that's finally working well, it works well on my machine. But if I try to pass that off to somebody else, there's a very high likelihood that between all of the environment variable configurations and system requirements and run times, there's a good chance it's not going to run on that person's machine. And they're not composable. So even if I get a really good chatbot working for my university,

0:20

SPEAKER_00

single turn of your agent, so that way you can diagnose it and fine tune it and make sure things aren't going off the rails. That is very hard to do. Agents are also not portable. So if I do get a good agent working, if I've managed to climb to the top of this mountain and I've got a good agent that's finally working well, it works well on my machine. But if I try to pass that off to somebody else, there's a very high likelihood that between all of the environment variable configurations and system requirements and run times, there's a good chance it's not going to run on that person's machine. And they're not composable. So even if I get a really good chatbot working for my university, the chances that I'm going to then be able to reuse that for another thing is very, very low. I can't just easily share that. So what often happens is after a short pursuit towards agents, people back away and say, okay, fine, no more agents, no agents. Instead, we're going to do the MCP thing. We've heard about this, it works. And sure enough, model context protocol does work. And really, it works pretty well to take your corporate information like Zillow's information and then shove that into one of these really large pre-existing agents, something like Claude or ChatGPT, which I would consider a large general purpose agent. And it sort of works like that. And it works okay. But if you take a look, this is actually from the MCP website. And if you take a look at what is supported in MCP clients around the world, you will notice that only one of these columns is actually filled out all the way down. And that, of course, is tools. So MCP has become a de facto tool distribution mechanism for agents. So if I need to get my company's tools into that other agent, then MCP is a good way to do that. It has not proven to be great at providing other value yet. And frankly, tools are just not enough. I like to joke that we didn't land a man on the moon by giving one guy a ton of tools. That's not a realistic way to get a really large project done. So maybe MCP is not the way. But, aha, we have skills. We have skills. And skills are great. I actually do enjoy skills. I'm sure you do too. We install them all the time for all kinds of things. And fundamentally, what a skill is, is a markdown file, which works as documentation. Now, interestingly, there's lots of research out there that shows that if you use very many of these, it actually makes your agent substantially worse. But they do work as documentation for various complex things. So, back to the analogy of a man getting to the moon, it's a little bit like just giving this guy a ton of documentation. And the documentation is going to help, but it's not the fundamental problem. So what's the fundamental problem? Okay. Let's build up a basic agent stack here. Let's start with a model. All agents start with a model. Big one, small one, doesn't matter. They start with a model. Then you have something like a system prompt on top of that, which tells the model what its role in the grand universe is, sort of its life objective. Then we have tools, the things that it can actually do, the effects it can take. And then skills would be layered on top of that. And then MCP would be layered on top of that. And then finally, you have all the messages from the conversation. That is roughly the stack of information that gets passed along within the runtime of an agent. And if you take a look here, almost all of it is context. Basically everything—the system prompt, tools, skills—all of that is stuff that ends up in the context of the agent. And so basically people are trying to solve the integration problem by working on the context or the model. These are the two areas where we constantly see new advances. We also see new things come out like skills and MCP, new technologies, new protocols. They are all coming out in the area of the context and the model.

0:27

SPEAKER_00

So how does it actually work then? Well, basically you work at a company, you occasionally need to do some business travel. So you've got a couple of travel MCPs installed. You've also got Figma and Playwright installed on yours. And all of these are building up in that context layer. And then you've got some Gmail MCPs to go check your mail for you and some Google Sheets to go fill out some expense reports or something like that. And then you've got skills. You're a developer. So you've got some React fixers and linters. This is actually, I think, the number one or the number two most popular MCP server that's out there. Maybe you've got Matt's grill me skill or maybe you've got the GitHub skill. And basically what you're doing is you are inflating that context layer. And we have a term for this in engineering. It's called inheritance. The idea of inheritance is you take an object and then you add more attributes to it to allow that one object to have other properties. And that's exactly what we are doing here with an agent. We're saying this agent is pretty good, but if we add all of these additional extra layers, then the agent can do stuff that it previously couldn't do before. That is exactly what inheritance is. And the truth about inheritance is it works. It does work. That's why these things are out there and they are working. But there's an old saying: composition over inheritance. And it turns out this is as old as time. Eventually inheritance starts to break down. Imagine, okay, I've got five skills on ChatGPT or on Claude. And that works pretty well. Now, what if I have 100 skills? What if I have 1000 skills? There's some point at which I get diminishing returns from adding additional context. That's just obvious. We all kind of understand that implicitly. So is there an alternative? Well, composition is the alternative to inheritance. It looks something like this. So imagine we have another little agent. And again, we're trying to provide Figma as a thing that can be done by our primary agent. Well, what we could do is have a tiny little agent where the actual system prompt of the agent is written specifically to be a Figma agent. It knows everything about Figma. It knows all of its context, all of its API, all of the right places to click and the things to do and mouse movements to make and everything like that. And then it has these precise tools that it needs to perform all of those actions and nothing more, just that. And then it has a very small message history, which just has to do with the Figma portion of this. And then you can have more of these. You can still have your Gmail and your travel and your Google sheets and all of that kind of stuff. But each of them is a separate isolated agent, a full agent, not just a little server with tools on it. It's a full agent with its own message history, its own agentic loop. And then above these, you have a

0:32

SPEAKER_00

the things to do and mouse movements to make and everything. And then it has these precise tools that it needs to perform all of those actions and nothing more, just that. And then a very small message history, which just has to do with the Figma portion of this. And then you can have more of these. You can still have your Gmail and your travel and your Google sheets and all of that stuff. But each of them is a separate isolated agent, a full agent, not just a little server with tools on it. It's a full agent with its own message history, its own agentic loop. And then above these, you have a coordinator and the communication mechanism for all of these small agents speaking to the larger agent above it is just English. They just talk to each other the way humans do. So if the primary agent is saying, oh, I should check my mail to see if there's anything about going on a trip. Well, it knows to go ask Gmail for any new emails about a trip. Those funnel their way back up says, oh yeah, actually, there's a trip coming up to Los Angeles this weekend. And then it can go to our travel agent and start to make bookings. That's a rough idea of how something like this could work. And the reality is it does work. And we know it works because this is actually how we got to the moon. There were teams of experts, teams of experts with faces that looked like that and faces that looked like that, each of them with different skills and capabilities and faces that looked like that. This is the Apollo 11 launch day. And look right here, there's an agent. I just found an agent sitting right there. That brain of his is that's his LLM. And here's his tools right there on the dashboard. Those are the tools. Now he didn't have all the tools. He just had those tools. And he was really, really, really good at that. And then look at that mouth. That's the messages. We are used to this. We can understand this. It implicitly works. It's almost a form of biomimicry for the agentic world. It works. And I call them domain-specific agents. I don't think I was the first person to utter the words domain-specific agents. Certainly not the first person to have this idea. But that is what I want to talk to you about agents that are just targeted to very specific domain. And we over here at standard agents have been building this ecosystem for quite some time. So we've gotten to have a really good inside look at how they actually work. And I'm not ready to come out here and announce a product or anything like that. But I can give you a little bit of a peek. First of all, they are far more token efficient, far more token efficient. We regularly see over 80% token efficiency for any given task. Now, it's a little more complicated because you have to define those tasks a little bit more ahead of time. But if you can have an agent portability where I can take that Gmail agent, squeeze it up, and then send it to somebody else, we can create an ecosystem where we don't have to create every one of these skills and capabilities. But within that domain, you're going to get dramatic efficiency. And part of the reason is if you think about the way that the context works, I don't need to have the entire context of the conversation when I make a choice to do something. Instead, my primary coordinator level can just ask the Gmail, hey, get that last email from Debbie. And that is the totality of the context. It literally just has the system message, its tools, and that message that came in. And so it is then able to perform this very targeted, very specialized, tiny little thing without all of the surrounding context. It's also far more practical with small language models. If you look at the difference in two models like DeepSeek V4 Flash and Claude 3.5, the cost difference is mind-boggling. It is 137 times cheaper than Claude per task. 137 times. Now granted, if DeepSeek V4 Flash fails over and over again to do the job, then not only is it going to be not that much cheaper, it's also going to be much more annoying to use it. But that's why domain-specific agents are so great. Because you don't need to have the V4 Flash do everything. Instead, it only needs to do the tasks that have been specifically picked for it to do. And with a very minimal context, it can execute those very faithfully. So you get these dramatic cost reductions, not only with the token efficiency, but also because you can use much smaller language models and even non-language models. You can use image generation and diffusion models. You can use all kinds of other models for smaller tasks. You can also enforce really strict limits on the capabilities. And I think you know what I'm talking about. I'm talking about this. We are all flying awfully close to the sun nowadays. Everybody's just bypassing permissions left and right. And of course you have to because a coding agent with a big model can do anything. And so we use it to do everything. In a world that would be powered by smaller domain-specific agents, those agents can't do everything. They can only do the things that are already explicitly approved for them to do. It doesn't mean that you still can't have permissions and permission dialogues, but you are opting into a much more controlled ecosystem. And I promise you when you explain that to Doug in IT, he puts his heart at ease understanding the difference between those two. And fourth, these have excellent scaling characteristics. Because each of these agents is its own small little execution environment, you can parallelize them. You can put them on the cloud very easily without needing a giant VPC up there. You can run thousands of instances all at the same time in all kinds of regions of the world. They don't actually need to be geographically co-located or anything like that. So they have very, very good scaling characteristics. Unfortunately, they don't exist. That's the downside. These domain-specific agents don't really exist. Not in a big public way. Like I said, here at Standard Agents, we have them, we are working with them on a daily basis, but they are not out there in public very much yet. However, that's changing. That is going to change very quickly. We're about halfway through 2026. And I'm here to make a public prediction that I think as we roll on from this point to the end of 2026, we are going to see a dramatic uptick in people talking about building domain-specific agents, frameworks around them. All kinds of things are coming down the pipe. And it's not going to be a small trickle. It's going to accelerate rapidly. And this will become one of the main players in the agentic ecosystem. And 2027, I would say, is basically the year of multi-agent orchestration. That's another word you'll start to hear a lot. So that's my big, bold public prediction. I was really excited just a few days ago when Vercel released Eve. This is the first time I actually saw the term that I had been blasting out into the void come back and hit me in my own face. The framework for building agents, build a company brain, personal assistant, or domain-specific agent. So there we go. About halfway through the year.

0:38

SPEAKER_00

accelerate rapidly. And this will become one of the main players in the agentic ecosystem. And 2027, I would say, is the year of multi-agent orchestration. That's another word you'll start to hear a lot, I think. So that's my big, bold public prediction. I was really excited just a few days ago when Vercel released Eve. This is the first time I actually saw the term that I had been blasting out into the void come back and hit me in my own face. The framework for building agents, build a company brain, personal assistant, or domain-specific agent. So there we go. About halfway through the year, we're going to start picking up steam. That's my prediction. And there's a number of reasons. One of them is something that most people believe right now is that the cost of intelligence is going down. That trend reversed in 2026, actually. We track this on a website. Tokens are not getting cheaper anymore. They are actually going up even when adjusted for IQ. They're up 29% when you adjust for IQ just this year. Halfway through the year, we're already up 30%. And that can be caused by lots of different things. Of course, we've got this memory crunch. And probably the long-term trend over a 10-year cycle or something is that intelligence will go down. But that does not mean that we need to be paying 137 times the cost for something that can be done just as effectively. The problem is it's harder to break those things apart. Now, if you don't account for IQ, tokens are up 76% this year. Almost 100% increase in tokens just this year. And we're not even halfway through it. So we are really trending upwards on token costs. So anything we can do, especially with large businesses, to bring that down is going to be really important. The other use case to really consider is putting AI in front of customers. You can't put fable in front of a customer unless that customer has a massive lifetime value. It's just too expensive. So you need to find a way to create great efficacy while being efficient. And domain-specific agents are going to be the way to do that. So I'm going to leave you here momentarily. But before I do, let's just dream a little bit. Let me dig in a little bit deeper to how an agent could be orchestrated and what an ideal agent would actually look like. And then I promise to leave you alone. Here we go. So remember, we got that model and we got the system prompt. And then at the tool layer, let's break that apart a little bit. On one hand, we have these functions. This would be like an actual function that can get executed, like write a file to the file system. Then we have prompts. Prompts are a lot like the system prompt, but they are smaller individual prompts that can get injected. And sub-prompts that can run a function that actually calls an LLM. So let's say I have a main agent running, but I want to use a nano banana just to generate an image when I'm using GPT-4 5.2 as my primary. Well, you can just have a tool that's a prompt. That would be really cool if you could do that. And then another type of tool could be another full-blown agent, like a complete other domain-specific agent could just be one of the tools. So that's the tool layer. And then you have hooks. What are hooks? Well, in this ideal world, a hook might be something that can harness or change or mutate or perform side effects. So let me give you an example. LLMs have no idea what time it is at any given point in time. Turns out a really great way to tell them what time it is, is you inject an artificial message or an artificial tool call in the message history. So it looks like somebody just said, hey, what time is it? And the other person replied, oh, it's 6:45 PM Pacific time. Pretty simple. You can do that with a hook, or you could fire off some side effect using a hook. So this is an important piece of an agent. And then finally, there's these agent rules. And agent rules are kind of complicated. It's like, how many times should one side have a turn? Can it go on for 10,000 turns or 10,000 steps before its turn is up? There's all kinds of interesting little rules. When it calls a tool, is it required to validate the whole thing or not? There are all kinds of very specific tools or rules that belong to a specific agent. And altogether, if we bundled all that up, we would call that an agent. But it's kind of missing a couple of things. One, every agent should really have a file system. If you've ever done this with ChatGPT or Claude or Codex, if you just ask it, hey, can you make me a PDF for my son's birthday party? Well, it'll do it. And it'll store it in its own little file system. So the big labs have already realized that in order to create an effective chat interface, not to mention a big agent, it needs some sort of file system. So every agent should have its own little sandbox file system. And also every agent should have a sandboxed code execution location. So it can write files, it can run those files, and it can do that safely without exfiltrating anything, without interacting with an OS at a higher level. That needs to be baked in as a primitive to every single domain-specific agent. Okay, so let's say that that's our ideal agent. And now let's talk about that little agent tool there. What is that? Well, those can be sub-agents, recursive sub-agents even. You could have an agent that calls a sub-agent that calls sub-agents that call sub-agents. And there could be one or there could be many of these at different levels. So for example, you could have this coordinator agent that's at the very top, and then you could have a Salesforce agent. And that agent knows Salesforce inside and out. It knows all of its APIs, it has all the credentials to communicate with your Salesforce instance in all the appropriate ways. And then it needs to communicate with a Google Workspace agent. So it can do all kinds of stuff in there. It can run spreadsheets. I can say, hey, what are all my top salespeople this year? And boom, it can look in Salesforce, it can coordinate with the sub-agent, create a sheet for you, send that back. Perfect. But maybe then you need to generate some assets. So the Salesforce agent actually has another sub-agent that it can talk to at any time it wants. And it's amazing at generating assets. Maybe that sub-agent doesn't just have codex image generation, maybe it has nano banana, maybe it has an SVG generator, all kinds of stuff. So that way it is an amazing asset generator and performs some of its own reflection and QA. And then our primary agent might need a whole legal team agent just so it can check the work that's coming out of these other ones. And maybe the legal team agent really needs a GDPR compliance agent just for those European customers. You know,

0:45

SPEAKER_00

Send that back. Perfect. But maybe then you need to generate some assets. So the Salesforce agent actually has another sub-agent that it can talk to at any time it wants. And it's amazing at generating assets. Maybe that sub-agent doesn't just have Codex image generation, maybe it has Nano Banana, maybe it has an SVG generator, all kinds of stuff. So that way it is an amazing asset generator and performs some of its own reflection in QA. And then our primary agent might need a whole legal team agent just so it can check the work that's coming out of these other ones. And maybe the legal team agent really needs a GDPR compliance agent just for those European customers. The main one doesn't have all of it—we don't want to have 45 megabytes of context just on GDPR. So we make that a separate sub-agent. And then maybe the legal team also needs an OSHA compliance agent, which is also very complicated. And so it has a separate one for that. You kind of get the idea. You can end up with all kinds of highly efficient, small little agents that are all working together, but maintaining small minimal context windows all the way through. That's the idea behind domain specific agents. So thank you very much. I appreciate you listening to my talk. Again, Standard Agents is where we're working—standardagents.ai. You can actually sign up on there for early access. We are slowly starting to roll this out to a few people. If your business is super ambitious and really wants to try out small domain specific agents, then you can write me at info at standardagents. And of course, I'd appreciate a follow. Thank you so much. Bye.

0:51

SPEAKER_00

more that I won't get into, but maybe check them out if you're interested. Okay, I think we can all agree that the moment in time that we are in is very similar to the Industrial Revolution. In fact, it might be like an accelerated Industrial Revolution. Maybe it's a bigger deal, but it's certainly not smaller. I probably don't need to convince you of that if you're listening to one of these talks. But that is the moment we find ourselves in. So I actually think it's helpful to go back and sort of look at what was the key catalyst of the Industrial Revolution. And ultimately, it was that we learned how to harness energy with machines.

1:29

SPEAKER_00

We learned how to harness energy with machines. And what's interesting is that in this next era, we are essentially learning to harness intelligence with agents. And agents, I think, can be thought of a little bit like the machine of yesteryear. It's the thing that is going to use the intelligence. Not so much us, but the agents. What's interesting about that is, I bet if I was in an actual room with you guys and we all put up our hands, I bet a lot of you when I say what is an agent instantly have examples that pop into mind, but also probably can't pull out a definition immediately. Some of you

2:12

SPEAKER_00

the agents are not going to be able to do that. But I think that's kind of interesting. But the reality is that we haven't even coalesced on a definition of what an agent is, even though we're well into the agentic era at this point. And I think that's kind of interesting. Here's my definition. You can feel free to agree with it or not. But agents are deterministic software that harness the non-deterministic results produced by models in pursuit of some desired objective. Now, deterministic software might make you think more like a harness. And I actually think the distinction between an agent and a harness is really pedantic, not very helpful. And for the most part,

2:53

SPEAKER_00

in most cases, you can just conflate the two. A harness is an agent and an agent is a harness. Okay. And for the purposes of this talk, we're going to go ahead and just move forward with that. I think you could probably make some good arguments for why one is the other and vice versa, but really not important right now. Now, if you did have some examples pop to mind, they might've been like Claude or Codex, you know, OpenClaw, Hermes. But you know what's interesting is I bet if you went out onto, you know, the streets of corporate America in any city, maybe not San Francisco, but any city in America, and you asked somebody just in an office building,

3:34

SPEAKER_00

could you name an agent by name? I think some people are going to get Claude. Some people might get Codex and that's about it. I don't think hardly anybody's going to be getting OpenClaw or Hermes. And really even Claude, I don't know that people would even know that that's an agent. These things are not well understood. And yet what's so crazy is everybody is building agents. I have a real estate agency down the street that's building agents. I know in like independent private insurance brokers building their own agents. I know Fortune 500 companies, lots of them building their own custom agents. Everybody is trying to build their own custom agents. And I know people

4:22

SPEAKER_00

don't believe me, but go talk to them. Just go talk to people. They are trying to build custom agents. And I can't help but wonder why. Nobody seems to be asking this question why. There's already AI everywhere. You can get on ChatGPT, all the way down to some open source model from China, on some, you know, rickety website. There's everything in between, but still people want to build custom agents. And ultimately it comes down to integration. Businesses want their data properly integrated into AI. They believe and are probably right that if they appropriately leverage AI, they're going to

5:00

SPEAKER_00

have these dramatic gains in their business and so on and so forth. So they need to figure out how to get integrated and build their own custom agents is obviously a way to do that. And it's one of the first ways that they discover as a mechanism for doing it. The problem though is that agents are really hard. You have to take very, very careful care of the agentic loop and make sure that it's properly orchestrated. There are a ton of different provider abstractions you need to think about. Fortunately, there's some good tools coming out around that, you know, like the the Vercel AI SDK is

5:37

SPEAKER_00

great. Durable execution, you need to make sure if there's faults, we can pick back up. These are relatively hard problems, especially if you're thinking about it at scale. And the reality is there's just tons more. There's all kinds of validations and stop conditions and so on and so forth. And so what often happens is people do try to build their own custom agents and they sort of work as a demo, but not much more than that. And really, it turns out that it's an absolute nightmare for people. Building robust agents is just hard. And if you go talk to anybody in an IT department,

6:13

SPEAKER_00

they are pulling their hair out because there are so many different concerns. There's no defined way to build an agent right now, like actually no defined way. The closest thing maybe is Eve that just came out from Vercel is maybe like the closest thing. But in reality, everybody's kind of coming up with their own way to do it. Telemetry and observability on these agents is unbelievably hard, especially at scale. Like if you want to know exactly what is getting transmitted on every single step of every single turn of your agent, so that way you can diagnose it and fine tune it and make sure things

6:48

SPEAKER_00

aren't going off the rails. That is very hard to do. Agents are also not portable. So if I do get a good agent working, if I've managed to climb to the top of this mountain and I've got a good agent that's finally working well, well, it works well on my machine. But if I try to pass that off to somebody else, there's a very high likelihood that between all of the environment variable configurations and system requirements and run times, there's a good chance it's not going to run on that person's machine. And they're not composable. So even if I get a really good chatbot working for my university,

7:26

SPEAKER_00

the chances that I'm going to then be able to reuse that for another thing is very, very low. I can't just easily share that. So what often happens is after a short pursuit towards agents, people kind of back away say, okay, fine, no more agents, no agents. Instead, we're going to do the MCP thing. We've heard about this, it works. And sure enough, model context protocol, it does work. And really, it works pretty well to take your corporate information like Zillow's information and then shove that into one of these really large pre-existing agents, something like Claude or ChatGPT, which I would consider a large general

8:08

SPEAKER_00

purpose agent. And it sort of works like that. And it works okay. But if you take a look, this is actually from the MCP website. And if you take a look at what is supported in MCP clients around the world, you will notice that only one of these columns is actually filled out all the way down. And that, of course, is tools. So MCP has become a de facto tool distribution mechanism for agents. So if I need to get my company's tools into that other agent, then MCP is a good way to do that. It has not proven to be great at providing other value yet. And frankly, tools are just not enough. You know, I like to joke that we

8:59

SPEAKER_00

didn't land a man on the moon by giving one guy a ton of tools. That's not a realistic way to get a really large project done. So, you know, maybe MCP is not the way. But, aha, we have skills. We have skills. And skills are great. I actually do enjoy skills. I'm sure you do too. We install them all the time for all kinds of things. And fundamentally, what a skill is, is a markdown file, which basically works as documentation. Now, interestingly, there's lots of research out there that shows that if you use very many of these, it actually makes your agent substantially worse. But they do work as documentation

9:41

SPEAKER_00

for various complex things. So, you know, back to the analogy of a man getting to the moon, it's a little bit like just giving this guy, you know, a ton of documentation. And the documentation is going to help, but it's not the fundamental problem. So what's the fundamental problem? Okay. Let's build up a basic agent stack here. Let's start with a model. All agents start with a model. Big one, small one, doesn't matter. They start with a model. Then you have something like a system prompt on top of that, which tells the model what its role in the grand universe is, sort of like its life objective.

10:19

SPEAKER_00

Then we have tools, the things that it can actually do, the effects it can take. And then skills would be layered on top of that. And then MCP would be layered on top of that. And then finally, you have all the messages from the conversation. That is roughly the stack of information that gets passed along within the runtime of an agent. And if you take a look here, almost all of it is context. Basically everything, the system prompt tools, skills, all of that is stuff that ends up in the context of the agent. And so basically people are trying to solve the integration problem by working on the context or the model. These are the two areas where we constantly see

11:08

SPEAKER_00

new advances. We also see new things come out like skills and MCP, new technologies, new protocols. They are all coming out in the area of the context and the model. So how does it actually work then? Well, basically you work at a company, you occasionally need to do some business travel. So you've got a couple of travel MCPs installed. You've also got Figma and Playwright installed on yours. And all of these are building up in that context layer. And then you've got some Gmail MCPs to go check your mail for you and some Google Sheets to go fill out some other some other expense reports or something like that. And then you've got skills. You're a developer. So

11:55

SPEAKER_00

you've got some React fixers and linters. This is actually, I think like the number one or the number two most popular MCP server that's out there. Maybe you've got Matt's grill me skill or maybe you've got the GitHub skill. And basically what you're doing is you are inflating that context layer. And we have a term for this in engineering. It's called inheritance. The idea of inheritance is you take an object and then you add more attributes to it to allow that one object to have other properties. Right? And that's exactly what we are doing here with an agent. We're saying this agent is pretty good, but if we add all of

12:37

SPEAKER_00

these additional extra layers, then the agent can do stuff that it previously couldn't do before. That is exactly what inheritance is. And the truth about inheritance is it works. It does work. That's why these things are out there and they are working. But there's an old saying, composition over inheritance. And it turns out this is as old as time. Eventually inheritance starts to break down. Imagine like, you know, okay, I've got five skills on chat GPT or on or on Claude, excuse me. And that works pretty well. Now, what if I have 100 skills? What if I have 1000 skills? There's some point at

13:21

SPEAKER_00

which I get diminishing returns from adding additional context. That's, that's just obvious. We all kind of understand that implicitly. So is there an alternative? Well, composition is the alternative to inheritance. It looks something like this. So like imagine we have another little agent. And again, we're trying to provide Figma as an, as a, as a thing that can be done by our primary agent. Well, what we could do is have a tiny little agent where the actual system prompt of the agent is written specifically to be a Figma agent. It knows everything about Figma. It knows all of its, all of its context, all of its API, all of the right places to click and

14:03

SPEAKER_00

the things to do and mouse movements to make and everything like that. And then it has these precise tools that it needs to perform all of those actions and nothing more, just that. And then a very small message history, which just has to do with the Figma portion of this. And then you can have more of these. You can still have your Gmail and your travel and your Google sheets and all of that kind of stuff. But each of them is a separate isolated agent, a full agent, not just a little server with tools on it. It's a full agent with its own message history, its own agentic loop. And then above these, you have a

14:41

SPEAKER_00

coordinator and the communication mechanism for all of these small agents speaking to the larger agent above it is just English. They just talk to each other the way human does. So if the primary agent is saying, oh, I should, I should check my mail to see if there's anything about going on a trip. Well, it knows to go ask Gmail for any new emails about a trip. Those funnel their way back up says, oh yeah, actually, there's a trip coming up to Los Angeles this weekend. And then it can go to our travel agent and start to make bookings. That's kind of a rough idea of how something like this could work. And the reality is it does work.

15:28

SPEAKER_00

And we know it works because this is actually how we got to the moon. There were teams of experts, teams of experts with faces that looked like that and faces that looked like that, each of them with different skills and capabilities and faces that looked like that. This is the Apollo 11 launch day. And look right here, there's an agent. I just found an agent sitting right there. That brain of his is, that's his LLM. And here's his tools right there on the dashboard. Those are the tools. Now he didn't have all the tools. He just had those tools. And he was really, really, really good at that. And then look

16:07

SPEAKER_00

at that mouth. That's the messages. We are used to this. We can understand this. It implicitly works. It's almost a form of biomimicry for the agentic world. It works. And I call them domain specific agents. I don't think I was the first person to utter the words domain specific agents. Certainly not the first person to have this idea. But that is what I want to talk to you about agents that are just targeted to very specific domain. And we over here at standard agents have been building this ecosystem for quite some time. So we've gotten to have a really good inside look at how they actually

16:46

SPEAKER_00

work. And I'm not ready to come out here and announce a product or anything like that. But I can give you a little bit of a peek. First of all, they are far more token efficient, far more token efficient. We regularly see over 80% token efficiency for any given task. Now, it's a little more complicated because you have to define those tasks a little bit more ahead of time. But if you can have an agent portability where I can take that Gmail agent, squeeze it up, and then send it to somebody else, we can create an ecosystem where we don't have to create every one of these skills and

17:22

SPEAKER_00

capabilities. But within that domain, you're going to get dramatic efficiency. And part of the reason is if you think about the way that the context works, I don't need to have the entire context of the conversation when I make a choice to do something. Instead, my primary coordinator level can just ask the Gmail, hey, get that last email from Debbie. And that is the totality of the context. It literally just has the system message, its tools, and that message that came in. And so it is then able to perform this very targeted, very specialized, tiny little thing without all of the surrounding

18:03

SPEAKER_00

context. It's also far more practical with small language models. If you look at the difference in two models like DeepSeq V4 Flash and Fable 5, the cost difference is mind-boggling. It is 137 times cheaper than Fable per task. 137 times. Now granted, if DeepSeq V4 Flash fails over and over and over again to do the job, then not only is it going to be not that much cheaper, it's also going to be much more annoying to use it. But that's why domain-specific agents are so great. Because you don't need to have the V4 Flash do everything. Instead, it only needs to do the tasks that have been specifically picked for it to do. And with a very minimal context,

19:07

SPEAKER_00

it can execute those very faithfully. So you get these dramatic cost reductions, not only with the token efficiency, but also because you can use much smaller language models and even non-language models. You can use image generation and diffusion models. You can use all kinds of other models for smaller tasks. You can also enforce really strict limits on the capabilities. And I think you know what I'm talking about. I'm talking about this. We are all flying awfully close to the sun nowadays. Everybody's just bypassing permissions left and right. And of course you have to because a coding agent with a big model

19:49

SPEAKER_00

can do anything. And so we use it to do everything. In a world that would be powered by smaller domain-specific agents, those agents can't do everything. They can only do the things that are already explicitly approved for them to do. It doesn't mean that you still can't have permissions and permission dialogues, but you are opting into a much more controlled ecosystem. And I promise you when you explain that to Doug in IT, he puts his heart at ease understanding the difference between those two. And fourth, these have excellent scaling characteristics. Because each of these agents is its own small little execution environment, you can parallelize them.

20:35

SPEAKER_00

You can put them on the cloud very easily without needing like a giant VPC up there. You can run thousands of instances all at the same time in all kinds of regions of the world. They don't actually need to be geographically co-located or anything like that. So they have very, very good scaling characteristics. Unfortunately, they don't exist. That's the downside. These domain-specific agents don't really exist. Not in a big public way. Like I said, here at Standard Agents, we have them, we are working with them on a daily basis, but they are not out there in public very much yet. However, that's

21:21

SPEAKER_00

changing. That is going to change very quickly. We're about halfway through 2026. And I'm here to make a public prediction that I think as we roll on from this point to the end of 2026, we are going to see a dramatic uptick in people talking about building domain-specific agents, frameworks around them. All kinds of things are coming down the pipe. And it's not going to be a small trickle. It's going to accelerate rapidly. And this will become one of the main players in the agentic ecosystem. And 2027, I would say, is basically the year of multi-agent orchestration. That's another word you'll start to

22:03

SPEAKER_00

hear a lot, I think. So that's my big, bold public prediction. I was really excited just a few days ago when Vercel released Eve. This is the first time I actually saw the term that I had been blasting out into the void come back and hit me in my own face. The framework for building agents, build a company brain, personal assistant, or domain-specific agent. So there we go. About halfway through the year, we're going to start picking up steam. That's my prediction. And there's a number of reasons. One of them is something that most people believe right now is that the cost of intelligence is going down.

22:44

SPEAKER_00

That trend reversed in 2026, actually. We track this on a website. Tokens are not getting cheaper anymore. They are actually going up even when adjusted for IQ. They're up 29% when you adjust for IQ just this year. Halfway through the year, we're already up 30%. And that can be caused by lots of different things. Of course, we've got this memory crunch. And probably the long-term trend over a 10-year cycle or something is that intelligence will go down. But that does not mean that we need to be paying 137 times the cost for something that can be done just as effectively. The problem is it's harder to

23:27

SPEAKER_00

break those things apart. Now, if you don't account for IQ, tokens are up 76% this year. Almost 100% increase in tokens just this year. And we're not even halfway through it. So we are really trending upwards on token costs. So anything we can do, especially with large businesses, to bring that down is going to be really important. The other use case to really consider is putting AI in front of customers. You can't put fable in front of a customer unless that customer has a massive lifetime value. It's just too expensive. So you need to find a way to create great efficacy while being efficient.

24:12

SPEAKER_00

And domain-specific agents are going to be the way to do that. So I'm going to leave you here momentarily. But before I do, let's just dream a little bit. Let me dig in a little bit deeper deeper to how an agent could be orchestrated and what an ideal agent would actually look like. And then I promise to leave you alone. Here we go. So remember, we got that model and we got the system prompt. And then at the tool layer, let's break that apart a little bit. On one hand, we have these like functions. This would be like an actual function that can get executed, like write a file file to the file system. Then we have prompts. Prompts are a lot like the system prompt, but they

24:57

SPEAKER_00

are smaller individual prompts that can get injected. And sub prompts that can, you know, you can run a function that actually calls an LLM. So let's say I have a main agent running, but I want to use a nano banana just to generate an image when I'm using GLM, you know, 5.2 as my primary. Well, you can just have a tool that's a prompt. That would be really cool if you could do that. And then another type of tool could be another full-blown agent, like a complete other domain-specific agent could just be one of the tools. So that's the tool layer. And then you have hooks. What are hooks? Well, in this ideal

25:40

SPEAKER_00

world, a hook might be something that can kind of harness or change or mutate or perform side effects. So let me give you an example. LLMs have no idea what time it is at any given point in time. Turns out a really great way to tell them what time it is, is you inject an artificial message or an artificial tool call in the message history. So it looks like somebody just said, hey, what time it is? And the other person replied, oh, it's 6.45 PM Pacific time. Pretty simple. You can do that with a hook, or you could fire off some side effect using a hook. So this is an important piece of an agent.

26:23

SPEAKER_00

And then finally, there's these agent rules. And agent rules are kind of complicated. It's like, how many times should one side have a turn? Can it go on for 10,000 turns or 10,000 steps before its turn is up? There's all kinds of interesting little rules. When it calls a tool, is it required to validate the whole thing or not? There are all kinds of very specific tools or rules that belong to a specific agent. And altogether, if we bundled all that up, we would call that an agent. But it's kind of missing a couple of things. One, every agent should really have a file system. If you've

27:05

SPEAKER_00

ever done this with ChatGPT or Claude or Codex, if you just ask it, you know, not inside of a project or anything like, hey, can you make me a PDF for my son's birthday party? Well, it'll do it. And it'll store it in its own little file system. So the big labs have already realized that in order to create an effective chat interface, not to mention a big agent, it needs some sort of file system. So every agent should have its own little sandbox file system. And also every agent should have a sandboxed code execution location. So it can write files, it can run those files, and it can do that safely without

27:47

SPEAKER_00

exfiltrating anything, without interacting with an OS at a higher level. That needs to be baked in as a primitive to every single domain specific agent. Okay, so let's say that that's our ideal agent. And now let's talk about that little agent tool there. What is that? Well, those can be sub-agents, recursive sub-agents even. You could have an agent that calls a sub-agent that calls sub-agents that call sub-agents. And there could be one or there could be many of these at different levels. So for example, you could have this coordinator agent that's at the very top, and then you could

28:23

SPEAKER_00

have a Salesforce agent. And that agent knows Salesforce inside and out. It knows all of its APIs, it has all the credentials to communicate with your Salesforce instance in all the appropriate ways. And then it needs to communicate with a Google Workspace agent. So it can do all kinds of stuff in there. It can run spreadsheets, I can say, hey, what are all my top salespeople this year? And boom, it can look in Salesforce, it can coordinate with the sub-agent, create a sheet for you, send that back. Perfect. But maybe then you need to generate some assets. So the Salesforce agent

28:56

SPEAKER_00

actually has another sub-agent that it can talk to at any time it wants. And it's amazing at generating assets. Maybe that sub-agent doesn't just have like, you know, codex image generation, maybe it has nano banana, maybe it has an SVG generator, all kinds of stuff. So that way it is an amazing asset generator and performs some of its own reflection in QA. And then our primary agent might need a whole legal team agent just so it can check the work that's coming out of these other ones. And maybe the legal team agent really needs a GDPR compliance agent just for those European customers. You know,

29:32

SPEAKER_00

the main one doesn't have all, you know, we don't want to have 45 megabytes of context just on GDPR. So we make that a separate sub-agent. So, you know, and then maybe the legal team also needs like an OSHA compliance agent, which is also very complicated. And so it has a separate one for that. You kind of get the idea. You can end up with all kinds of highly efficient, small little agents that are all working together, but maintaining small minimal context windows all the way through. That's the idea behind domain specific agents. So thank you very much. I appreciate you listening

30:10

SPEAKER_00

to my talk. Again, standard agents is where we're working standard agents.ai. You can actually sign up on there for early access. We are slowly starting to roll this out to a few people. If your business is super ambitious and really wants to try out small domain specific agents, then you know, you can write me info at standard agents. And of course, I'd appreciate a follow. Thank you so much. Bye.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note