Open Reader

Tribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk

completed 12:02 Aug 29, 2026 Watch on YouTube

Current Status

completed

Video ID

dQ-_i1tZiws

RAG / Chat

Enabled
Tribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk
Description

Maersk's standard operating procedures were screenshots. A sequence of images showing what a person sees and where they click, which is a perfectly good record for a human and useless to an agent. Dmitry Buykin calls the gap tribal dungeons: the knowledge exists, just not in a form anything can execute safely. An agent version of the same procedure needs preconditions, decisions, identifiers, backend calls, validation, recovery and evidence that it actually worked. Most of the project was that translation, negotiated with the people who own the process, because experts own the what and agents own the how. His sharpest point is about where the engineering actually lives. The agent loop is not the system. The refining loop around it is, and the corpus of procedures outweighs the runtime roughly twenty to one, because the same shipping step means different things in different countries. Accuracy was not designed up front in a diagram, it was earned through more than 100,000 corrections over nine months, with heat maps turning traces into priorities and a single cell often costing the team a month or two. A correction only counts once it becomes an executable change, which is the line between an opinion and a production fix. Discovery needs agent freedom, production needs a cage, and a harness exists to make the dumb mistakes impossible rather than to give the model more room. Speaker info: - https://x.com/tzakus - https://www.linkedin.com/in/buykin/ Timestamps: 0:00 - The long tail is the expensive part 2:04 - What an SOP has to become for an agent 3:56 - The refining loop is the system 4:50 - Running 200 instances against legacy backends 5:47 - Triage, traces, and shared evidence 6:44 - Where vibe coding and spec driven work run out 7:39 - A hundred thousand corrections, and heat maps 8:36 - Please be careful is not a guardrail 9:31 - Five moves, and compounding improvement 10:31 - Composite tools, and why they skip MCP

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Production-scale AI agents in global shipping succeed not through a better agent loop or larger model, but through an adaptive operating system that turns tribal operational knowledge into bounded, observable, replay-tested executable procedures.
  • Why it matters: This is a concrete production account of agent orchestration over legacy systems, where exceptions, human-expert feedback, write controls, and compounding process memory—not model capability—determine reliability.
  • Best use: Use it as a design reference for building high-stakes agent systems: formalize work before automating it, constrain execution by failure mode, and make every correction become a testable production change.

Executive Summary

Dmitry Buykin describes Maersk's agent challenge as the unautomated long tail of shipping exceptions. A single shipment spans multiple parallel state machines and legacy applications; once any one system fails or diverges from its happy path, resolving the case requires expert coordination across incomplete and inconsistent systems. The central obstacle is "tribal" process knowledge that exists in people, screenshots, and loosely written SOPs but is not yet safe for an agent to execute.

The proposed architecture has three components: an SOP corpus as process memory, an execution runtime, and a mechanism for capturing subject-matter-expert feedback. The key distinction is that the agent loop is only a component; the real system is the iterative refinement loop around it. SOPs must be expanded beyond UI click instructions into preconditions, decisions, identifiers, backend calls, validations, recovery paths, and proof that work completed correctly.

Maersk operationalizes quality through traces, expert triage, replay evaluation with production write access disabled, and fixes that are encoded as executable changes rather than informal feedback. Buykin reports more than 100,000 corrections over nine months, using failure heat maps to prioritize clusters of scenarios. In this framing, an agent failure is the start of diagnosis: each failure must map to a preventive control, such as a workflow/classifier evaluation, a write gate, or a required human review.

The strategic asset is therefore not rented model intelligence but the adaptive architecture and reusable process memory it accumulates. Successful repeated sequences are consolidated into composite tools and reusable snippets, enabling rollout across countries despite local variation. Buykin also argues against tool-attachment: Maersk does not treat MCP as a default, preferring distilled, purpose-tuned function interfaces that provide tighter quality control over bloated legacy-system responses.

Key Takeaways

  • Claim: The economically important agent opportunity is the exception-heavy operational long tail, not the already-automated happy path. | Evidence: Buykin says each shipment is an orchestration of many parallel state machines; when one cannot complete, the process requires expert orchestration across multiple incomplete systems. He characterizes the remaining work as a long tail with more exceptions than existing systems were built to handle. | Implication: Ken should evaluate agent opportunities by the quality and representability of exception handling, rather than by apparent success on routine workflows. | Caveat: Exception automation is intrinsically harder because behavior varies by scenario, system state, and local operating conditions; it cannot safely be approached as a generic chat or loop-agent problem.
  • Claim: An organization cannot safely automate a process that it cannot represent in executable form. | Evidence: The talk contrasts a legacy SOP made of screenshots and click sequences with an agent-ready SOP containing preconditions, decisions, identifiers, backend calls, validation, recovery, and evidence of successful execution. Buykin summarizes the division as experts owning the "what" and agents owning the "how." | Implication: Process representation should be treated as a first-class asset and implementation workstream, not as documentation to be inferred on the fly by an agent. | Caveat: Creating agent-ready procedures requires substantial translation and negotiation with SMEs to align on operational common sense; this is described as most of the effort.
  • Claim: The production system is the feedback-and-refinement loop around the agent, composed of SOP memory, execution runtime, and SME feedback capture. | Evidence: Buykin explicitly identifies these three architectural parts and says the refining loop—not the agent loop—is the most complex part. He describes the SOP corpus as company process memory that must be modified and aligned to each country's conditions, visually portraying it as far larger than runtime at roughly a 20:1 proportion. | Implication: For multi-region deployments, build versioned process knowledge and feedback flows as the control plane; do not expect a single prompt, workflow, or tool definition to generalize unchanged. | Caveat: Country-specific process differences create many variations, so central process memory must preserve local conditions rather than assume one universal workflow.
  • Claim: Reliable agent quality is earned through trace-driven replay and executable corrections, not through intuition or simply upgrading models. | Evidence: Maersk uses traces as shared evidence for SME and engineering review, runs replay against real examples with rights disabled to protect production, and only counts a correction when it becomes an executable change. Buykin reports more than 100,000 corrections over nine months and uses heat maps of traced scenario clusters to prioritize work. | Implication: Instrument every run, retain replayable cases, and require fixes to land as tests, rules, tool changes, or workflow updates; subjective expert feedback alone does not compound into system quality. | Caveat: Improving a red scenario cluster can take one to two months of effort from the combined engineering and AI-agent team, indicating that meaningful reliability gains are operationally expensive.
  • Claim: Safety requires preventive constraints matched to the specific failure mode, with human approval retained for critical paths. | Evidence: Buykin distinguishes a wrong workflow requiring classifier/workflow evaluation, a wrong write requiring a write gate, and a wrong assumption requiring a review step. He argues that a harness should make dumb mistakes impossible rather than merely give an agent more latitude, and states that review and approval remain in the loop on critical paths. | Implication: Design agent control planes around explicit failure taxonomy: routing/evaluation controls for wrong task handling, permission gates for irreversible actions, and mandatory review for ambiguous assumptions or critical outcomes. | Caveat: A generic guardrail is insufficient when the underlying workflow itself is wrong; controls must eliminate the unsafe path rather than only detect or warn about it afterward.
  • Claim: The durable asset is adaptive process architecture that compounds reusable successful behavior, while foundation models are interchangeable rented intelligence. | Evidence: Maersk aggregates repeatable successful step sequences and scenarios into larger composite tools and reusable snippets for other agents, with the aim of rolling them from one country to hundreds. Buykin says his team does not use MCP as a default because legacy-system interfaces can be bloated; instead it distills responses and tunes tools through function coding. | Implication: Prioritize proprietary workflow memory, tested composite tools, and tightly specified tool contracts over dependency on any one model or tool-interface standard. | Caveat: The anti-MCP conclusion is an implementation choice driven by Maersk's legacy-system and quality-control needs, not evidence that MCP is universally unsuitable.

Detailed Brief

Production scale and operational bottlenecks

  • Claims: The system operates at more than 200 instances, with latency ranging from a few minutes to as much as 10 minutes.; Legacy-system dependency, rather than the agent loop itself, is identified as the primary driver of latency.; Expert time is the binding operational constraint, so failure triage must convert large volumes of traces into actionable clusters.
  • Evidence: Buykin says an SME bench performs triage and returns actionable failure groupings rather than raw cases for review.; Heat maps group tracked scenarios so experts and engineers can focus on the highest-benefit work.
  • Caveats: The talk provides no baseline error rate, throughput, cost, or quantified business impact, so scale claims should not be read as a full ROI case.; The precise meanings of the 200-plus instances and the stated corpus-to-runtime 20:1 ratio are not defined in the transcript.
  • Implications: Capacity planning for enterprise agents must include legacy-system latency, SME review capacity, and remediation lead time—not only model inference cost and response time.; A triage layer that prioritizes failure clusters is necessary once agent traces exceed what individual experts can inspect manually.

Notable Concepts & Terms

  • Tribal dungeons: Buykin's term for operational knowledge that exists in experts and fragmented artifacts but has not been formalized into agent-executable procedures.
  • SOP corpus: The company process-memory layer containing standardized, localized, and continuously updated procedures; presented as a larger strategic asset than runtime.
  • Parallel state machines: The multiple systems and workflows governing a shipment simultaneously; inconsistency among them creates exception work.
  • Trace: Shared evidence of what happened in an agent case, enabling SMEs and engineers to inspect the same facts and agree on corrective action.
  • Executable correction: A fix encoded into behavior, tooling, workflow, or evaluation such that it can be replayed and validated, rather than remaining an expert opinion.
  • Write gate: A control that prevents unsafe or incorrect changes to production systems when an agent attempts a write action.
  • Composite tools: Larger reusable tools assembled from proven repeatable action sequences, intended to let lessons from one geography scale across many.
  • Rented intelligence: The view that frontier models are externally sourced and replaceable, whereas the enduring proprietary value is the adaptive architecture and accumulated process knowledge.

Operator Notes / Why Ken Should Care

  • Require every high-impact workflow candidate to pass a representability review: explicit states, preconditions, allowed actions, identifiers, validations, recovery paths, and completion evidence.
  • Build a trace-to-fix pipeline in which every material failure is clustered, assigned an owner, replayed in a no-write environment, and closed only through a versioned executable change.
  • Create a control matrix that maps wrong routing, wrong write, and wrong assumptions to distinct preventive mechanisms; prohibit relying on a single generic guardrail.
  • Measure SME triage capacity and legacy-system latency as core production constraints before committing to scale targets.
  • Treat MCP and other agent interface conventions as implementation options; benchmark them against distilled, task-specific tool contracts for payload size, observability, permissioning, and error containment.

Source/Metadata

  • Title: Tribal Dungeons of Global Shipping: AI Agents at Global Scale — Dmitry Buykin, Maersk
  • Transcript words: 1262
  • Duration seconds: 722
  • Timestamp note: No timestamps or chapters were present in the supplied transcript.

Transcript

1234 words en Processed in 69.3s

Hello, everyone. This is a practitioner report from real production work. So let's get into it. And I'll skip the generic yet another loop agent intro. This is about the hard part most agents must skip. And about turning this operational knowledge into something an agent can execute safely. This comes from real work in my company I'm working for, supporting global shipping operations, and ground that in production. On paper, it's one workflow usually, but in reality, every shipment is an orchestration of many parallel state machines. While they agree the happy paths work, the moment one reads, you get exception work. The easy majority is already automated in many companies. What's left is the long tail and more exceptions than systems built to handle them. That tail is the expensive part. And then there's my favorite category. It comes with a special plate. Yeah. See? For AI Builder Dreams and their laptops. This is what you can find outside of the AI bubble in San Francisco. The signal process depends on many systems being coherent at once. If any step can't complete, the happy path breaks, and then it takes expert orchestration across multiple incomplete systems. All these variations and pathways should be captured in SOPs. SOPs are the standard operating procedure, common and regulated industries. So an expert and the model read them the same way. That gap is the hard part. Stable intent detection. Tool calls you can guarantee are safe, integrating with legacy backends, and results evaluated with experts. I call these tribal dungeons. The knowledge exists but not in a form an agent can execute. And you can't safely run a process. You can't safely run a process the organization cannot represent. Standard legacy is a piece. A bunch of screenshots organized in sequence. But screenshots are not a process. A legacy is a piece. Explain what the person sees and clicks. And an agent SOP needs a more complex setup: Preconditions, decisions, identifiers, back-end calls, validation, recovery, and evidence of successful execution. Experts own the what. Agents own the how. An exception becomes a guardrail. Most of the effort is the translation and negotiation between them to align on common sense. Three parts here in this architecture. It's SOP memory organized as SOP corpus. Execution runtime and SME feedback capture. The agent loop is not the system. The refining loop around the agent is the system. And it's the most complex part. Oops, sorry. SOP is, okay. It's this slide for UK. This is the correct one. So, and it's a good illustration why the same thing means different, describing different countries. And it's creating a lot of variations between a piece and each country. And that corpus is the asset, the company's process memory, modified and aligned with every country's conditions. And far bigger than runtime. You can see the proportion 20 to 1. And this is currently operating system. And this is the scale we run in production today. Over 200 instances and spikes. And latency deviates from a few minutes to up to 10 minutes. And mainly, the main reason for it is that we depend on many legacy systems, which cannot be faster than the agent loop itself. Expert time is the bottleneck. So, the SME bench does the triage for us. It clusters the failures and hands back something you can act on, not just look at it. The trace is the shared evidence that lets an expert and an engineer review the same case and agree on what happened. A correction only counts when it becomes an executable change. And that's the line between an opinion and a production fix. And this is where quality comes from. Not from vibes. Not from a bigger model. From replaying real examples with disabled rights to protect the production systems. And checking whether behavior improved. You can see here on the cognitive proportion of this effort ratio between each activity in our project. So, usually, vibe coding ends here. Here ends spec-driven development. Because it cannot grow and improve accuracy more than this stage on this scale. And this is where the real work starts. Nothing exotic. It's common engineering sense applied at scale. So, if you don't know all this terminology which developed over the last 30 years in software development, I recommend to check because this is what every AI coding agent should know to help you develop reliable production systems. And accuracy wasn't designed in one diagram up front. It was earned one small correction at a time at the scale you see here. So, we have over 100,000 corrections over the last nine months in the system when we're developing it. And these heat maps turned thousands of traces into priorities. It's how we keep experts and engineers looking at the same problems and prioritize where the most beneficial work for them. Every cell is a group of tracked scenarios we have. And usually, to turn one block in red, it's around one, two months of force for the whole team. All team of engineers and also AI agents. The agent failing is where the investigation starts, not where it ends. Each failure maps to a specific fix. Discovery needs agent freedom and production needs a cage. The harness isn't there to give the agent more room. It's there to make the dumb mistakes impossible. So, on this scale, please be careful. It's not a guardrail. If we have wrong workflow, then classifier eval. If it's wrong write, then write gate. If it's wrong assumption, then it's a mere view. A preventive measure eliminates the unsafe path. On critical paths, review and approval stay in the loop. The engineering focus is to build safe handoffs and a trail you can trust. The real outcome wasn't the agent in the system. It was the methodology we built around it. If you want the blueprint, then it's these five moves. Make work representable. Make execution bounded. Make behavior observable for every agent. Make correction cheap. And last thing is make improvement compound. So, gradually, systematically improve the quality of the system. AI-native operation is more than agency and workflow. It's a system that learns from what works and folds it back into code as new composite tools, adapting to the applications and the people around it. The best AI models are rented intelligence for us. The adaptive architecture we built is the asset, the final asset. And we're aggregating all repeatable sequences of steps, successful scenarios, and merging them into bigger tools which combine the disproven scenarios into the reusable snippets by other agents. So, and then it's possible to roll out then not only for one country, but for hundreds of countries in one go. So, this is all for the talk. A little time for questions, and I'll be around afterwards. And the final reminder, if you are an AI builder, if you emotionally attach the tools, not MCPs. We're not using MCPs because for us it's always not the best choice. Because all systems are usually really bloated, and we have to distill responses and tune the tools through function coding to our agents. Then we can control quality of our software and ensure that it's correctly processing assigned tasks. Thank you. Any questions? Okay. Then, thanks for your attention. Then, I will be around, so you can ask me questions if you want. Then, thanks for your attention. Then, I will be around, so you can ask me questions if you want.