AI Engineer

Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra

1673 summary words 7 min summary Watch video

Start with the signal

7 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: For large, structured physical systems, LLMs should translate ambiguous language into a bounded retrieval plan while deterministic code executes exact entity resolution, filtering, and set operations.
  • Why it matters: This is a reusable production architecture for preventing context-window, recall, hallucination, cost, and reproducibility failures when agents must operate over hundreds of thousands of near-identically named entities.
  • Best use: Use it as a design reference for agent control planes: keep intent parsing and final explanation in the model, but move searchable structure, identity resolution, and exact operations into deterministic services.

Executive Summary

Phaidra describes “semantic blindness”: an LLM can handle equipment-name resolution in a small demo, but becomes unreliable when asked to identify or enumerate assets across a gigawatt-scale data center containing hundreds of thousands of GPUs and more than a million total equipment nodes. Customer naming conventions are inconsistent and names often differ by only one character, making both raw-context prompting and embedding-based retrieval unreliable.

Their replacement architecture exploits the physical hierarchy already present in an AI factory: data center, hall, aisle, row, rack, and device. Rather than expose every asset name to the model, Phaidra gives it a compact summary of hierarchy types and asks it to produce a structured plan: entity type, subtree/scope, and filter. A deterministic resolver then executes indexed subtree retrieval and exact set intersections to produce the result set.

The reported result is a shift from scaling with equipment count to scaling roughly with tree depth. Phaidra says its prior approach fell from 80% correctness at 64 GPUs to about 30% at 460,000 GPUs, while the new resolver achieved 100% accuracy in its test suite and zero failures across 66 cases on six production systems. It also reduced a one-gigawatt evaluation pass from 116 million tokens to 390,000, while per-query cost remained about 9,000 tokens across both small and large systems.

The broader engineering argument is that AI-native products should often begin as prompt-driven Software 3.0 prototypes, then progressively move known, structured, reproducible work back into deterministic Software 1.0 components. The LLM remains responsible for ambiguous interpretation, planning, and answer synthesis—not for exhaustive search, counting, deduplication, or identity-critical retrieval.

Key Takeaways

  • Claim: Passing large inventories of asset names directly to an LLM fails at production scale because finite context and near-identical identifiers create silent omission, hallucination, and enumeration failures. | Evidence: Phaidra cites facilities with 400,000+ GPUs and over one million equipment nodes; names may differ only by a single character, such as chiller 6 versus chiller 7. The speakers also say repeated output of similar tokens can trigger model frequency penalties or guardrails, causing long lists to stop prematurely. | Implication: Do not make an LLM the source of truth for exhaustive entity enumeration or exact asset identity when the inventory is large or identifiers are highly similar. | Caveat: The failure modes are described from Phaidra's data-center use case; exact model behavior, including output guardrails and repetition penalties, can vary by provider and deployment.
  • Claim: Vector RAG is not a sufficient fallback for identity-sensitive retrieval when entity names are semantically similar but operationally distinct. | Evidence: The speakers argue that a 20-character equipment name differing by one character produces embeddings too similar to reliably distinguish devices such as adjacent chillers or CDUs, leading to poor accurate recall. | Implication: Separate semantic retrieval from canonical entity resolution: use a structured identifier, graph, or index for the latter rather than relying on embedding proximity. | Caveat: RAG can still assist with semantic documentation retrieval or broader discovery; the critique is specifically about exact resolution of dense, near-duplicate operational identifiers.
  • Claim: A hierarchical representation lets query-planning cost scale with tree depth rather than the number of physical instances. | Evidence: Their infrastructure hierarchy runs from data center through data hall, aisle, row, rack, and GPU; a 64-GPU system and a 460,000-GPU system can therefore yield similarly sized structural summaries because the number of hierarchy layers changes little while the number of leaves grows rapidly. | Implication: For agent systems over large structured domains, expose the ontology and navigable topology to the model, not the full instance-level inventory. | Caveat: This approach depends on having a usable hierarchy or graph schema; poorly modeled, cross-cutting, or rapidly changing relationships require additional indexing and data-model work.
  • Claim: The model should create a structured retrieval plan, while deterministic backend services should resolve the actual entities. | Evidence: For the request “GPUs ... running hot in data hall 11,” the LLM outputs the target type (GPUs), scope (the data-hall-11 subtree), and condition (running hot). The backend retrieves the indexed GPU subset for that hall and intersects it with the set matching the temperature condition. | Implication: Implement a typed planner/resolver boundary: validate model output against a schema, execute retrieval through controlled tools, and return a canonical result set before asking the model to explain it.
  • Claim: When user references remain vague, ask the LLM to generate search patterns rather than asking it to inspect every possible matching name. | Evidence: Phaidra says the model can infer patterns in naming conventions or data and return an executable pattern; the backend runs that pattern against the inventory, so the LLM never needs to process a large token volume of raw names. | Implication: Use the LLM as a compiler from natural language into constrained query expressions, not as the runtime engine that scans and selects from the full corpus. | Caveat: Pattern generation still needs deterministic validation, constrained syntax, and safe failure behavior, particularly where a broad pattern could select an incorrect operational asset group.
  • Claim: Phaidra reports that the planner-plus-deterministic-resolver design materially improved correctness and made token cost nearly independent of system size. | Evidence: With the same model and data and three runs per case, the prior method reportedly declined from 80% correctness at 64 GPUs to roughly 30% at 460,000 GPUs. The new method reported 100% accuracy in those tests and zero failures across 66 cases on six production systems. At one-gigawatt scale, it used 390,000 tokens per evaluation pass versus 116 million previously—approximately a 297x reduction—and about 9,000 tokens per query at both sizes. | Implication: The key architectural bet is not merely lower cost: it is that deterministic execution converts an otherwise probabilistic agent failure surface into one that can be tested for exactness. | Caveat: These are vendor-reported evaluation results; the transcript does not define the full case distribution, correctness rubric, latency, model version, or independent replication.
  • Claim: AI-native products should prototype with broad LLM behavior, then migrate any structured, rule-governed, reproducibility-critical functions into conventional software. | Evidence: Referencing Karpathy's Software 1.0 versus 3.0 framing, Phaidra says it began with a near-pure prompting solution for a demo, then moved bulk retrieval, exact set logic, counting, deduplication, and other known engineering problems into deterministic code as scale exposed weaknesses. | Implication: Treat early all-LLM workflows as discovery mechanisms, then explicitly identify and harden the steps that can be represented as data models, rules, schemas, or deterministic tools.

Detailed Brief

Production architecture boundary

  • Claims: The end-to-end flow is deliberately limited to two or three stages rather than an open-ended agentic loop.; The model is retained for ambiguity handling and response synthesis, but is prevented from performing high-volume retrieval or exact set construction.
  • Evidence: The described flow is user query to planning LLM to deterministic resolver to final result set.; The resolver uses pre-indexed trees based on physical location and interactions between equipment.
  • Caveats: The transcript does not describe authorization enforcement, freshness/streaming-index handling, conflict resolution when a query maps to multiple valid scopes, or what the system does when the planner produces an invalid plan.
  • Implications: The useful implementation unit is a query-planning contract with explicit entity types, scopes, predicates, and a validated execution grammar.; This design creates natural observability points: planner output, resolver inputs, selected canonical entities, applied predicates, and the final result-set cardinality.

Notable Concepts & Terms

  • Semantic blindness: Phaidra's term for an LLM losing reliable entity-level understanding when faced with a very large inventory of similarly named physical assets.
  • Linearizer: A compact summary of the system graph or hierarchy that gives the LLM enough structural context to plan a lookup without seeing every equipment instance.
  • Pre-indexed trees: Deterministic subsets of equipment organized by location and relationships, enabling efficient scope retrieval such as all GPUs in a given data hall.
  • Result set: The canonical set of entities produced by the deterministic resolver after indexed retrieval and set operations; it is the factual basis for the agent's final answer.
  • Software 1.0: Deterministic code used here for exact retrieval, set logic, counting, deduplication, and other reproducibility-critical operations.
  • Software 3.0: LLM-prompted behavior used here for interpreting novel phrasing, planning where and how to search, and writing a human-readable response.
  • Frequency penalty: The speakers' explanation for why models may stop or be suppressed when generating long lists of highly repetitive identifiers.

Operator Notes / Why Ken Should Care

  • Audit agent workflows for steps where the model is currently enumerating, counting, deduplicating, or selecting canonical IDs from large datasets; move those steps behind deterministic tools.
  • Define a constrained planner schema with fields equivalent to entity type, scope/subtree, predicates, and optional validated name-pattern expressions.
  • Build evaluation sets that scale entity counts while preserving query intent, and track exact recall, precision, silent omissions, hallucinated entities, token cost, and latency separately.
  • Require result-set provenance in operational agents: retain the resolved IDs, index version, applied filters, and query plan so output can be audited and reproduced.
  • Do not accept the reported 100% result as general proof; test the pattern against your own graph quality, naming entropy, permissions model, and cross-hierarchy relationships.

Source/Metadata

  • Title: Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra
  • Transcript words: 2938
  • Duration seconds: 984
  • Timestamp note: No timestamps or chapter markers were provided in the transcript.
Full transcript 2910 words · 13 min read
0:00

Welcome everyone. My name is Rahul Singh. I'm a Staff AI Research Engineer at Phaedra. And I'm Vansh Leistik. I'm a Senior Engineering Manager, also at Phaedra. And today we want to talk about the time when we gave an LLM 500,000 sensor names and it got confused. We call this problem semantic blindness. At Phaedra, we build AI agents for AI factories. This includes agents which allow our customers to talk about their data centers and explain to themselves and understand how the data centers are working and what problems they are facing on a day-to-day basis. User queries can be anything from what chiller is running hot

0:37

to analyze the distribution of temperatures across my data halls to is any of my GPUs facing any problems. Now, from this variety of queries, you can see that these include talking about specific equipment, is chiller 6 alright, to talking about groups of equipment, GPUs in data hall 1. The industry has not really figured out a common name pattern yet, and every single customer can have their own things. From simple names like racks with GPUs, data halls, which give you an exact idea of where the printings are, to something that is more difficult to comprehend like CH3 something something 6. We've seen all kinds of names in the industry going forward.

1:24

Now, when you're building a demo system, this works because you're working at a small scale. A simple LLM can look at all the names of your equipment and figure out what the user is talking about. But this problem really becomes intractable as you go for scale. For example, at 1 gigawatt-scale factories, you will see 400,000-plus GPUs. And to support those GPUs, you have power meters, you have chillers, you have other equipment. LLM context windows are finite, and you will very quickly saturate them, and it just becomes a problem. We say a product is something that works for all scenarios and does not fail silently. A demo just has to work for one.

2:07

In addition to having the LLM figure out these names, we could also have embedded them in a RAG database, a vector embedding approach. But the problem is oftentimes the names are so similar that semantics just fail. There's very little difference between a vector of, sorry, a string name 20 characters long, which differs by, let's say, one character, chiller 6 instead of chiller 7, or CDU something versus something else. It's very small. So you get a lot of problems with getting accurate recall. Also, LLMs suffer from what we call a frequency penalty. If you keep on outputting very similar names over and over again,

2:50

or very similar tokens, more accurately, over and over again, there are internal penalties in the LLM which just shut off their output. So if the user says list all the names of, let's say, GPUs in IL7N, there are, let's say, a hundred. Just by listing through them, the LLM's guardrails would think that it's spiraling, and it would just shut the system up. So we can't really have these two approaches. RAG would not work. Naive LLMs would not work. As we move into production systems, we need something that can scale as the system scales. Now, there are naive solutions which we shall discuss going forward.

3:27

A naive approach to solve this problem would be just to divide and conquer. Take your different equipment, branch them in different shards, and pass them through LLMs. Parallel recall should work, right? Well, that's what we thought. The problem is you get horrible recall and hallucinations. You will see LLMs invent phantom equipment that do not exist, and also silently drop things that do exist. Now, this creates a problem for mission-critical systems where the users need to know exactly what is happening with their systems. Any problems will quickly erode user trust. At the same time, you will miss specific problems which can cascade into bigger problems going forward.

4:09

Something that we realized here is that as the size of this physical infrastructure grows, LLM-based solutions cannot grow with the size of individual components or instances or nodes. We have to find something that grows sublinearly with increasing equipment count. And this is what we figured out. So, we should not grow with instances, we should grow with tree depth. Now, what do we mean by tree depth here? We realized that there is a hierarchical structure in which an AI factory is arranged. You will have data centers, then each data center will have different data halls, each of them will have different aisles,

4:48

then you will have different rows and racks, and then GPUs. Similarly, a chiller plant will have rooms where chillers are arranged, then you will have pumps, there will be a separate cooling tower unit, and it's kind of like a tree. The depth of the tree grows very slowly, the width grows extremely fast. In other words, you will have a hierarchy that only adds new equipment very rarely, but it adds a lot of them when it does. You will have a lot of GPUs, but without GPUs you won't have a lot of things. Now, we realized that this could be used to solve our problems. And there are four insights that really come into the picture for this. One is a linearizer.

5:29

What I mean by the linearizer is, the LLM has to figure out where each of these things are arranged and how to map from a vague user query to specific equipment or groups of equipment. Here, a summarized representation of our system's graph can be really useful. For example, a one gigawatt-scale factory can have over a million nodes, and each node represents a unique equipment. But because you want to go from the root to the leaf, all you have to do is describe all the parts, and that's a very small finite list. For example, here, you can see that to get to a GPU, you just need four different layers to get to it. To get to a chiller, similarly four layers.

6:15

To get to switches, similarly four layers. With this, a 64 GPU system and a 460,000 GPU system produce roughly the same size of summaries. This consolidated context gives your LLM all it needs to know to figure out how a plant is arranged and how different equipment are distributed in that plant. The second insight that we had was LLMs are good for planning but not good for searching. This is what we realized when we saw very poor recall with our sharded solutions. Instead of making the LLM sift through all the different fuzzy names that we can get from our users, we asked the LLM to give us exactly how to look for them. For example, here, the query says,

7:07

this is giving me all the GPUs that are running hot in data hall 11. The LLM does not need to go through all these names of all our GPUs to figure out which ones are in data hall 11 and then figure out how they are running hot. All it needs to do is structured outputs, which tells us that, well, we need to collect GPUs. The scope under which we need to collect is a subtree which is just data hall 11. And the filter that we need to apply to finally figure out what exactly we need is GPUs that are running hot. And this can be different things depending on the context, and you can implement different filters.

7:41

All we need to know is how do we create the set that we want to get to. The third thing that we realized is that once you have a structured output to figure out what exactly you need, building the back end for it is relatively simple and straightforward. Now, all we need to do is create different subsets of our equipment, or pre-indexed trees as we like to call them, based on location and based on how they interact with other equipment around them to get an idea of what you need to look at. So, for example, from the previous slide, we saw that we wanted to look at GPUs in data hall 11.

8:19

We could just get all the GPUs in data hall 11 in one pre-indexed subtree, or other data halls, for example, or other racks, for example. And then to get to the final result, all we need to do is run a query to find all GPUs that are running hot and take their intersection. Set operations ensure that we have perfect recall and accuracy irrespective of how we want to filter the queries and what fuzziness the user may have in their query. Next slide, please. Yes. And finally, what if you have something that is very vague, in which case you have to get the LLMs to search?

9:00

But instead of searching directly over names, we figured out that getting the LLMs to give us patterns to look for, this could be patterns in the data, patterns in the names, can be much more useful to figure out what the user is talking about instead of just passing the entire name list to the LLM. The LLM never has to see a very large number of tokens. All it needs to do is see some patterns in the naming convention, figure out what exactly the user is talking about, create a pattern, and then we can execute it on the back end ourselves. This makes sure that the LLM has a constant or relatively constant cost of operation.

9:55

Whereas if we had to read through everything, we would have been scaling linearly with increasing equipment count, which itself grows exponentially as the size of the system increases. Finally, to process a user query end to end, we go from the user's query, which could be anything from one piece of equipment or a group of equipment, to a plan that LLM which figures out user intent and gives us a search pattern. Based on that search plan, we have a deterministic resolver which indexes, does set operations, figures out exactly what we need to do, and creates the final set that maps to the user's query.

10:33

And this is what we call the result set. All of this is a two- or three-step process instead of a multi-step agentic loop which can keep on running over and over again. And this keeps our total cost also relatively flat and constant. Cool. If Rahul's job is to design the architecture, my job is production readiness. So we need to make sure that whatever we designed, the elegant solutions we came up with, actually hold up under real load, real customer data, and all the messy edge cases you only ever see in production. So before any of this went near a customer, we put it through some extensive tests and evals.

11:13

We measured the new system head to head against the old ones with the same LLM model, same data, and we did three runs per case just to make sure. Our goal was simple. Prove that Rahul's solutions are ready for production. And they definitely are. If you look at some stats here, the old approach degraded pretty fast with scale. So we got 80% correctness at 64 GPUs, and that dropped to about 30% when the GPU count grew to 460,000. On the other hand, our new approach has maintained correctness with 100% accuracy across all of those tests that we gave it. And this is not just a synthetic test. This is also our real data.

11:55

So we had 66 cases on six real production systems, and they produced zero failures as well. And not just correctness, it's also dramatically lighter. So when we talk about a one gigawatt-scale data center, the old approach burned 116 million tokens for just a single evaluation pass while still having a lot of errors. If you look at the new one, it's 390,000 tokens, which comes out to around 300 fewer tokens. But the part that matters the most, the cost is flat. As Rahul was talking about before, the system grows in size, but the cost of the query was 9,000 tokens a query, whether the system was 64 GPUs or 460,000.

12:43

So instead of having the cost grow exponentially when the customers grow their systems, it stays the same. That's the impact. Now there's something about what we learned and the part I want you to take back to your own work. It starts with a lens from Karpathy. Most of you have probably seen his presentation or his framing about this before. He talks about different kinds of software. Software 1.0 is deterministic code you write. It's predictable, exact, but a lot less flexible. On the other hand, we have software 3.0, which is basically behavior you prompt out of an LLM. In our case, that might be show me the GPUs in data hall 11 that are running hot.

13:21

This is very flexible, smart, but it's a bit fuzzy. Now his observation was that in legacy software, 3.0 is steadily eating 1.0. More of what used to be deterministic code becomes a prompt.

13:41

Just hold that picture because the lesson for new-built AI-native systems runs a little bit the other way. But the real skill is knowing which work the LLM is not best for. And you always want to keep the things that LLMs should be used for to the things that they do well. It's great at parsing an ambiguous request, judging where to look for data and what to look for, handling phrasing we've never seen from a new user that has a different query, and at the end also synthesizing and writing the final human-readable answer. But everything you can data-model, you should move into code. This is the key one, especially for large systems.

14:19

If your data has structure, call it a hierarchy, graph, or a schema, a language model scanning it token by token is definitely the wrong tool. Bulk retrieval, exact set logic, counting, dedupe across near-identical names, which happens a lot in the data center land. Anything that must be 100% reproducible, it should be deterministic code. The simple heuristic that usually works, if you can write down the structure or the rules, it's a 1.0 job. And pure LLM is weakest exactly when the system is large and well-structured, which is precisely where we operate and where our customers operate. So in a sense, we ran Karpathy's trend backwards. We started at almost pure 3.0.

15:06

We threw everything in a context window because that is the fastest way to find out what's even worth building. And as we said, it worked pretty well on a simple demo at the start. Then, as it met real scale, we moved the parts that can be treated as known engineering problems into 1.0. And we still kept the hard judgment in the LLM. So that's the inversion. Legacy software drifts from 1.0 towards 3.0, and new AI-native software starts at 3.0 and matures towards 1.0, for the use cases that earn it, of course. We are not here to replace model judgment. We want to feed it with the 1.0 tools, as we call it.

15:47

So every 1.0 function you add is more reliable ground for the LLM to stand on. So to finish up, let the LLM keep the hard decisions. Everything it shouldn't be guessing on, add it in as code. And hand that structure back to the model to work with. So demo with software 3.0, productionize by adding in software 1.0. So thanks for watching, everyone. Any follow-up questions at all, you can reach us by email. Rahul, RussellGhul at Fydera.ai, Pants.Listik at Fydera.ai, or you can drop them in the video comments below. Thank you for watching. Thanks.

16:23

Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple Couple

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note