AI Engineer

Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora

3553 summary words 16 min summary Watch video

Start with the signal

16 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Legora scaled from hundreds of thousands to billions of legal documents by migrating from Elasticsearch/Postgres to turbopuffer, whose object-storage-native architecture with per-namespace isolation and memory-hierarchy optimization delivered order-of-magnitude latency improvements, dramatically lower costs, and operational simplicity for multi-tenant enterprise workloads.
  • Why it matters: The architecture patterns and tradeoffs demonstrated here—writes-to-blob, tree-based vector indexes optimized for few S3 round trips, per-namespace encryption/residency, cache-hierarchy tuning—are directly applicable to Ken's agent orchestration systems, especially when scaling retrieval for multi-tenant or regulated environments.
  • Best use: Watch the full talk for the complete architecture evolution narrative; extract the turbopuffer design principles (object-storage-first, namespace isolation, memory hierarchy, tree vs graph indexing) and apply them to OpenClaw's retrieval layer, especially if Ken is considering vector search at scale or enterprise multi-tenancy requirements.

Executive Summary

Legora is a collaborative AI platform for legal work (contract review, legal research, litigation support) used by law firms and in-house legal teams. Jacob Lauritzen (Legora engineer) and Simon Eskildsen (turbopuffer CEO/co-founder) describe Legora's journey scaling search from hundreds of thousands to billions of legal documents across two workloads: project search (confined to tens to millions of documents per project) and legal research (deep research across laws, cases, and regulations at 10 billion vectors and growing).

Legora's search infrastructure evolved through five generations: (1) single Elasticsearch cluster; (2) multi-region Elasticsearch (EU/US/APAC) for data residency; (3) per-enterprise Postgres with pg_vector/disk_ann/ts_vector and aggressive partitioning (4,000 partitions, hash-binned projects) to meet full physical isolation and customer-managed encryption key (CMEK) requirements; (4) migration to turbopuffer at 400 million documents; (5) turbopuffer for legal research nearing 10 billion vectors. The Postgres phase delivered isolation but exploded in cost and latency (p99 spiked from 100ms to 20s) due to cache thrashing when hot and cold projects landed in the same partitions. Turbopuffer solved this by treating each project as a namespace, keeping cold projects in object storage (S3/GCS/Azure Blob) and hot projects in memory/NVMe cache, delivering order-of-magnitude latency improvements and much lower cost.

Turbopuffer's architecture is object-storage-native: writes go directly to S3 as a write-ahead log (hundreds of milliseconds write latency, acceptable for search), and indexes (vector, text, columnar) are built asynchronously in the background. At query time, turbopuffer routes to the node with the highest cache affinity, checking memory → NVMe SSD → object storage in sequence. The entire design minimizes S3 round trips (target: ~3 round trips per query, since S3 p99 is ~200ms per 1MB blob). Vector search uses a tree of clusters (not a graph) to avoid random access; text search (true BM25, not ts_vector) uses compressed inverted indexes and early termination. The memory hierarchy is explicit: root centroids live in DRAM, leaves on SSD, cold namespaces on blob.

For Legora, turbopuffer's per-namespace isolation enabled full physical separation, per-namespace encryption keys, per-customer buckets, and multi-region data residency without operational overhead. Even after disabling the NVMe cache for multi-tenant encryption constraints, memory-only performance was sufficient. Legal research benefits especially: EU law (hot) stays in cache, Danish law (cold, rarely queried) stays on blob with 500ms fetch latency, which is acceptable for deep research. The result: Legora went from managing separate Postgres/Elasticsearch instances per enterprise to a single turbopuffer cluster with 70+ tenants (growing to 200+), focusing engineering effort on product rather than infra, and making the CFO happy about cost.

Key Takeaways

  • Claim: Turbopuffer's write-to-object-storage-first architecture trades write latency (hundreds of milliseconds) for read adaptability and cost efficiency, making it ideal for search workloads where slow writes are acceptable. | Evidence: Writes go directly to S3/GCS/Azure Blob as JSON-like write-ahead log entries (1.json, 2.json, etc.), with no disk replication or Paxos. Indexes are built asynchronously in the background. Simon explicitly contrasts this with inventory reservation systems (e.g., Shopify Kylie Jenner flash sale) where this would not work, but for search it excels. | Implication: For Ken's agent systems, if retrieval reads vastly outnumber writes and sub-100ms write latency is not required, an object-storage-native design like turbopuffer's can dramatically reduce cost and operational complexity compared to traditional disk-replicated databases.
  • Claim: Per-namespace isolation (where each namespace is a separate S3 directory/bucket) enables full physical separation, per-customer encryption keys (CMEK), multi-region data residency, and simplified multi-tenancy without multiplying infrastructure. | Evidence: Legora required full physical isolation and CMEK for enterprises (banks, top law firms). Turbopuffer allows every namespace to be encrypted with a different key, stored in a different bucket (some customers have thousands of buckets), and located in different regions. Legora went from separate Postgres/Elasticsearch per enterprise to a single turbopuffer cluster with 70+ tenants growing to 200+. | Implication: If Ken is building multi-tenant agent systems with enterprise customers requiring data residency, CMEK, or physical isolation, per-namespace storage is a proven pattern that avoids the operational nightmare of per-customer infrastructure.
  • Claim: Aggressive partitioning in Postgres (4,000 partitions, hash-binned projects) led to cache thrashing and catastrophic latency degradation (p99 from 100ms to 20s) when hot and cold projects coexisted in the same partition. | Evidence: Legora used pg_vector (disk_ann), ts_vector (not BM25), and 4,000 partitions with hash-based bin packing. Cold projects (never queried after close) and hot projects (constantly active) landed in the same partitions. Postgres would pull entire partitions into memory, thrash the cache across partitions, and spike p99 to 20 seconds. | Implication: For Ken, aggressive partitioning is a false solution to multi-tenancy at scale when workload heat is highly variable. Blob-based storage that naturally segregates cold data is architecturally superior to in-database partitioning for mixed hot/cold access patterns.
  • Claim: Turbopuffer's tree-based vector index (clusters of clusters) is optimized for object storage and disk, avoiding the random-access penalty of graph-based indexes (HNSW) which require many S3 round trips with 200ms p99 per trip. | Evidence: Simon explains that graph-based vector search (HNSW, etc.) navigates from the center of the graph, incurring 200ms p99 per node traversal on S3. Turbopuffer instead organizes vectors into a tree of clusters (conceptually a very complicated B-tree on vector geometry), so queries traverse the tree in ~3 round trips, with root centroids in DRAM and leaves on SSD. | Implication: For Ken, if OpenClaw or agent systems require vector search at scale (especially over blob storage or NVMe), tree/clustering approaches (like Faiss IVF, ScaNN, or turbopuffer's design) will be cheaper and faster than graph indexes. Graph indexes are optimized for memory/register random access, not storage hierarchies.
  • Claim: True BM25 full-text search is more computationally expensive and difficult at web scale than vector search, requiring careful minimization of round trips, list compression, and early termination during set intersection. | Evidence: Simon describes BM25 as a hash map (token → document IDs) with set intersection and TF-IDF-style scoring. The art is minimizing round trips (download dictionary, then lists), compressing lists, and stopping intersection early once high-scoring documents are found (e.g., documents with 'York' and 'population' beat documents with only 'New'). Text search at web scale is harder than vector search. | Implication: Ken should not underestimate the complexity of hybrid search (text + vector). If building retrieval systems, invest in real BM25 implementations (not ts_vector, which Legora abandoned) and understand the memory bandwidth and round-trip optimization required. Turbopuffer's approach (download compressed lists, intersect in memory, early termination) is the gold standard.
  • Claim: Legal research at 10 billion vectors with heavy filtering (hierarchical jurisdictions, temporal validity, regulatory graphs) requires fan-out to many namespaces, where a long tail of cold jurisdictions (e.g., Danish law) can stay on blob with 500ms latency and a few hot jurisdictions (e.g., EU law) stay in cache. | Evidence: Legora's legal research fans out queries across jurisdictions (cities, counties, states, federal, international) and respects hierarchical authority and temporal overrides (judges overruling prior decisions). Danish law is rarely queried and stays on blob; EU law is hot and cached. Deep research workloads tolerate 500ms cold-blob fetch latency. | Implication: For Ken's agent systems, if retrieval involves hierarchical or graph-like filtering (e.g., regulatory compliance, multi-jurisdiction policy, or temporal validity), turbopuffer's namespace model (cold on blob, hot in cache) is a proven pattern. Accept higher latency for cold data in deep-research-style agent workflows.
  • Claim: Even without NVMe cache (disabled for multi-tenant encryption constraints), turbopuffer's memory-only performance was sufficient for Legora's enterprise workloads, demonstrating the efficiency of the object-storage + memory-hierarchy design. | Evidence: Legora's enterprise customers did not accept unencrypted NVMe cache. Turbopuffer was going to implement encryption for the disk cache but instead just disabled it and tested performance. Memory cache alone delivered acceptable latencies, so they kept it that way. Simon notes turbopuffer will support encrypted disk cache in the future. | Implication: For Ken, if security/compliance requires disabling disk cache, turbopuffer's architecture can still deliver acceptable performance with memory alone. This is a materially different cost/performance envelope than traditional vector databases that assume full memory or disk residency. | Caveat: This demonstrates sufficiency for Legora's use case (legal search with tolerance for some latency), not that memory-only is universally optimal. Future encrypted disk cache will expand applicability.

Detailed Brief

Legora's search infrastructure evolution: five generations

  • Claims: Generation 1: Single Elasticsearch cluster for all tenants, simple setup, worked initially for hundreds of thousands of documents.; Generation 2: Multi-region Elasticsearch (EU/US/APAC) to meet data residency requirements (Americans want US processing, Europeans want EU, Australians want Australia). Setup multiplied by 3-4x, annoying overhead but met requirements.; Generation 3: Per-enterprise Postgres with pg_vector (disk_ann), ts_vector (not BM25), 4,000 partitions, hash-binned projects, to meet full physical isolation and CMEK requirements. Expensive, search performance degraded, cache thrashing at scale.; Generation 4: Migration to turbopuffer at ~400 million documents, one namespace per project. Gained true BM25, much better relevancy, order-of-magnitude latency improvement, lower cost, single cluster across regions.; Generation 5: Turbopuffer for legal research approaching 10 billion vectors, one namespace per jurisdiction (EU hot, Danish cold, etc.), fan-out queries with hierarchical/temporal filtering.
  • Evidence: Jacob lists the evolution explicitly: 'hundreds of thousands to billions', starting with single Elasticsearch, then multi-region, then Postgres, then turbopuffer.; Postgres phase: pg_vector disk_ann, ts_vector (not BM25), 4,000 partitions, hash-based bin packing. P99 latencies spiked from 100ms to 20s due to partition thrashing.; Turbopuffer phase: One namespace per project, BM25, blob-based storage, median latencies improved by order of magnitude, p99 even better.
  • Caveats: Timestamps not provided for the five-generation timeline, so Ken cannot determine how long each phase lasted or how quickly Legora scaled.
  • Implications: The five-generation narrative shows that initial simple architectures (single Elasticsearch) break under multi-tenancy, data residency, and enterprise isolation requirements, and that traditional database solutions (Postgres) are not cost-effective at scale for mixed hot/cold workloads. The turbopuffer migration is the inflection point.

Turbopuffer's vector search: tree of clusters vs graph indexes

  • Claims: Vector search has two fundamental approaches: graph-based (HNSW, etc.) which navigates from the center of the graph, or tree-based (clustering) which organizes vectors into a tree of clusters.; Graph-based indexes incur 200ms p99 per S3 round trip per node traversal, making them poorly suited to object storage or disk.; Turbopuffer uses a tree of clusters: vectors are organized into clusters, then clusters of clusters, etc., forming a tree (like a B-tree on vector geometry). Root centroids stay in DRAM, leaves on SSD.; This tree approach minimizes round trips (~3 per query) and respects the memory hierarchy: hot data (root) in DRAM, cold data (leaves) on SSD/blob.
  • Evidence: Simon: 'The problem with a graph... you have to navigate from the center of the graph. And then every time you navigate through these nodes, you're doing 200 millisecond p99 to S3... fundamentally you're at odds with the fact that graph is about a random sequential trade-off that you have in memory and in registers, but not further down to memory hierarchy.'; Simon: 'Turbo Puffer then creates clusters of clusters and clusters of clusters of clusters to essentially organize all of the vector data in a tree. You can basically think of Turbo Puffer as a very complicated B-tree.'; Simon: 'This is fundamentally the cheapest way that you can run a database, period. So for something like Legora or even Web Search, which is in the hundreds of billions or tens of billions... this is fundamentally the cheapest way to do it.'
  • Implications: For Ken's agent systems, if vector search is required at scale (especially over object storage or NVMe), tree/clustering indexes (Faiss IVF, ScaNN, or turbopuffer's approach) are architecturally superior to HNSW/graph indexes. Graph indexes are optimized for in-memory random access, not storage hierarchies.

Turbopuffer's full-text search: BM25, inverted index, early termination

  • Claims: Text search is a hash map: token → set of document IDs. Query 'New York population' finds three sets, intersects them, and scores documents (York is rarer than New, so documents with York score higher).; BM25 is the scoring algorithm. The art is minimizing round trips (download dictionary, then lists), compressing lists, and early-terminating intersection once high-scoring documents dominate.; At web scale, text search is more computationally expensive and difficult than vector search, counter-intuitively.
  • Evidence: Simon: 'The way that text search works is essentially you can think of it as a hash map. You have a big document, and then you take every single one of the tokens, and you put them into the key in the hash map. The value in the hash map is some set with all of the document IDs that has that term.'; Simon: 'A document that has York in it is probably more valuable than a document that has New in it, because York is a more rare word. When people say BM25, this is the scoring that they're referring to.'; Simon: 'The art of full text search is, one, we want to minimize the number of round trips... And then the second round trip is to get these massive lists. Try to make the list as small as possible by compressing them. But also while you're doing the text search, you're trying to minimize the amount, again, of memory bandwidth... At some point... documents that just have New in it are irrelevant anymore.'; Simon: 'Counter-intuitively to most people, text search at web scale is more difficult and more computationally expensive than doing vector search.'
  • Implications: Ken should not treat BM25 as trivial. If building hybrid search, invest in real BM25 (not ts_vector, which Legora abandoned) and understand the memory bandwidth and compression optimization required. Turbopuffer's early-termination approach is gold standard.

Legal research workload: hierarchical jurisdictions, temporal validity, graph-like filtering

  • Claims: Legal research corpus is hierarchical: cities, counties, states, federal law, international law. Must respect authoritative hierarchy.; Temporal validity: judges overrule prior decisions, new regulations have exemptions or special cases of old regulations. If you find one regulation, you must find related ones.; This creates a graph-like explosion of queries. Legora fans out legal research queries across many jurisdictions and filters aggressively.; Some jurisdictions are hot (EU law, queried constantly), some are cold (Danish law, rarely queried). Turbopuffer's namespace model (hot in cache, cold on blob) handles this naturally. 500ms cold-blob latency is acceptable for deep research.
  • Evidence: Jacob: 'It's hierarchical. You have cities and you have counties and you have states and you have federal law and it's the same all around the world. And so you need to respect that authoritative hierarchy.'; Jacob: 'There's also some temporal validity. So one judge might overrule a decision that's been made somewhere else and you need to also respect that and figure that out. And then sometimes there's even a new regulation that has exemptions or special cases of an old regulation. And so if you're finding this one, you need to find all the other ones as well.'; Jacob: 'If you do a legal research query, we will fan it out to a bunch of different queries and we'll keep going... it explodes the search.'; Jacob: 'Some of them, here's an example where the EU gets queried all the time, that's super hot. And some of them, let's say Danish law, because we're Danish, no one cares really... So it doesn't really get queried. And so that can just stay on blob and that's fine. And because it's a deep research style workload, if there's 500 milliseconds latency to fetch that cold blob, that's okay.'
  • Implications: For Ken's agent systems, if retrieval involves hierarchical or graph-like filtering (e.g., multi-jurisdiction policy, regulatory compliance, temporal validity), turbopuffer's namespace model is proven. Accept higher latency for cold data in deep-research-style workflows.

Notable Concepts & Terms

  • turbopuffer: Object-storage-native vector+text search engine (S3/GCS/Azure Blob) designed around minimizing S3 round trips (~3 per query, S3 p99 ~200ms/1MB), using tree-based vector indexes (not graphs) and memory-hierarchy optimization (DRAM for root centroids, SSD for leaves, blob for cold namespaces). Per-namespace isolation enables full physical separation, CMEK, multi-region data residency.
  • namespace (turbopuffer): Turbopuffer's concept of a table, implemented as a separate S3 directory/bucket. Every namespace can be encrypted with a different key, stored in a different bucket/region, and cached independently. Legora uses one namespace per project (project search) or per jurisdiction (legal research).
  • BM25: Best Match 25, a probabilistic ranking function for full-text search that scores documents based on term frequency (TF) and inverse document frequency (IDF). Turbopuffer implements true BM25; Legora's Postgres phase used ts_vector (not BM25), which hurt retrieval quality.
  • CMEK (Customer-Managed Encryption Keys): Enterprise security requirement where the customer (e.g., bank, law firm) controls encryption keys in their own key vault. Vendor (Legora) must request access to decrypt data at rest. Customer can revoke access, rendering data unreadable. Turbopuffer supports per-namespace CMEK.
  • Full physical separation: Enterprise requirement for data to be stored in separate databases, buckets, and encryption contexts, not just logically partitioned within a shared database. Legora's Postgres phase required separate Postgres per enterprise; turbopuffer achieves this via per-namespace buckets and keys.
  • Memory hierarchy (DRAM → NVMe SSD → object storage): Simon's core design principle: data should be pushed as far down the hierarchy as economically optimal. Hot data (root centroids, frequently accessed namespaces) in DRAM, warm data on NVMe SSD, cold data (rarely queried projects/jurisdictions) on blob. Turbopuffer is architected to respect this hierarchy, unlike traditional databases.
  • Tree-based vector index (clustering): Turbopuffer organizes vectors into clusters, then clusters of clusters, forming a tree (conceptually a B-tree on vector geometry). Root centroids (always queried) stay in DRAM, leaves (document vectors) on SSD. Avoids random-access penalty of graph-based indexes (HNSW) on object storage.
  • Cache thrashing (Postgres phase): Catastrophic performance degradation when Postgres pulled many large partitions into memory, then evicted them to load others, repeatedly. Hot and cold projects coexisted in the same partitions, so queries across projects thrashed the cache, spiking p99 from 100ms to 20s.
  • Write-to-object-storage-first: Turbopuffer writes directly to S3/GCS/Azure Blob as write-ahead log (no disk replication, no Paxos), accepting hundreds of milliseconds write latency. Indexes are built asynchronously. This tradeoff is acceptable for search workloads where reads vastly outnumber writes.
  • Fan-out queries (legal research): Legora's legal research workflow fans out a single user query into many sub-queries across jurisdictions, hierarchies, and temporal relationships, then aggregates results. High QPS spikes, but turbopuffer's namespace model (cold on blob, hot in cache) handles this efficiently.

Operator Notes / Why Ken Should Care

  • Action: If Ken is building retrieval for multi-tenant agent systems with enterprise customers, adopt turbopuffer's per-namespace isolation model (separate S3 buckets/keys per tenant or project) to meet data residency, CMEK, and physical separation requirements without multiplying infrastructure.
  • Action: For vector search at scale over object storage or NVMe, use tree/clustering indexes (Faiss IVF, ScaNN, or turbopuffer's approach) rather than graph indexes (HNSW). Graph indexes are optimized for memory/register random access, not storage hierarchies.
  • Decision: If OpenClaw or agent orchestration requires hybrid search (text + vector), invest in real BM25 implementations (not Postgres ts_vector) and understand the memory bandwidth, compression, and early-termination optimization required. Turbopuffer's approach (download compressed inverted lists, intersect in memory, stop early) is gold standard.
  • Risk: Aggressive partitioning in traditional databases (Postgres, Elasticsearch) is a false solution to multi-tenancy at scale when workload heat is highly variable. Hot and cold tenants in the same partition cause cache thrashing. Blob-based storage (turbopuffer's model) naturally segregates cold data.
  • Watch: Turbopuffer's write-to-object-storage-first architecture (no disk replication, no Paxos, writes to S3 write-ahead log) trades write latency (hundreds of milliseconds) for read adaptability and cost efficiency. This is acceptable for agent systems where retrieval reads vastly outnumber writes, but not for transactional workloads (e.g., inventory reservations).
  • Decision: If security/compliance requires disabling disk cache (e.g., unencrypted NVMe is unacceptable for multi-tenant enterprises), turbopuffer's memory-only performance can still be sufficient. Test this configuration for OpenClaw if disk encryption is required.
  • Action: For agent workflows involving hierarchical or graph-like filtering (e.g., multi-jurisdiction policy retrieval, regulatory compliance, temporal validity), adopt turbopuffer's namespace model (cold on blob, hot in cache) and accept higher latency (e.g., 500ms) for cold data in deep-research-style workflows.
  • Watch: Simon's claim that 'text search at web scale is more difficult and more computationally expensive than vector search' is counterintuitive but important. Do not underestimate BM25 complexity when designing hybrid retrieval for OpenClaw.

Source/Metadata

  • Title: Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora
  • Transcript words: 6230
  • Duration seconds: 1236
  • Timestamp note: Timestamps not present in transcript; unable to provide MM:SS/HH:MM:SS navigation.
Full transcript 3764 words · 28 min read
0:00

Okay. Hi, everyone. Welcome to this 20-minute talk about connecting AI to loads of legal documents.

0:12

My name is Jacob. I'm an engineer at Legora. Yeah, and I've got to step into frame here. I'm Simon. I'm the CEO and co-founder of Turbo Puffer, a search engine that we work with Legora and others on. Super quickly, introduction to Legora. We're a collaborative AI platform for legal work. And so that means we have law firms that are clients and we have in-house legal teams that are clients. And they use Legora to do reviews of contracts. They use it to go through an absurd amount of contracts and make sure that they all look good. They use them to create new contracts. They do legal research, which means looking over all potential law. And they collaborate inside Legora.

1:00

So you can think of Legora as a linear slash figma slash notion slash GitHub for legal work. It's a lot. We are one of the fastest growing companies right now. We work extremely fast.

1:17

Yeah. Tons of numbers on the screen. I'll just skip through that. What we really want to talk about is search today. So at Legora, there's two types of search that we do. There is project search and legal research. Project search is basically projects in Legora is like the unit of work that you have. So if, let's say, you are SpaceX and you want to acquire Cursor, then that would be one project with your law firm. And so they would go into Legora and they would upload all these documents. And the law firm that helped you would go through all of the employment agreements, all of the contracts with suppliers.

1:50

I know Cursor is using TurboPuffer, so maybe there's a contract there they'd look at. But basically you do the search confined to a project. And projects can be tens of documents to millions of documents. The other use case is legal research. And legal research is a deep research style workload where we'll search across tons of laws, previous cases, regulations, et cetera. And people use this to answer questions such as how do we handle this specific thing? And they'll also use it for litigation. Maybe they want to sue someone or maybe they are getting sued and they'll use legal research to support and help their case. So if we start at number one, project search.

2:29

We've been through a little bit of a ride here on how we do search. Starting at hundreds of thousands of documents all the way into two billions of documents. And we've tried a lot of different things. So first we started with the very simple one, which is just a single elasticsearch clause. And we started with the cluster for all of our search workflows. That worked relatively well. It was a simple setup. All of the tenants, our clients, our users, would be on one big blob storage where we store the raw documents. And on one big elasticsearch where we would do all of their searching, the indexing and the searching. Super simple. Worked relatively well initially.

3:15

Then we wanted to enter the land of the free. And we got some new requirements. Americans only want processing to happen within the US. And Europeans only want it to happen within the EU. And Australians only want it to happen within Australia. And so we had to basically move to multiple elasticsearch instances. And what we actually did was we took the entire setup and we just iterated over the set. That is EU, US, and Asia Pacific. And so we just had this multiplied by three or four. Annoying. Lots of overhead. But it got us to where we needed to be. Then the next iteration of the story is enterprise. So really big banks, the biggest law firms in the world.

3:51

They have really annoying requirements. And number one they have is they'll ask for full physical isolation of all of their data.

3:57

There's probably a little bit of a debate on what physical isolation actually means. But essentially it means they want their own database. They also want customer managed encryption keys. And what that means is they basically have a key vault thing where they have an encryption key. And they give us access to read the key. And we then use that key to encrypt and decrypt all of their data at rest. And what that gives them is they can just revoke our access to their key. And then we can't decrypt their data anymore and so it's safe. And so in a way that gives enterprises a lot of control over all of their data. Because they control the keys to reading it.

4:33

So we moved from Elasticsearch to Postgres. And I imagine a bunch of you guys are wondering why would you ever put your vectors into Postgres? Yes. It actually works surprisingly well.

4:47

And the reason that we did this was we were already using Postgres for OLTP workloads. And so we already had to do this split of multiple Postgres and multiple blobs. And so it was really easy for us to try to shift all of our search into Postgres as well because then we only have one system. So the setup here was pg vector, specifically disk ANN, TS vector for the search, not BM25, which meant we lost a little bit of retrieval performance there. And what we'd do is we would partition the table where we would store all the document chunks. We'd partition it aggressively, like 4,000 partitions.

5:09

And then each project, we'd basically hash the project key and we'd bin pack them into the partitions. That actually worked relatively well. But it was expensive. And search performance wasn't super good. And what happened was when we scaled a lot, everything just broke and exploded. And so what happened was you can imagine that you have a bunch of projects. And some of them, you spin up a project, you work on it, and then you close it and you basically never go back to it again. And we have a bunch of those where they never get queried and we have a bunch that get queried all the time because they're super active projects.

5:39

And when we pack them into partitions, the cold ones and the hot ones would land on the same ones and the partitions would get really big. And so when we queried them, Postgres would pull the partition, put it into memory, we'd do the stuff and then we'd query another partition and another partition and it would essentially thrash the cache all the time. And what that meant was our latencies would spike. So we went from search and ingestion p99 of 100 milliseconds into 20 seconds, which you can imagine is a really bad user experience. So then we went to Turbo Puffer at about 400 million documents. And what we did with Turbo Puffer was we did one namespace per project.

6:04

And the advantages of moving to Turbo Puffer is we got BM25, real BM25, much better relevancy, much better latencies, and it was extremely simple to operate because we could just have a single Turbo Puffer cluster. We didn't have to have a bunch of different ones like with Postgres and it would, since it's blob based, it could just query the blobs that we had anyway. And much lower cost and it was extremely simple to operate and we didn't have this problem with the partitions because if a project is not used, it's just in blob. And so it's really easy. And Simon can talk a bit more about why that works so well.

6:20

Yeah, so Legora has some of, and legal in general has, by the way, if Jake and I have similar accents and maybe even look a bit similar, it's because we're both Danish. Turbo Puffer has a particular architecture that supports these kinds of very regulated environments really well. But in order to understand that, we have to understand what kind of search engine is Turbo Puffer. Why is it different than the ones that they used in the past? Since the very beginning of Turbo Puffer, the design has more or less been the same.

6:45

There may be changes in the future, but the design has stood the test of time. When you do a write to Turbo Puffer, we write directly to object storage. There is no disk replication. There's no paxos. There's none of that. Direct to S3. That's the fundamental tradeoff in Turbo Puffer, right? Hundreds of milliseconds. If you're Shopify and doing inventory reservations for a Kylie Jenner flash sale, not going to work. Very good for search.

7:28

Because generally when you're doing search, doing a slow write is fine as long as the read performance is adaptable and good. So that's what happens on write. It just goes into write-ahead log. You can imagine you write 1.json, 2.json, 3.json. Obviously, it's a database, so it's not JSON. But for illustrative purposes, that's what happens. And in the background, we build the vector indexes, the text indexes, the columnar indexes, and so on to satisfy the queries that Jake and other customers have. So then at query time, we can go in and then the query reaches some namespace. And namespace is our concept of a table.

8:09

Direct to S3. That's the fundamental tradeoff in Turbo Puffer, right? Hundreds of milliseconds. If you're Shopify and doing inventory reservations for a Kylie Jenner flash sale, not going to work. Very, very good for search. Because generally when you're doing search, doing a slow write is fine as long as the read performance is adaptable and good.

8:11

So that's what happens on write. It just goes into write-ahead log. You can imagine you write 1.json, 2.json, 3.json. Obviously, it's a database, so it's not JSON. But for illustrative purposes, that's what happens. And in the background, we build the vector indexes, the text indexes, the columnar indexes, and so on to satisfy the queries that Jake and other customers have.

8:12

So then at query time, we can go in and then the query reaches some namespace. And namespace is our concept of a table. You can think of it as a directory on S3 that's isolated from everything else. We go to the node that is most likely to have it. It could go to any node, right? It could go to every single node, and they're all read replicas. But it would go with some affinity to the node that has the highest probability of having it in cache.

8:14

We check the memory cache for any objects, NVMe, SSD cache, and then finally to object storage. Everything in Turbo Puffer is optimized around doing as much work in as few round trips as possible, right? S3 has a P99 on a one megabyte blob size of around 200 milliseconds, so you want to do as few round trips as possible, right? Ideally, you do around three. And everything in Turbo Puffer in the database is, I'm going to need your fingerprint. You got it. Everything in Turbo Puffer is designed around minimizing the number of round trips. This is also amazing for modern disks. If you do a lot of concurrency in few round trips, you utilize them optimally, and everything in Turbo Puffer is designed around this.

8:17

So why is this so good for a company like Legora? Well, object storage native, if you design it around the atomic unit of separation being the namespace or the table, every single table could be encrypted with a different key. Every single namespace could be in a different bucket. We have customers that have thousands of buckets that they have namespaces in so that their customers get the warm IT fuzzies of having the bucket in their own cloud account. They can also be encrypted with their own keys. You can share buckets. You can do whatever configuration that you need at the namespace level. You can reencrypt with different keys. You can move them around, and you can reencrypt with other keys.

8:18

For Legora in particular, this was really important for this full physical separation, right? An encryption separation. All of the namespaces needed to be physically at rest with different keys and as separated as possible. S3, GCS, Azure Blob Storage, they passed that. And the other parts of the hierarchy also. Except the NVMe SSD cache, because in the SSD cache we consider that to be volatile like memory. But your customers did not. So what we did was that we thought we were going to implement encryption into the disk cache. But instead we just disabled the disk cache and saw how it fared. And the performance of Turbo Puffer even without the disk cache, but just the memory cache, was so good that we just kept it that way for some of the Legora workloads where we couldn't have the disk cache for multi-tenancy. Turbo Puffer will support that in the future. But it just goes to show the natural point where Turbo Puffer allows these encryption and storage and separation to become fully multi-tenancy native. I'll hand it back to you on what happened then.

8:24

And then, drum roll please, latencies look like this. Is my mic working? No? Could I, or I'll start screaming really loudly. It speaks for itself if you can't hear me. Okay. Latencies improved in order and magnitude basically. And these are median latencies, so P99 were even better. So obviously this is a huge thing when you're doing, I mean one thing is if you're doing a single rack style thing. But if you have an agent that does 20 queries, 100 queries, these really add up. So that was on the project side.

8:25

And then a more recent thing is legal research. So legal research is a difficult problem. And the reason it's difficult is that the corpus is extremely big. So we're racing towards 10 billion vectors and we're growing extremely fast. We also have quite high read. So QPS can spike a lot because we do a lot of fan out. If you do a legal research query, we will fan it out to a bunch of different queries and we'll keep going. And the reason we do that is we need this heavy filtering because essentially it's a graph for a few different reasons.

8:32

Firstly, it's hierarchical. You have cities and you have counties and you have states and you have federal law and it's the same all around the world. And so you need to respect that authoritative hierarchy. There's also some temporal validity. So one judge might overrule a decision that's been made somewhere else and you need to also respect that and figure that out. And then sometimes there's even a new regulation that has exemptions or special cases of an old regulation. And so if you're finding this one, you need to find all the other ones as well. So you can imagine that it explodes the search.

8:33

And so we started on Elasticsearch for this, but also moving to Turbo Puffer. And Elasticsearch just got extremely expensive because we have to have everything there. But with Turbo Puffer, we can basically take different jurisdictions and we can make them namespaces in Turbo Puffer. And that means some of them, here's an example where the EU gets queried all the time, that's super hot. And some of them, let's say Danish law, because we're Danish, no one cares really. It's such a small country. So it doesn't really get queried. And so that can just stay on blob and that's fine. And because it's a deep research style workload, if there's 500 milliseconds latency to fetch that cold blob, that's okay. It's fine. It's not really a big problem.

8:35

So the way that Turbo Puffer is designed lends itself super well to this super long scale of cold, weird namespaces and a few that are really, really hot. And Simon wants to talk more about that. Yeah. So I was talking about why the company is called Turbo Puffer, another talker earlier today. But one of the other explanations of the name of Turbo Puffer is that it's about puffing into the different memory hierarchies and really mastering when data should be in particular memory hierarchies.

8:43

So you can think about it here, right, of something the EU law might be more or less part of almost every one of the legal research queries, right? So that probably sits closer to NVMe SSDs than memory, right? The economics change as you move up and down this hierarchy. In memory, you want things that are queried a lot, right? Then the economics of memory are great. NVMe SSDs, you can do a lot of things directly on them. But the economics change as you move up and down this boundary, the latency changes, and the way that the database is architected to take advantage of it in terms of round trips versus random versus sequential, all changes as you navigate this hierarchy.

8:53

Turbo Puffer is a database that is really designed around the memory hierarchy, and all of the smarts in Turbo Puffer is that all of these namespaces are puffed in and out of the cache. You can think of this as we want to spend as much time, have as much data pushed as far down in this hierarchy as possible to get the best performance cost ratios.

8:59

So how does that apply to search? Well, for something like vector search, for example, there's two fundamental ways to do vector search. One is to navigate it, basically design a graph. The problem with a graph on something like object storage or disk, again, we want to have things as far down that memory hierarchy as possible. The problem with a graph, this is not a graph, this is a tree, but in a graph, you have to navigate from the center of the graph. And then every time you navigate through these nodes, you're doing 200 millisecond P99 to S3, right? And so you're trying to shrink the diameter of the graph, you're trying to do all these tricks to make the graph, but fundamentally you're at odds with the fact that graph is about a random sequential trade-off that you have in memory and in registers, but not further down to memory hierarchy.

9:02

The way Turbo Puffer does it is organize it into clusters, right? Vectors you can think of in two dimensions just as points in a massive coordinate system, and we can organize them into clusters. Turbo Puffer then creates clusters of clusters and clusters of clusters of clusters to essentially organize all of the vector data in a tree. You can basically think of Turbo Puffer as a very complicated B-tree, right?

9:06

And then every time you navigate through these nodes, you're doing 200 millisecond p99 to S3, right? And so you're trying to shrink the diameter of the graph, you're trying to do all these tricks to make the graph, but fundamentally you're at odds with the fact that graph is about a random sequential trade-off that you have in memory and in registers, but not further down to memory hierarchy. The way Turbo Puffer does it is organize it into clusters, right? Vectors you can think of in two dimensions just as point is a massive coordinate system, and we can organize them into clusters. Turbo Puffer then creates clusters of clusters and clusters of clusters of clusters to essentially organize all of the vector data in a tree. You can basically think of Turbo Puffer as a very complicated B-tree, right? Because it's a tree on this geometry of this entire space and the clustering of it in an approximate way. Now, the root centroids further up the tree, you can imagine, are part of every single time you search, right? We're always trying to figure out which clusters that we're in, and we're always looking at the upper levels of the tree. So they're going to be further up the memory hierarchy, right? Closer to the registers, almost all in DRAM. Now, the leaves that have all of the actual legal cases or whatever long document it could be, it could be images, all of that is probably going to be on SSDs with that single one millisecond roundtrip at the end. It doesn't make sense to have all that puffed into DRAM. This is fundamentally the cheapest way that you can run a database, period. So for something like Legora or even Web Search, which is in the hundreds of billions or tens of billions, depending on how much of the web you've scraped, this is fundamentally the cheapest way to do it, and we have customers that are indexing massive parts of the entire web into Turbo Puffer, which is really also a part of what legal research is. Full text is also really respectful of the memory hierarchies. The way that text search works is essentially you can think of it as a hash map. You have a big document, and then you take every single one of the tokens, and you put them into the key in the hash map. The value in the hash map is some set with all of the document IDs that has that term. So then if you search for New York population, you're finding those three places in the hash map, and then you're taking the three sets and doing an intersect on the sets. While you're intersecting, you're also trying to do some kind of scoring, right? A document that has York in it is probably more valuable than a document that has New in it, because York is a more rare word. When people say BM25, this is the scoring that they're referring to. The art of full text search is, one, we want to minimize the number of round trips. So first you download the parts of the dictionary that are relevant, round trip one, maybe a round trip one before that to index into where the parts of the terms are. And then the second round trip is to get these massive lists. Try to make the list as small as possible by compressing them. But also while you're doing the text search, you're trying to minimize the amount, again, of memory bandwidth that you want to intersect these lists. You can probably imagine that at some point there's a point where you've seen so many documents with population in York that have much higher scores, that documents that just have New in it are irrelevant anymore. This is a mega crash course in how text search works. And counter-intuitively to most people, text search at web scale is more difficult and more computationally expensive than doing vector search. I'll hand it over to you.

9:08

Cool. So key learnings from what you heard today, retrieval is extremely important to Legora. It's key to legal reasoning. Turbo Puffer really excels when, for us, because it makes it extremely easy to operate. We have 70 plus tenants. We have 100. We have 200 tenants. If we had to have separate Elastic Search databases for each of these, it would be just hell. But we can do this natively with Turbo Puffer with data residency and CMEG, et cetera. And then it's extremely cost efficient generally when you have these types of workflows or workloads that we do where there's a long tail of cold indices, basically, that you don't need to query so much. And you're okay paying the small latency cost for it. So now with Turbo Puffer and four seconds to go, now we can focus on making Legora. We can focus on the product, making it really great, and not on scalability and infra. And also David, our CFO, is really happy about the cost. So it's great. Thanks, everyone.

9:09

It could go to every single node, and they're all read replicas. But it would be go with some affinity to the node that has the highest probability of having it in cache. We check the memory cache for any objects, NVMe, SSD cache, and then finally to object storage. Everything in Turbo Puffer is optimized around doing as much work in as few round trips as possible, right? S3 has a P99 on a, like, one megabyte blob size of around 200 milliseconds, so you want to do as few round trips as possible, right? Ideally, you do around three. And everything in Turbo Puffer in the database is, oh, I'm going to need your fingerprint. You got it.

9:42

Everything in Turbo Puffer is designed around minimizing the number of round trips. This is also amazing for modern disks. If you do a lot of concurrency in few round trips, you utilize them optimally, and everything in Turbo Puffer is designed around this. So why is this so good for a company like Legora? Well, Oblix Storage Native, if you design it around the atomic unit of separation being the namespace or the table, every single table could be encrypted with a different key. Every single namespace could be in a different bucket.

10:11

We have customers that have thousands of buckets that they have namespaces in so that their customers get the warm IT fuzzies of having the bucket in their own cloud account. They can also be encrypted with their own keys. You can share buckets. You can do whatever configuration that you need at the namespace level. You can reencrypt with different keys. You can move them around, and you can reencrypt with other keys. For Legora in particular, this was really important for this full physical separation, right? An encryption separation. All of the namespaces needed to be physically at rest with different keys and as separated as possible.

10:48

S3, GCS, Azure Bobstories, they passed that. And the other parts of the hierarchy also. Except the NVMe SSD cache, because in the SSD cache we consider that to be volatile like memory. But your customers did not. So what we did was that we thought we were going to implement encryption into the disk cache. But instead we just disabled the disk cache and saw how it fared. And the performance of TurboPuffer even without the disk cache, but just the memory cache, was so good that we just kept it that way for some of the Legora workloads where we couldn't have the disk cache for multi-tenancy. TurboPuffer will support that in the future.

11:24

But it just goes to show the natural point where TurboPuffer allows these encryption and storage and separation to become fully multi-tenancy native. I'll hand it back to you on what happened then. And then, drum roll please, latencies look like this. Is my mic working? No? Could I, or I'll start screaming really loudly. It speaks for itself if you can't hear me. Okay. Latencies improved in order and magnitude basically. And these are median latencies, so P99 were even better. So obviously this is a huge thing when you're doing, I mean one thing is if you're doing a single sort of rack style thing.

12:06

But if you have an agent that does 20 queries, 100 queries, these really, really add up. So that was on the project side. And then a more recent thing is legal research. So legal research is a kind of a difficult problem. And the reason it's difficult is that the corpus is extremely big. So we're racing towards 10 billion vectors and we're growing extremely fast. We also have quite high read. So QPS can spike a lot because we do a lot of fan out. Like if you do a sort of legal research query, we will fan it out to a bunch of different queries and we'll keep going.

12:44

And the reason we do that is we need this heavy filtering because essentially it's kind of like a graph for a few different reasons. Firstly, it's hierarchical. You know, you have cities and you have counties and you have states and you have federal law and it's the same all around the world. And so you need to respect that authoritative sort of hierarchy. There's also some temporal validity. So one judge might overrule a decision that's been made somewhere else and you need to also respect that and figure that out. And then sometimes there's even like a new regulation that has exemptions or special cases of an old regulation.

13:18

And so if you're finding this one, you need to find all the other ones as well. So you can imagine that it sort of explodes the search. And so we started on Elasticsearch for this, but also moving to Turbo Puffer. And Elasticsearch just got extremely expensive because we have to have everything there. But with Turbo Puffer, we can basically take different jurisdictions and we can make them namespaces in Turbo Puffer. And that means some of them, here's an example where like you have the EU that gets queried all the time, that's super hot. And some of them, let's say Danish law, because we're Danish, no one cares really. It's such a small country.

13:53

So like it doesn't really get queried. And so that can just stay on blob and that's fine. And because it's sort of a deep research style workload, if there's 500 milliseconds latency to fetch that cold blob, that's okay. That's fine. It's not really a big problem. So the way that Turbo Puffer is designed lends itself super well to this super long scale of like cold, weird namespaces and a few that are really, really hot. And Simon wants to talk more about that. Yeah. So I was talking about why the company is called Turbo Puffer, another talker earlier today.

14:28

But one of the other explanations of the name of Turbo Puffer is that it's about puffing into the different memory hierarchies and really mastering when data should be in particular memory hierarchies. So you can think about it here, right, of something like the EU law might be more or less part of almost every one of the legal research queries, right? So that probably sits closer to NVMe SSDs than memory, right? The economics kind of change as you move up and down this hierarchy. In memory, you want things that are queried a lot, right? Then the economics of memory are great. NVMe SSDs, you can do a lot of things directly on them.

15:05

But the economics change as you move up and down this boundary, the latency changes, and the way that the database is architected to take advantage of it in terms of round trips versus random versus sequential, all changes as you navigate this hierarchy. Turbo Puffer is a database that is really designed around the memory hierarchy, and all of the smarts in Turbo Puffer is that all of these namespaces are puffed in and out of the cache. You can think of this as we want to spend as much time, have as much data pushed as far down in this hierarchy as possible to get the best performance cost ratios. So how does that apply to search?

15:37

Well, for something like vector search, for example, there's two fundamental ways to do vector search. One is to navigate it, basically design a graph. The problem with a graph on something like object storage or disk, again, we want to have things as far down that memory hierarchy as possible. The problem with a graph, this is not a graph, this is a tree, but in a graph, you have to navigate from the center of the graph. And then every time you navigate through these nodes, you're doing 200 millisecond p99 to S3, right? And so you're trying to shrink the diameter of the graph, you're trying to do all these tricks to make the graph,

16:10

but fundamentally you're at odds with the fact that graph is about a random sequential trade-off that you have in memory and in registers, but not further down to memory hierarchy. The way Turbo Puffer does it is organize it into clusters, right? Vectors you can think of in two dimensions just as point is a massive coordinate system, and we can organize them into clusters. Turbo Puffer then creates clusters of clusters and clusters of clusters of clusters to essentially organize all of the vector data in a tree. You can basically think of Turbo Puffer as a very, very complicated B-tree, right?

16:42

Because it's a tree on this geometry of this entire space and the clustering of it in an approximate way. Now, the root centroids further up the tree, you can imagine, are part of every single time you search, right? We're always trying to figure out which clusters that we're in, and we're always looking at the upper levels of the tree. So they're going to be further up the memory hierarchy, right? Closer to the registers, almost all in DRAM. Now, the leaves that have all of the actual legal cases or whatever long document it could be, it could be images, all of that is probably going to be on SSDs with that single one millisecond roundtrip at the end.

17:15

It doesn't make sense to have all that puffed into DRAM. This is fundamentally the cheapest way that you can run a database, period. So for something like Legora or even Web Search, which is in the hundreds of billions or tens of billions, depending on how much of the web you've scraped, this is fundamentally the cheapest way to do it, and we have customers that are indexing massive parts of the entire web into Turbo Puffer, which is really also a part of what legal research is. Full text is also really respectful of the memory hierarchies. The way that text search works is essentially you can think of it as a hash map.

17:49

You have a big document, and then you take every single one of the tokens, and you put them into the key in the hash map. The value in the hash map is some set with all of the document IDs that has that term. So then if you search for New York population, you're finding those three places in the hash map, and then you're taking the three sets and doing an intersect on the sets. While you're intersecting, you're also trying to do some kind of scoring, right? A document that has York in it is probably more valuable than a document that has New in it, because York is a more rare word. When people say BM25, this is the scoring that they're referring to.

18:24

The art of full text search is, one, we want to minimize the number of round trips. So first you download the parts of the dictionary that are relevant, round trip one, maybe a round trip one before that to index into where the parts of the terms are. And then the second round trip is to get these massive lists. Try to make the list as small as possible by compressing them. But also while you're doing the text search, you're trying to minimize the amount, again, of memory bandwidth that you want to intersect these lists. You can probably imagine that at some point there's a point where you've seen so many documents with population in York

18:57

that have much higher scores, that documents that just have New in it are irrelevant anymore. This is like a mega crash course in how text search works. And counter-intuitively to most people, text search at web scale is more difficult and more computationally expensive than doing vector search. I'll hand it over to you. Cool. So key learnings from what you heard today, retrieval is extremely important to Legora. It's key to legal reasoning. Turbo Puffer really excels when, for us, because it makes it extremely easy to operate. We have 70 plus tenants. We have 100. We have 200 tenants.

19:37

You know, if we had to have separate Elastic Search databases for each of these, it would be just hell. But we can do this natively with Turbo Puffer with data residency and CMEG, et cetera, et cetera. And then it's extremely cost efficient generally when you have these types of workflows or workloads that we do where there's a long tail of cold indices, basically, that you don't need to query so much. And you're okay paying the small latency cost for it. So now with Turbo Puffer and four seconds to go, now we can focus on making Legora. We can focus on the product, making it really, really great, and not on scalability and infra.

20:11

And also David, our CFO, is really happy about the cost. So it's great. Thanks, everyone.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note