Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA
Description
Ask your own company what it actually contracted for with a given vendor and nobody can tell you. The number exists, negotiated carefully, sitting in a pricing table inside a signed PDF no system ever read back. Hiral Shah puts a figure on the aggregate: roughly two trillion dollars of negotiated value locked in agreements that organizations never return to, because recovering it means human reading, human review, and manual work across disconnected systems. The scale on Docusign's side is its own engineering problem. Close to two million paying customers, a billion users, and about a million agreements flowing through each day, all of it needing to come out structured, queryable, and useful. Agreements also refuse to be flat. One governs another, which amends a third, so answering a simple question often means traversing twenty years of a business. The specific thing that breaks is tables. Pricing tiers, rate cards, SKUs and service levels are exactly the terms people need, and they are exactly what generic extraction handles worst, because reading a page line by line destroys a merged cell or a nested column. Shah and Sean Sodha describe the model they built together for that one job, an unusually small vision language model at roughly 900 million parameters, designed as an extractor rather than a generator. It replaces a stack of separate layout and table models with a single pass that returns reading order, semantic structure and preserved tables. Their headline lesson is about restraint: a purpose built small model ran table extraction around twenty times faster than the general alternatives they tested, and lower context meant lower latency and lower cost at their volume. Speaker info: - https://www.linkedin.com/in/shahhiral/ - https://www.linkedin.com/in/sean-sodha/ Timestamps: 0:00 - Why agreement data is an engineering problem 1:42 - A million agreements a day 2:19 - Two trillion dollars nobody goes back for 3:59 - Why tables break generic extraction 5:1
Summary
Generated by claude-sonnet-4-5-20250929At-a-Glance
- Verdict: Watch fully
- Core thesis: DocuSign partnered with NVIDIA to solve petabyte-scale agreement data extraction using a purpose-built 850M-parameter VLM (NemoTron Parse) that extracts complex tables from unstructured PDFs 20x faster than open-source alternatives, enabling queryable structured data for downstream enterprise systems.
- Why it matters: Demonstrates production-scale document processing architecture with concrete model selection tradeoffs (purpose-built vs. generic LLM, FP16 vs. quantization, batch OCR vs. on-demand), real performance benchmarks, and hybrid pipeline design for handling billions of documents—directly relevant to Ken's agent orchestration, data ingestion, and control plane work.
- Best use: Study for document processing architecture patterns, model efficiency tradeoffs, hybrid OCR/VLM pipeline design, and scale considerations when building agent systems that ingest unstructured enterprise data.
Executive Summary
Hiral Shah (DocuSign) and Sean Sodha (NVIDIA) presented their partnership solving agreement data extraction at massive scale: DocuSign processes 1 million agreements daily for 1.9 billion users, with $2 trillion in negotiated value trapped in unstructured PDFs. Traditional OCR and generic LLMs fail on complex tables with merged cells, nested structures, and pricing tiers because they read text line-by-line, breaking semantic context. DocuSign built Intelligent Agreement Management (IAM) using NVIDIA's NemoTron Parse, a purpose-built 850M-parameter VLM that extracts layout, reading order, and table structure in a single shot.
The technical architecture uses a hybrid approach: NemoTron Parse handles table extraction while separate OCR modules process clauses and metadata fields. The team prioritized accuracy first, then performance optimization via quantization (FP16 → FP8/FP4), multi-token generation, and encoder-decoder techniques. The model runs 20x faster than open-source alternatives on table extraction benchmarks (RD TableBench), with latency/cost benefits from lower context windows. The system pre-processes documents at ingest (batch OCR for high-throughput queryability) rather than on-demand extraction, storing results in a proprietary agreement data model that feeds downstream procurement/finance systems like Coupa.
Key architectural decisions include: (1) Purpose-built models for specific jobs rather than general LLMs; (2) Upfront compute investment for batch processing to enable low-latency retrieval later; (3) Hybrid pipelines splitting table extraction (NemoTron Parse) from text/metadata extraction (OCR); (4) API and CSV export for downstream integration. The demo showed Agreement Manager uploading an order form, extracting metadata and pricing tables in seconds, and exporting structured data. Future roadmap includes NemoRetriever for document search (find the right document before extracting data) and NVIDIA agent toolkit integration for production-scale agents.
Key Takeaways
- Claim: Generic LLMs and traditional OCR fail on enterprise agreement tables because they read text line-by-line, breaking semantic structure in merged cells, nested tables, and pricing tiers. | Evidence: DocuSign tried generic LLMs and traditional extraction tools on complex order forms with pricing tables, SLAs, and rate cards; all failed because line-by-line reading destroyed table context, requiring manual legal/procurement team intervention. | Implication: For agent systems ingesting structured data from PDFs (pricing, specs, SLAs), Ken should use purpose-built VLMs or extraction models rather than general-purpose LLMs; the semantic layout understanding matters more than raw text extraction.
- Claim: NemoTron Parse is an 850-900M parameter VLM designed as a single-shot extraction model that outputs semantic layout, text, reading order, and preserved table structure without generation. | Evidence: Sean described it as a 'tiny C-radio VLM' at ~850M parameters, serving via NVIDIA NIM or vLLM; it extracts rather than generates, designed for all-in-one document processing without separate YOLOX models for table/page elements. | Implication: Ken should consider small, purpose-built VLMs (<1B params) for agent data ingestion pipelines when accuracy and structure preservation matter more than generative capabilities; these can run efficiently on VLLM without GPU overkill. | Caveat: Model currently runs on FP16; quantization to FP8/FP4 and multi-token generation optimizations are planned for next few months to further improve performance.
- Claim: NemoTron Parse extracted tables 20x faster than open-source alternatives in DocuSign's benchmarks, with lower latency and cost due to smaller context windows and parameter count. | Evidence: Hiral stated 'we ran this against a lot of the other open source models. When you think about how many tables can it extract per second, NemoTron was 20x faster'; Sean added that lower parameter count and FP16 precision enabled cost/latency wins at DocuSign's scale (millions/billions of docs). | Implication: For Ken's agent control planes processing high-volume document ingestion, prioritize model efficiency metrics (tables extracted/sec, tokens/sec, cost per extraction) over general leaderboard scores; 20x throughput gains enable real-time or near-real-time data availability.
- Claim: DocuSign uses a hybrid pipeline: NemoTron Parse for table extraction and separate OCR modules for clauses/metadata, with upfront batch processing to enable low-latency downstream queries. | Evidence: Hiral explained 'we do a lot of hybrid approach and purpose built for the needs... from a table piece, [NemoTron] does break down the whole layout along with extracting. We still do OCR from a lot of other fields and metadata and the clauses.' Sean added 'you are heavy on the compute at the upfront side with all the OCR so you don't have to worry about it later on.' | Implication: Ken should architect agent ingestion pipelines with separate modules for different data types (tables, text, metadata) rather than one-size-fits-all models; pre-process high-volume data at ingest time to enable fast retrieval later, but support on-demand extraction for interactive use cases. | Caveat: Batch processing requires upfront compute investment and storage for pre-processed results; not suitable for low-latency ad-hoc document uploads (e.g., Q&A on a single newly uploaded contract).
- Claim: DocuSign's IAM stores extracted data in a proprietary agreement data model and exposes it via API/CSV for downstream procurement, finance, and sales systems like Coupa. | Evidence: Demo showed Agreement Manager extracting order details, pricing tables, and metadata into a structured format downloadable as CSV; Hiral noted 'businesses want to use this data to do a lot of downstream work... a procurement team wants to put it into Coupa to make sure when I'm paying that works.' | Implication: Agent orchestration systems should output structured data in formats/APIs that integrate with enterprise systems of record (ERPs, CRMs, procurement tools); extraction is only valuable if it feeds operational workflows, not just search/Q&A.
- Claim: Purpose-built models for specific jobs (table extraction, document retrieval, metadata extraction) outperform general LLMs for enterprise scale and accuracy, enabling faster time-to-market. | Evidence: Hiral stated 'a purpose-built model for the job you're trying to do is a big, big part of how we've been thinking about, and that's where we've been able to accelerate, bring things to market much faster.' Sean added that NemoTron Retriever builds specialized embedding, re-ranking, and extraction models rather than using general VLMs for all tasks. | Implication: For Ken's OpenClaw and agent systems, compose specialized models (embedding for retrieval, small VLM for extraction, LLM for reasoning) rather than defaulting to frontier models for every task; this reduces cost, latency, and time-to-production.
- Claim: DocuSign and NVIDIA are extending the partnership to include NemoRetriever (finding the right document before extraction) and NVIDIA agent toolkit for production-scale agents. | Evidence: Sean described next steps: 'we started with how do I extract as much possible information from a page, and now we'll scale to how do I now find that page to begin with... we'll talk about the NVIDIA agent toolkit with them over the next few months, and then actually start scaling into more production-scale agents.' | Implication: Agent system builders should plan for multi-stage pipelines: document retrieval → page selection → extraction → reasoning; retrieval/re-ranking models are critical before extraction models, especially at petabyte scale. | Caveat: These capabilities are roadmap items (next few months), not yet in production at DocuSign's scale.
Detailed Brief
Architecture and Scale Details
- Claims: DocuSign processes 1 million agreements daily for 1.9 billion users across 1.9 million paying customers.; Deloitte study found $2 trillion in negotiated value trapped in agreements due to human reading, disconnected systems, and manual workflows.; Agreements are hierarchical (one agreement governs others), requiring multi-document queries to answer simple questions like 'what did we contract for total tokens with Claude?'; Agreement Manager serves as a central repository with AI-powered extraction, search, and API access for downstream systems.
- Evidence: Hiral showed demo uploading an order form, processing in seconds, extracting pricing tables and metadata, and exporting to CSV.; Sean noted NemoTron Retriever leads leaderboards (Vidori V1/V2/V3, MTab, MMTab) and open-sources models, techniques, blueprints, and agent skills.; NemoTron Parse model is ~850M params, runs on FP16, serves via NVIDIA NIM or vLLM, and is designed for single-shot extraction without generation.
- Caveats: Batch processing requires upfront compute and storage; not optimal for real-time ad-hoc document uploads (e.g., single contract Q&A).; Model quantization (FP8/FP4) and multi-token generation are planned but not yet deployed; current accuracy focus precedes performance push.
- Implications: Ken should design agent ingestion with batch processing for high-throughput queryability and on-demand extraction for interactive use cases, choosing based on latency vs. throughput needs.; Smaller, specialized models (sub-1B params) can achieve production-scale performance if purpose-built and quantized; avoid default reliance on multi-billion parameter VLMs.
Model Selection and Optimization Tradeoffs
- Claims: DocuSign prioritized accuracy before performance: 'Right now we're focusing on the accuracy side. Are we adding value to the system? And then from there, we'll push out that Pareto curve on the performance side.'; Quantization paths (FP16 → FP8 → FP4), multi-token generation, and encoder-decoder optimizations are planned for next phase.; Lower context windows and parameter count reduced latency and cost at DocuSign's scale.; NemoTron is open-source with published datasets, techniques, quantization/distillation/pruning approaches, and blueprints.
- Evidence: Sean described Pareto curve strategy: accuracy first, then performance optimization via quantization, multi-token generation, and encoder-decoder techniques.; DocuSign tested multiple open-source models; NemoTron Parse was 20x faster on table extraction (tables extracted/sec metric).; Sean noted 'we're trying to move everyone to Blackwell' for performance gains; model currently runs on FP16 but can be further optimized.
- Caveats: Multi-token generation and FP4 quantization not yet deployed; current performance is FP16 baseline.; Encoder-decoder optimizations are VLM-specific; not all techniques apply to other model architectures.
- Implications: Ken should adopt phased optimization: deploy for accuracy first, then iterate on quantization and performance after validating value; avoid premature optimization.; Use open-source models with published optimization blueprints (like NemoTron) to accelerate production deployment and cost reduction.; Benchmark models on task-specific metrics (e.g., tables extracted/sec) rather than general leaderboard scores when building agent pipelines.
Batch vs. On-Demand Processing Tradeoffs
- Claims: DocuSign uses batch processing for high-throughput queryability: pre-process all documents at ingest, store structured data, enable fast retrieval later.; On-demand processing (ad-hoc Q&A on newly uploaded contracts) requires low latency but lower throughput; different compute allocation strategy.; Upfront compute investment in batch OCR enables petabyte-scale queryability without per-query extraction cost.
- Evidence: Sean explained: 'you have petabytes of documents that you want to be queryable at some point, right? So you are heavy on the compute at the upfront side with all the OCR. So you don't have to worry about it later on.'; Hiral added: 'when I am querying at that scale of thousands, I do need to have pre-processed and identified... businesses want to use this data to do a lot of downstream work.'; Q&A scenario: 'people may upload a contract to begin with for Q&A, and that's a very low latency use case... your different batch sizes, your concurrences, your different techniques on how you process the document will be different.'
- Caveats: Batch processing not suitable for real-time interactive use cases (e.g., user uploads contract and asks immediate questions).; Requires storage infrastructure for pre-processed structured data; upfront cost higher than on-demand extraction.
- Implications: Ken's agent control planes should support both modes: batch processing for high-volume corpus ingestion (pre-process everything) and on-demand extraction for interactive agent sessions (process one doc at a time).; Choose processing mode based on use case latency/throughput needs: batch for queryable data lakes, on-demand for real-time agent interactions.
Notable Concepts & Terms
- NemoTron Parse: NVIDIA's 850M-parameter VLM designed for single-shot document extraction (layout, tables, reading order, structure) without generation; 20x faster than open-source alternatives for table extraction at DocuSign's scale.
- NemoTron Retriever: NVIDIA's portfolio of embedding models, re-ranking models, and document extraction models for enterprise-scale retrieval; leads leaderboards (Vidori, MTab, MMTab); roadmap for DocuSign to enable 'find the right document before extraction.'
- Intelligent Agreement Management (IAM): DocuSign's AI-first platform for entire agreement lifecycle: creation (GenAI), negotiation (redlining), signing, storage, and insights extraction; processes 1M agreements/day for 1.9B users.
- Agreement Data Model: DocuSign's proprietary schema for structuring extracted agreement data (metadata, pricing tables, clauses, SLAs) at document and organization level; enables API/CSV export to downstream systems like Coupa.
- Hybrid OCR/VLM Pipeline: DocuSign's architecture: NemoTron Parse for table extraction, separate OCR modules for clauses/metadata; splits processing by data type for better accuracy and efficiency than one-size-fits-all models.
- Pareto Curve Optimization: Strategy of optimizing accuracy first, then pushing performance (quantization, multi-token generation, encoder-decoder techniques) after validating value; NVIDIA and DocuSign's phased approach.
- RD TableBench: Benchmark for table extraction accuracy; NVIDIA uses it to compare NemoTron Parse against open-source models and track progress toward enterprise-grade performance.
- NVIDIA NIM: NVIDIA's inference microservices for serving models like NemoTron Parse; alternative to vLLM for deploying VLMs in production.
Operator Notes / Why Ken Should Care
- Evaluate purpose-built VLMs (<1B params) for OpenClaw's document ingestion pipeline instead of defaulting to frontier LLMs; NemoTron Parse's 20x throughput gain at DocuSign's scale shows small specialized models can outperform general models for structured extraction.
- Design agent control planes with dual processing modes: batch for high-volume corpus ingestion (pre-process at ingest, store structured data) and on-demand for interactive sessions (low-latency extraction); choose based on use case latency vs. throughput needs.
- Implement hybrid pipelines splitting extraction by data type (tables, text, metadata) rather than one-size-fits-all models; DocuSign's NemoTron Parse + OCR architecture achieved better accuracy and efficiency than generic LLMs.
- Adopt phased optimization: deploy for accuracy first, then iterate on quantization (FP16 → FP8 → FP4) and performance (multi-token generation, encoder-decoder techniques) after validating value; avoid premature optimization.
- Benchmark agent system models on task-specific metrics (e.g., tables extracted/sec, cost per extraction) rather than general leaderboard scores; DocuSign measured tables/sec, not just accuracy, to validate NemoTron Parse.
- Plan multi-stage agent pipelines: document retrieval (NemoRetriever/embeddings) → page selection → extraction (NemoTron Parse) → reasoning (LLM); retrieval/re-ranking precedes extraction at petabyte scale.
- Output structured data in formats/APIs that integrate with enterprise systems of record (ERPs, CRMs, procurement tools); DocuSign's CSV/API export to Coupa shows extraction is only valuable if it feeds operational workflows.
- Use open-source models with published optimization blueprints (like NemoTron) to accelerate production deployment and cost reduction; NVIDIA open-sources models, datasets, quantization/distillation/pruning techniques, and agent skills.
Source/Metadata
- Title: Your Agreements Are a Database You Can't Query — Hiral Shah, Docusign & Sean Sodha, NVIDIA
- Transcript words: 4446
- Duration seconds: 1101
- Timestamp note: Timestamps unavailable in transcript; could not provide MM:SS navigation.
Transcript
Welcome everyone to our session. I'm Haral, I'm a senior director of products at DocuSign, and I'm joined by Sean. Hello everyone, I'm a product manager at NVIDIA. So today Sean and I are going to talk about a massive problem that every enterprise faces, which is agreement data, large-scale agreement data. Agreements are a big part of any relationship, any B2B organization goes through, day in, day out, and a lot of that data is captured inside agreements, and it's very critical, whether it's pricing tables, and it's all in a lot of different unstructured format, and that's what we're going to show is how DocuSign partnered with NVIDIA are fixing that on making that data available, readable, usable for a lot of our organizations. So here is a quick map of our talk today. We'll start with just the stakes. Why does this matter? Why is the scale so large? And then we'll dive deep into the technical architecture of how we are approaching it, how we have tackled this thing, especially document processing at scale. And then finally we'll cover what we have learned from our evaluation of all of the different models we've tried for different purposes and share our learnings with you. So why, just to understand why this is a big problem. When you think about DocuSign, raise your hands, how many of you have used DocuSign? Anyone who's employed probably the HR docs, right? So it's a massive scale. Everyone uses DocuSign. For us, it's a massive engineering problem as well, because just look at the scale. We have 1.9 million customers who are paying us and a billion users. What does that imply? We process a million agreements a day that needs to now be structured, made readable, made queryable, usable. And in the past we've worked with Deloitte on a study and it says that there's $2 trillion captured in this agreement negotiated value that no one capitalizes, no one goes back and gets that data back. And why? It's because they have to do a lot of human reading, human reviews, there's disconnected systems, a lot of manual workflows that are there. So that's why DocuSign built IAM, an intelligent agreement management platform that takes the entire approach in an AI first way, the entire agreement life cycle. Whether you are creating agreements, Gen AI helps a lot with that, whether you're negotiating to understanding and redlining, all the way to after signing, storing, and making a lot of insights from this data. So when you think about the challenges involved, in an agreement, there is the unstructured data. It could be a PDF, it could be a PNG. And what are people wanting to do is simple questions. They can't get that. That is the data that is trapped inside one agreement, but also the whole corpus of millions of agreements, 10 years, 20 years of business has put into that. So agreements aren't flat. They're also hierarchical in nature. One agreement governs the other. The other does something else. So you always are needing a lot of things to answer a question. Simple thing. I'm sure you all are using a lot of Claude and all, especially at your company's organization. Simple thing. What did we contract for the total tokens with Claude? No one knows. That's captured inside this agreement in different forms and fashion. So you need to be extracting this data to find things, but also insights and push it downstream where you're tracking doing more things. And when you analyze an enterprise contract, a big set of things are captured, what I call vital terms, like pricing tiers, the SKUs, the information, SLAs, rate cards, they're all in table format. Now, traditional document extraction tools or a generic LLMs completely fail here. We've tried, we've definitely done this, because they're reading text line by line, which breaks a lot of that concept within the table. Merged cells. Some things are not bounded. So this makes a massive operational overhead for the downstream legal teams, procurement teams, sales teams to get that queries and get that answers done. And they spend hours and hours digging through this to even just locate a basic thing. And that's where we partnered with NVIDIA and leveraged a purpose-built model, which is really making our architecture for table extraction take us to make and solve these complex use cases. So with this, we are making things scalable. We can understand it with the layout, but also deliver really accurate results. And to share more about how we're leveraging the NemoTron, I'm going to hand it to Sean. Sean Huston: All right, hello everyone. So real quick on the NemoTron Retriever initiative. So for those who know about NemoTron, you can raise your hand real quick. Awesome. So NemoTron is all about building world-class open source models and publishing the datasets, the techniques, the quantization approaches, distillation approaches, pruning approaches, every technique possible, blueprints to go with that, you name it. Throughout the NemoTron portfolio, we have specifically NemoTron Retriever, which is building embedding models, re-ranking models, and document extraction models. So real quick here, our first initiative is, if you're a large scale enterprise that deals with petabytes scale data, our first initiative is how do we make sure that you find the right document given a certain query, your agent sends a set of queries to the corpus afterwards. Once you find those top five documents, whatever it may be, then we say, okay, you found the right document, now how do you then find the right information within the document? And this is where the work with the DocuSign team has gone really great, where we've worked with them to build the NemoTron parse model to focus specifically on table extraction, which is a really complicated technique. If you think about it, the number of permutations of tables are quite vast when you think about nested tables, merged cells, merged columns, merged rows, whatever it may be, and that can get really complex and really hairy of a problem. So real quick, as I mentioned before, our team is responsible for leading a lot of leaderboards in the retrieval space. So Vidori, V1, V2, V3, MTab, MMTab. So our team knows how to build world-class retrieval models, given a lot of leadership, given a lot of leaderboard winnings that we've had in the last year or so. And then of course, as I mentioned before, we open source everything. So we share the open source model weights, the techniques, and then we release with those blueprints and skills that agents can use afterwards. So touch a little bit on the actual model that we are working with with DocuSign was the NemoTron parse model. So when you think VLM, you generally think a multi-billion parameter model. It's very heavy. It's high latency. This is a very small, tiny C-radio VLM. It's about 850, 900 million parameter model designed to be that all-in-one package model where you deploy it. And instead of having small, let's say, YOLOX models that do table extraction or page element extraction, wherever it may be, this is a single-shot model that you can feed a document in, and out comes the semantic formatting layouts, the text, the reading order, the preserved structure of the table, et cetera. This can be served via the NVIDIA NIM or via VLLM as well. And so it's a tiny, small model that you can use. It's not a generator. It's more of an extractor at the end of the day. So real quick as well, we always want to make sure that we're building towards benchmarks that matter most to the enterprise space. So we want to make sure that both on the Pareto curve of accuracy versus performance, we'll make sure that we're going to be releasing world-class models to the ecosystem. So you'll see here generally is just a very standard benchmark of table extraction. I believe this one was RD TableBench. And we compare some popular open source models here, and then we compare how our NemoTron parse model does compared to that industry. And we continue to strive to improve this as time goes on. So I think we believe we have a demo as well. Yeah. Is that as well? Just press one. Yeah. There we go. Let me show you how easy it is to turn any agreement into structured, usable data with Agreement Manager, which is a central repository of every agreement an organization has ever signed. Let's look at this. So when we look at the Agreement Manager view here, we have an ability to see the entire list of agreements, but also go and upload a new agreement. So I'm uploading a new order form into Agreement Manager. As you can see, I can select from a computer, import from other places. The moment I select the agreement, it starts uploading. of table extraction. I believe this one was RD TableBench. And we compare some popular open source models here, and then we compare how our NemoTron parse model does compared to that industry. And we continue to strive to improve this as time goes on. So I think we believe we have a demo as well. Yeah. Is that as well? Just press one. Yeah. There we go. Let me show you how easy it is to turn any agreement into structured, usable data with Agreement Manager, which is a central repository of every agreement an organization has ever signed. Let's look at this. So when we look at the Agreement Manager view here, we have an ability to see the entire list of agreements, but also go and upload a new agreement. So I'm uploading a new order form into Agreement Manager. As you can see, I can select from a computer, import from other places. The moment I select the agreement, it starts uploading and starts processing with AI. And just like that, you can see that the jobs engine has processed it. Let's take a closer look at this agreement. So when you go into the action, you can go and browse the file. Within seconds, Agreement Manager has extracted a rich set of metadata. Everything from key terms to commercial details are automatically structured, highlighted. And immediately, you can jump to that section where the details are found. Built-in goes deeper. This is where the NVIDIA's model comes in, that it's extracted all the structured pricing data around this agreement. It goes in, breaks down these complex tables into order details. And as you can see, we can break it down. We can download all of this data. This is powered by the advanced parsing, leveraging NVIDIA's NemoTron model, turning even dense tables into something that is instantly usable. And of course, you can take this data with you. You can see when we've downloaded into CSV, how we've structured all of it for your finance team, procurement team, even further analysis. All of it is also available through API. And that's how Agreement Manager has transformed agreements into actionable insights in seconds, leveraging NVIDIA. So I think what you saw there from a demo perspective, we've tried to shorten it. We have a whole repository. What you see a list, we get customers which has thousand agreements to all the way millions of agreements within. But the big piece is how do we understand and get that data that makes it very valuable to an end business user, right? A legal person, a procurement person, a salesperson who's doing a lot of the deals, or even a leader, right? A business unit. The CTO goes and asks, what did we do? This is how we have making each of the things a lot more structured, so we have our own proprietary agreement data model, which we are structurizing each agreement but also at a whole organization level, and leveraging a lot of the NVIDIA things. We've been able to do a really good job, especially with all of those tables, pricing, SLAs, and then make that available. And then we also have a robust search that is on top of it. So when you think about what have we learned, right? When you think from a NemoTron plus DocuSign, one of the biggest things for us, we definitely have done a lot of different models for different purposes. So a purpose-built model for the job you're trying to do is a big, big part of how we've been thinking about, and that's where we've been able to accelerate, bring things to market much faster. The second big piece around the model efficiency, so as Sean was talking about, the number of parameters, context and all that matters in a different environment for different things. For us, the lower context also meant lower latency, lower cost, to deliver the scale that we are talking about. Last, around the faster extraction. So we ran this against a lot of the other open source models. When you think about how many tables can it extract per second, NemoTron was 20x faster, which helps us when we're talking about the millions and billions of scale that we are serving for all of our customers. So a lot of it is having that smaller purpose-built things is the way for an enterprise as an organization to go and leverage, and then serve that from an end user perspective. And then what's next? So let Sean talk through those. Yeah, so working with the DocuSign team has been awesome so far, and we're going to continue to deepen that partnership as well over the next few months. So with them, we started with, how do I extract as much possible information from a page, and now we'll scale to, how do I now find that page to begin with? So we'll start a little bit with the NemoTron, NemoRetriever effort, and then of course we'll talk a little bit about the NVIDIA agent toolkit with them over the next few months, and then actually start scaling into more production-scale agents then. So, perfect. I think that's what we had. We have time for a couple questions. Anyone in the room? Okay, I have someone there. So just to recap for everybody, if you didn't hear, the question is, right, OCR is always a thorn in the whole process, so are we thinking about letting that go and starting from agentic from the get-go? I can talk from my perspective. So I think for us, right, there are different use cases at different points in time. Many times, if you are reactive, you have a question and you're coming, some of that can work dynamically at a smaller scale. The question is the latency. When I am querying at that scale of thousands, I do need to have pre-processed and identified, so that's one. I think the second big part of the use case for us, a lot of times businesses want to use this data to do a lot of downstream work. So an example is a procurement team. This is my pricing table. I want to put it into Coupa to make sure when I'm paying that works. At that time, the agent is helping, but I can't do that on a one document by document basis. That said, there are ways that we are compressing. That's why NemoTron worked for us. It's how do you do it from a layout understanding just for that purpose? But I would like to let you add. Yeah, I think it depends on the use case a little bit. I think for this specific instance, right, you have petabytes of documents that you want to be queryable at some point, right? So you are heavy on the compute at the upfront side with all the OCR. So you don't have to worry about it later on, right? And there are some instances where people may upload a contract to begin with for Q&A, and that's a very high, that's a very low latency use case, right? So you have a high throughput versus a latency use case. And in that scenario, your different batch sizes, your concurrences, your different techniques on how you process the document will be different. And where you spend that compute in that cycle will be changing between the different use cases. One more there. Yeah. We do a lot of what I call hybrid approach and purpose built for the needs and the use cases. So from a table piece, it does break down the whole layout along with extracting. We still do OCR from a lot of other fields and metadata and the clauses, all of the text type of thing. So I think we had an architecture where we have our pipeline going through two different routes for that. As a follow-up, we have a blog out there. How are we really solving this at scale across? And if you look at that, there are a lot of different piecemeal modules and things together. Yeah, one more. Not last. At scale, what's the modification? What do you mean to do with that model? So we, this model is currently on FP16, but there are paths towards going to FP8 and FP4 in the next few months as well. What about... We do use that. And then we are also using some of the older ones. And that's the journey as a partnership is to tweak as you get more of the customers. Yeah. So for this, there are many techniques on how to improve the performance side, right? So quantization, right? We're trying to move everyone to Blackwell, right? That's why NVFB is the big thing now, as well as multi-token generation for this. It's a VLM architecture, right? So your encoder decoder techniques can definitely be further optimized. Right now this model just generates one token at a time. You can do multi-token generation, of course. So there's plenty of performance things. Right now we're focusing on the accuracy side. Are we adding value to the system? And then from there, we'll then push out that Pareto curve on the performance side. So what's up? Are you guys using the catalogs on the video GPUs, or are you primarily running through the port? I believe they just deploy the VLM directly. So there are many techniques on how to improve the performance side, right? Quantization, right? We're trying to move everyone to Blackwell, right? That's why NVF before is the big thing now, as well as multi-token generation for this. It's a VLM architecture, right? Your encoder decoder techniques can definitely be further optimized. Right now this model just generates one token at a time. You can do multi-token generation, of course, too. So there's plenty of performance things. Right now we're focusing on the accuracy side. Are we adding value to the system? And then from there we'll push out that Pareto curve on the performance side. Are you guys using the catalogs on the video GPUs, or are you primarily running through the port? I believe they just deploy the VLM directly. I've done a lot on recall side. When you do that demo, you ask the question, is that graph, or is this a new port? This is just an extraction. This is now a retrieval. It's coming from that agreement data that we've extracted and stored. Maybe I can chat with you offline about how our architecture works fully as well. We're almost coming up on time there. I think that's all we have. I'm happy to hang around in the back with more questions. Good luck with a lot of your challenges with AI. Thank you. Thank you. I would like to catch up in November. I think that's what we had. We have time for a couple questions. Anyone in the room? Okay, I have someone there. So just to recap for everybody, if you didn't hear, it was, the question is, right, like OCR is always a thorn in the whole process, so are we thinking about letting the go of that and starting from agentic from the get-go? I can talk from my perspective. So I think for us, right, like there are different use cases at different points in time. Many times, if you are reactive, you have a question and you're coming, some of that can work dynamically at a smaller scale. The question is the latency. When I am querying at that scale of thousands, I do need to have pre-processed have identified, so that's one. I think the second big part of the use case for us, a lot of times businesses want to use this data to do a lot of downstream work. So an example is a procurement team. This is my pricing table. I want to put it into Coupa to make sure when I'm paying that works. At that time, like, you know, the agent is kind of helping, but I can't do that on a one document by document. That said, there is ways that we are compressing. That's kind of why NemoTron worked for us. It's like, how do you do it from a layout understanding just for that purpose? But I would like to let you add. Yeah, I think it depends on the use case a little bit. I think for this specific instance, right, you have petabytes of documents that you want to be queryable at some point, right? So you are heavy on the compute at the upfront side with all the OCR. So you don't have to worry about it later on, right? And there's some instances where people may upload a contract to begin with for Q&A, and that's a very high, that's a very low latency use case, right? So you have a high throughput versus a latency use case. And in that scenario, your different batch sizes, your concurrences, your different techniques on how you process the document will be different. And where you spend that compute in that cycle will be changing between the different use cases. One more there. Yeah. We do a lot of like more of what I call hybrid approach and a purpose built for like the needs and the use cases. So from a table piece, it does kind of, you know, do the whole layout along with extracting. We still do OCR from a lot of other fields and metadata and the clauses like all of the text kind of thing. So I think we had an architecture where we have our pipeline going through two different routes for that. As a follow, we have a blog out there. How are we really solving this at scale across? And if you look at that, there's a lot of different piecemeal modules and stuff together. Yeah, one more. Not last. At scale, what's the modification? What do you mean to do with that model? So we, this model is currently on FP16, but there are paths towards going on to FP8 and FP4 in the next few months as well too. What about . We do use that. And then we are also kind of using some of the older ones. And that's the journey as a partnership is to kind of go tweak as you get more of the customers. Yeah. So for this, there are many techniques on how to improve the performance side, right? So quantization, right? So we're trying to move everyone to Blackwell, right? That's why NVF before is the big thing now, as well as multi-token generation for this. It's a VLM architecture, right? So your encoder decoder techniques can definitely be further optimized. So right now this model just generates one token at a time. You can do multi-token generation, of course, too. So there's plenty of performance things. Right now we're focusing on the accuracy side. Like, are we adding value to the system? And then from there, we'll then push out that Pareto curve on the performance side. So what's up? Are you guys using, like, the catalogs on the video GPUs, or are you, like, primarily running through the port? I believe they just deploy the VLM directly. No, I've done a lot on recall side. When you, like, when you do that demo, you ask the question, is that graph, or is this a new port? Oh, this is just an extraction. This is now a retrieval. Yeah. Yeah. Yeah. It's coming from that agreement data that we've kind of extracted and stored. Yeah, maybe I can chat with you offline and how our architecture kind of works fully as well. No, we're almost coming up on time there. But I think that's kind of all we have. I'm happy to hang around in the back with more questions. And good luck with a lot of your challenges with AI. Thank you. Thank you. I would like to catch up on November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November November Thank you.