Open Reader

Structuring the Unstructured - Cedric Clyburn, Red Hat

completed 20:40 Jun 28, 2026 Watch on YouTube

Current Status

completed

Video ID

-x5GEVnkuRw

RAG / Chat

Enabled
Structuring the Unstructured - Cedric Clyburn, Red Hat
Description

Modern organizations generate vast amounts of data stored in diverse and often unstructured formats, such as PDFs, scanned documents, and proprietary file types. For engineers working with AI, the challenge isn’t just about extracting text but also about preserving the structure, context, and relationships within the data. Whether fine-tuning models or building retrieval-augmented generation (RAG) pipelines, effective document processing is essential for creating AI applications that bring value. This live demo session is all about the techniques and open source tools needed to transform unstructured documents into structured formats like JSON or Markdown, ready for AI workflows. You’ll learn how to handle challenges like multi-page tables, image-heavy layouts, and scanned documents using context-aware methods with Docling, part of the Linux AI & Data Foundation. Speakers: - Cedric Clyburn (Red Hat): Cedric Clyburn (@cedricclyburn) is a Senior Developer Advocate at Red Hat and open-source contributor (vLLM, Podman) who helps developers adopt emerging technologies through speaking, workshops, and community leadership (as an organizer of Kubernetes Community Day New York). X/Twitter: https://x.com/cedricclyburn LinkedIn: linkedin.com/in/cedricclyburn/ GitHub: github.com/cedricclyburn

Summary

Generated by claude-sonnet-4-5

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Dockling (Linux Foundation open-source tool) solves enterprise unstructured data processing with local, fast, cheap OCR+layout analysis to extract accurate markdown/JSON from PDFs/images without proprietary APIs, enabling RAG/agent workflows at 50x cost savings vs vision models
  • Why it matters: Most enterprise data is unstructured (PDFs, scanned docs, images); poor extraction causes LLM hallucinations, incorrect outputs, and research contamination—Dockling provides local, air-gapped, cost-effective extraction for tables/images/text with structured output
  • Best use: Watch fully for operational blueprint: local document processing pipeline, chunkless RAG pattern, scaling via REST API/MCP server, air-gapped deployment; reference demos for table extraction, image annotation, structured output from invoices

Executive Summary

Cedric Clyburn (Red Hat) presents Dockling, a Linux Foundation open-source document processing tool that addresses the core bottleneck in enterprise AI: converting unstructured data (PDFs, images, scanned docs, tables) into formats LLMs can reliably use. The talk demonstrates that naive PDF parsers produce truncated/merged text, frontier VLMs are prohibitively expensive ($30/M tokens) and non-deterministic at scale, while Dockling runs locally on CPU with OCR+layout models to extract markdown/JSON/Pydantic objects with preserved structure. Red Hat uses it internally for thousands of product documentation PDFs.

The technical implementation spans multiple deployment patterns: pip-installable CLI/library for single documents, REST API microservice (docling-serve) for batch processing, and MCP server integration for agentic workflows (Cursor, Claude Code). Cedric shows live Jupyter demos extracting tables from 8-page PDFs, annotating images with local Granite VLM via Ollama, and implementing 'chunkless RAG'—where the LLM navigates a markdown outline of document sections (418 sections in IBM annual report example) instead of embedding thousands of vectors. This pattern answered queries in one iteration by referencing structured section IDs.

The cost/scale argument is compelling: Loandro at Hugging Face processed Common Crawl PDFs with Dockling at 50x cost savings vs naive VLM/OCR, running on CPU without GPU. The tool handles edge cases like removing PII via layout visualizer bounding boxes, structured output for invoice fields (bill number, total), and prevents the documented research contamination where 20 scientific papers cited a nonsensical term created by OCR merging words across PDF columns. For air-gapped/compliance environments, everything runs locally without external API calls.

Operational takeaway: Dockling enables enterprise RAG/agent systems to scale document processing without vendor lock-in, GPU dependency, or data exfiltration. The hybrid chunker, Pydantic exports, and MCP server support make it integration-ready for existing AI harnesses. Ken should consider this for any workflow involving legacy document conversion, especially for clients with compliance constraints or high document volumes where VLM costs would be prohibitive.

Key Takeaways

  • Claim: Naive PDF parsers and frontier VLMs both fail at enterprise scale—parsers lose table structure/images, VLMs cost $30/M tokens and produce non-deterministic output across model versions | Evidence: Demo showed simple parser outputting truncated/merged text from table in linear format; VLM markdown was high-quality but speaker noted version drift (5.1 vs 5.2) creates inconsistent structured output at scale; 20 scientific papers now cite nonsensical term from OCR misinterpreting scanned PDF that merged words across columns | Caveat: Speaker works at Red Hat (vendor interest), but cost/quality tradeoffs are documented with third-party Hugging Face case study (50x savings); no benchmarks provided for accuracy vs frontier VLMs on complex documents | Implication: For Ken's agent systems: local document processing eliminates API costs and non-determinism risk; prioritize Dockling for compliance/air-gapped clients or high-volume workflows where VLM costs compound; watch for accuracy limits on highly complex layouts | Timestamp: 03:20
  • Claim: Dockling achieves 50x cost savings vs VLMs/OCR for large-scale PDF processing, running on CPU without GPU requirement | Evidence: Loandro at Hugging Face processed Common Crawl PDFs for FineWeb dataset (thousands of documents); comparison showed CPU-based Dockling processing at 1/50th the cost of GPU VLM+OCR naive approach; tool uses OCR + layout analysis models instead of frontier vision models | Caveat: No absolute throughput numbers provided (pages/sec, docs/hour); 'thousands of documents' scale mentioned but not quantified; CPU performance likely depends on document complexity and OCR backend choice | Implication: Operational win for batch document processing in agent pipelines; Ken should benchmark Dockling CPU throughput vs VLM API for his use cases; consider hybrid approach (Dockling for bulk, VLM for edge cases requiring semantic understanding) | Timestamp: 08:15
  • Claim: Chunkless RAG pattern uses document outline as retrieval index instead of vector embeddings, enabling single-iteration query answering on 418-section documents | Evidence: Demo processed IBM 2025 annual report (418 sections) and Red Hat revenue query; LLM navigated markdown outline with section summaries, retrieved relevant section in 1-5 iterations without vector database; compared to 'thousands of vectors' in traditional RAG semantic similarity approach | Caveat: Pattern assumes document has logical section structure and outline; unclear how it handles cross-sectional queries or when answer spans multiple non-adjacent sections; no comparison metrics vs traditional chunking (accuracy, retrieval time) | Implication: Novel RAG architecture for Ken's agent workflows—eliminates embedding model, vector DB, and chunking strategy decisions; best for structured reports/docs where outline is meaningful; test retrieval accuracy vs semantic search for your use cases | Timestamp: 16:45
  • Claim: Dockling extracts structured data from tables and invoices as Pydantic objects, enabling field-specific extraction without pulling headers/titles | Evidence: Demo showed extracting 8 tables from research paper PDF to dataframes; invoice example extracted bill_number and total_invoice_price as structured fields; output is Pydantic data type usable in Python workflows, not just markdown strings | Caveat: No error handling shown for malformed tables or OCR failures; unclear how it handles tables spanning pages or nested table structures; no accuracy benchmarks vs manual extraction provided | Implication: For Ken's content/business ops: enables automated data extraction from contracts, invoices, reports without manual parsing; Pydantic output means type-safe integration with agent tools; consider for document classification/routing workflows | Timestamp: 12:30
  • Claim: Dockling integrates with MCP (Model Context Protocol) to give AI agents document processing capabilities through standardized tool interface | Evidence: Demo showed MCP server providing conversion/generation/manipulation tools to Qwen 3.6 model in VS Code; agent could 'convert this document, give summary, create document with action items from another PDF'; runs via uvx and config.yaml in Cursor/Claude Code | Caveat: MCP adoption is early-stage; config shown but not debugging/error scenarios; unclear how MCP server scales for concurrent agent requests or handles processing timeouts | Implication: For Ken's agent platform: MCP provides standardized document processing interface across agent frameworks; enables multi-step document workflows (extract→summarize→export) in single agent interaction; consider Dockling MCP as building block for document-centric agents | Timestamp: 19:20
  • Claim: Dockling-serve deploys document processing as REST API microservice for batch processing hundreds/thousands of PDFs via Kubernetes | Evidence: Showed pip install docling-serve, CLI serving on specific port, POST requests with options (OCR backend, image annotation flags); mentioned container/Kubernetes deployment for scale; contrasted with single-document pip install docling | Caveat: No throughput benchmarks, queue management details, or horizontal scaling guidance provided; unclear how it handles processing failures, retries, or partial results for corrupted documents | Implication: Operational deployment pattern for Ken: REST API enables async document processing in agent workflows; Kubernetes deployment supports multi-tenant or high-volume scenarios; consider this for client environments where documents arrive via webhook/upload rather than batch | Timestamp: 18:10
  • Claim: Layout visualizer with bounding boxes enables PII removal and structure validation before extraction | Evidence: Demo showed visualizing section headers, text blocks, images with bounding box overlays from layout model; speaker noted this allows removing 'personally identifiable information from a customer' before extracting to application | Caveat: No PII detection/redaction automation shown—appears to be manual inspection tool; unclear if layout model can auto-detect PII fields vs just visualizing structure | Implication: For Ken's compliance-focused clients: layout visualization enables manual PII audit before document ingestion; consider building auto-redaction layer on top using layout model output; useful for regulated industries (healthcare, finance) where document inspection is required | Timestamp: 14:50
  • Claim: Local vision model integration (Granite via Ollama) enriches images/diagrams with detailed descriptions for RAG context without external API calls | Evidence: Demo configured PDF pipeline to call local Ollama Granite model at OpenAI-compatible endpoint; generated detailed caption describing DocLean pipeline diagram beyond original image caption; runs fully locally without internet | Caveat: No comparison of local VLM quality vs GPT-4V/Claude; processing time not shown; unclear which Granite model version or size used; may require fine-tuning for domain-specific diagrams | Implication: For Ken's air-gapped deployments: local VLM enrichment adds semantic context to images/charts without cloud dependency; consider for technical documentation, architecture diagrams, or charts where visual context matters; benchmark local model accuracy for your image types | Timestamp: 15:30

Detailed Brief

The Document Processing Problem and Dockling's Solution

  • Claims: Most enterprise data is unstructured (PDFs, scans, images, tables) and needs transformation for LLM consumption; Simple parsers lose structure (tables→linear text, missing images), frontier VLMs are expensive ($30/M tokens) and non-deterministic; Dockling is Linux Foundation open-source tool using OCR+layout analysis to extract markdown/JSON locally on CPU
  • Evidence: 20 scientific papers cite nonsensical term from OCR merging PDF columns—demonstrates downstream contamination risk; Demo compared naive parser (truncated table output) vs VLM (good but costly/inconsistent) vs Dockling (structured, local); Red Hat uses Dockling internally for thousands of product documentation PDFs; Hugging Face's Loandro achieved 50x cost savings processing Common Crawl PDFs with CPU-based Dockling vs GPU VLM/OCR
  • Caveats: Speaker is Red Hat employee presenting company-adjacent tool (vendor interest); No absolute accuracy benchmarks vs frontier VLMs on complex documents provided; CPU performance likely varies with document complexity and OCR backend choice
  • Implications: Enterprises with legacy document archives can now process at scale without cloud API costs or data exfiltration; Local processing enables air-gapped/compliance deployments where VLM APIs are prohibited; Non-determinism of VLMs at scale (version drift 5.1→5.2) makes Dockling's consistent output valuable for production systems; Ken should evaluate Dockling for any client with high document volumes, compliance constraints, or GPU budget limits

Chunkless RAG: Document Outline as Retrieval Index

  • Claims: Traditional RAG embeds thousands of vectors for semantic similarity search across document chunks; Chunkless RAG uses markdown outline of document sections as retrieval index—LLM navigates outline to find relevant section; Pattern answered query on 418-section IBM annual report in 1 iteration vs typical multi-hop retrieval
  • Evidence: Demo processed IBM 2025 annual report (418 sections) with query 'What was Red Hat's revenue growth in 2025 and contribution to software segment?'; LLM iterated max 5 times checking section relevance, retrieved answer in single iteration by referencing section IDs; Outline showed section summaries (e.g., 'docling can be installed from PyPI' for getting started section); Contrast made to traditional RAG: 'thousands of vectors in database' vs 'markdown outline of document'
  • Caveats: Assumes document has logical section structure and meaningful outline—may fail on unstructured narratives; No metrics comparing retrieval accuracy vs traditional semantic chunking provided; Unclear how pattern handles queries requiring cross-sectional synthesis or non-adjacent section answers; Single-iteration claim may be optimistic for complex queries spanning multiple domains
  • Implications: Eliminates need for embedding model selection, chunking strategy, and vector database infrastructure; Reduces RAG system complexity and operational overhead—fewer moving parts to tune/maintain; Best suited for structured reports, technical docs, regulatory filings where outline is semantically meaningful; Ken should test retrieval accuracy vs semantic search for his use cases before replacing traditional RAG; Consider hybrid: outline-first retrieval with fallback to semantic search for cross-sectional queries

Deployment Patterns: CLI, REST API, and MCP Server

  • Claims: Dockling supports three deployment modes: pip-install CLI/library, REST API microservice (docling-serve), and MCP server for agents; REST API enables batch processing hundreds/thousands of PDFs via Kubernetes with configurable OCR/annotation options; MCP server integration gives agents standardized document processing tools (convert/generate/manipulate) via Model Context Protocol
  • Evidence: CLI demo: pip install docling, process single PDF to markdown in ~seconds for 8-page document; REST API: pip install docling-serve, serve on port, POST with options (OCR backend, image annotation flags); MCP demo: Qwen 3.6 in VS Code using MCP server tools to 'convert document, summarize, export action items from another PDF'; Mentioned container/Kubernetes deployment for scale, uvx for MCP server launch, config.yaml for agent integration
  • Caveats: No throughput benchmarks (pages/sec, docs/hour) or horizontal scaling guidance provided; Error handling, retries, and partial result scenarios not addressed in demos; MCP adoption is early-stage—unclear how many agent frameworks support it beyond Cursor/Claude Code; No queue management or concurrent request handling details for REST API at scale
  • Implications: For Ken's agent platform: MCP provides clean abstraction for document processing tools across frameworks; REST API enables async processing in production workflows—documents arrive via webhook/upload, get processed, results cached; Kubernetes deployment supports multi-tenant SaaS scenarios or high-volume enterprise clients; CLI/library mode best for prototyping or single-tenant local deployments; Consider building job queue (Celery, BullMQ) around REST API for production reliability and observability

Structured Output and Table Extraction

  • Claims: Dockling extracts tables as dataframes and Pydantic objects, preserving structure lost in linear parsing; Supports field-specific extraction from invoices/forms (bill number, total) without pulling headers/boilerplate; Output formats include markdown, HTML, JSON, Pydantic—usable across different downstream systems
  • Evidence: Demo extracted 8 tables from research paper PDF, rendered as dataframes in Jupyter notebook; Invoice example showed extracting bill_number and total_invoice_price as structured fields; Pydantic output demonstrated: document.pages, document.tables, programmatic access to page-level content; Compared to naive parser which output table rows linearly: 'merged and isn't decipherable even by me as a human'
  • Caveats: No error handling shown for malformed tables, OCR failures, or tables spanning pages; Accuracy benchmarks vs manual extraction not provided; Unclear how it handles nested tables, rotated text, or complex multi-column layouts; Invoice extraction may require training/fine-tuning for custom form layouts
  • Implications: For Ken's business ops: automate data extraction from contracts, invoices, reports without manual parsing teams; Pydantic output enables type-safe integration with Python agent tools and validation logic; Consider for document classification/routing workflows: extract key fields to determine processing path; Use table extraction to populate structured databases from legacy PDF reports/statements; Test accuracy on your specific document types before replacing manual extraction workflows

Image Annotation and PII Handling

  • Claims: Local vision models (Granite via Ollama) can enrich images/diagrams with detailed descriptions without external APIs; Layout visualizer with bounding boxes enables PII identification and removal before extraction; Supports scenarios where SME who created diagram has left organization—VLM reconstructs knowledge
  • Evidence: Demo configured PDF pipeline to call local Ollama Granite model at OpenAI-compatible endpoint for image captioning; Generated detailed description of DocLean pipeline diagram beyond original caption; Layout visualizer showed bounding boxes around section headers, text blocks, images for structure inspection; Speaker noted: 'remove personally identifiable information from a customer' using layout model
  • Caveats: No comparison of local VLM quality vs GPT-4V/Claude for image understanding; Processing time for VLM enrichment not shown—may be slow on CPU; PII detection appears manual (via visualization) rather than automated—no redaction API shown; Unclear which Granite model version/size used or if fine-tuning needed for domain-specific diagrams
  • Implications: For Ken's air-gapped clients: local VLM enrichment adds semantic context without cloud dependency; Use for technical documentation, architecture diagrams, charts where visual understanding adds RAG value; Consider building automated PII redaction layer using layout model output + regex/NER for common patterns; Benchmark local VLM accuracy for your image types—may need domain-specific fine-tuning; Image enrichment particularly valuable for legacy docs where original context is lost

Notable Concepts & Terms

  • Dockling: Linux Foundation open-source document processing tool using OCR+layout analysis to extract structured data (markdown/JSON/Pydantic) from PDFs/images locally on CPU; main subject of talk
  • Chunkless RAG / Agentic RAG: Pattern where LLM navigates document outline (section summaries) as retrieval index instead of embedding chunks in vector DB; speaker claims single-iteration answers on 418-section documents vs traditional semantic search
  • docling-serve: REST API microservice mode of Dockling for batch processing hundreds/thousands of documents via Kubernetes with configurable OCR/annotation options; enables production deployment at scale
  • MCP (Model Context Protocol): Standardized protocol for agent-tool communication; Dockling MCP server exposes conversion/generation/manipulation tools to AI agents (Cursor, Claude Code) via config.yaml integration
  • Layout visualizer / bounding boxes: Feature showing structural elements (headers, text blocks, images) with position overlays; enables PII identification and document structure validation before extraction
  • Pydantic document type: Dockling's typed output format enabling programmatic access to pages, tables, sections; provides type-safe integration with Python agent tools vs unstructured markdown strings
  • FineWeb PDFs: Hugging Face dataset of processed Common Crawl PDFs (thousands of documents) used in cost comparison case study showing 50x savings with Dockling vs VLM/OCR approach
  • Hybrid chunker: Dockling feature mentioned for chunking exported documents; details not provided but implies alternative to naive token-based or sentence chunking strategies

Operator Notes / Why Ken Should Care

  • Dockling solves high-volume document processing economics: 50x cost savings vs VLM APIs makes it viable for batch processing thousands of legacy PDFs in agent knowledge bases
  • Chunkless RAG pattern eliminates embedding model, vector DB, and chunking strategy decisions—reduces system complexity and operational overhead for document-based agents
  • MCP server integration provides standardized document processing interface across agent frameworks—consider as building block for multi-step document workflows (extract→validate→export)
  • Local/air-gapped deployment critical for compliance-heavy clients (healthcare, finance, government) where document exfiltration to VLM APIs is prohibited
  • Pydantic output enables type-safe tool integration—extracted invoice fields, table data, section IDs usable directly in agent decision logic without parsing markdown strings
  • For Ken's content/GTM workflows: automate data extraction from contracts, reports, technical docs to populate CRMs, knowledge bases, or feed content generation pipelines
  • REST API deployment via Kubernetes supports multi-tenant SaaS scenarios—clients upload documents, async processing, results cached for RAG/agent access
  • Consider hybrid approach: Dockling for bulk extraction, frontier VLM for edge cases requiring deep semantic understanding or complex visual reasoning
  • Layout visualizer useful for document audit workflows—show clients structure extraction accuracy before committing to full pipeline
  • Watch for accuracy limits on highly complex layouts (nested tables, rotated text, multi-column academic papers)—benchmark against your specific document types

Watch Map

  • 00:00: Problem intro: unstructured data (PDFs, images) as AI context bottleneck; naive parsers vs VLMs vs Dockling positioning
  • 03:20: Viral tweet case study: 20 scientific papers cite nonsensical term from OCR column-merging error; importance of extraction accuracy
  • 05:15: Comparison demo: simple parser (truncated), frontier VLM (expensive/non-deterministic), Dockling (local/structured)
  • 08:15: Scale/cost evidence: Hugging Face FineWeb 50x savings with CPU-based Dockling vs GPU VLM/OCR
  • 10:30: Live demo begins: Jupyter notebook, pip install docling, document converter basics
  • 12:30: Table extraction demo: 8 tables from research paper PDF to dataframes; Pydantic access patterns
  • 14:50: Layout visualizer demo: bounding boxes for structure inspection, PII removal use case
  • 15:30: Image annotation demo: local Granite VLM via Ollama enriching diagram descriptions
  • 16:45: Chunkless RAG demo: IBM 2025 annual report (418 sections), markdown outline as retrieval index, single-iteration answer
  • 18:10: Deployment patterns: docling-serve REST API for batch processing, Kubernetes scale-out
  • 19:20: MCP server demo: Qwen 3.6 in VS Code using document processing tools via Model Context Protocol
  • 20:15: Wrap-up: open source benefits, integration with RAG frameworks, LinkedIn connect CTA

Source/Metadata

  • Title: Structuring the Unstructured - Cedric Clyburn, Red Hat
  • Transcript words: 8143
  • Duration seconds: 1240
  • Timestamp note: Timestamps estimated from transcript flow and demo transitions; actual chapter markers not present in source

Transcript

3912 words en Processed in 138.4s

Hey, hey, welcome. My name is Cedric Laburn. I'm an open source engineer here at Red Hat, and I think we can all agree that context is the most important aspect to building an AI application or an agent, right? It's the reason that harnesses have become so popular in order to manage the LLM's context. But the thing is, no matter what model or agent that you're using, there is so much data that we're not able to use properly because it's in unstructured formats. I'm talking everything from PDFs to presentations to contracts and technical docs, even meeting notes, scan documents, diagrams, tables, images, and more. And sorry, I know that's a lot, but you understand what I mean, right? All this data needs to be transformed into something that an LLM can actually understand. And that's why by the end of this session, you'll understand how to extract structure between raw enterprise documents and use it to power better downstream AI systems like RAG and agents. So let's get started. Now, I think Jensen from NVIDIA made this point super clear at his keynote that unstructured data is becoming this new context layer for AI. And the reality, though, for many teams, and I know this personally working at Red Hat, that PDFs and data are spread across dozens of different systems. So we've got a lot to cover today. As you might know, a large majority of the world's data is unstructured. And so no matter what model you're using, if you're working with data like PDFs and unstructured types of formats, this is a bit tricky to work with. Because there are solutions out there, but they might be proprietary or require you to send your private data to someone else's server. And for not just text, how do we take documents and their graphs or tables and images to formats that LLMs can understand like Markdown or JSON? And I'm going to show you how in the session today, because we're going to be using an open source tool, part of the Linux foundation that is called Dockling, and learn about extraction, parsing, chunking, and much more. And I've got some live demos for you. So we're going to have some fun. And just in case you'd like it, we have the session slides here on the right, and a little overview of the specifics I'll be showing you in today's session. But without further ado, let's get started. So why is there a need for advanced document processing? As I briefly mentioned before, you might have a lot of technical documentation or meeting minutes or different types of documents and invoices that you need to use and maybe RAG or different type of applications where the context is provided to an LLM. So whether it's RAG or retrieval augmented generation to answer questions based on this data, or you're using this to fine tune a new specialized model, well, data is this key ingredient behind those applications. And it doesn't matter if you're using Nvidia acceleration or an open source or proprietary model, that data and the way you process it is the key determining factor in whether your answer is going to be correct or incorrect for the user or customer at the end of the day. And that's what's most important. And how important is it? Well, I have this viral tweet from earlier where 20 scientific papers now feature a new nonsensical term that doesn't exist because AI misinterpreted a very old article that was scanned and taken to a PDF, merging two different words from two different columns in this PDF. And because researchers are using these models in order to help them write, now we have different types of scientific papers that all feature this word and are even being cited by other people. And so that's how important it is to make sure that the data that we're processing is processed in a way that's accurate and not hallucinated and able to be used confidently in our applications that we're delivering to users and customers. So it's quite important. Now, if we were to use a tool like dockling that I can run on my own machine, you could see that these two words are quite far away from each other and shouldn't have been combined in the first place. But that's how we're going to learn about extracting this text here in a second. So if we were to try a simple PDF parser for a PDF like this, that includes a table here, it has an image, there's captions, and there's regular sections of text. Well, we might get an answer like this here on the right in Markdown. You know, this might be very fast and cheap to run, even on CPU. But the issue is, is that a lot of this text has been truncated, has been merged, and isn't decipherable even by me as a human. And if I sent this to a model, I don't think I could trust that the model could extract specifics from, say, for example, this table, because the table has been sort of just spit out linearly. And this information isn't fit for most use cases where I need to ask questions or have an agent do validation and extraction on this source data. So this isn't going to cut it, right? There's undesired page headers. We don't understand the table. And where's the content from the image, right? It's not even there. When we're using frontier models, this is not bad, right? But quite expensive. I'm sending this to a model that's maybe $30 per million output tokens. You can see how this can get quite expensive as I scale this up to dozens or hundreds or, in a lot of cases, thousands of PDFs that organizations have to work through to use an AI application. And the differences between maybe a 5.1 of a model that was depreciated and a 5.2 version of a model make it tricky to have structured output that's consistent every single time. And so while it might be good quality, and I can see that most of this looks accurate in the exported markdown, we might be susceptible to hallucinations because models are non-deterministic. And this is really tricky at scale. And so what is the middle ground? Well, that's where Dockling comes in. It's a fast and cheap and most importantly, local CLI and library that I can use to take various types of input sources and convert this to markdown, JSON, and a pydantic data type that I can use in my applications and that I can scale up if I have thousands of different types of formats that need to be used or translated to something like markdown. And so we rendered this as HTML, but you can see here we've got this export of the specific table that we had here previously, because as a PDF, this data type is spread out and it's proprietary. So it's hard to extract this type of data from the source PDF here. And I'll show you how Dockling does it by using a combination of OCR and specific vision models that extract the format and allow me to do things like structured output if I only say, for example, want a specific column to be outputted from a content source. So it's really cool. And I've been using it at Red Hat, my team uses it because we have thousands of PDFs that we need to work with specifically from product documentation. But also we have content like images that we want to extract using vision models. And what Dockling does with a here previously, because as a PDF, this data type is spread out and it's proprietary. So it's hard to extract this type of data from the source PDF here. And I'll show you how Dockling does it by using a combination of OCR and specific vision models that extract the format and allow me to do things like structured output if I only want a specific column to be outputted from a content source. So it's really cool. And I've been using it at Red Hat, my team uses it because we have thousands of PDFs that we need to work with specifically from product documentation. But also we have content like images that we want to extract using vision models. And what Dockling does with a pip install Dockling is allow me to convert single documents or websites or anything else to markdown or whatever file type I want, and be able to work with the page layout so I don't lose the consistency in the structure of the source document. And there's a lot of other integrations and for situations where maybe you need to run this locally, or you don't want to pay for a service or you have an air-gapped environment. So you're able to do this using this open source project that is part of the Linux Foundation. So before I show you the quick demo that I want to highlight the project, I want to talk a little bit about scale and cost. As I just mentioned it there earlier, but there is a public use case I want to show you from Landro at Hugging Face where he compared a source of common crawl PDFs and where he did a little bit of pre-work on them and actually extracted the structure using OCR and using Dockling in order to remove certain parts and clean this up for the FineWeb PDFs export, which is thousands of tokens from PDFs around the web that you could use for training a model, etc. But the two comparisons he did using GPU and using CPU for Dockling allow him to do this at 50 times cost savings compared to VLMs and OCR naively. And so this is what's really cool is that you can really scale this up. And he did this on CPU and not needing GPU, which is really cool. And it's not just document conversion, right? So his example was getting things ready from PDFs. But let's say we have images, right? This image with a vision language model was able to be described and annotated, right from this specific image right here. Now we have all of this really important context that can be used in a RAG application as an example, but also just for the end user to be able to understand what's happening in this image. And say, for example, the SME doesn't work there anymore at the organization. Now we have a way to understand what this is with the help of an LLM, whether proprietary or local. And then finally, it's not just document conversion and annotation, but also structured output. Say for example, I have this invoice and I need to extract specifically the bill number, the total invoice price, the name of the sender, we're able to get that in a format that is structured. So I don't have to worry about it pulling in the headings and titles. No, I just want the total invoice price and the bill number. And I can get that in a format that's Pydantic, but also just super simple if I'm just trying to get a couple things out of a huge document. So let's hop over to the demo. Let me show you what I'm talking about. And by the way, we're going to be using this docling workshop repository for the demos today. So feel free to check it out. In this first example, we're going to be using docling to convert a popular file type like a PDF, because remember that data is the foundation for all AI systems. And in order to leverage that data, we have to properly ingest different file formats with accuracy. And without doing that, we can lose information or information can become unreadable from tables and diagrams and images. So with docling, what we're able to do with a simple pip install down here is start processing some of these different file types. So I'll show you how we do that. First off, we're going to be importing essential components like the document converter, as well as some other dependencies. And I'll show you the simplest way, which is to just start out with a PDF that we have online, which is Dockling's own research paper. And here you can see we have different titles and subtitles, we have components such as images and captions that we have here. And at the bottom of this PDF, we also have something like a table here that we need to extract as well as images that might be helpful for our AI application or agent. So let me come back to the notebook here. And I'll show you how we do this simple conversion by using the Dockling document converter, and exporting this to markdown. So here you can see is a rough example of how fast this can be for an eight page PDF, to be able to export this in a way that an LLM can start using right. But at the same time, this is a Pydantic data type. So you can see we can explore this PDF, the number of pages and tables and see what is on which page of this PDF, and export this into a variety of different formats, like markdown, HTML, dictionary, and much more. But the real value here is not just with basic text and columns, but working with tables. So this PDF here has a variety of different tables that we want to be able to extract. And so by doing a converter for this specific document, we can then extract these tables here, and be able to export this to a data frame. So we can render this out in our Jupyter notebook here. So as I run this cell, you can see that we've exported eight different tables from that source PDF. And we can list these, but also be able to get these in a format ready to use in a RAG application, or just to query with our LLM. So we've taken a look at how to pull text and tables from a PDF. But what about visualizing and extracting images from that source document as well. Now for this specific example, what we'll do is set up a PDF pipeline that will allow us to scale up the images and push this into a document converter. So now when I go to inspect the images and picture content in the cells, we can see this nicely mapped out where we have the picture. So the source image, the caption of that, and all of the embedded text elements that we could use in some type of retrieval augmented generation application to ask questions about what's happening in these different photos in our source PDF. I think what would help here is also to be able to visualize the document layout by using the bounding boxes provided by the layout visualizer here. So here, what we're going to do is visualize all of the different elements and components that can be extracted from, say, for example, that source PDF. So section headers or text or subtitles or different components, such as that photo that we just pulled and extracted from the PDF. So this is one of the models that Dockling provides, which can be used for situations where you might have personally identifiable information from a customer that you want to remove from a source document type before you extract that into your application. We can also use vision About what's happening in these different photos in our source PDF. I think what would help here is also to be able to visualize the document layout by using the bounding boxes provided by the layout visualizer here. So here, what we're going to do is visualize all of the different elements and components that can be extracted from, say, for example, that source PDF. So section headers or text or subtitles or different components, such as that photo that we just pulled and extracted from the PDF. So this is one of the models that DocLean provides, which can be used for situations where you might have personal identifiable information from a customer that you want to remove from a source document type before you extract that into your application. We can also use vision language models in order to enrich the source images and diagrams that might be in these document types using something like OLAMA or a third party LLM. So what we're going to do here is set up a PDF pipeline that's going to use a local running granite model and say, hey, give us a detailed description of what's happening in this image. And with that document converter, we're going to go ahead and display that enriched document by calling the open AI endpoint with OLAMA that we have running locally at the completions endpoint. And so here we have an annotated caption of what's happening with that DocLean pipeline image that we were taking a look at earlier, where originally we were just pulling the caption. But now we can use a vision language model to describe what's happening in these photos. And with all this additional information, this can help us to build a solid RAG pipeline to where we can do questioning and answering over our source data. And for this example, I want to show you what's known as chunkless RAG or agentic RAG using docling. And now by starting off with our document outline, so processing a PDF with docling like we just did, we can allow an LLM to be able to pick the most relevant part of the document that is related to the user's question and pull that full text from the docling document itself to try to answer that question for the user. And this can run in an agentic loop. And what's really important here is that we're doing RAG but without having to use a chunker or embedding model or vector database, etc. So the index ends up being the markdown outline of the document. Now when the user asks a question, the entire retrieval index typically would be thousands of vectors in a database where we would do semantic similarity to make sure these sections are similar to the user's question. But for us, what we're going to do is be able to see a markdown outline of the document with each section summary outline. So if the LLM is looking for something about getting started with docling, well, it can just pull from this reference of text to see that docling can be installed for PyPI. And that's the entire retrieval index. So let's say we have a query, what are the main AI models using docling? Well, we're setting up a RAG agent here to be able to iterate on that specific question about five times. So we see that there are 20 sections available when the users ask the question. And we're going to search for that specific part of text that talks about the AI models and be able to determine, is this relevant to answering the question. And you can see here the final answer in one iteration was pulled from that source material without having to go in a vector database, but instead search that docling document structure for the specific text. So it's a quite interesting way to be able to answer users questions through this chunkless retrieval augmented generation pattern. And while this first example might have been simple, what we can do is also pull the IBM 2025 annual report into the context here, which has 418 sections. So it's much larger and ask a question like, what was Red Hat's revenue growth in 2025? And how did it contribute to overall software segment? And so here we're going to be iterating multiple times to figure out, is this section relevant to the user's question? And if not, we need to pull in more information. And so that is how chunkless RAG can work in a situation where you're using a tool like docling. But what happens when we have hundreds or hundreds of thousands of PDFs that we want to have processed? Well, this is where we can deploy docling as a REST API service using something that's known as docling serve. This allows us to scale things up and to run this as a microservice as a container or through Kubernetes. So when we set things up here, we're going to do a pip install docling serve. And we're going to be able to serve this from a CLI with docling serve on a specific port. Or once we had that server started, be able to send things to that specific endpoint with different types of options and arguments, do we want OCR? Do we want a specific backend? Or do we want images annotated? And so that's how we can scale things up and allow this API endpoint to be able to handle hundreds or thousands of different document types at a time. And let's say for example, that you're trying to build an AI agent, you can also use the docling MCP server. So this allows us to automate things with our AI agent and give it the capabilities that docling has through that model context protocol and allow us to standardize the communication between, say, for example, cloud code or continue in our developer CLI to the MCP server, which can handle the document processing for us without us having to know all of those different arguments and commands. So it makes it quite easy. And the example I show here is by using one of the client models on my own Mac itself and connecting this to my VS code instance. So let's understand the tools that are available. For the MCP server, we have conversion tools, generation tools, say, for example, if I want to process a specific part of a PDF and manipulation tools. And this is all provided to the LLM and the agent that we're going to be using with the MCP server. So here, I'm just checking that my local MLX server to run an LLM is running. And it looks like we've got Qwen 3.6 here. And we're going to verify that the docling MCP server is also running, which would be done using UVX here. And so we'll do that in this cell here to make sure that the MCP server is available. Now, in cloud code or codex or another type of AI application, we're going to install the extension. So for us, this means adding an MCP server in the config.yaml. And now at this point, we can use a model and an MCP server to do things such as, convert this document, give me a summary, or create a document with a section of action items and pull in a list from another PDF and export that as markdown. So we can use all of those docling components through the MCP server in order to agentically process and parse these documents using an AI agent like cursor or cloud code, or one of the many open source options that are out there. Now let's head back to the slides and wrap things up. So let's put it all together with docling. We've seen that we can take a PDF into a format like markdown or JSON with a fully local operation from our own machine without even needing model and an MCP server to do things such as convert this document, give me a summary, or create a document with a section of action items and pull in a list from another PDF and export that as markdown. So we can use all of those docking components through the MCP server in order to agentically process and parse these documents using an AI agent like cursor or cloud code, or one of the many open source options that are out there. Now let's head back to the slides and wrap things up. So let's put it all together with docking. We've seen that we can take a PDF into a format like markdown or JSON with a fully local operation from our own machine without even needing a GPU. It's fast, it's cheap, and most importantly, it's open source. So I encourage you to check it all out. Behind the scenes, there's different pipelines that use a combination of OCR and layout analysis to structure everything together, put it together as a Pydantic dockling document that then you can use to export to different types of formats, create datasets, chunk that using the hybrid chunker, and do so much more. It integrates with a lot of different RAG frameworks and agentic systems and harnesses. So feel free to try it out. And I want to give a big thank you to the AI engineering team for having me on. This is something that I'm really passionate about with processing documents, but also doing this in an open source way because we're here at Red Hat and we love open source. So feel free to check out the slides, connect with me on LinkedIn. I appreciate the opportunity. Enjoy the conference and keep up with the AI engineering ecosystem. Things are looking super bright right now for the open source world and for AI engineers in general. So see you next time. And thanks for watching. you understand what I mean, right? All this data needs to be transformed into something that an LLM can actually understand. And that's why by the end of this session, you'll understand how to extract structure between raw enterprise documents and use it to power better downstream AI systems like RAG and agents. So let's get started. Now, I think Jensen from NVIDIA made this point super clear at his keynote that unstructured data is becoming this new context layer for AI. And the reality, though, for many teams, and I know this personally working at Red Hat, that PDFs and data are spread across dozens of different systems. So we've got a lot to cover today. As you might know, a large majority of the world's data is unstructured. And so no matter what model you're using, if you're working with data like PDFs and unstructured types of formats, this is a bit tricky to work with. Because there are solutions out there, but they might be proprietary or require you to send your private data to someone else's server. And for not just text, how do we take documents and their graphs or tables and images to formats that LLMs can understand like Markdown or JSON? And I'm going to show you how in the session today, because we're going to be using an open source tool, part of the Linux foundation that is called Dockling, and learn about extraction, parsing, chunking, and much more. And I've got some live demos for you. So we're going to have some fun. And just in case you'd like it, we have the session slides here on the right, and a little overview of the specifics I'll be showing you in today's session. But without further ado, let's get started. So why is there a need for advanced document processing? As I briefly mentioned before, you might have a lot of technical documentation or meeting minutes or different types of documents and invoices that you need to use and maybe RAG or different type of applications where the context is provided to an LLM. So whether it's RAG or retrieval augmented generation to answer questions based on this data, or you're using this to fine tune a new specialized model, well, data is this key ingredient behind those applications. And it doesn't matter if you're using, you know, Nvidia acceleration or an open source or proprietary model, that data and the way you process it is the key determining factor in whether your answer is going to be correct or incorrect for the user or customer at the end of the day. And that's what's most important. And how important is it? Well, I have this viral tweet from earlier where 20 scientific papers now feature a new nonsensical term that doesn't exist because AI misinterpreted a very old article that was scanned and taken to a PDF, merging two different words from two different columns in this PDF. And because researchers are using these models in order to help them write, now we have different types of scientific papers that all feature this word and are even being cited by other people. And so that's how important it is to make sure that the data that we're processing is processed in a way that's accurate and not hallucinated and able to be used confidently in our applications that we're delivering to users and customers. So it's quite important. Now, if we were to use a tool like dockling that I can run on my own machine, you could see that these two words are quite far away from each other and shouldn't have been combined in the first place. But that's how we're going to learn about extracting this text here in a second. So if we were to try a simple PDF parser for a PDF like this, that includes a table here, it has an image, there's captions, and there's regular sections of text. Well, we might get an answer like this here on the right in Markdown. You know, this might be very fast and cheap to run, even on CPU. But the issue is, is that a lot of this text has been truncated, has been merged, and isn't decipherable even by me as a human. And if I sent this to a model, I don't think I could trust that the model could extract specifics from, say, for example, this table, because the table has been kind of just spit out linearly. And this information isn't fit for most use cases where I need to ask questions or have an agent do validation and extraction on this source data. So this isn't going to cut it, right? There's undesired page headers. We don't understand the table. And where's the content from the image, right? It's not even there. When we're using frontier models, this is kind of not bad, right? But quite expensive. I'm sending this to a model that's maybe $30 per million output tokens. You can see how this can get quite expensive as I scale this up to dozens or hundreds or, in a lot of cases, thousands of PDFs that organizations have to work through to use an AI application. And the differences between maybe a 5.1 of a model that was depreciated and a 5.2 version of a model make it tricky to have structured output that's consistent every single time. And so while it might be good quality, and I can see that most of this looks accurate in the exported markdown, we might be susceptible to hallucinations because models are non-deterministic. And this is really tricky at scale. And so what is the middle ground? Well, that's where Dockling comes in. It's a fast and cheap and most importantly, local CLI and library that I can use to take various types of input sources and convert this to markdown, JSON, and a pydantic data type that I can use in my applications and that I can scale up if I have thousands of different types of formats that need to be used or translated to something like markdown. And so we rendered this as HTML, but you can see here we've got this export of the specific table that we had here previously, because as a PDF, this data type is spread out and it's proprietary. So it's hard to extract this type of data from the source PDF here. And I'll show you how Dockling does it by using a combination of OCR and specific vision models that extract the format and allow me to do things like structured output if I only say, for example, want a specific column to be outputted from a content source. So it's really cool. And I've been using it at Red Hat, my team uses it because we have thousands of PDFs that we need to work with specifically from product documentation. But also we have content like images that we want to extract using vision models. And what Dockling does with a pip install Dockling is allow me to convert single documents or websites or anything else to markdown or whatever file type I want, and be able to work with the page layout so I don't lose the consistency in the structure of the source document. And there's a lot of other integrations and for situations where maybe you need to run this locally, or you don't want to pay for a service or you have an air-gapped environment. So you're able to do this using this open source project that is part of the Linux foundation. So before I show you the quick demo that I want to highlight the project, I want to talk a little bit about scale and cost. As I just mentioned it there earlier, but there is a public use case I want to show you from Landro at Hugging Face where he compared a source of common crawl PDFs and where he did a little bit of pre-work on them and actually extracted the structure using OCR and using Dockling in order to remove certain parts and clean this up for the find PDFs export, which is thousands of tokens from PDFs around the web that you could use for training a model, etc. But the two comparisons he did using GPU and using CPU for Dockling allow him to do this at 50 times of a cost savings compared to VLMs and OCR naively. And so this is what's really cool is that you can really scale this up. And he did this on CPU and not needing GPU, which is really cool. And it's not just document conversion, right? So his example was getting things ready from PDFs. But let's say we have images, right? This image with a vision language model was able to be described and annotated, right from this specific image right here. Now we have all of this really important context that can be used in a RAG application as an example, but also just for the end user to be able to understand what's happening in this image. And say, for example, the SME doesn't work there anymore at the organization. Now we have a way to understand what this is with the help of an LLM, whether proprietary or local. And then finally, it's not just document conversion and annotation, but also structured output. Say for example, I have this invoice and I need to extract specifically the bill number, the total invoice price, the name of the sender, we're able to get that in a format that is structured. So I don't have to worry about, hey, it's going to pull in like the headings and titles. No, I just want the, you know, total invoice price and the bill number. And I can get that in a format that's pydantic, but also just super simple if I'm just trying to get a couple things out of a huge document. So let's hop over to the demo. Let me show you what I'm talking about. And by the way, we're going to be using this docling workshop repository for the demos today. So feel free to check it out. In this first example, we're going to be using docling to convert a popular file type like a PDF, because remember that data is the foundation for all AI systems. And in order to leverage that data, we have to properly ingest different file formats with accuracy. And without doing that, we can lose information or information can become unreadable from tables and diagrams and images. So with docling, what we're able to do with a simple pip install down here is start processing some of these different file types. So I'll show you how we do that. First off, we're going to be importing essential components like the document converter, as well as some other dependencies. And I'll show you the simplest way, which is to just start out with a PDF that we have online, which is doclings own research paper. And here you can see we have different titles and subtitles, we have components such as images and captions that we have here. And at the bottom of this PDF, we also have something like a table here that we need to extract as well as images that might be helpful for our AI application or agent. So let me come back to the notebook here. And I'll show you how we do this simple conversion by using the doclings document converter, and exporting this to markdown. So here you can see is a rough example of how fast this can be for a eight page PDF, to be able to export this in a way that an LLM can start using right. But at the same time, this is a Pydantic data type. So you can see we can explore this PDF, the number of pages and tables and see what is on which page of this PDF, and export this into a variety of different formats, like markdown, HTML, dictionary, and much more. But the real value here is not just with basic text and columns, but working with tables. So this PDF here has a variety of different tables that we want to be able to extract. And so by doing a converter for this specific document, we can then extract these tables here, and be able to export this to a data frame. So we can render this out in our Jupyter notebook here. So as I run this cell, you can see that we've exported eight different tables from that source PDF. And we can list these, but also be able to get these in a format ready to use in a RAG application, or just to query with our LLM. So we've taken a look at how to pull text and tables from a PDF. But what about visualizing and extracting images from that source document as well. Now for this specific example, what we'll do is set up a PDF pipeline that will allow us to scale up the images and push this into a document converter. So now when I go to inspect the images and picture content in the cells, we can see this nicely mapped out where we have the picture. So the source image, the caption of that, and all of the embedded text elements that we could use in some type of retrieval augmented generation application to ask questions about, hey, what's happening in these different photos in our source PDF. I think what would help here is also to be able to visualize the document layout by using the bounding boxes provided by the layout visualizer here. So here, what we're going to do is visualize all of the different elements and components that can be extracted from, say, for example, that source PDF. So section headers or text or subtitles or different components, such as that photo that we just pulled and extracted from the PDF. So this is one of the models that DocLean provides, which can be used for situations where you might have personal identifiable information from a customer that you want to remove from a source document type before you extract that into your application. We can also use vision language models in order to enrich the source images and diagrams that might be in these document types using something like OLAMA or a third party LLN. So what we're going to do here is set up a PDF pipeline that's going to use a local running granite model and say, hey, give us a detailed description of what's happening in this image. And with that document converter, we're going to go ahead and display that enriched document by calling the open AI endpoint with OLAMA that we have running locally at the completions endpoint. And so here we have an annotated caption of what's happening with that DocLean pipeline image that we were taking a look at earlier, where originally we were just pulling the caption. But now we can use a vision language model to describe what's happening in these photos. And with all this additional information, this can help us to build a really solid rag pipeline to where we can do questioning and answering over our source data. And for this example, I want to show you what's known as chunkless rag or agentic rag using docling. And now by starting off with our document outline, so processing a PDF with docling like we just did, we can allow an LLN to be able to pick the most relevant part of the document that is related to the user's question and pull that full text from the docling document itself to try to answer that question for the user. And this can run in an agentic loop. And what's really important here is that we're doing rag but without having to use a chunker or embedding model or vector database, etc, etc. So the index ends up being the markdown outline of the document. Now when the user asks a question, the entire retrieval index typically would be thousands of vectors in a database where we would do semantic similarity to make sure hey, these sections are similar to the user's question. But for us, what we're going to do is be able to see a markdown outline of the document with each section summary outline. So if the LLN is looking for something about getting started with docling, well, it can just pull from this reference of text to see that docling can be installed for PyPy. And that's the entire retrieval index. So let's say we have a query, what are the main AI models using docling? Well, we're setting up a rag agent here to be able to iterate on that specific question about five times. So we see that there are 20 sections available when the users ask the question. And we're going to search for that specific part of text that talks about the AI models and be able to determine hey, is this relevant to answering the question. And you can see here the final answer in one iteration was pulled from that source material without having to go in a vector database, but instead search that docling document structure for the specific text. So it's a quite interesting way to be able to answer users questions through this chunkless retrieval augmented generation pattern. And while this first example might have been simple, what we can do is also pull the IBM 2025 annual report into the context here, which has 418 sections. So it's much larger and ask a question like, hey, what was Red Hat's revenue growth in 2025? And how did it contribute to an overall software segment? And so here we're going to be iterating multiple times to figure out, hey, is this section relevant to the user's question? And if not, we need to pull in more information. And so that is how chunkless RAG can work in a situation where you're using a tool like docling. But what happens when we have hundreds or hundreds of 1000s of PDFs that we want to have processed? Well, this is where we can deploy docling as a REST API service using something that's known as docling serve. This allows us to scale things up and to run this as a microservice as a container or through Kubernetes. So when we set things up here, we're going to do a pip install docling serve. And we're going to be able to serve this from a CLI with docling serve on a specific port. Or once we had that server started, be able to send things to that specific endpoint with different types of options and arguments like, hey, do we want OCR? Do we want a specific backend? Or do we want images annotated? And so that's how we can scale things up and allow this API endpoint to be able to handle hundreds or thousands of different document types at a time. And let's say for example, that you're trying to build an AI agent, you can also use the docling MCP server. So this allows us to automate things with our AI agent and give it the capabilities that docling has through that model context protocol and allow us to standardize the communication between, say, for example, cloud code or continue in our developer CLI to the MCP server, which can handle the document processing for us without us having to know all of those different arguments and commands. So it makes it quite easy. And the example I show here is by using one of the clint models on my own Mac itself and connecting this to my VS code instance. So let's understand the tools that are available. For the MCP server, we have conversion tools, generation tools, say, for example, if I want to process a specific part of a PDF and manipulation tools. And this is all provided to the LLM and the agent that we're going to be using with the MCP server. So here, I'm just checking that my local MLX server to run an LLM is running. And it looks like we've got Quinn 3.6 here. And we're going to verify that the docling MCP server is also running, which would be done using UVX here. And so we'll do that in this cell here to make sure that the MCP server is available. Now, in cloud code or codex or another type of AI application, we're going to install the extension. So for us, this means adding an MCP server in the config.yaml. And now at this point, we can use a model and an MCP server to do things such as, hey, convert this document, give me a summary, or create a document with a section of action items and pull in a list from another PDF and export that is markdown. So we can use all of those docking components through the MCP server in order to agentically process and parse these documents using an AI agent like cursor or cloud code, or one of the many open source options that are out there. Now let's head back to the slides and wrap things up. So let's put it all together with docking. We've seen that we can take a PDF into a format like markdown or JSON with a fully local operation from our own machine without even needing a GPU. It's fast, it's cheap, and most importantly, it's open source. So I encourage you to check it all out. Behind the scenes, there's different pipelines that use a combination of OCR and layout analysis to structure everything together, put it together as a Pydantic dockling document that then you can use to export to different types of formats, create datasets, chunk that using the hybrid chunker, and do so much more. It integrates with a lot of different RAG frameworks and agentic systems and harnesses. So feel free to try it out. And I want to give a big thank you to the AI engineering team for having me on. This is something that I'm really passionate about with processing documents, but also doing this in an open source way because, you know, we're here at Red Hat and we love open source. So feel free to check out the slides, connect with me on LinkedIn. I appreciate the opportunity. Enjoy the conference and keep up with the AI engineering ecosystem. Things are looking super bright right now for the open source world and for AI engineers in general. So see you next time. And thanks for watching.