Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI
Description
Full implementation is open-source: https://github.com/iusztinpaul/ai-research-os-workshop Agent Engineering: Building Multi-Agent Systems Course: https://academy.towardsai.net/courses/agent-engineering Turning thousands of notes, videos, documents, and repositories into usable AI context requires more than a bigger context window. It requires memory and context engineering: organizing sources, indexing what matters, and loading only what the model needs. The talk shows how the authors turn an Obsidian vault of 10,000+ notes, documents, videos, and repositories from a passive knowledge archive into live context for AI agents. A system that the authors use daily as their personal AI research OS for writing code and creating content. You'll learn to design a deep research algorithm that runs across both the open web and your personal "Second Brain", then store what it finds as a research memory you and your agents can maintain, visualize, and grow. All based on the authors' 3 major iterations of the system over the past 18 months. You'll walk away knowing how to: - Choose the right tool for the job: Codex/Claude Code vs. NotebookLM vs. RAG vs. a personalized research memory - Design deep-research pipelines across Obsidian, NotebookLM, GitHub, Google Drive, Readwise, and YouTube - Build a token-efficient memory layer from plain files, not a vector or graph DB - Plug tens of thousands of personal notes into an LLM knowledge (or wiki) base that scales - Implement the memory layer between your Second Brain and any agent harness (Codex, Claude Code, or your own) For: Engineers ready to move beyond hoarding notes and turn their "Second Brain" into a context that their AI agents can research, maintain, and grow. Speaker info: Paul Iusztin - X: https://x.com/pauliusztin_ - in: https://www.linkedin.com/in/pauliusztin/ - youtube: https://www.youtube.com/@itsdecodingai Louis-François Bouchard - X: https://x.com/Whats_AI - in: https://www.linkedin.com/in/whats-ai/ - yo
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: Build a personal AI research OS by combining a deep research algorithm with a file-based, LLM-queryable wiki layer that evolves with your questions, avoiding the complexity of vector databases while maintaining context across projects.
- Why it matters: Solves the core problem of losing research context between work sessions and enables compound learning from your second brain (10,000+ notes) without manually curating links or rebuilding context every time you switch projects.
- Best use: Watch the architecture explanation (V1→V3 evolution), skip setup sections, focus on the file-based indexing strategy and how the wiki layer works—then adapt the open-source code to your own workflow.
Executive Summary
Paul Iusztin and Louis-François Bouchard built an AI Research OS to solve a universal problem: they have 10,000+ notes across Obsidian, Readwise, Notion, and other tools, growing by ~250 files/month, but lose research context between projects. Existing tools (ChatGPT, Codex, Notebook LM) either lack memory or lock you into proprietary systems. Their solution: a three-layer system that runs deep research queries across your second brain + public web, stores raw sources immutably, maintains a queryable index, and generates a living wiki that evolves with every question you ask.
The system progressed through three versions. V1 used manually curated golden links + a deep research algorithm (multi-agent query rounds) to generate static research.md files—useful for their AI course but not reusable. V2 replaced golden links by querying the second brain directly (Obsidian, Readwise, Notebook LM, GitHub repos) using the same deep research loop, making the system self-referential. V3 added the wiki layer: instead of compiling research into a single static file, it stores raw sources, builds an index.yaml catalog with summaries/metadata, and generates wiki derivatives (concepts, entities, comparisons, open questions) that agents traverse hierarchically to save tokens.
The file-based design deliberately avoids vector databases and knowledge graphs. An agent reads index.yaml (catalog of all sources + summaries), then drills down: first checks the source wiki page (executive summary), then wiki derivatives (concept/entity notes), and only reads the raw source if needed. Every question leaves a trace—new concept files, comparisons, logged queries—so the wiki grows organically. The system is scoped per project (not your entire second brain) and integrates with Claude Code/Codex via custom skills.
Demos show three use cases: (1) researching agentic engineering by ingesting hand-picked links + running light/deep/fast query modes (1–3 rounds, 3–6 queries/round); (2) ingesting three GitHub repos (OpenCode, Poe, Hermes) to extract architecture notes, memory systems, permission flows, then auto-generating comparisons; (3) ingesting arbitrary URLs with no setup beyond Git/curl. The repo is open-source and designed to teach memory/context management techniques—not to be a polished product. Next improvements: better source provenance, memory compaction, and linting.
Key Takeaways
- Claim: The bottleneck isn't providing more context to LLMs—it's how you manage context and memory over time so agents can leverage past research without reingesting everything. | Evidence: Louis-François notes: 'With an agent, the context window becomes everything. The database, the file system, the memory, the reasoning space. It has to do it all. And when you stop the conversation, it loses everything.' Paul adds that rebuilding context for every new project is expensive (high token cost, 10–20 min wait times). | Caveat: This assumes you have a second brain worth preserving (5,000+ notes). For ad-hoc queries or one-off tasks, ChatGPT/Codex is faster. | Implication: Ken should think of AI systems as stateful collaborators, not stateless tools. The value compounds when memory persists and evolves across sessions—critical for agent-driven workflows or long-term research projects. | Timestamp: 03:15
- Claim: A file-based index (YAML + markdown) with hierarchical traversal is more practical for personal wikis than vector databases or knowledge graphs. | Evidence: Paul: 'Forget the infrastructure you think you need—vector databases, knowledge graphs, semantic search. All that is beautiful but adds a lot of complexity, especially for personal research OSs you want to use lightly.' The agent reads index.yaml → source wiki page (executive summary) → wiki derivatives (concepts/entities) → raw source only if needed, saving tokens at each step. | Caveat: This approach doesn't scale to production systems with thousands of concurrent users or real-time updates. Paul explicitly says it's 'not a polished product' and prioritizes inspectability (Obsidian graph view) over performance. | Implication: For Ken's agent systems, consider file-based memory for prototyping or personal use cases where inspectability and simplicity matter more than query speed. Defer to vector DBs for production scale. | Timestamp: 18:45
- Claim: The wiki evolves with every question you ask—not just when you ingest new sources—so it becomes a living reflection of your learning gaps and thought process. | Evidence: Paul: 'Every question leaves a trace. The LLM can create a new concept file, a new node file, a new comparison file, and every question is tracked into a log. The wiki doesn't evolve only when you ingest new data or do a deep research round—it evolves as you start talking with it.' | Caveat: No mention of pruning or versioning stale concepts. Over time, the wiki could accumulate noise or outdated thinking without manual curation. | Implication: This is a meta-learning tool: Ken could use it to audit his own knowledge gaps, spot patterns in his questions, or identify areas where he's repeating research. Also useful for content/course creators to avoid duplicating topics. | Timestamp: 22:10
- Claim: The deep research algorithm uses multi-agent query rounds (orchestrator + sub-agents) to generate 40–50 sources per topic, then ranks them by relevance to avoid context bloat. | Evidence: Paul: 'The orchestrator creates multiple questions based on the initial topic and scraped context. Each agent manages its own question, uses Gemini grounded in Google to gather resources, creates executive summaries, then passes info back to the main agent. We did this for three rounds—six queries per round—ending up with 40–50 links. We applied a ranking algorithm to find high-signal notes, fully scraped the top sources, and kept summaries for the rest.' | Caveat: Token cost scales with query depth (light = 1 round/3 queries; deep = 3 rounds/6 queries). Paul warns this is 'expensive' and time-consuming (10–20 minutes even on light mode). | Implication: Ken should reserve deep research mode for high-stakes projects (e.g., new product strategy, major content series) and use light/fast modes for incremental learning. The ranking step is critical—without it, you drown in noise. | Timestamp: 13:20
- Claim: The system deliberately doesn't touch your raw second brain (Obsidian vault)—it creates project-scoped wikis that reference but never mutate your notes. | Evidence: Paul: 'Obsidian is an immutable snapshot that the LLM never touches. This is my data. I don't want the LLM to touch my personal notes that I manually write. Whenever we start working on a new project, we reference this second brain and scope it down to our own project.' | Caveat: This means the system doesn't automatically backfill insights into your main vault. You manually decide what to promote from project wikis to your master notes. | Implication: Ken's agent workflows should separate immutable source of truth (personal notes, code repos) from mutable derived artifacts (wikis, summaries). This prevents runaway agent edits and preserves human curation as the ultimate arbiter. | Timestamp: 23:50
- Claim: The system works for three distinct use cases: (1) deep research on a topic using your second brain + web, (2) ingesting and analyzing GitHub repos to extract architecture/design patterns, (3) quick wiki creation from arbitrary URLs with zero setup. | Evidence: Demo 1: Paul ingests hand-picked links + runs deep research on agentic engineering, generating concepts (agent loop, tool registry, context compaction) and comparisons (agentic RAG vs. file systems). Demo 2: Clones OpenCode/Poe/Hermes repos, extracts notes on memory systems, permission flows, generates cross-repo comparisons. Demo 3: Ingests three random URLs using only Git/curl—no Obsidian or Readwise required. | Caveat: Setup friction varies: demos 2–3 need only Git/curl; demo 1 requires Obsidian, Readwise, Notebook LM CLIs and authentication. The README lists dependencies but Paul skips setup in the video. | Implication: Ken can start with demo 3 (URL ingestion) to prototype quickly, then layer in second-brain connectors as needed. The repo-ingestion use case (demo 2) is especially powerful for competitive analysis or understanding open-source tools. | Timestamp: 26:30 / 30:10 / 32:45
- Claim: The repo is designed to teach memory/context management techniques for AI engineering—not to be a production-ready product—so customization is expected and encouraged. | Evidence: Louis-François: 'The core goal is that we want you to adapt it for your needs. Our goal is to teach AI engineering—it's not to build the next best product. It's a builder workflow through Claude Code and Codex in the terminal. We didn't add Google Drive, Notion, Slack connectors because they're not useful to us yet, but you can easily implement them—YouTube transcript took one prompt.' | Caveat: This means ongoing maintenance and feature additions are not guaranteed. The repo is a teaching artifact first, tool second. Source provenance, memory compaction, and linting are acknowledged weaknesses. | Implication: Ken should treat this as a reference architecture, not a dependency. Fork it, strip out what you don't need, and extend it where your workflow diverges. The educational value is in understanding the file-based indexing + wiki evolution pattern. | Timestamp: 35:20
Detailed Brief
The problem: losing research context between projects
- Claims: Both speakers have 10,000+ notes across fragmented tools (Obsidian, Readwise, Notion, Google Drive) growing by ~250 files/month; Existing tools fail: ChatGPT/Codex lose context between sessions; Notebook LM isn't agent-native or code-friendly; vector DB solutions require heavy infrastructure; The real bottleneck isn't providing more context—it's managing context and memory so agents can reuse past research without reingesttion
- Evidence: Paul: 'Whenever I actually want to start working on something, I never recall what I have in my second brain. Or I have to spend a ton of time finding meaningful notes.'; Louis-François: 'With an agent, the context window becomes everything—the database, file system, memory, reasoning space. When you stop the conversation, it loses everything.'; Paul notes that rebuilding context is 'expensive' (high token cost, 10–20 min wait times)
- Caveats: This problem only matters if you have a large, active second brain (thousands of notes); For quick queries or one-off tasks, ChatGPT/Codex is still faster and simpler
- Implications: Ken should think of AI systems as stateful collaborators that compound value over time; Memory and context management are the real engineering challenges in agent workflows—not just prompt engineering or model selection
V1: Deep research with golden links → static research.md
- Claims: First version used manually curated 'golden links' as seed context for a multi-agent deep research algorithm; Orchestrator agent generates multiple questions; sub-agents query Gemini grounded in Google, scrape sources, create executive summaries; Runs for 3 rounds × 6 queries = ~40–50 sources, applies ranking to identify high-signal sources, fully scrapes top results, keeps summaries for the rest; Output: single static research.md file—worked for their AI course (35 lessons) but not reusable
- Evidence: Paul: 'We used golden links as seed for context. During query rounds, each agent managed its own question, used Gemini grounded in Google to gather resources, passed summaries back to main agent which aggregated everything.'; Ranking algorithm compares each source against the initial topic to find highest-signal content
- Caveats: Golden links must be hand-picked, adding manual labor; Static output means you must rerun the entire expensive process to update research
- Implications: Multi-agent query rounds with ranking is a reusable pattern for Ken's agent systems—orchestrator delegates, sub-agents specialize, main agent aggregates and dedupes; Static outputs are useful for one-time deliverables (reports, course lessons) but don't compound over time
V2: Self-referential research querying your second brain
- Claims: V2 replaces golden links by querying the second brain directly (Obsidian, Readwise, Notebook LM, GitHub repos); Same deep research algorithm but now targets personal notes + public web instead of just public web; Eliminates manual link curation—your second brain becomes the source of golden links organically
- Evidence: Paul: 'Now we plug in all our second brains—Obsidian, Readwise, Notebook LM, GitHub. We target our queries from the deep research algorithm to our second brain plus the public web.'; Still outputs static research.md, still expensive to rerun
- Caveats: Requires connectors for each second-brain tool (Obsidian, Readwise, Notebook LM CLIs + auth); Still outputs static research.md—doesn't solve the reusability problem
- Implications: Self-referential querying is the key innovation: your second brain becomes the filter for signal, reducing reliance on generic web search; Ken should think about 'agent memory' as bidirectional—agents query your notes, but your notes also guide what agents retrieve
V3: Living wiki layer with file-based indexing
- Claims: V3 adds a wiki layer: stores raw sources immutably, builds index.yaml catalog with summaries/metadata, generates wiki derivatives (concepts, entities, comparisons, open questions); Agent traverses hierarchically: index.yaml → source wiki page (executive summary) → wiki derivatives → raw source (only if needed); Wiki evolves with every question—new concept files, comparisons, logged queries—not just when ingesting new data; File-based design avoids vector databases, knowledge graphs, semantic search infrastructure
- Evidence: Paul: 'We create an index out of all files, generate a wiki on top, which we can query. You should forget the infrastructure you think you need—vector DBs, knowledge graphs. We create this system just based on files and references.'; index.yaml contains 10 sources, 38 wiki pages (example shown); includes link to raw file, origin, title, authors, publication date, summary; Wiki derivatives: comparisons (agentic RAG vs. file systems), concepts (tool registry, context compaction, sandboxing), entities (OpenCode, Claude Code, MCP); Paul: 'Every question leaves a trace. The wiki evolves as you talk with it. You can see a true reflection of yourself, what you haven't understood, all your questions from the past.'
- Caveats: No pruning/versioning of stale concepts mentioned—wiki could accumulate noise over time; Doesn't scale to production systems with thousands of users or real-time updates; Paul explicitly says it's 'not a polished product'—prioritizes inspectability (Obsidian graph view) over performance
- Implications: File-based memory with hierarchical traversal is a practical alternative to vector DBs for personal/prototype use cases; Ken should consider this pattern for agent systems where inspectability, simplicity, and human oversight matter more than query speed; The 'living wiki' concept is powerful for meta-learning: audit your knowledge gaps, spot question patterns, avoid duplicating research
Immutable second brain + project-scoped wikis
- Claims: The system never mutates your raw second brain (Obsidian vault)—it's an immutable snapshot; For each new project, you create a scoped wiki that references your second brain but stores derivatives separately; Projects can be: writing articles, creating videos, courses, managing codebases, tracking features
- Evidence: Paul: 'Obsidian is an immutable snapshot that the LLM never touches. This is my data. I don't want the LLM to touch my personal notes. Whenever we start working on a new project, we reference this second brain and scope it down.'; Paul uses PARA method (Projects, Areas, Resources, Archive) coined by Tiago Forte—resources are flat lists; projects/areas reference them; Example projects: Paul used the system to create his presentation slides
- Caveats: Manual decision required: what insights from project wikis get promoted back to your master vault; No automatic backfill of learnings into your main second brain
- Implications: Ken's agent workflows should separate immutable source of truth from mutable derived artifacts—prevents runaway edits, preserves human curation; Scoped wikis prevent context pollution: each project has its own knowledge base that doesn't clutter your global notes
Three demo use cases and token economics
- Claims: Demo 1: Deep research on agentic engineering—ingest hand-picked links, run light/deep/fast query modes (light = 1 round/3 queries; deep = 3 rounds/6 queries), generates concepts (agent loop, tool registry), comparisons (agentic RAG vs. file systems), takes 10–20 min; Demo 2: Ingest three GitHub repos (OpenCode, Poe, Hermes)—clone repos, extract architecture notes (permission flow, memory system), auto-generate cross-repo comparisons; Demo 3: Ingest arbitrary URLs with zero setup (no Obsidian/Readwise)—just Git/curl required; Token cost warning: deep mode is expensive (token cost + time), reserve for high-stakes projects
- Evidence: Paul runs demo 1 on auto mode: 'I use this hundreds of times, I know it won't delete anything.' Obsidian graph shows concepts like 'agent loop resources', 'compaction vs recursive LMs', 'agentic RAG vs file systems.'; Demo 2 output: individual repo notes on memory systems, permission flows; wiki comparisons show architectural differences across harnesses; Demo 3: Paul says 'You can run examples 2 and 3 without any setup beyond Git/curl—not dependent on Obsidian, Readwise, Notebook LM.'; Paul: 'Light or fast mode is more than enough because this process consumes a lot of tokens. Do deep mode only when you want to look over tons of notes.'
- Caveats: Setup friction varies: demo 1 requires Obsidian, Readwise, Notebook LM CLIs + auth; demos 2–3 are minimal; README lists dependencies but Paul skips setup in video ('I don't want to waste your time')
- Implications: Ken should start with demo 3 (URL ingestion) to prototype quickly, layer in second-brain connectors later; Repo-ingestion use case (demo 2) is especially powerful for competitive analysis or understanding open-source tools—applicable to Ken's investing/GTM research; Token cost vs. value tradeoff: use light mode for incremental learning, deep mode for high-stakes projects (product strategy, major content series)
Limitations, future improvements, and educational intent
- Claims: Repo is designed to teach memory/context management—not to be a production-ready product—customization expected; Missing connectors (Google Drive, Notion, Slack) because they're not useful to Paul/Louis yet—but easily implemented ('YouTube transcript took one prompt'); Known weaknesses: hard to rank source quality/trust, no versioning for outdated sources, memory compaction needs work; Next improvements: stronger linting, better source provenance, memory compaction (state of the art progressing fast)
- Evidence: Louis-François: 'The core goal is for you to adapt it for your needs. Our goal is to teach AI engineering—not to build the next best product. It's a builder workflow through Claude Code/Codex in the terminal.'; Paul: 'You should forget the infrastructure you think you need—vector DBs, knowledge graphs. This is personal, not production scale.'; Louis-François: 'YouTube transcript connector—honestly, a few seconds, just one prompt. Super easy for Codex to implement.'; Acknowledged weaknesses: 'Hard to know which sources are outdated or weak compared to other systems we build. It's not really the priority here.'
- Caveats: Ongoing maintenance and feature additions not guaranteed—it's a teaching artifact first, tool second; Not polished UX/UI—terminal-based workflow ('that's by design, we don't care')
- Implications: Ken should treat this as a reference architecture, not a dependency: fork it, strip what you don't need, extend where your workflow diverges; Educational value is in understanding the file-based indexing + wiki evolution pattern—not in using the tool as-is; For Ken's agent systems/investing workflows, prioritize adding connectors that matter (e.g., Google Drive for client docs, Slack for team context)
Notable Concepts & Terms
- Golden links: Manually curated high-quality sources used as seed context for the deep research algorithm in V1—replaced by self-referential second-brain queries in V2/V3
- Deep research algorithm: Multi-agent system: orchestrator generates questions, sub-agents query Gemini/Google, scrape sources, create executive summaries; main agent aggregates across 3 rounds × 6 queries = ~40–50 sources, then ranks by relevance
- index.yaml: File-based catalog of all sources + wiki pages, containing summaries, metadata (origin, title, authors, date), and references—serves as entry point for agent traversal instead of vector database
- Wiki derivatives: LLM-generated artifacts on top of raw sources: concepts (tool registry, context compaction), entities (OpenCode, Claude Code), comparisons (agentic RAG vs. file systems), open questions
- Source wiki page: Executive summary of each raw source (article, paper, video, repo)—agent checks this before reading the full raw source to save tokens
- Hierarchical traversal: Token-efficient query pattern: agent reads index.yaml → source wiki page → wiki derivatives → raw source only if needed, drilling down incrementally
- Living wiki: Wiki that evolves with every question asked—not just when ingesting new data—by creating new concept/comparison files and logging queries, reflecting your learning process over time
- Immutable snapshot: Your raw second brain (Obsidian vault) that the LLM never modifies—project wikis reference it but store derivatives separately to prevent runaway edits
- PARA method: Note organization framework (Projects, Areas, Resources, Archive) by Tiago Forte—Paul uses it to keep resources as flat lists, projects/areas reference them
- Light/Deep/Fast query modes: Controls depth of research: light = 1 round/3 queries; fast = 2 rounds/3 queries; deep = 3 rounds/6 queries. Deep mode is expensive (tokens + time), reserve for high-stakes projects
- Gemini grounded in Google: Sub-agents use this to gather web sources during deep research rounds—combines Gemini LLM with Google Search grounding for factual retrieval
- Context compaction: Technique to reduce context size while preserving signal—explicitly mentioned as a known weakness ('memory compaction needs work') and state-of-the-art challenge in agent systems
Operator Notes / Why Ken Should Care
- File-based memory with hierarchical traversal (index → summary → derivatives → raw) is a practical pattern for agent systems where inspectability and human oversight matter—consider for prototyping Ken's agent workflows before scaling to vector DBs
- Self-referential second brain querying (agents query your notes, your notes guide retrieval) is the key innovation here—applicable to any long-running agent system where context needs to persist across sessions
- Token economics lesson: deep research (3 rounds × 6 queries) is expensive and slow (10–20 min); use light/fast modes for incremental learning, deep mode for high-stakes projects like product strategy or major content series
- The 'living wiki' that evolves with questions is a meta-learning tool—Ken could use this to audit knowledge gaps, avoid duplicating research, or track question patterns over time (useful for content creation, investing research, GTM strategy)
- Immutable source of truth + mutable derived artifacts pattern prevents runaway agent edits—apply this to Ken's agent systems by separating core knowledge base (never touched by LLM) from project-scoped wikis (LLM can modify)
- Repo-ingestion use case (clone GitHub repos, extract architecture/memory/permission notes, generate cross-repo comparisons) is powerful for competitive analysis or understanding open-source tools—directly applicable to Ken's investing/GTM research workflows
- Educational repo design: treat as reference architecture, not dependency—fork, customize connectors (e.g., Google Drive for client docs, Slack for team context), strip unneeded features; value is in learning the pattern, not using the tool as-is
- Known weaknesses to improve if adapting: source provenance/trust ranking, versioning for outdated sources, memory compaction (state of the art progressing fast); these are active AI engineering challenges Ken should track
Watch Map
- 00:00: Problem: 10,994 notes across fragmented tools, losing research between projects
- 03:15: Why existing tools fail: ChatGPT/Codex lose context, Notebook LM not agent-native, vector DBs too complex
- 05:30: Decision tree: when to use ChatGPT vs. Codex vs. Notebook LM vs. build your own
- 08:45: V1 architecture: golden links + deep research algorithm → static research.md
- 13:20: Deep research algorithm explained: orchestrator + sub-agents, 3 rounds × 6 queries, ranking
- 15:50: V2 architecture: replace golden links by querying second brain (Obsidian, Readwise, Notebook LM, GitHub)
- 18:45: V3 architecture: add living wiki layer with file-based indexing (no vector DB)
- 20:30: File-based index design: index.yaml catalog + hierarchical traversal (summary → derivatives → raw)
- 22:10: Living wiki: evolves with every question, leaves traces (concepts, comparisons, logged queries)
- 23:50: Immutable second brain + project-scoped wikis (LLM never touches raw vault)
- 26:30: Demo 1: Deep research on agentic engineering (light mode, 10–20 min, generates concepts/comparisons)
- 30:10: Demo 2: Ingest GitHub repos (OpenCode, Poe, Hermes), extract architecture notes, generate cross-repo comparisons
- 32:45: Demo 3: Ingest arbitrary URLs with zero setup (Git/curl only)
- 35:20: Limitations and future improvements (source provenance, memory compaction, linting)
- 37:00: Educational intent: repo designed to teach memory/context management, not be a product; customization expected
Source/Metadata
- Title: Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI
- Transcript words: 10590
- Duration seconds: 2372
- Timestamp note: Timestamps reconstructed from transcript structure and video duration (39:32); chapter markers not explicitly provided in transcript
Transcript
I spent 18 months turning my second brain into my living research memory. Let me explain. So within my second brain, I currently have over 5,000 notes in Obsidian and another 5,000 notes in Readwise and some scattered in Notion and Google Drive. And all of this is growing on average with 250 files per month. And this is what I want. On the left, you can see my whole Obsidian vault, this huge mess. And whenever I start working on something such as an article, a new project, a new code base, a new feature or whatever, I want to actually pull high signal notes that are actually useful for my current work. And you would ask yourself, why not use directly Codex, Cloud or Notebook LM? And the thing is that I am. But you need a system that sits between those harnesses and your second brain. Okay, so let's go back to the root of my problem, which is that I'm always losing my research. For example, my reading list is a graveyard. When I'm scrolling social media and I see that cool X post, a new article, a new YouTube video, a GitHub repository, it doesn't matter. Whenever I actually want to start working on something, I never recall what I have in my second brain. Or I have to spend a ton of time actually finding meaningful notes that I can use in my work, right? And another problem that I have is that I want the system to actually be anchored into my personal notes, into my personal values, into my personal face. I want the system to be personal, to reflect my own thoughts, right? And that's why in today's video, Luis Francois and I will teach you how to build your own AI research OS. This also comes with code. So you can also try it out yourself. And I'm Paul Justine. I'm the founder and CEO of Decoding AI, where I do a ton of content on courses on how to ship AI products. And I'm also the co-author of the LM Engineers Handbook bestseller. And the system, the AI research OS that I will teach you in this video, is the system that I use in my daily work. And now I will pass the torch to Luis Francois. [SPEAKER_00] Thanks, Paul. [SPEAKER_00] So I'm Luis Francois Bouchard. [SPEAKER_00] I'm the co-founder and CEO of Towards AI, where we build educational courses. [SPEAKER_00] And I'm also the creator of What's AI, a YouTube channel where I explain AI engineering techniques. [SPEAKER_00] I used to explain AI research before, now focusing on AI engineering. [SPEAKER_00] I'm also the author of the book Building L&Ms for Production. [SPEAKER_00] And before that, I was a PhD student. [SPEAKER_00] So I honestly make research for a living. [SPEAKER_00] I used to do a PhD, as I said, in AI and doing tons of research and research work. Now I build courses, I write videos, I research for videos, I build trainings for companies for a living. And all of these things that I do start with very good research. And also leveraging tons of knowledge and insights that we get at Towards AI from building for clients. So I have tons of notes as well, just like Paul. And we try to leverage them the best possible. And as you'll see, we build some sort of tool to leverage our second brain, where, as you'll see, there will be some differences between how I use it and how Paul uses it. And that's the core goal of the repository that we built on this project, is that we want you to adapt it for your needs. The whole goal is how can we make research better, but more specifically, how can we better leverage what we have? So let's dive into it. And first, we need to figure out which tool to use and when, because this whole research system that we built is not for every query. If you just need a fast answer, like a few quick questions or just something that you would just Google, well, obviously, just Google it or ask ChatGPT, Cloud, whichever system you want. But the problem when doing that is that if you have a lot of following up questions or it's a bigger project that you need to build on and have a very long context or tons of information to share, relying on ChatGPT isn't ideal. And it also means that you are fully dependent on the architecture that OpenAI or ChatGPT's team built. So the next step here is to ask yourself for a more complex problem. Do you need to act quickly or do you want to build some next feature and do something very difficult? If you just have a small repo for a quick change or write one article, just do one thing that you know won't be repeatable that much, definitely use Codex or Cloud Code or some agent that you trust. Sometimes you need to keep on digging to make it better, to improve efficiency, to optimize it more. And so typically when you have to do that, you want your research sources, your research to stick and to be able to refer to them in the future. So if you want a process like this where the sources that you find, the notes that you take, stick around in time and have an agent be able to leverage that efficiently and being able to come back to these information, to ask follow up questions, to digest content even more. And right now, for instance, when I make a new video, I want also the agent of the system to understand the previous videos I made, to not duplicate content, to not repeat myself and to refer to some other content. In this case, there are some tools that are very interesting that you might have tried before, like Notebook LM, that is super powerful to do research, to digest content efficiently and to come back to it. But the problem with Notebook LM is that it's, well, first, the main problem is that you don't own it. You cannot do anything you want with it. You cannot personalize as much as possible. It's not agent native. And it's obviously weak for coding tasks since it's just browser based. So it's far from ideal from something that Paul and I needed and that most AI engineers need in general. So if you need your agents to be able to leverage all you do, whether it is a big research, a new video, whatever you write, you do, you code, you typically want your other agents, your other projects to be able to leverage what you learn from what you just did. And one thing that we advise, especially for production, obviously for a product, is to build some sort of retrieval rack pipeline with vector databases. But this needs an infrastructure. It's not really human friendly to be able to digest quickly, to check notes, to make edits. It's hard to inspect by hand. You need to build everything around it. It's definitely far from ideal for just something I want to use on a daily basis. Obviously, it's super powerful at scale. Very interesting, especially in a product. But as I said, this project is for us. And I don't want something live, super professional as a product. I just want something I will use and that my agents and different projects can leverage as best as possible. But this needs an infrastructure. It's not really human friendly to be able to digest quickly, to check notes, to make edits. It's hard to inspect by hand. You need to build everything around it. It's definitely far from ideal for just something I want to use on a daily basis. Obviously, it's super powerful at scale. Very interesting, especially in a product. But as I said, this project is for us. And I don't want something live, super professional as a product. I just want something I will use and that my agents and different projects can leverage as best as possible. So the last question to ask ourselves here is that if you want everything there but more personalization, So a personalized research assistant that builds some sort of Wikipedia that compounds over time and is easily inspectable and usable, where you have tons of sources, documents, videos, comparisons, implementations, new research, new topics that you keep on adding and that you keep on wanting to leverage and review easily, this is where you may want to build something yourself. And in our case, we build the personalized research OS that we will share in this talk with exactly what we built and how. But the downside is that it definitely needs a bit more setup than just opening cloud code. Right now, the main problem with using cloud code and other agentic tools is that you give codecs links, PDFs, and different information. For example, my most recent loop engineering video. And then the next session you use codecs, you have to paste it all again or ask it to use skills. And whatever structure that codecs or chat GPT, whatever tool that you use, built on the fly to leverage what you did, the scripts it ran, the scripts it had, you all lost it or kept it inside a skill that you have to ask it to reuse and that usually isn't ideal and just grows and grows over time. And the problem is that all this information that you give to the model is not the bottleneck. The bottleneck is how can you leverage it in the future? Meaning that with an agent, the context window becomes everything. The database, the file system, the memory, the reasoning space. It has to do it all. And when you stop the conversation, it loses everything. And the thing is that we don't need necessarily to provide more and more context for better research. You need a proper memory and context management and ideally some personality with it, especially in my case when I do videos. So what we did is that we decided to build a system with plain files, mostly markdown files that we can leverage easily and that agents can leverage easily. I won't detail it very much here because Paul will talk about it in depth. And as I said, Paul has 5,000 or something notes. I have just a few hundred, but that's just to say that we need to consider that we didn't start from nothing. We already both had some sort of large database. In my case, I made hundreds of videos and I take many notes. So I still need to leverage these years of content that I already made and tons of meetings that I have with my team, with clients when we build for them that I want to leverage as well because we learn a lot by building for people. We have highlights from interesting tweets that I see, interesting articles that I see. And I want all my projects to be able to leverage my agent skills. So I decided to pivot and instead of having a folder for cloud code skills and having all my meeting recaps in Granola and having years of notes on Apple notes and the tweets on a saved Chrome tab, instead, I moved everything automatically into Obsidian. It's just a note reader, obviously, so you don't have to use that. You can just save it locally. But I used the codex to set up everything so that Granola is automatically saved there. My notes are now on Obsidian just because it's a nice UI. I like it and I can use it from my phone, my computer, my Windows, Mac, everything. So anyways, I moved everything to Obsidian, which means that it's saved locally in my file system, which means it's basically my companion for researching and building everything I build now. And what we built, obviously, leverages that. We built a repo called AI Research OS for this workshop, where it's basically just skills for cloud code and codex with plugins to be able to do a very deep research about a topic or a simpler search or destination, different tools that you can use. The most useful and complete one will be the research tool that I use, for example, when I kick off a new video topic. And the goal of this repo is to have you implement it, install the cloud plugins from it, and tune it to your needs. Right now, it can connect to, as I said, Obsidian with my local notes. It can use Readwise, Notebook LM, your GitHub repos, any links that you send for GitHub or YouTube videos, and web links, obviously, and documents. But there are tons of things missing, as we will discuss in the end, that you can easily implement just asking cloud code or codex to do so. Like, for example, I implemented the YouTube video transcript in, honestly, a few seconds, just one prompt. It's super easy for codex to implement it. So the thing is that this whole repository and this whole project is a very useful companion for my own work. But as I said, it implements tons of features and state-of-the-art context management and memory management techniques that I believe AI engineers need to know. And now, Paul will dive into all this three-layer system that we built with the raw content, the index that I mentioned, and the wiki-like synthesized version of all your notes, all your research, all your work. So he'll cover everything we did, how it ended up, and show how to use it. [SPEAKER_01] Okay, so now I want to go over three versions of our system and how it progressed over time, and most importantly, why we added more complexity. So in the first version, we wanted to scope it just to create lessons for our agent engineering course. So we wanted to keep it super simple, where we had as input a topic and a research MD as output. So within the input, we had the topic plus a set of golden links, which were manually hand-picked by us. We applied this deep research algorithm, and we had as output a static research MD file. And if we go and zoom into the architecture, we first scraped the links of the golden links, right, because we already know them, and we used them as seed for context for the deep research algorithm, which was a really powerful technique So we wanted to keep it super simple, where we had as input a topic and a research MD as output. So within the input, we had the topic plus a set of golden links, which were manually hand-picked by us. We applied this deep research algorithm, and we had as output a static research MD file. And if we go and zoom into the architecture, we first scraped the links of the golden links, right, because we already know them, and we used them as seed for context for the deep research algorithm, which was a really powerful technique because we had more context on how to frame our questions. And during the query rounds, we used a very classic deep research algorithm where we had one main agent, the orchestrator, which created multiple questions based on the initial topic and the scraped context. And each agent managed its own question and used Gemini grounded in Google to gather multiple resources. And each agent gathered these resources, which returned multiple links and created some executive summaries of each link. And then it passed all this information back to the main agent, where the main agent aggregated all this information in a summarized way so it did not explode in the context. And we did this for three rounds. So after three rounds of generating six queries per round, we ended up with 40, 50 links in total. So you can imagine that there's a lot of noise over there. So that's why we also applied a ranking algorithm where we wanted to find the information with the highest signal. And we compared each source against the topic, the initial topic of the user. And that way, we fully scraped only the top key elements based on the ranking score. And for the rest of the links, we just kept the summaries. And then we compiled everything into this ResearchMD file as a single flat file, which we used for each lesson of our course in our particular use case. But as you can imagine, it was pretty limited. For the course, it worked great, right? We generated 35 lessons really quick, but we wanted more. So we started to aim this deep research loop to the second brain as well, right? Before it was targeting only the public web, which made this pretty generic and we had to manually find all those golden links. So by aiming this deep research loop to the second brain, where we basically organically keep track of all the information, all the research that we really want and is filtered by us, we can organically gather all those golden links into our deep research algorithm. So let's look at how this new algorithm looks like. It's the same loop, right? But now we target our own sources instead of just the public web. Now for input, we have only the topic because we don't need the golden links. As I said, the golden links are actually a reflection of our second brain system. In theory, you can also add them if you really want to, but that's the beauty of this new strategy because you can just put as input some topic and we'll find everything that it needs. And then we use this topic only as seed for the context to generate the queries. And now we do the same deep research algorithm, right? The same query rounds, but instead of targeting only the public web, now we plug in all our second brains, such as our Obsidian, our Readwise, our Notebook LM, our GitHub. You can also use, for example, Gemini Deep Research for this, similar to how we use Notebook LM or you can extend this with whatever you want, for example, YouTube, Google Drive, Notion, or whatever makes sense on your infrastructure. The idea is that now we target our queries from the deep research algorithm to our second brain plus the public web. And after, we apply the same algorithm, such as ranking, fully scraping summaries, and compile everything into this research MD file. But now we have another problem, right? This research MD file is static. It's a pile of static data. And research is not static, right? So after you end up with this file, you most often realize that you want to ask another question or some information is outdated and you don't need it anymore. Or you want more out of this research MD file, which means that you need to start all of this from scratch. And the operation that I showed you above is an expensive operation. It consumes a lot of tokens and it takes a lot of time. So you don't want to run it from scratch. And that's why you need to add a wiki layer on top of it. And that's why V3 of this system is actually a deep research algorithm plus an LM knowledge base on top of it, aka the wiki layer. So the new algorithm looks like this. So we have sources in and a wiki out. And the sources, as I said before, can be Obsidian, Notebook LM, Google Drive, or YouTube, Notion, even custom URLs, right? That's also powerful as well, where you use tools such as Bright Data to parse basically any single page application, any type of site, any type of public information that's out there, we can put it in. aka the wiki layer. So the new algorithm looks like this. So we have sources in and a wiki out. And the sources, as I said before, can be Obsidian, Notebook LM, Google Drive, YouTube, Notion, or even custom URLs, right? That's also powerful as well, where you use tools such as Bright Data to parse any single page application, any type of site, any type of public information that's out there, we can put it in. And then you apply the same deep research algorithm, you store everything into raw files, right? Instead of compiling everything into a research MD file, now we store each file individually. And we create an index out of all these files. And ultimately, we generate a wiki on top of it, which we can query. We can query the wiki plus the index. Okay, so this is just the high level architecture of the new system. Let's zoom into it. So what I want to start with is that you should forget the infrastructure you think you need, such as vector databases, knowledge graphs, semantic search, text search. All that is beautiful, but adds a lot of complexity, especially for personal wikis, personal research operating systems that you want to use very lightly. So I want a system just based on files, right? A simple mechanism that's very rooted into how your computer works. And that's why we'll create all this system just based on files and just based on references. So no database, just a simple index based on references. And how this works? We have an agent that reads an index.yaml file that's a catalog of all your data plus the summaries of each source and some metadata around it. For example, here on the right, you can see part of an index.yaml file that contains 10 sources and 38 wiki pages as derivatives of these sources where we can see into the sources list of the YAML file, the first source, for example. And as you can see, it has the link to the original file plus metadata such as the origin, the title, the authors, the publication date, summary, and things like this which can be flexible, right? And the next step is that based on this index.yaml file we need to point to all the wiki pages, to all the wiki derivatives, to all the raw sources. So basically, this index.yaml file is an entry point for our agent, right? It's what we will give to our agent to actually reason on how to find our data. It's an index ultimately, right? So the next step is to understand how the wiki actually looks like. So on the left, you can see the high-level structure of the wiki where we have the raw folder, the wiki folder, and the index. In the raw folder, we actually just have the raw data which is immutable. You don't want to touch that. And the index points to everything that we need. And in the wiki, we actually have derivatives created by the LLM which contains things such as comparisons between multiple concepts, entities, or just simple notes as a reflection of our questions or repositories that we ingested and we can create multiple notes based on a repository, right? Or open questions that, based on our questions, that LLM couldn't answer yet. And everything that you can analyze on top of your raw data. And on the right, based on Obsidian, we can see the subgraph reflected just based on this index. And this is just the first iteration but as the wiki grows, you can see connections made between entities and concepts. For example, concepts are things such as tool registry, context compaction, sandboxing, or entities are Open Code, Cloud Code, MCP, right? So as you can see, you can beautifully start visually and practically create connections. Now, the next question is how do we actually query this wiki? So as I said, the agent will have as input this index.yaml file which contains summaries and metadata about Are open code, cloud code, MCP, right? So as you can see, you can beautifully can start visually and practically create connections. Now, the next question is how do we actually query this wiki? So as I said, the agent will have as input this index.yaml file which contains summaries and metadata about each source. But what happens next, right? The next step is actually to look into the source wiki page where the source wiki page is an executive summary of each page which is not just a summary but a more expanded summary of each source. And sometimes the agent just looks into this, gets what it needs and goes back which is very token efficient, right? And if it doesn't find within this source wiki page, we also need links into the wiki derivatives such as concepts, entities, nodes, comparisons and so on and so forth. And only if it doesn't find the necessary information up to this point, it needs and it actually reads the whole raw source, right? Which contains the whole article, the whole paper, the whole video or whatever. And this makes just through pure referencing and creating this simple hierarchy, this makes everything very token efficient. Now, the beautiful part is that this wiki is actually alive, right? For example, every question leaves a trace into your wiki. So for every question, the LLM can create a new concept file, a new node file, a new comparison file and every question is tracked into a log. So the wiki doesn't evolve only when you ingest new data or do a deep research round, it actually evolves as you start talking with it, right? That's the beautiful part. And you can see a true reflection of yourself, of what you haven't understood, of all your questions from the past. And the beautiful part is that the wiki is never frozen, right? Similar to the research empty files. At any point, you can ingest a new custom link that you think you need into the wiki or even run a new deep research round. Or as I said previously, the wiki keeps evolving just purely based on your questions. And another important thing to understand is that this wiki doesn't sit on top of your entire second brain, right? For example, in my particular use case, I use the paramethod coined by Tiago Forte where all my data is structured between project areas, resources, and archive. Where all my notes, resources that I save, sources that I save, PDFs, articles, or whatever are just piped directly into the resource, a flat list. And whenever I need something, I just reference them into projects and areas, right? And like this, Obsidian is just an immutable snapshot that the LLM never touches, right? So this is my data. I don't really want the LLM to touch my personal notes that I manually write, right? So then, how can we actually put this wiki to use, right? So, as I said, we have the big Obsidian snapshot, which is our global second brain. And then, whenever we start to work on a new project, we reference this second brain from this deep research algorithm that I explained and we scope it down to our own project, right? So whenever we want to start working on something, we run this deep research loop or we start ingesting some particular repositories, articles, notes, and so on and so forth. And we usually do that through a set of skills plugged into a harness. this second brain from this deep research algorithm that I explained and we scope it down to our own project, right? So whenever we want to start working on something, we run this deep research loop or we start ingesting some particular repositories, articles, notes, and so on. And we usually do that through a set of skills plugged into a harness. And a project can be anything such as writing a new article, doing a new video, doing a set of slides. I apply this technique doing this slide. Or you can even apply it for something more complex such as writing a book, doing a course, or keeping track of the whole code base, right? You can also use it for that. So a project can be anything where you want, as I said initially, to transform research into work. The project is the work and your second brain is the research. So now I want to show you a few demos. So what you need to do is go to the AI Research OS Workshop repository. And here you can find all the skills required to run what we presented into this presentation. And everything is packed as a cloud code plugin, but you can very easily tweak it and install it with any other harness. And also in the ReadMe, you can find details on how to install all the other dependencies because this system is dependent on tools such as Obsidian, Readwise, Notebook LM, and so on. So you need to set up specific CLIs or authentication issues. But I don't really want to waste any of your time with setup issues and I want to go straight directly into the examples. So I prepared here three examples. The first one is a research on one of my previous articles on agentic engineering. And within these files, I have a brain dump of everything that I knew I wanted to talk on this subject. And on top of that, I also added a few references that I knew 100% that I want to add into this wiki. And what we need to do to actually trigger the algorithm on top of these files is to open up a cloud session and then just call the skill, the research skill and pointed it to this file. And that's it. Everything else is baked directly into the skill. It will understand my intent that I want to create a wiki on agentic hardness engineering, just looking at these files and looking at the topic. And it will know that before starting the deep research algorithm, it actually needs to scrape this information to use it as context when it frames the questions for the deep research algorithm. Now, we need to wait a bit for the agent to reason on top of it. And I will actually just put it on auto mode to speed this up, right? I use this hundreds of times, so I know it won't delete anything from my computer or it won't do anything weird. Okay, so now this is the most important part, right? So it asks me how deep I want the deep research algorithm to be. We have light, deep, fast. This mostly controls how many questions you want to run per one round and how many rounds you want to run. And usually light or fast is more than enough because remember, this process consumes a lot of tokens. So you need to do the deep one only when you really want to look over tons and tons of notes. And for this use case, I will just pick the light one to speed this up. And in this use case, it just does one round of three queries, right? And for the fast one, it does two rounds of three queries. So I will just keep it around that spectrum. And now the process will take around 10 to 20 minutes to actually look around my Obsidian, to look around my Readwise, my Notebook LM and run those queries on top of this. of three queries, right? And for the fast one, it does two rounds of three queries. So I will just keep it around that spectrum. And now the process will take around 10 to 20 minutes to actually look around my Obsidian, to look around my Readwise, my Node.LM and run those queries on top of this. And I actually run this, right? And now let's open this wiki that we created based on the prompt before in Obsidian. And let's look what's inside the wiki. So we have three big objects, the row files, which are a row copy of what we found, the index, which contains all the references towards the wiki, right? We can also in Obsidian have this beautiful subgraph where we can very quickly understand what is going on. This is created purely based on this file. And then the most interesting part is inside the wiki where we have comparisons, concepts, entities, and sources. The sources contain the executive summary of our raw sources. So the LLM doesn't really need to every time when it reads them and it needs them to read the raw sources, but it needs just the executive summaries, which are computed just one time during the ingestion. And for example, for the comparisons, it understood out of the box that it needs to do comparisons between a genetic rag versus file systems or compaction versus recursive language models or anything of interest based on the sources. And the most interesting part is actually the concepts. So it automatically extracted all the concepts that we need to understand from this pile of resources. For example, if we open the agent loop resources, we automatically can look and get this beautiful summary containing graphs, text, and explaining everything that we need on this topic. And we can do the same on all the concepts from this wiki or on the entities from the wiki and so on and so forth. Okay, so now let's go to the second example. It's a simple example where I want to learn more on harness engineering and how harnesses work. So in this prompt over here, I just want to ingest the three open source repositories on open code, PI, and Hermes. And I don't want to do deep research at all, right? I just want to ingest those repositories and explore topics such as the general architecture, agents, architecture, subagents, memory system and the agent permission flow. And let's run the research on top of these prompts. And now what the research will do, will clone automatically all these repositories and will explore all the repositories on the topics that I gave here. And it will create notes at the individual level of each repository on how they work, on how the architecture works at the repository level. And then we can create higher level nodes within the wiki derivatives and compare other architectures or create Notes at the individual level of each repository on how they work, on how the architecture works at the repository level. And then we can create higher level nodes within the wiki derivatives and compare other architectures or create aggregate architectures on the general trends of all those harnesses and explore and learn everything that we want about harnessing engineering directly from the code. And again, usually I just do auto mode and let it do its own GIS, but I already run this. So here is the wiki for this GitHub repository. And again, we have the raw files and we have the index and the wiki. And here probably within the repos, we can see all three repositories. And for example, inside the OpenCode repositories, we can see empty files explaining everything that we need on particular topics, such as the permission flow, the memory system, and so on so forth. And we can see this in all the other repositories. And the most interesting part here is actually that we have comparisons on all of these, right? So we can actually understand the differences within the architecture within these harnesses. Or we also have all the concepts extracted from these repositories, repositories. And we can understand what are the key architectural decisions from these repositories. So you can go crazy with this. And this is super useful if you want to, for example, write your own harness. And the third example is the simplest one in reality, which is just based on ingesting some simple links, right? And again, I will just exit the previous run and I want to run this example from scratch. And here, I just want to show you that you can use this also with a very basic setup where I just want to ingest three custom random links. And as before, I just passed this prompt and it will start the research process. And I want to highlight that you can run example two ending on the GitHub repositories and this example without any other setup, without setting up Obsidian Readwise or anything else. You can just install this plugin and run these examples because here is not dependent on any other service than Git and using core to get these URLs. And again, here, if we go into Obsidian and explore the wiki created out of this example three, on any other service than Git and using core to get these URLs. And again, here, if we go into Obsidian and explore the wiki created out of this example three, we can see the index, we can see the wiki itself with all the sources, right? One executive summary for its source and all the concepts extracted entities and so on and so forth. And the idea is that as you start asking questions on top of this, everything starts to get more interesting. So now let's assume that we want to ask, for example, a question on harness engineering based on the wiki created out of the GitHub repositories. So what we have to do is just again hit the research skill, point it to the wiki that we just created and then just ask our question. And let's assume that I want to learn more on sandboxing, more exactly how remote sandboxing works and how is plugged into the heart. This can be any other question. And now what it will happen, it will query this wiki, it will give you an answer and based on this you can also start creating notes, comparisons or maybe it will extract and find new entities that you care about. It will start updating the wiki. [SPEAKER_00] All right. So now where is this project going? What is it? What should you be needs more connectors as you saw. We need to add Google Drive, Notion, Slack, tons of other connectors that could be useful to you. But to be honest, it's not really useful to me and my current workflow or to Paul. So we didn't add them yet because the core of this project is to be useful for us and for you to take over and add whatever you need. And the other main goal of this project is [SPEAKER_00] This project is to be useful for us and for you to take over and add whatever you need. And the other main goal of this project is to teach memory and context management. So all these extra features aren't really useful towards that purpose. There are some weaknesses, like it's hard to know which sources are outdated or weak or strong compared to some other system that we build. So we know we can improve this, but again it's not really the priority here. And lastly, it's still obviously a builder workflow. You use it through cloud code and Codex, and it's just to me in the terminal and I really like it. So it's not a final polished product with a nice UI, nice UX, and honestly that's by design. So we don't really care about this because our goal is to teach AI engineering. It's not to build the next best product. Still, we have a few next improvements we want to do very shortly, from having a stronger linting to a better memory compaction because that's a big issue and it's just very complicated in general to manage memory correctly. And the state of the art is always progressing there. We, as I said, need better source provenance to trust the sources and be able to rank them properly and reuse them properly if needed, and be able to know, access quickly as a user if this source is relevant or not. And we have other next improvements to do, but those are mostly for optimization and for the future. And the thing is that we actually built all of that into another product that you [SPEAKER_00] But those are mostly for optimization and for the future. The thing is that we actually built all of that into another product that you can even build yourself because we created a course called Agent Engineering where we built a similar deep research system with a writing and research agent. We build a system with the same goal to be able to learn best AI engineering practices. It's a very in-depth course where I assume it takes around 60 hours to complete with a final project being the multi-agent system I just described that you'll build for yourself. So if this presentation and the demo repo that you saw was interesting, please consider checking out the Towards AI Academy with our courses on there, including the Agent Engineering course to learn more on the best practices when building around and with agents. The same query rounds, but instead of targeting only the public web, now we plug in all our second brains, such as our Obsidian, our Readwise, our Notebook LM, our GitHub. You can also use, for example, Gemini Deep Research for this, like similar to how we use Notebook LM or you can extend this with whatever you want, for example, YouTube, Google Drive, Notion, or whatever makes sense on your infrastructure. The idea is that now we target our queries from the deep research algorithm to our second brain plus the public web. And after, we apply the same algorithm, such as ranking, fully scraping summaries, and compile everything into this research MD file. But now we have another problem, right? This research MD file is static. It's a pile of static data. And usually research is not static, right? So after you end up with this file, you most often realize that you want to ask another question or some information is still and you don't need it anymore. Or basically, you want more out of this research MD file, which means that you need to start all of this from scratch. And the operation that I showed you above is an expensive operation. It consumes a lot of tokens and it takes a lot of time. So you don't want to run it from scratch. And that's why you need to add a wiki layer on top of it. And that's why V3 of this system is actually a deep research algorithm plus an LM knowledge base on top of it, aka the wiki layer. So the new algorithm looks like this. So we have sources in and a wiki out. And the sources, as I said before, can be like Obsidian, Notebook LM, Google Drive, or YouTube, Notion, even custom URLs, right? That's also powerful as well, where you use tools such as a bright data to parse basically any single page application, any type of site, any type of public information that's out there, we can put it in. And then you apply the same deep research algorithm, you store everything into raw files, right? Instead of compiling everything into a research MD file, now we store each file individually. And we create an index out of all these files. And ultimately, we generate a wiki on top of it, which we can query. We can query basically the wiki plus the index. Okay, so this is just the high level architecture of the new system. Let's zoom into it. So what I want to start with is that you should forget the infrastructure you think you need, such as vector databases, knowledge graphs, semantic search, text search. All that is beautiful, but adds a lot of complexity, especially for like this personal wikis, personal research operating systems that you want to use very lightly. So I want a system just based on files, right? A simple mechanism that's very rooted into how your computer works. And that's why we'll create all this system just based on files and just based on references. So no database, just a simple index based on references. And how this works? We have an agent that reads an index.yaml file that's basically a catalog of all your data plus the summaries of each source and some metadata around it. For example, here on the right, you can see part of an index.yaml file that contains 10 sources and 38 wikipages as derivatives of these sources where we can see there into the sources list of the YAML file, the first source, for example. And as you can see, it has like the link to the original file plus a metadata such the origin, the title, the authors, the publication date, summary, and things like this which can be flexible, right? And the next step is that based on this index.yaml file we need to point to all the wiki pages, to all the wiki derivatives, to all the raw sources. So basically, this index.yaml file is an entry point for our agent, right? It's what we will give to our agent to actually reason on how to find our data. It's an index ultimately, right? So the next step is to understand how the wiki actually looks like. So on the left, you can see the high-level structure of the wiki where we have the raw folder, the wiki folder, and the index. In the raw folder, we actually just have the raw data which is immutable. You don't want to touch that. And the index points to everything that we need. And in the wiki, we actually have derivatives created by the LLM which contains things such as comparisons between multiple concepts, entities, or just simple notes as a reflection of our questions or repositories that we ingested and we can create multiple notes based on a repository, right? Or open questions that, based on our questions, that LLM couldn't answer yet. And everything that you can analyze on top of your raw data. And on the right, based on Obsidian, we can see like the subgraph reflected just based on this index. And this is just like the first iteration but as the wiki grows, you can see connections made between entities and concepts. For example, concepts are things such as tool registry, context compaction, sandboxing, or entities are open code, cloud code, MCP, right? So as you can see, you can beautifully can start visually and practically create connections. Now, the next of us question is how do we actually query this wiki? So as I said, the agent will have as input this index.yaml file which contains summaries and metadata about each source. But what happens next, right? The next step is actually to look into the source wiki page where the source wiki page is like an executive summary of each page which is basically not just a summary but a more expanded summary of each source. And sometimes the agent just looks into this, gets what it needs and goes back which is very token efficient, right? And if it doesn't find within this source wiki page, we also need links into the wiki derivatives such as concepts, entities, nodes, comparisons and so on and so forth. And only if it doesn't find the necessary information up to this point, it needs and it actually reads the whole raw source, right? Which basically contains the whole article, the whole paper, the whole video or whatever. And this makes just through pure referencing and creating this simple hierarchy, this makes everything very token efficient. Now, the beautiful part is that this wiki is actually alive, right? For example, every question leaves a trace into your wiki. So for every question, the LLM can create a new concept file, a new node file, a new comparison file and every question is tracked into a log. So basically, the wiki doesn't evolve only when you ingest new data or do a deep research round, it actually evolves as you start talking with it, right? That's the beautiful part actually. And like that, you can see a true reflection of yourself, of what you haven't understood, of all your questions from the past. And the beautiful part is that the wiki is never frozen, right? Similar to the research empty files. At any point, you can ingest a new custom link that you think that you need into the wiki or even run a new deep research round. Or as I said previously, the wiki keeps evolving just purely based on your questions. And another important thing to understand is that this wiki doesn't sit on top of your entire second brain, right? For example, in my particular use case, I use the paramethod coined by Tiago Forte where all my data is structured between project areas, resources, and archive. Where all my notes, resources that I save, sources that I save, PDFs, article, or whatever are just piped directly into the resource, a flat list. And whenever I need something, I just reference them into projects and areas, right? And like this, Obsidian is just an immutable snapshot that the LLM never touches, right? So this is my data. I don't really want the LLM to touch my personal notes that I manually write, right? So then, how can we actually put this wiki to use, right? So, as I said, we have the big Obsidian snapshot, which is our global second brain. And then, whenever we start to work on a new project, we reference this second brain from this deep research algorithm that I explained and we scope it down to our own project, right? So basically, whenever we want to start working on something, we run this deep research loop or we start ingesting some particular repositories, articles, notes, and so on and so forth. And we usually do that through a set of skills plugged into a harness. And a project can be basically anything such as writing a new article, doing a new video, doing a set of slides. I apply this technique doing this slide. Or you can even apply it for something more complex such as writing a book, doing a course, or keeping track of the whole code base, right? You can also use it for that. So basically, a project can be anything where you want, as I said initially, to transform research into work. The project is the work and your second brain is the research. So now I want to show you a few demos. So what you need to do is go to the AI Research OS Workshop repository. And here you can find all the skills required to run what we presented into this presentation. And everything is packed as a cloud code plugin, but you can very easily tweak it and install it with any other harness. And also in the ReadMe, you can find details on how to install all the other dependencies because the thing is that this system is dependent on tools such as Obsidian, Readwise, Notebook LM, and so on and so forth. So you need to set up specific CLIs or authentication issues. But I don't really want to waste any of your time with setup issues and I want to go straight directly into the examples. So I prepared here three examples. The first one is a research on one of my previous articles on agentic engineering. And within these files, I have a brain dump of everything that I knew I wanted to talk on this subject. And on top of that, I also added a few references that I knew 100% that I want to add into this wiki. And what we need to do to actually trigger the algorithm on top of these files is to open up a cloud session and then just call the skill, the research skill and pointed it to this file. And that's it. Everything else is baked directly into the skill. It will understand my intent that I want to create a wiki on agentic hardness engineering, just looking at these files and looking at the topic. And it will know that before starting the deep research algorithm, it actually needs to scrape this information to use it as context when it frames the questions for the deep research algorithm. Now, we need to wait a bit for the agent to reason on top of it. And I will actually just put it on auto mode to speed this up, right? I use this hundreds of times, so I know it won't delete anything from my computer or it won't do anything weird. Okay, so now this is the most important part, right? So it asks me how deep I want the deep research algorithm to be. We have light, deep, fast. fast. This mostly controls how many questions you want to run per one round and how many rounds you want to run. And usually light or fast is more than enough because remember, this process consumes a lot of tokens. So you kind of need to do the deep one only when you really want to look over tons and tons of notes. And for this use case, I will just pick the light one to speed this up. And in this use case, it just does one round of three queries, right? And for the fast one, it does two rounds of three queries. So I will just keep it around that spectrum. And now the process will take around 10 to 20 minutes to actually look around my Obsidian, to look around my Readwise, my Node.LM and run those queries on top of this. And I actually run this, right? And now let's open this wiki that we created based on the prompt before in Obsidian. And let's look what's inside the wiki. So we have three big objects, the row files, which are basically a row copy of what we found, the index, which contains all the references towards the wiki, right? We can also in Obsidian have this beautiful subgraph where we can very quickly understand what is going on. This is created purely based on this file. And then the most interesting part is inside the wiki where we have comparisons, concepts, entities, and sources. The sources contain the executive summary of our raw sources. So the LLM doesn't really need to every time when it reads them and it needs them to read the raw sources, but it needs just the executive summaries, which are computed just one time during the ingestion. And for example, for the comparisons, it understood out of the box that it needs to do comparisons between like a genetic rag versus file systems or compaction versus recursive language models or anything of interest based on the sources. And the most interesting part is actually the concepts. So it automatically extracted all the concepts that we need to understand from this pile of resources. For example, if we open the agent loop resources, we automatically can look and get this beautiful summary containing like graphs, text, and explaining everything that we need on this topic. And we can do the same on all the concepts from this wiki or on the entities from the wiki and so on so forth. Okay, so now let's go to the second example. It's a simple example where I want to learn more on harness engineering and how harnesses work. So in this prompt over here, I just want to ingest the three open source repositories on open code, PI, and Hermes. And I don't want to do deep research at all, right? I just want to ingest those repositories and explore topics such as the general architecture, agents, architecture, subagents, memory system and the agent permission flow. And let's run the research on top of these prompts. And now what the research will do, will clone automatically all these repositories and will explore all the repositories on the topics that I gave here. And it will create notes at the individual level of each repository on how they work, on how the architecture works at the repository level. And then we can create higher level nodes within the wiki derivatives and compare other architectures or create aggregate architectures on the general trends of all those harnesses and explore and learn everything that we want about harnessing engineering directly from the code. And again, usually I just do auto mode and let it do its own GIS, but I already run this. So here is the wiki for this GitHub repository. And again, we have the raw files and we have the index and the wiki. And here probably within the repos, we can see all the three repositories. And for example, inside the OpenCode repositories, we can see empty files explaining everything that we need on particular topics, such as the permission flow, the memory system, and so on so forth. And we can see this in all the other repositories. And the most interesting part here is actually that we have comparisons on all of these, right? So we can actually understand the differences within the architecture within these harnesses. Or we also have all the concepts extracted from these repositories, repositories. And we can understand what are the key architectural decisions from these repositories. So you can go crazy with this. And this is super useful if you want to, for example, write your own harness. And the third example is the simplest one in reality, which is just based on ingesting some simple links, right? And again, I will just exit the previous run and I want to run this example from scratch. And here, I just want to show you that you can use this also with a very basic setup where I just want to ingest three custom random links. And as before, I just passed this prompt and it will start the research process. And I want to highlight that you can run example two ending on the GitHub repositories and this example without any other setup, like without setting up Obsidian Readwise or anything else. You can just install this plugin and run these examples because here is not dependent on any other service than Git and using core to get these URLs. And again, here, if we go into Obsidian and explore the wiki created out of this example three, we can see the index, we can see the wiki itself with all the sources, right? One executive summary for its source and all the concepts extracted entities and so on and so forth. And the idea is that as you start asking questions on top of this, everything starts to get more interesting. So now let's assume that we want to ask, for example, a question on harness engineering based on the wiki created out of the GitHub repositories. So what we have to do is just again hit the research skill, point it to the wiki that we just created and then just ask our question. And let's assume that I want to learn more on sandboxing, more exactly how remote sandboxing works and how is plugged into the heart. This can be basically any other question. And now what it will happen, it will basically query this wiki, it will give you an answer and based on this you can also start creating notes, comparisons or maybe it will extract and find new entities that you care about. Basically it will start updating the wiki. All right. So now where is this project going? What is it? What should you be needs more connectors as you saw. We need to add Google Drive, Notion, Slack, tons of other connectors that could be useful to you. But to be honest, it's not really useful to me and my current workflow or to Paul. So we didn't add them yet because the core of this project is to be useful for us and for you to take over and add whatever you need. And the other main goal of this project is to teach memory and context management. So all these extra features aren't really useful towards that purpose. There are other some weaknesses like it's hard to know which sources are outdated or weak or strong compared to some other system that we build. So we know we can improve this but again it's not really the priority here. And lastly it's still obviously a builder workflow you use it through cloud code and codex and it's just to me in the terminal and I just really like it. So it's not a final polished product with a nice UI nice UX and honestly that's by design. So we don't really care about this because our goal is to teach AI engineering it's not to build the next best product. Still we have a few next improvements we want to do very shortly from having a stronger linting to a better memory compaction because that's a big issue and it's just very complicated in general to manage memory correctly and the state of the art is always progressing there. We as I said need better source provenance to trust the sources and be able to rank them properly and reuse them properly if needed and be able to know access quickly as a user if this source is relevant or not and we have other next improvements to do but those are mostly for optimization and for the future and the thing is that we actually built all of that into another product that you can even build yourself because we created a course called agent ! engineering where we built a similar deep research system with a writing and research agent where we build a system with the same goal to be able to learn best AI engineering practices it's a very in-depth course where I assume it takes around 60 hours to complete with a final project being the multi-agent system I just described that you'll build for yourself so if this presentation and the demo repo that you saw was interesting please consider checking out the Towards AI academy with our courses on there including the agent engineering course to learn more on the best practices when building around and with agents done