Open Reader

Lessons from Studying Every Memory System — Shlok Khemani, Independent

completed 19:30 Aug 12, 2026 Watch on YouTube

Current Status

completed

Video ID

5ZGyKWjQDr0

RAG / Chat

Enabled
Lessons from Studying Every Memory System — Shlok Khemani, Independent
Description

A profile ChatGPT keeps on Shlok Khemani says he travelled to Turkey in 2025. He never has. The memory came from conversations where he was choosing between Turkey and Thailand, he went to Thailand, and the profile kept both with overlapping dates. What bothers him is not the mistake but the incuriosity: nothing notices the conflict, and the evidence to settle it was sitting in his email as flight and hotel bookings. He calls that a product problem, not a technology one. The rest is a year of reverse engineering how consumer memory systems are built, all of it his reading from the outside rather than anything documented. By his account ChatGPT went from a user managed list of facts to a running profile rebuilt in the background, roughly 4,000 tokens of dense keyword clues he could only inspect by jailbreaking his way to it. Claude started opposite, with no profile and two retrieval tools over past conversations, then added one about a quarter the size, in full sentences, refreshed daily and visible in settings. His frame for the difference is that memory is a function of compute: a profile costs to maintain and costs again in every context window it enters, so a large profile updated rarely and a small one updated daily are two answers to one budget question. Speaker info: - https://x.com/shloked - https://www.linkedin.com/in/shlokkhemani/ - https://shloked.com Timestamps: 0:00 - What memory means here: consumer personalization 2:09 - ChatGPT memory v1, a list of facts 4:08 - The running profile arrives 6:04 - The trip that never happened 6:45 - A profile you cannot look at 7:23 - Claude's first version, tools instead of a profile 8:03 - Publishing that they were opposites 9:18 - Three years of convergence 11:14 - There is no single way to do memory 12:30 - Memory is a function of compute 13:47 - Continual learning is already here 15:39 - The context problem no architecture solves 17:34 - Products that each know you separately

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI memory is not a generic RAG feature but a product-specific, compute-constrained control loop that must synthesize, retrieve, expose, and reconcile a user's evolving context.
  • Why it matters: The talk provides concrete comparative architecture patterns from ChatGPT and Claude, plus failure modes directly relevant to building durable agent and personalized-AI memory systems.
  • Best use: Use it to pressure-test memory architecture decisions: profile size and refresh cadence, always-on context versus on-demand retrieval, user editability, source coverage, and conflict handling.

Executive Summary

Shlok Khemani traces consumer-AI memory from ChatGPT's 2024 explicit fact list to its later background-maintained user profile, then contrasts that approach with Claude's initially retrieval-only system and subsequent hybrid design. His central observation is convergence at a high level: both products now combine a running user profile placed into new-chat context with tools that retrieve material from past conversations on demand. Their implementation choices remain materially different.

ChatGPT's early memory design imposed management work on users and accumulated stale facts. Its later profile system moves maintenance into asynchronous background synthesis, but can still confidently preserve incorrect inferences. Khemani's example: ChatGPT interpreted discussions weighing Thailand versus Turkey as evidence that he traveled to both, despite only traveling to Thailand. The important failure is not merely imperfect information, but a lack of awareness that the information is contradictory or unresolved.

The practical architecture lesson is that no universal memory stack exists. Neither ChatGPT nor Claude appears to use the simplistic default of embedding every conversation chunk in a vector store and semantically retrieving it. Instead, memory needs to be built in-house alongside the product because its right form depends on the product's interaction model, data sources, latency and cost envelope, and user-control requirements.

Khemani frames profile memory as a compute-allocation problem: frequent, high-quality profile synthesis costs maintenance compute, while longer profiles raise the cost of every served conversation. He also argues that current products face a larger context-coverage problem: their memories are isolated by application and generally cannot reason across the user's richest sources—email, calendar, photos, real-world conversations—or actively ask to resolve gaps. That cross-source, conflict-aware context layer is where personal AI remains immature.

Key Takeaways

  • Claim: The dominant consumer-memory pattern is converging on a hybrid of an always-available running profile plus on-demand retrieval over prior conversations. | Evidence: ChatGPT V2 maintains a background-updated profile that is inserted into each new conversation and later added past-conversation lookup; Claude V1 began with keyword/topic and time-based search tools only, then Claude V2 added a running profile while retaining retrieval. | Implication: For agent systems, separate persistent operating context from episodic-history retrieval rather than forcing all memory into either a static prompt or a retrieval index. | Caveat: Architectural convergence does not mean implementation parity: ChatGPT's profile is roughly 4,000 tokens and dense, while Claude's is roughly 1,000 tokens, sentence-based, refreshed daily, and directly editable.
  • Claim: Explicit fact-list memory creates user burden and staleness; asynchronous profile synthesis reduces management burden but does not solve truth maintenance. | Evidence: ChatGPT Memory V1 required users to say things such as "remember that I'm vegetarian" and exposed a list of memories. A stale entry such as "Shlok is going to Bengaluru" could remain injected long after it stopped being true. V2 automatically synthesizes a profile from conversations every few days. | Implication: A production memory layer needs provenance, temporal validity, uncertainty or competing hypotheses, and a correction path—not just extraction and summarization. | Caveat: Automatic summarization can produce durable false inferences: ChatGPT recorded both Thailand and Turkey as 2025 travel after the speaker merely discussed choosing between them.
  • Claim: Memory architecture must be treated as a core product capability and built in-house, not outsourced as a generic component. | Evidence: Khemani found that ChatGPT, Claude, Gemini, and agent products such as Claude Code, OpenClaw, and Hermes use materially different patterns, including profiles, timing logs, Markdown files, heartbeats, knowledge bases, and skills. He says leading consumer products all have memory but build it themselves. | Implication: Ken should define memory semantics—what is remembered, when it is retrieved, who can revise it, and how it expires—as proprietary control-plane behavior rather than accepting a vendor's default RAG abstraction. | Caveat: The talk does not claim external storage or retrieval vendors are useless; the point is that ownership of the memory policy and product integration cannot be delegated.
  • Claim: Profile memory is fundamentally a compute tradeoff between update quality/frequency and per-request serving cost. | Evidence: Khemani distinguishes maintenance cost, driven by update frequency and compute per synthesis, from serving cost, driven by profile tokens injected into every conversation. ChatGPT uses an approximately 4,000-token profile updated every few days; Claude uses an approximately 1,000-token profile updated every 24 hours. | Implication: Set explicit memory budgets by class: small, frequently refreshed operational state; selectively retrieved episodic detail; and costly deeper reconciliation jobs triggered only when the expected value warrants them. | Caveat: The idealized design—very frequent updates, powerful subagents, and a vastly larger profile—is economically impractical in a GPU-constrained environment.
  • Claim: Running-profile systems already resemble continual learning, but the loop occurs outside model weights. | Evidence: A profile influences each interaction; new interactions add information; a periodic "dreaming" process synthesizes that information back into the profile; and the revised profile shapes future interactions. | Implication: Near-term personalization strategy should focus on external, inspectable learning loops rather than assuming per-user fine-tuning will be the practical solution. | Caveat: Whether individual users will receive genuinely self-learning, weight-updated models is unresolved because personal training is expensive and lacks the enterprise economics of amortizing cost across many users.
  • Claim: The limiting problem is not only memory quality but incomplete and fragmented context acquisition, compounded by weak conflict detection. | Evidence: The speaker's actual Thailand decision happened offline and left corroborating evidence in flight and hotel emails, but ChatGPT neither reasoned over that email data nor recognized the contradiction between its Thailand and Turkey travel inference. His chatbot, assistant, vertical-app, agent, and device memories are also isolated from one another. | Implication: The differentiated opportunity is a permissioned cross-source context layer that tracks evidence, identifies ambiguous facts, and requests confirmation rather than silently cementing uncertain conclusions. | Caveat: Broader source access raises significant privacy, permissions, data-quality, and user-consent requirements, none of which the talk resolves.

Detailed Brief

Product-level comparison and user-control differences

  • Claims: ChatGPT V1 exposed individual memories directly, but its later raw running profile was initially hidden; users could inspect it only through a jailbreak-style prompt described by the speaker.; Claude V2 made its raw profile visible in settings and allowed explicit edit requests, including review and deletion of prior edits.; ChatGPT later exposed a generated summary of the already generated profile and allowed edit requests, while deprecating the original V1 fact list.
  • Evidence: Claude's profile refreshes every 24 hours and is written as complete sentences rather than ChatGPT's dense keyword-like contextual clues.; Gemini also uses a running profile but attaches detailed timing metadata, including when a memory was created and last updated.
  • Caveats: Visibility does not automatically make a memory system legible: a generated summary of a generated profile can obscure what the actual serving context contains.; Direct editing can correct known errors but does not ensure the system detects unknown errors or inconsistencies.
  • Implications: Design memory UX around inspectability of the actual operational representation, not merely a friendly paraphrase.; Treat timestamps, edit history, source links, and revision history as trust primitives for user-facing memory.

Future model and ecosystem questions

  • Claims: Khemani believes personalized AI is still early despite rapid progress; he characterizes AI memory as only about three years old.; He identifies a future possibility of one self-learning model per person, but leaves the required bootstrap data, learning mechanism, and economic model as open questions.
  • Evidence: He recommends Gwern's essay "Guardian Angels" for a detailed exploration of the one-model-per-person future.; His critique is product-oriented: he argues that models could be designed to notice missing or conflicting personal information, but current products do not prioritize that behavior.
  • Caveats: The talk is a reverse-engineering and product-analysis perspective, not a verified disclosure of providers' complete internal implementations.
  • Implications: The near-term strategic moat is likely better context governance and reconciliation behavior, not claims of permanent, omniscient personal memory.; Monitor whether platforms open interfaces for portable, consented personal context; fragmentation is both a user pain point and a platform-control issue.

Notable Concepts & Terms

  • Running profile: A periodically synthesized persistent representation of a user that is added to new conversations, distinct from a list of isolated facts.
  • Dreaming: The informal name for asynchronous background synthesis that reviews recent conversations and updates the user profile.
  • Always-on memory versus on-demand retrieval: The core architecture split: inject compact standing context every turn, while querying prior episodes only when relevant.
  • Truth maintenance: The unresolved need to detect stale, contradictory, speculative, or superseded memories rather than treating extracted assertions as durable facts.
  • Memory as a function of compute: Profile refresh frequency and synthesis depth consume maintenance compute; profile length consumes recurring serving cost.
  • Continual learning outside the weights: A feedback loop in which external profile state learns from interactions and changes future model behavior without retraining model parameters.
  • Context problem: Memory quality is bounded by the sources the system can access and reason over; isolated chat logs miss real-world decisions and rich personal data.
  • RAG: Retrieval-augmented generation is presented as an insufficient default for memory; embedding and searching conversation chunks alone does not define a complete memory product.

Operator Notes / Why Ken Should Care

  • Establish a memory schema that separates durable preferences, time-bounded state, hypotheses or plans, episodic evidence, and derived summaries; do not store all extracted statements as equivalent facts.
  • Require each material memory to carry source provenance, timestamp or validity window, confidence, and a supersession mechanism.
  • Implement a conflict queue: when sources disagree or evidence only supports an intention rather than a completed event, ask the user or retain uncertainty instead of writing a definitive profile assertion.
  • Budget memory as a tiered compute system: compact context injected per run, event-driven or scheduled profile refreshes, and selective deep reconciliation over high-value sources.
  • Make the actual agent-facing profile inspectable and editable, with revision history; avoid exposing only a separately generated summary.
  • Evaluate a consented connector strategy for email, calendar, documents, and operational systems, but gate it with least-privilege permissions and explicit rules for what may update durable memory.
  • Read Gwern's "Guardian Angels" as input to longer-horizon investment or product-thesis work on persistent personal agents.

Source/Metadata

  • Title: Lessons from Studying Every Memory System — Shlok Khemani, Independent
  • Transcript words: 4009
  • Duration seconds: 1170
  • Timestamp note: No timestamps or chapters were present in the supplied transcript; the latter portion repeats material from the earlier talk.

Transcript

2891 words en Processed in 103.9s

Shlok Sulek Reviewer Okay. Hi, everyone. I'm Shlok, and I've spent the past year studying different memory systems. Now, before I get started, one thing I've realized, speaking to people over the last two days, is that memory is a very overloaded term now. It can mean a lot of different things. So when I talk about memory today, it is going to be in the context of personalization, especially for consumer AI applications. Now, a little bit about me. My claim to fame, the reason I get to speak to you here, is that I've spent the past year trying to reverse engineer how products like ChatGPT, Claude, Gemini, and Poke implement their memory systems. And I've then worked with multiple teams across different domains in helping them design their memory. I'm going to break the talk down into two parts. First, we're going to look at how memory has evolved over the past three years, especially in the context of ChatGPT and Claude. And then in part two, I'm going to discuss some of the lessons I've learned, maybe a rant, and where I think all of this is going. To kick things off, we go back to ancient times, which in our industry is 2023. This is ChatGPT just after the launch of GPT-4. Now, you could have back-and-forth conversations within a single thread, and context was maintained inside that thread. But as soon as it started a new conversation, nothing was carried over. Now, for early adopters, this wasn't a problem. GPT-4 was such an amazing model that if we ever had the need to carry context, we would do so by hand. But as ChatGPT started becoming more popular, as regular people started using it for things like learning, cooking, as a companion, the need for some sort of memory system became really apparent. So in February of 2024, we got ChatGPT Memory V1. And what you could do is ask ChatGPT to remember things about you. So you could say things like, "Hey, remember that I'm vegetarian." And ChatGPT would extract what it thought was a fact, which is that the user is vegetarian, store it in a list of memories, and this list was then added to the context window for every single conversation. You could also then go into settings and view this list of memories. And if you thought that something didn't apply anymore, you could delete a memory. Now, as the first serious memory implementation within our industry, I think this was a really decent effort, but there were also some fundamental flaws with it. The biggest one was that, as a user, because you could see every time a memory was created, it felt like you were responsible for both creating memories while you were just trying to have a conversation. So the burden of memory management fell to the user. Also, if you notice this list of memories here, these held true at the time they were being created, but that doesn't necessarily hold true over time. For example, it says that Shlok is going to Bengaluru. Now, I obviously am in SF right now. I am not going to Bengaluru, but this fact, this memory, is still added to my context window today. So staleness was another huge problem with this version of ChatGPT's memory. A little more than a year later, in April of 2025, ChatGPT released V2 of its memory, and this was a little more sophisticated. The most important addition was this thing called user knowledge memories. I am just going to call it a running profile for the rest of this talk. And what a running profile really is, is that every few days, ChatGPT looks at all the conversations you have had with it, extracts anything it thinks is important for it to know about you, and updates this profile that it maintains on you. Now, this update process is also what a bunch of folks now call dreaming. How many of you all were there for Lance Martin's talk yesterday? Okay, not many, but he did a great talk on this. So every few days, ChatGPT looks at the new conversations you're having, updates your profile, and then this updated profile is added to the context window for every single new conversation. These are two excerpts from my running profile. I want you to notice a few things. First, these are extremely dense memories. So ChatGPT tries to pack in as much context as it can within every single memory. What's essentially happening here is that they're trying to put in keywords almost like clues, and because LLMs, especially the frontier models today, are so good at inferring context from limited information, when you're having a conversation, it connects these clues to what you're talking about. Also, these are just two of 16 different sections in my profile. Other sections include my personal life and things I'm working on. In total, my profile is almost 4,000 tokens long. And because these updates are happening asynchronously, they're happening in the background, this new version does away with the flaw we discussed in V1, which is the burden of memory management was taken away from the user. But I want you to notice the highlighted memory. This is about places I traveled to in 2025. But if you pay attention, it says Thailand and Turkey, but the dates are overlapping. And that's because the source of this memory was conversations I was having with ChatGPT, deciding between where to go among these two places. Now, I did end up going to Thailand. I've never been to Turkey. But ChatGPT still says that I've been to Turkey in 2025. So the staleness problem with V2 didn't completely go away. Another very important thing is that if you go to your settings, ChatGPT doesn't let you view this raw profile. So you could view your memories from V1. Your raw profile is not visible to you. Now you may ask, "Shlok, how did you see your profile then?" And that's because this prompt works really well if you want to jailbreak ChatGPT and view your raw profile. You might have to attempt it a few times, try different thinking modes. But Claude enough and you shall receive. In August of 2025, Claude released its first version of memory. This surprised me a bit because if you compare ChatGPT and Claude, they are very similar applications. You have a chat box, you have back-and-forth conversations, you have a list of previous conversations, you can start a new conversation. My assumption going into studying Claude was that the memory systems would also be similarly designed. Not the case, at least for V1. So in V1 of Claude, you had no user profile, you had no list of facts. Instead, the model was given two tools. It was given a tool to search over previous conversations by keyword or topic, and it was given another tool to search over conversations by time period. So queries like, "What did we discuss last week?" or "What did we discuss at the start of November of 2025?" So in V1, every single conversation starts fresh, with no context on the user. And when the model thinks that it needs to retrieve something, it can do so on demand. On September 11 of last year, I released a blog post saying, "Claude's memory architecture is the opposite of ChatGPT's." This hit the Hacker News front page. Funnily, on that very day, Claude released V2 of its memory, and they added a running profile similar to ChatGPT, but with a few differences. First, Claude made this profile visible to users. So you could go to settings, and you could view your raw profile. Second, this profile was 1,000 tokens. So it was much smaller than ChatGPT's 4,000 tokens. And also, if you notice, these are complete sentences rather than the dense keyword approach of ChatGPT. So less dense and smaller. Claude's profile updates every 24 hours, versus ChatGPT's every few days. And Claude also let the user make explicit edits to this profile. So you could request an edit, and that edit would lead to a re-synthesis of the profile. It gave you an interface to manage previous edits, and you could delete the things that no longer held true. And this is how Claude's memory works even today. So it hasn't changed since September of last year. We have seen two updates within ChatGPT's memory this year, though. The first was that it added a tool to look over past conversations, like we just saw with Claude. So the model can retrieve summarized context based on queries it makes. And then a month ago, at the start of June, ChatGPT finally made the user profile visible to users, somewhat. So what you can see is an LLM-generated summary of your profile, which is weird because your profile is already an LLM-generated summary of your conversations. It's all a bit confusing. I've written about it, but it is visible in some sense. You could also request explicit edits to your profile. And with this update, ChatGPT deprecated the V1 fact list from its memory system. So what we've seen here is a convergence after three years of each of these products evolving independently, where they both now have a running profile. This profile is visible and editable, again, somewhat. And the model has tools to look over past conversations. Okay, so what can we learn from this evolution, and where are things going? I think the biggest lesson for me is that there is no single way to do memory. It wasn't too long ago that everyone, including me, assumed that RAG was the way to go about memory, where you would take conversations, chunk them, create embeddings, put them in a vector store, and then, as user queries came in, do some sort of semantic search. But as we saw, neither ChatGPT nor Claude really do this. Instead, they both evolved independently using different approaches. And while the general architectures have converged, the specific implementation details are still very different. And then if you look at Gemini, it also has a running profile, but each memory comes with detailed timing logs. So when was it created? When was it last updated? And then if you look at agents like ClaudeCode, OpenClaw, Hermes, they have completely different memory systems, with markdown files, heartbeat, knowledge bases, skills. The point being that there is no one way to do memory. The implication of this is that memory cannot be outsourced. If you're a serious team, you do not outsource memory. It is something that you build alongside your product. Your memory system evolves with your product, and it cannot be thought of as an afterthought. And there is plenty of evidence for this. So if you look at all of the top consumer products today across different categories, each of these has some form of memory, yet none of them outsource it. All of them build memory in-house. Lesson two: memory is a function of compute. What does that mean? Let's look at the costs associated with a running profile. So there are two types of costs. There is a cost to maintain a profile. And that depends on how frequently you update it and how much compute you apply to each update. And then, because these profiles are part of the context window for every single conversation, there's a cost of serving, which is the longer the profile, the more it costs to serve. Now, thought experiment: if you were to design the ideal memory system with no restraints, what would you do? You might want to update your profile every hour, or maybe after every conversation. You might want to task a bunch of Opus sub-agents with the update itself. And why stop at 4,000 tokens? Why not make it 400,000 tokens, store every single thing you would want about the user? Unfortunately, we live in a GPU-constrained world, and tradeoffs have to be made. And you can see that happening here. So ChatGPT, the profile length is 4,000 tokens. It updates every few days. So they have a higher serving cost for a lower update cost. And for Claude, it's 1,000 tokens, updates every 24 hours, so they make the exact opposite tradeoff. And this is what I mean by memory as a function of compute. You have to really think about how much compute you want to put into memory. Third, we had a bunch of talks about continual learning today. I'm not an expert here, but what I would say is that continual learning is already here. Going back to running profiles, what exactly is happening here? Your running profile starts with something that the model knows about you. This is then applied to every single conversation. Each of these conversations brings in new information. This new information is then synthesized through the dreaming process back into the profile. And then this profile dictates future conversations. And this loop keeps repeating itself again and again and again. And what you have is a continual learning process. Now, obviously, this learning loop is happening outside the weights. And a big question, particularly for consumer AI, is: will this process ever make its way into the weights? Now, obviously, updating weights, training models, is an expensive process. Continual learning does make sense at an enterprise level because the costs of these models are amortized across different employees and different customers. But that's not the case at an individual level. So big, open questions that I don't yet know the answers to are: will each of us get our own self-learning model? What data do we need to kick the CL process off? And how do we generate it? And finally, who's going to pay for this? How would the economics for this work? Gwern recently wrote an essay called Guardian Angels, where he explores this topic in beautiful detail. And if you're interested in what the future for one model per person looks like, I would recommend reading this. Finally, my rant is that we have a massive context problem. You could have the best memory architecture in the world. You could pour infinite amounts of compute into it. You could have continuous learning working at an individual level, where every single data point you bring up is somehow perfectly integrated into the model weights. Yet your memory system is capped by how much context it can gather about you. Let's go back to the example we discussed earlier, which was the conflict between where I traveled in the summer of 2025. These are the two source conversations. Again, I was trying to use ChatGPT to decide between which of these two countries to go to. Now, the decision to go to Thailand was actually made in a conversation I had with my partner in person. And ChatGPT couldn't reason over this or listen to this. But there were also traces of this conversation in my emails because I had flight and hotel bookings for Thailand. But even if ChatGPT is connected to my email, it doesn't reason over my email and it doesn't update my profile from my email, so it couldn't resolve this conflict. And I think that's okay. It's understandable. But what really bothers me is that ChatGPT today doesn't realize that there is a conflict. It's not curious about trying to fill in gaps in the information it knows about me. And this is particularly interesting and also infuriating because it's not a technology problem. It's a product problem. There is no fundamental reason at an LLM level that these things can't be solved. It's just that our products today are not designed to help us with this. So my personal stack today is a bunch of chatbots, assistants, vertical-specific applications, agents, and even hardware devices. Each of these products is trying to build its own memory of me. None of these memories are shared with each other. So I have to rebuild context within every single product from scratch every time. When something in my life changes, I have to individually update all of them. And then I have a bunch of very rich existing context sources, like my email, calendar, and photos. None of these products are able to reason over my existing, very rich context sources. So for me, none of this feels like 2026. And what I keep asking myself every day is, when will personal AI feel like personal AI? All of that frustration aside, I still think we're very early. Memory for AI is just a three-year-old field. Memory is also foundational to how humans interact with AI. And because I hope to be talking to AI all my life, and I know that's going to be the case for every single one of us here today, memory is something that's going to be important for the rest of human history. And there's so much left to build. That's it from me. Thank you so much. You can find my website. You can find me on Twitter. Have a great rest of the conference. Thank you. Thanks. and it cannot be thought of as an afterthought. And there is plenty of evidence for this. So if you look at all of the top consumer products today, across different categories, each of these has some form of memory, yet none of them outsource it. All of them build memory in-house. Lesson two, memory is a function of compute. What does that mean? Let's look at the costs associated with a running profile. So there are two types of costs. There is a cost to maintain a profile. And that depends on how frequently you update it, and how much compute you apply to each update. And then because these profiles are part of the context window for every single conversation, there's a cost of serving, which is the longer the profile, the more it costs to serve. Now, third experiment. If you were to design the ideal memory system with no restraints, what would you do? You might want to update your profile every hour, or maybe after every conversation. You might want to task Fable with a bunch of Opus sub-agents for the update itself. And why stop at 4,000 tokens? Why not make it 400,000 tokens, store every single thing you would want about the user? Unfortunately, we live in a GPU-constrained world, and tradeoffs have to be made. And you can see that happening here. So chat GPT, the profile length is 4,000 tokens. It updates every few days. So they have a higher serving cost for a lower update cost. And for Claude, it's 1,000 tokens, updates every 24 hours, so they make the exact opposite tradeoff. And this is what I mean by memory as a function of compute. You have to really think about how much compute you want to put into memory. Third, we had a bunch of talks about continual learning today. I'm not an expert here, but what I would say is that continual learning is already here. Going back to running profiles, what exactly is happening here? Your running profile starts with something that the model knows about you. This is then applied to every single conversation. Each of these conversations bring in new information. This new information is then synthesized through the dreaming process back into the profile. And then this profile dictates for the conversations. And this loop keeps repeating itself again and again and again. And what you have is a continual learning process. Now, obviously, this learning loop is happening outside the weights. And a big question, particularly for consumer AI, is will this process ever make its way into the weights? Now, obviously, updating weights, training models is an expensive process. Continuous learning does make sense at an enterprise level because the costs of these models are amortized across different employees, different customers. But that's not the case at an individual level. So big, big open questions that I don't yet know the answers to, which is, will each of us get our own self-learning model? What data do we need to kick the CL process off? And how do we generate it? And finally, who's going to pay for this? How would the economics for this work? Gwern recently wrote an essay called Guardian Angels, where he explores this topic in beautiful detail. And if you're interested in what the future for one model a person looks like, I would recommend reading this. Finally, my rant is that we have a massive context problem. You could have the best memory architecture in the world. You could pour infinite amounts of compute into it. You could have continuous learning working at an individual level where every single data point you bring up is somehow perfectly integrated into the model weights. Yet, your memory system is capped by how much context it can gather about you. Let's go back to the example we discussed earlier, which was the conflict between where I traveled to in the summer of 2025. These are the two source conversations. Again, I was trying to use ChatGPT to decide between which of these two countries to go to. Now, the decision to go to Thailand was actually made in a conversation I had with my partner in person. And ChatGPT couldn't reason over this or couldn't listen to this. But there were also traces of this conversation in my emails because I had flight and hotel bookings for Thailand. But because even if ChatGPT is connected to my email, it doesn't reason over my email and it doesn't update my profile over my email, it couldn't resolve this conflict. And I think that's okay. It's understandable. But what really bothers me is that ChatGPT today doesn't realize that there is a conflict. It's not curious about trying to fill in gaps in the information it knows about me. And this is particularly interesting and also infuriating because it's not a technology problem. It's a product problem. There is no fundamental reason from an LLM level that these things can't be solved. It's just that our products today are not designed to help us with this. So my personal stack today is a bunch of chatbots, assistants, vertical-specific applications, agents, and even hardware devices. Each of these products is trying to build its own memory of me. None of these memories are shared with each other. So I have to rebuild context within every single product from scratch every time. When something in my life changes, I have to individually update all of them. And then I have a bunch of very rich existing context sources. Like my email, calendar, photos. None of these products are able to reason over my existing very rich context sources. So for me, none of this feels like 2026. And what I keep asking myself every day is when will personal AI feel like personal AI? All of that frustration aside, I still think we're very early. Memory for AI is just a three-year-old field. Memory is also foundational to how humans interact with AI. And because I hope to be talking to AI all my life, and I know that's going to be the case for every single one of us here today, memory is something that's going to be important for the rest of human history. And there's so much left to build. That's it from me. Thank you so much. You can find my website. You can find me on Twitter. Have a great rest of the conference. Thank you. Thanks.