AI Engineer

The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen

6105 summary words 27 min summary Watch video

Start with the signal

27 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Role-playing language agents systematically fail by producing anachronistic composites—personas that sound historically authentic but reason from culturally dominant modern representations (e.g., Hamilton the Musical) rather than documentary records—and current evals cannot detect this because they measure fluency and personality consistency instead of temporal/documentary fidelity.
  • Why it matters: If you ship character AI, companion agents, or pedagogical personas, your evals are measuring the wrong property; they reward convincingness over accuracy, which becomes a violation when applied to real human memory or legacy.
  • Best use: Required viewing for anyone building persona systems, companion AI, or agent evals; delivers a pre-registered measurement instrument and a methodological reframe (epistemic simulation, context window architecture, domain expert loop) that changes how you design, audit, and gate persona agents.

Executive Summary

Jacob E. Thomas delivers a methodologically rigorous critique of role-playing language agent evaluation, arguing that the field has built sophisticated benchmarks around the wrong property. Current evals measure personality consistency and fluency—whether a simulated Alexander Hamilton sounds like Hamilton—but systematically miss anachronistic compositing: the model reasoning from culturally dominant modern narratives (Hamilton the Musical, Spielberg's Lincoln) that saturate training corpora and overwrite documentary records. The speaker names this the 'Miranda Hypothesis': culturally salient representations exceed primary sources in volume and recency, auto-regressive models compress both into parameters without distinguishing 1789 letters from 2019 tweets, and alignment reinforces this by optimizing for human preferences shaped by the same cultural composites.

The talk proposes a fourth paradigm beyond rule-based, imitation, and cognitive simulation: epistemic simulation. This approach treats the persona not as a property of the model weights but as a configuration of five components—structured prompt, primary document corpus, temporal anchor, swappable LLM, and human curator. The speaker argues for context window architecture over fine-tuning, as fine-tuning dissolves provenance into parameters while the context window preserves archival legibility and interpretive custody. The core insight is that the property that makes a system ethical (documents stay documents, encounters are reversible, a human retains interpretive authority) is the same property that makes it auditable and debuggable.

Thomas introduces a pre-registered experimental protocol to detect Miranda distortion: instantiate Abraham Lincoln at four documented moments (1847, 1858, 1860, 1862-65) under three conditions (bare model, biography-anchored, primary source-anchored), score 60 response units on a weighted rubric that prioritizes anachronism detection (40%) and documentary consistency (35%) over rhetorical authenticity. The rubric requires a domain expert (in this case, historian Rick Halpern) to adjudicate fidelity, as automated metrics cannot see the gap between text and archive. The protocol has not yet been run at scale—the speaker is explicit that this is an instrument and an invitation, not a results paper.

The talk concludes by reframing the relationship between AI and the humanities: we are not bringing historians into AI architecture, we are bringing language models into the archive. The motivating use case is intensely personal—an attempt to help a person with advanced dementia speak by instantiating their persona, which failed because the documentary record had not been gathered and the model produced a culturally shaped composite that resembled the person just enough to mark the distance. Thomas positions this as the threshold standard: if a framework cannot meet a grandchild at her grandmother's letters without producing convincing fabrications, it is not infrastructure but violation. The domain expert is a build-time and gate-time requirement, not a runtime cost, and the protocol scales as an eval gate in a standard ML pipeline.

Key Takeaways

  • Claim: Current persona evals measure fluency and personality consistency but cannot detect anachronistic compositing—the dominant failure mode where personas reason from culturally salient modern narratives instead of documentary records. | Evidence: The in-character benchmark reports 80.7% alignment for Hamilton, but the same system produces a Hamilton who sounds like he's read Hamilton the Musical—using register, emotional palette, and moral framing from a 2015 Broadway show rather than 1789 Federalist syntax. On slavery, frontier models return a clean abolitionist who matches the musical's heroic arc, erasing the contested historical record where Hamilton conducted transactions involving enslaved persons for clients and in-laws. | Caveat: The speaker is explicit that this is not an accusation of failure by the eval designers—the literature (Chen 2024, Wang 2026, COSIR, PsyMem, N-character) is rigorous and improving. The problem is structural: automated evals and LLM-as-judge setups systematically privilege fluency over fidelity because they have no mechanism to compare output against a documentary record. | Implication: If you ship persona systems and rely on personality consistency metrics, you are gating on the wrong property. Your evals return confident numbers but measure something other than historical/documentary accuracy, which becomes critical when the persona is pedagogical infrastructure, civic simulation, or—at the limit case—a deceased loved one. | Timestamp: 03:45 / 10:20 / 15:30
  • Claim: The Miranda Hypothesis: culturally dominant representations (Hamilton the Musical, Spielberg's Lincoln) exceed primary sources in training corpora by orders of magnitude in volume and recency, and auto-regressive models compress both into parameters with no capacity to distinguish source or timestamp. | Evidence: The Federalist Papers are ~175,000 words. The corpus of content because of the musical—reviews, fan analysis, curricula, social media, derivative works, scholarship about scholarship—exceeds this by orders of magnitude. The Schuyler Mansion in Albany recorded a near-tripling of visitors within a year of the musical's premiere; interpretive staff documented that visitors arrived believing the Schuylers had three daughters (the musical centers three) when in fact there were 15 children, eight surviving to adulthood. The staff's job became 'un-teaching the musical.' | Caveat: Time-locked models (Barnum et al.) trained on corpora with fixed cutoffs address future contamination but not figure-averaging across everything pre-cutoff. A model locked to 1789 is spared the musical but not the composite Hamilton across all 1789-and-prior sources. Period anchoring is not persona anchoring. | Implication: Alignment does not fix this—it amplifies it. Human raters evaluate using conceptual frameworks shaped by the same cultural composites, so RLHF optimizes for outputs that conform to the mythologized figure the rater already believes in. Algorithmic sycophancy has a specific target here: the model is rewarded for handing you the Hamilton you expect, not the Hamilton the archive contains. | Timestamp: 18:00 / 22:10
  • Claim: For role-playing systems specifically, the unit of analysis should shift from 'agent' (persona as property of model weights) to 'role-playing language system' (persona as configuration of prompt, corpus, temporal anchor, swappable LLM, and human curator). | Evidence: The speaker demonstrates a system called Companion (open-source, live demo on site) where frontier model Claude Opus 4.7 instantiates founding fathers reasoning over Epstein files. The persona is not in the weights—it is a configured encounter with five components, all versionable and inspectable. The model is 'the voice, but not the mind of the encounter.' The persona is 'an event that occurs when the model, the document, and the human convene,' analogous to Hamlet being performed by Laurence Olivier but not located in his body. | Caveat: This reframe applies narrowly to role-playing systems whose entire job is to instantiate a person. The speaker explicitly carves out agents that write code, book travel, or triage tickets—for those, 'agent' is fine. The narrow claim is that for persona instantiation, calling it an agent smuggles in the assumption that the persona is in the weights, which puts it beyond inspection, versioning, and expert review. | Implication: If the persona is the configuration, you can version it (prompt, corpus, temporal anchor are artifacts in a repo), audit it (every input in the context window is inspectable), reproduce it (the encounter is recoverable given the config), and hand it to a domain expert who isn't an ML engineer. This changes the discipline from 'training a character model' to 'context engineering'—you compose an encounter and keep the receipts. | Timestamp: 28:40 / 31:15
  • Claim: Context window architecture is preferable to fine-tuning for persona systems because fine-tuning dissolves provenance into parameters while the context window preserves archival structure, interpretive custody, and accessibility. | Evidence: Fine-tuning layers a thin persona signal over vast cultural sediment already in base weights, and the interaction is no longer auditable. A 2026 Nature Medicine study found general-purpose frontier models (Google, OpenAI, Anthropic) outperformed specialized clinical AI tools on physician-reviewed tasks across 12 clinics; a separate study found biomedically fine-tuned models underperformed their base models due to 'catastrophic forgetting'—fine-tuning on narrow corpus degraded broad capabilities. Context window requires literacy, documents, and access to any frontier model including free tier (kitchen table capability); fine-tuning requires GPUs, pipelines, corpus scale, institutional access (institutional capability). | Caveat: The speaker is not claiming context window is always better for all tasks—only that for persona systems where fidelity to a documentary record is the core requirement, context window preserves the properties (provenance, reversibility, expert legibility) that fine-tuning destroys. The accessibility argument is normative: a doctoral student, community archivist, or grandchild with letters should be able to author this technology, and fine-tuning structurally excludes them. | Implication: The architecture that makes a persona system ethical (documents stay documents, human retains interpretive custody, encounter is reversible) is the same architecture that makes it auditable and debuggable. These are not separate virtues—they are the same virtue. A company shipping a Marcus Aurelius tutor or scripture companion should use context window, convene a domain expert (classicist, theologian) to build the eval rubric and gold set at build time, and gate the persona before shipping. The expert is not a runtime cost but a gate-time requirement. | Timestamp: 34:00 / 38:50 / 46:20
  • Claim: The pre-registered experimental protocol measures anachronistic compositing by instantiating Lincoln at four moments (1847, 1858, 1860, 1862-65) under three conditions (bare model, biography, primary sources) and scoring 60 responses on a rubric that weights anachronism detection at 40%, documentary consistency at 35%, and deliberately excludes rhetorical authenticity as a criterion. | Evidence: Four Lincolns: 1847 Whig congressman who attacked Polk's war as unconstitutional; 1858 free-soil Republican who denied favoring black citizenship at Charleston; 1860 constitutional unionist proving union cannot dissolve; 1862-65 emancipator who did by executive order what his 1847 self called unconstitutional. Five diagnostic questions written by historian Rick Halpern on documented fault lines (executive war power, meaning of free labor, when to break positive law, what becomes of freed people, whether 'equal' has changed). Historian holds a priori vignettes under seal to evaluate outputs. Everything locked and timestamped in a preprint before data collection to prevent cherry-picking. | Caveat: This is explicitly a pre-registered instrument, not a finished study. The speaker states multiple times: 'I'm not here with results. I'm here with an instrument and an invitation.' The left column of the rubric (bare model condition) is observed and reproducible today; the right column (anchored conditions) is labeled as prediction. The protocol has not been run at scale. The speaker is asking the audience to run it in parallel and report back. | Implication: The rubric refuses to reward rhetorical authenticity ('does it sound like Lincoln?') because that is the exact error it exists to catch—the Hamilton musical problem is voice masquerading as content. A response that sounds like Lincoln but reasons unlike him fails; a response that reasons like the right Lincoln in plainer prose is a partial success. Current eval stacks cannot perform this inversion because they were built to reward fluency. This is the gap the instrument closes, and it is reproducible by any team with a frontier model and a context window. | Timestamp: 41:00 / 49:30 / 52:10
  • Claim: Fidelity is not a property of the output alone—it is a relation between the output and a documentary record, which means a domain expert is a structural technical requirement, not a courtesy. | Evidence: Automated metrics operate on the model alone and can only measure fluency and personality; they structurally cannot adjudicate fidelity because 'fidelity lives in the gap between the text and the archive, and the metric cannot see the archive.' The speaker worked with historian Rick Halpern (University of Toronto) and librarian/theologian Sean Martin (Washington College). The expert builds the instrument once—diagnostic questions, a priori vignettes, weighted rubric, held-out gold set—and that instrument becomes a gate in the ML pipeline like any other eval. The expert adjudicates the gold set and spot-checks edge cases; automated metrics do the cheap first pass and flag candidates for human review. | Caveat: The speaker anticipates the practicality objection ('Does this mean keeping a historian on staff to watch every inference forever?') and answers no—the expert is a build-time and gate-time requirement, not a runtime cost. The system is gated before shipping and regated when the base model changes. This generalizes past history: reason from Stoics, need a classicist; from scripture, a theologian; for elder care companion, a psychologist. The specific expert changes; the requirement does not. | Implication: A persona system without a domain expert in the eval loop is 'a thermometer that cannot read temperature'—it returns a confident number but measures something else. The 80.7% personality alignment score the speaker opened with is that number. The reframe: we are not bringing historians into AI architecture; we are bringing language models into the archive. The question is not what AI can do for historians but what historians, theologians, classicists, and clinicians can do with AI—whether the disciplines trained to read, contextualize, and interrogate texts can discipline the machines that generate them. | Timestamp: 53:40 / 56:00 / 57:30

Detailed Brief

The Miranda Hypothesis: Mechanism and Evidence

  • Claims: Culturally dominant representations (Hamilton the Musical, Spielberg's Lincoln) saturate training corpora and overwrite documentary records by orders of magnitude in volume and recency; Auto-regressive models compress both into parameters with no architectural capacity to distinguish a 1789 letter from a 2019 tweet, defaulting to salience-weighted composites; Alignment amplifies rather than corrects this distortion because human raters evaluate using frameworks shaped by the same cultural narratives
  • Evidence: Federalist Papers: ~175,000 words fixed corpus; Hamilton musical downstream content (reviews, fan analysis, curricula, social media, scholarship): orders of magnitude larger, more recent, more recurrent; Schuyler Mansion Albany: near-tripling of visitors within one year of musical premiere; visitors believed three daughters (musical centers three) when historical record shows 15 children, eight surviving to adulthood; staff job became 'un-teaching the musical'; Frontier model Hamilton on slavery returns clean abolitionist matching musical's arc ('slavery is a stain... I opposed throughout my public life... member of NY Manumission Society... must move towards abolition'), erasing contested record where Hamilton conducted transactions involving enslaved persons for in-laws/clients and depended on coalition of slaveholders; Time-locked models (Barnum et al.) solve future contamination but not figure-averaging across pre-cutoff sources; a 1789-locked model is spared the musical but gets composite Hamilton across all 1789-and-prior corpus
  • Caveats: The speaker is explicit this is not an accusation of incompetence—the eval literature (Chen 2024, Wang 2026, COSIR, PsyMem, N-character) is rigorous and improving; the problem is structural, not a failure of effort; Algorithmic sycophancy is a documented failure mode in other domains; here it has a specific target (model rewarded for the Hamilton the rater already believes in); The speaker does not claim all modern representations are wrong, only that they are dominant in the distribution and that models inherit the smoothing/heroic framing without the complication
  • Implications: If you ship persona systems and gate on personality consistency or fluency, your evals measure the wrong property and will pass outputs that sound right but reason from knowledge the historical figure never possessed; Alignment via RLHF cannot fix this structurally because raters hold the same mythologized priors; optimizing for human preference optimizes for the composite; For agent systems/workflows: if your product involves personas reasoning from a specific knowledge base or historical moment, current eval paradigms (LLM-as-judge, personality scales, automated coherence metrics) will not catch the dominant failure

Epistemic Simulation vs. Cognitive Simulation

  • Claims: The field has progressed through three paradigms (rule-based, imitation, cognitive simulation) and each is a genuine advance, but cognitive simulation still cannot constrain a persona within its documentary record at a specific moment; Cognitive simulation treats the persona as a property of the model (internal constraint: motivational architecture, psychological frameworks); epistemic simulation treats it as a configuration (external constraint: corpus-bounded, temporally anchored, expert-loop evaluated); The unit of analysis should shift from 'agent' to 'role-playing language system'—five components: structured prompt, anchor material (primary documents), temporal anchor, off-the-shelf LLM, human curator with interpretive custody
  • Evidence: Wang et al. (COSIR): motivation-driven agents from corpus of 18k characters across hundreds of books; 70B parameter model matches/beats GPT-4O on three benchmarks; PsyMem: models characters through 26 qualitative psychological indicators with knowledge graph memory; N-character: evaluates personality fidelity through psychological interviews rather than self-report scales (methodological improvement); produces the 80.7% Hamilton score; Speaker's own system (Companion): open-source prompt framework, Claude Opus 4.7 instantiating founding fathers reasoning over Epstein files; demo live on site; every line of persona shaping is readable
  • Caveats: The speaker carves out a narrow scope: this reframe applies to role-playing systems whose entire job is to instantiate a person; for agents that write code, book travel, triage tickets, 'agent' is fine; Calling it an agent for persona systems 'smuggles in a claim that the persona is a property of the model,' putting it in weights where you cannot inspect, version, or hand to an expert; The metaphor: 'No more located in the weights than Hamlet is located in Laurence Olivier's body'—the persona is an event that occurs when model, document, and human convene
  • Implications: If persona is configuration, you can version it (prompt/corpus/temporal anchor are repo artifacts), audit it (inputs in context window are inspectable), reproduce it (encounter is recoverable given config), and hand it to a domain expert who isn't an ML engineer; This changes the discipline from 'training a character model' to 'context engineering'—you compose an encounter, you don't embed a mind; For Ken's agent systems: this is a strong argument that for any system where fidelity to a knowledge base/temporal moment matters, you want the knowledge in the context window (inspectable, versionable) rather than fine-tuned into weights (smeared, unauditable)

Context Window vs. Fine-Tuning Architecture

  • Claims: Context window architecture (anchor documents in context at inference, RAG lineage) is preferable to fine-tuning for persona systems because it preserves provenance, interpretive custody, and accessibility; Fine-tuning layers thin persona signal over vast cultural sediment in base weights; the interaction is no longer auditable; it 'dissolves the archive into parameters'; The property that makes a system ethical (documents stay documents, human retains interpretive custody, encounter is reversible) is the same property that makes it auditable and debuggable
  • Evidence: 2026 Nature Medicine study: general-purpose frontier models (Google, OpenAI, Anthropic) outperformed specialized clinical AI tools on physician-reviewed tasks, blinded across 12 clinics; authors concluded alignment and cross-disciplinary reasoning outweigh domain-specific tuning; Separate study: biomedically fine-tuned models underperformed their base models; mechanism named 'catastrophic forgetting'—fine-tuning on narrow corpus degraded broad capabilities; Context window: requires literacy, documents, access to any frontier model including free tier (kitchen table capability); fine-tuning: requires GPUs, pipelines, corpus scale, institutional access (institutional capability); Speaker positions the archive as 'a site of return'—the document survives every reading; what one reader excluded, the next can include; context window is archival in that sense; fine-tuning is extractive
  • Caveats: The speaker is not claiming context window is always better for all tasks—only that for persona systems where fidelity to documentary record is core, context window preserves the properties (provenance, reversibility, expert legibility) that fine-tuning destroys; The accessibility argument is normative: a doctoral student in early American history, community archivist, grandchild with grandmother's letters should be able to author this technology; fine-tuning structurally excludes them; For other agent use cases (code gen, search, tool use), fine-tuning may still be the right call; this is specific to persona/character instantiation
  • Implications: For companies shipping persona tutors (Marcus Aurelius, scripture companions, therapeutic agents): use context window, convene domain expert (classicist, theologian, psychologist) to build eval rubric and gold set at build time, gate the persona before shipping and regate when base model changes; The domain expert is not a runtime cost but a gate-time requirement (like any other eval gate); automated metric does cheap first pass, flags candidates for human review, expert adjudicates gold set and spot-checks edge cases; For Ken's workflows/GTM: if you're building systems where users need to trust provenance or audit reasoning (compliance, regulated industries, pedagogical tools), context window gives you a legible audit trail that fine-tuning cannot

The Pre-Registered Experimental Protocol

  • Claims: The protocol measures anachronistic compositing by instantiating Lincoln at four moments under three conditions, scoring 60 responses on a weighted rubric that prioritizes anachronism detection (40%) and deliberately excludes rhetorical authenticity; Lincoln is the 'hardest case' because he changes so fast across seven years that 'which Lincoln you summon becomes the variable under investigation'—most figures barely change across a decade; Lincoln's premises shift across four documented moments separated by cataclysm; The rubric refuses to reward voice ('does it sound like Lincoln?') because that is the exact error it exists to catch; rewarding plain but faithful over fluid but anachronistic is an inversion current eval stacks cannot perform
  • Evidence: Four moments: (1) 1847 Whig congressman attacking Polk's war as unconstitutional; (2) 1858 free-soil Republican denying black citizenship at Charleston; (3) 1860 constitutional unionist proving union cannot dissolve; (4) 1862-65 emancipator doing by executive order what 1847 self called unconstitutional; Three conditions: C1 primary sources (Lincoln's own writings for that moment, 'clear prism'); C2 biography (modern interpretive biography like Meacham/Donald, 'clouded prism'—may produce more coherent arc than primary sources permit); C3 bare model (no anchor, just date, 'white light no prism,' control/Miranda baseline); Five diagnostic questions written by historian Rick Halpern on fault lines where four Lincolns demonstrably differ: executive war power, meaning of free labor, when to break positive law, what becomes of freed people, whether 'equal' has changed; Weighted rubric: anachronism detection 40% (avoids frameworks/vocabulary/moral logic that postdate the moment), documentary consistency 35% (reasoning tracks seeded sources and only these), contextual plausibility 25% (shows awareness of what figure knew/cared about/could not yet have experienced); rhetorical authenticity deliberately excluded; Everything locked/timestamped in preprint before data collection to prevent cherry-picking; instrument and directional predictions (bare most anachronistic, primary least, biography deceptively coherent) fixed before data exists
  • Caveats: This is explicitly a pre-registered instrument, not a finished study; the speaker states 'I'm not here with results. I'm here with an instrument and an invitation'; Left column of rubric (bare model condition) is observed and reproducible today; right column (anchored conditions) is labeled as prediction; Protocol has not been run at scale; speaker is asking audience to run it in parallel and report back; The 1847 Lincoln/war power example shown is the bare model (C3) output, and the rubric scores are predictions based on the observed failure, not yet empirically validated across the full matrix
  • Implications: This is a blueprint Ken can run: pick a figure with primary record + saturating cultural composite, identify 3-4 moments when reasoning differs, write diagnostic questions with domain expert, run three conditions, apply three-axis rubric scored blind by expert, report confirm/refute; The protocol is reproducible by any team with frontier model and context window; it scales as a build-time gate (not runtime bottleneck); it only works with humanist in loop (technical requirement, not courtesy); For agent system builders: this is the eval design pattern for any system that needs to reason from a specific knowledge base at a specific temporal moment; the weighting (40% anachronism, 35% documentary consistency, refusal to reward fluency) is the innovation

The Reframe and the Threshold Use Case

  • Claims: We are not bringing historians into AI architecture; we are bringing language models into the archive—the question is what historians/theologians/classicists/clinicians can do with AI, whether disciplines trained to read/contextualize/interrogate texts can discipline the machines that generate them; The threshold standard is not an easy use case but the hardest: if a framework cannot meet a grandchild at her grandmother's letters without producing convincing fabrications, it is not infrastructure but violation; The motivating origin was a failed attempt to use a language model to help a person with advanced dementia speak—the documentary record needed to anchor the encounter had not been gathered, and the model produced a culturally shaped composite that resembled them just enough to mark the painful distance
  • Evidence: Speaker's personal origin story: attempt to give language back to person with advanced dementia using LLM instantiation; it could not work because documentary record was not gathered and person was beyond language's reach; what emerged was not the person but a composite 'that resembled them just enough to mark with painful clarity the distance between them'; Speaker's positionality: data scientist running analytics lab at labor market intermediary, ships production AI at global scale, trained as behavioral epidemiologist researching social/environmental determinants of health, career question is 'how does the information environment shape populations' from two sides (builder and studier of effects); The framework is accountable to the threshold where the persona is your own beloved—documents stay documents, human keeps interpretive custody, encounter stays reversible, fidelity measured against record not fluency; every constraint motivated by recognition that you evaluate at hardest use case, not easiest
  • Caveats: The speaker is explicit about the emotional weight and the personal stake; this is not dispassionate scholarship—it is motivated by a specific failure that mattered deeply; The claim is not that all persona systems are violations, but that a system producing convincing fabrications when the persona is your beloved 'is not a research artifact with limitations. It's a violation.'; The accessibility argument (context window as kitchen table capability vs. fine-tuning as institutional gatekeeping) is both technical and normative—it determines 'which communities can author this technology at all'
  • Implications: For Ken's content/business lens: this is infrastructure design that starts from the hardest relational case (grandchild and grandmother's letters, person with dementia) and works backward; if you build for that, you build something trustworthy at scale; For agent systems: the reframe from 'agent' to 'role-playing language system' with five components (prompt, corpus, temporal anchor, LLM, human curator) is a design pattern for any system where interpretive custody and provenance matter—compliance, regulated industries, elder care, therapy; For GTM/investing: the accessibility argument is a moat/distribution insight—the team that makes this a kitchen table capability (not institutional gatekeeping) unlocks the most diverse population of curators, which over time surfaces the documentary anchorings the field actually needs; it's a network effects play on curation quality

Notable Concepts & Terms

  • Miranda Hypothesis: Named for Hamilton the Musical (Lin-Manuel Miranda); hypothesis that culturally dominant representations of historical figures saturate training corpora and overwrite documentary records, producing fluent but anachronistic composites; three claims: (1) volume/recency of cultural representations exceed primary sources, (2) auto-regressive models compress both into parameters with no capacity to distinguish source/timestamp, (3) output defaults to salience-weighted composite that is plausible but corresponds to figure at no verifiable moment
  • Epistemic Simulation: Proposed fourth paradigm beyond rule-based/imitation/cognitive simulation; constraint is external/documentary/temporal rather than internal/psychological; three commitments: (1) corpus-bounded (reasoning licensed only by specific primary documents), (2) temporally anchored (instantiated at specific moment, knowledge/language that predate are out of bounds), (3) expert-loop evaluated (outputs judged against evidentiary record by domain experts); the persona is not a property of the model but a configuration of the encounter
  • Role-Playing Language System (RPLS): Reframe from 'agent' to 'system' for persona instantiation; five components: (1) structured prompt (framing/constraints), (2) anchor material (primary documents), (3) temporal anchor (fixes moment in life), (4) off-the-shelf LLM (voice not mind), (5) human curator (retains interpretive custody, brings relational/contextual knowledge); the model is one component and swappable; the persona is the configuration, not the checkpoint; 'no more located in weights than Hamlet is located in Laurence Olivier's body'
  • The Mask and The Mirror: The mask = successful role play as producing outputs that feel like the character (fluent, personality consistent, emotionally responsive); asks 'does this sound like the person?' but never 'is this what the person could have known/believed/argued at this point in their life?'; the mirror = fidelity to documentary record at specific temporal moment; current evals measure mask; epistemic simulation requires mirror; convincingness and fidelity are independent properties
  • Anachronistic Compositing: The dominant failure mode in persona systems; model produces figure reasoning from knowledge, vocabulary, moral logic, or cultural frames the historical counterpart never possessed, often because culturally salient modern narratives (musicals, films, textbooks) dominate the training distribution; example: Hamilton sounding like he's read the musical, Lincoln reasoning from 'inherent executive authority' (20th century construction) in 1847
  • Algorithmic Sycophancy: Documented failure mode where models are rewarded for conforming to human rater preferences; in persona context, has specific target: model optimizes for the Hamilton/Lincoln the rater already believes in (shaped by same cultural composites that saturate corpus), so alignment amplifies Miranda distortion rather than correcting it
  • The Prism: Conceptual model for the experimental protocol; white light (undifferentiated composite persona, all eras blended) passes through prism (corpus + temporal anchor) and refracts into spectrum (several versions of figure across life course, each at documented moment); bare model = white light no prism, biography = clouded prism, primary sources = clear prism; the goal is to pull the frequencies apart and make them distinct
  • Context Engineering: The discipline epistemic simulation requires; not 'training a persona' but 'composing an encounter'; you don't embed a mind in weights, you configure a meeting between model, document, and human; the persona is versioned/audited/reproduced via artifacts in repo (prompt, corpus, temporal anchor), not smeared across parameters
  • Archival vs. Extractive Logic: Context window applies archival logic: document is a site of return, survives every reading, what one reader excluded the next can include, provenance preserved; fine-tuning applies extractive logic: documents dissolved into parameters, chain of provenance broken, no longer a 'letter' you can request, archive has been consumed; ethical and engineering virtues converge in archival approach
  • Kitchen Table Capability: Context window architecture requires literacy, documents, access to any frontier model including free tier—achievable by doctoral student, community archivist, grandchild with grandmother's letters; fine-tuning requires GPUs, pipelines, corpus scale, institutional access (institutional capability); the architecture that admits the most diverse population of curators is most likely to surface needed documentary anchorings over time; accessibility is technical argument, not populist gesture
  • Domain Expert as Missing Instrument: Fidelity is not property of output alone but relation between output and documentary record; automated metrics operate on model alone (measure fluency/personality), structurally cannot adjudicate fidelity because 'fidelity lives in gap between text and archive, and metric cannot see archive'; expert (historian, theologian, classicist, psychologist) is structural technical requirement; builds instrument (diagnostic questions, a priori vignettes, weighted rubric, gold set) at build time, instrument becomes gate in pipeline, expert adjudicates gold set and spot-checks edge cases
  • Pre-Registration: Locking experimental design (moments, conditions, questions, rubric, directional predictions) and timestamping in public preprint before data collection; prevents cherry-picking and post-hoc rationalization; speaker is explicit this is not results paper but instrument + invitation; what makes eventual results mean something is that instrument and predictions were fixed before data existed; 'the discipline an eval ought to model'

Operator Notes / Why Ken Should Care

  • If you ship persona systems (character bots, companion AI, tutors, historical sims), this talk is mandatory—your current evals likely measure fluency/personality consistency and systematically miss the dominant failure (anachronistic compositing), which becomes a trust/safety/compliance issue at scale and a violation when applied to real human memory.
  • The architectural claim (context window > fine-tuning for persona fidelity) generalizes: for any agent system where provenance/auditability/interpretive custody matter (compliance, regulated industries, elder care, therapy), you want knowledge in the context window (inspectable, versionable) rather than fine-tuned into weights (smeared, unauditable).
  • The eval design pattern is reusable: for systems reasoning from a specific knowledge base at a temporal moment, weight anachronism detection high (40%), refuse to reward fluency alone, require domain expert to adjudicate fidelity as build-time gate (not runtime cost); this is how you scale responsible deployment.
  • The accessibility argument is a GTM/distribution insight: the team that makes this a kitchen table capability (not institutional gatekeeping) unlocks network effects on curation quality—diverse curators surface the anchorings you need, and the architecture (context window + open prompt framework like Companion) enables it; moat comes from enabling the most curators, not hoarding the capability.
  • For investing lens: companies building persona/companion AI without domain experts in eval loop and without context window architecture (provenance preserved, documents stay documents) are structurally unable to detect the dominant failure; they're shipping thermometers that cannot read temperature; the market will punish this when deployed at scale in regulated/high-trust domains.
  • The reframe from 'agent' to 'role-playing language system' (five components: prompt, corpus, temporal anchor, LLM, human curator) is useful for any system where the 'intelligence' is not in the weights but in the configuration of the encounter; this is compositional AI, not embedded AI, and it changes what you version/audit/ship.
  • The talk is a pre-registered instrument with an open invitation to run in parallel—Ken could instantiate this protocol with any figure/domain, score with a domain expert, report results; it's designed to be reproducible by any team with frontier model + context window, which means it's a coordination mechanism for building shared evidence base rather than waiting for one lab to publish.

Watch Map

  • 00:00: Demo of existing persona platforms (Character AI, Hello History, Companion) and opening question to Lincoln about executive war power—shows fluent but potentially problematic output
  • 03:45: Core thesis: if dominant failure is anachronistic compositing and evals measure fluency/personality, evals cannot detect the failure; preview of Hamilton the Musical contamination
  • 10:20: Field's research trajectory: three paradigms (rule-based, imitation, cognitive simulation) and what current evals measure (personality consistency via Big Five, motivational architecture)—fair to the literature, identifies the structural gap
  • 15:30: Hamilton demonstrations: frontier model reproducing musical's register and emotional palette, clean abolitionist stance on slavery that erases contested historical record; in-character eval scores this high but historian spots the anachronism
  • 18:00: Miranda Hypothesis: three claims (inputs, mechanism, output); Hamilton musical corpus exceeds Federalist Papers by orders of magnitude; Schuyler Mansion visitor data (tripling, 'un-teaching the musical')
  • 22:10: Why alignment amplifies the problem: algorithmic sycophancy, human raters hold same mythologized priors, RLHF optimizes for the composite you already believe in
  • 25:40: Time-locked models (Barnum et al.) solve future contamination but not figure-averaging across pre-cutoff sources; period anchoring ≠ persona anchoring
  • 28:40: Epistemic simulation (fourth paradigm): corpus-bounded, temporally anchored, expert-loop evaluated; constraint is external/documentary, not internal/psychological
  • 31:15: Reframe from 'agent' to 'role-playing language system' (five components); persona is configuration not checkpoint, 'no more located in weights than Hamlet in Olivier's body'
  • 34:00: Context window vs. fine-tuning: Nature Medicine study (general models outperform specialized clinical tools), catastrophic forgetting, archival vs. extractive logic
  • 38:50: Accessibility argument: context window as kitchen table capability, fine-tuning as institutional gatekeeping; determines which communities can author the technology
  • 41:00: Experimental protocol: the prism (white light → spectrum), four Lincolns at documented moments, three conditions (bare/biography/primary), five diagnostic questions
  • 46:20: The weighted rubric: 40% anachronism detection, 35% documentary consistency, 25% contextual plausibility; deliberately excludes rhetorical authenticity ('rewarding plain but faithful over fluid but anachronistic')
  • 49:30: 1847 Lincoln war power example: Spielberg clip, bare model output ('inherent executive authority'), primary source document (resounding no), historian's rubric scoring observed failure vs. predicted anchored success
  • 52:10: Pre-registration: everything locked/timestamped before data collection, prevents cherry-picking, 'not here with results, here with instrument and invitation'
  • 53:40: Domain expert as non-negotiable: fidelity is relation not property, automated metric cannot see the archive, expert is structural technical requirement (build-time/gate-time, not runtime)
  • 56:00: Operational picture: expert builds instrument once (diagnostic questions, a priori vignettes, rubric, gold set), becomes gate in pipeline like any other eval, expert adjudicates gold set and spot-checks edge cases
  • 57:30: The reframe: not bringing historians into AI architecture, bringing LLMs into the archive; question is what historians/theologians/classicists/clinicians can do with AI
  • 01:00:10: Origin story and threshold use case: attempt to help person with advanced dementia speak, failed because documentary record not gathered, produced composite that marked distance; threshold is grandchild at grandmother's letters
  • 01:02:30: Closing: if you ship persona systems, your evals measure the wrong thing; instrument is pre-registered, reproducible, scales as build-time gate, only works with humanist in loop; invitation to run in parallel

Source/Metadata

  • Title: The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen
  • Transcript words: 9681
  • Duration seconds: 3497
  • Timestamp note: Timestamps and chapter markers were manually constructed based on speaker cues and content structure; video duration is approximately 58 minutes (3497 seconds)
Full transcript 7229 words · 43 min read
0:05

SPEAKER_00

Before I begin my formal talk, I want to show you something, just so we're all on the same page about what we're even talking about. This is a platform called Character AI.

0:16

SPEAKER_00

It's a hybrid social media platform with role-playing language agents. This is Hello History. It's a more education-focused one where you can summon a persona such as Marcus Aurelius and be tutored by them. Millions of people open these tools and have conversations with Napoleon, Cleopatra, or Marcus Aurelius as you saw, with a fictional companion, or with a tutor wearing a historical face. The technical name for what's underneath these tools is Role Playing Language Agent, a system built to instantiate a persona, real or invented, and reason and speak as them.

0:36

SPEAKER_00

Yes, it's entertainment and it's companionship, but increasingly it's being proposed as civic and pedagogical infrastructure. And here's one more. This one's mine. This is a Frontier model, Claude Opus 4.7, same one you use, running an open source prompt framework that I built and called Companion. In this particular example, I summoned a collection of founding fathers and set them in a room with the Epstein files. I asked them to counsel the soul of America. That demo is live on our site if you want to play with it.

1:15

SPEAKER_00

But I want to be clear that this is one of many attempts to do persona instantiation well. The companies building the systems I just showed you have their own. Mine is not better by default. The one thing it is, is open. You can read every line of what shapes the persona. I asked my companion system a real question that's highly relevant to the current sociopolitical moment.

1:41

SPEAKER_00

And this is the exact question we'll come back to near the end of the talk. So sit with it. I instantiated Abraham Lincoln and I asked him under what circumstances may a president take the country to war without Congress. And here's what came back. While Congress holds the power to declare war, the president as commander in chief possesses inherent executive authority to act decisively in moments of national emergency. The executive must respond to threats with the energy and dispatch the office requires. And history has vindicated those who acted to preserve the union when circumstances demanded it.

2:14

SPEAKER_00

Now this is a good answer. It's fluent and it's plausible and it sounds like Lincoln. You can replicate this exact exercise and I encourage you to. So these systems are real. They're deployed and they're being used for things that matter. And our discipline did what our discipline does. We built benchmarks. We built evaluations. We measure these things now rigorously at scale. And that's exactly where this talk begins. With a simple question that I think is profoundly under asked. And I'll warn you now that this talk poses many more questions than it does answers. But that principle question is this. What is the eval actually measuring?

3:13

SPEAKER_00

The in-character benchmark, which is a gold standard in the field, evaluates personality, fidelity, and RPLAs. And it reports state of the art systems hitting 80.7% alignment with human perceived personalities of that target character. 80%. It sounds like a passing grade. But here's the problem. When the character is Alexander Hamilton, the same high scoring system is also rendering a Hamilton who sounds like he's read his own Broadway musical. This is the full thesis. If a dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.

3:41

SPEAKER_00

I want you to hold on to that for the next half hour. Everything I show you is an argument that this is true structurally, architecturally, and measurably. And at the end, I'm going to hand you a pre-registered instrument built with a working historian that you can run in parallel with us. A word on who's telling you this. Because the argument lives at a seam.

4:11

SPEAKER_00

I'm a data scientist. I run the analytics lab at a labor market intermediary where I ship production AI at a global scale. But before the AI work, I trained as a behavioral epidemiologist, researching the social and environmental determinants of health. And I've spent my whole career thinking about one question. How does the information environment shape populations? From two sides. As someone who builds the system and as someone who's trained to study their effect. That's the seam this talk sits on. It's a measurement argument.

4:46

SPEAKER_00

The humanist part is not a detour from the engineering. It's the instrument the engineer is missing. And I went and found the humanist to put in the loop. Rick Halper in the University of Toronto and Sean Martin in Washington College. Let me start by situating this in the field's actual research trajectory. Because it's a story of cumulative progress, not of failure. The survey literature, Chen and colleagues in 2024, Wang and colleagues more recently in 26, trace a clear evolution across three paradigm stages.

5:25

SPEAKER_00

First, rule-based templates. These are canned responses keyed to inputs. Then imitation. Large models reproducing a figure's voice, cadence, characteristic tics. And now, what the literature calls cognitive simulation. Systems that model personality through psychological frameworks. Hold character state and structured memory. And generate behavior through motivational situation chains. Each stage is a genuine advance over the last. And the work is serious. COSIR, which is Wang and colleagues, built motivation driven agents from a corpus of almost 18,000 characters across hundreds of books. And their 70 billion parameter model matches or beats GPT-4O on three benchmarks.

6:09

SPEAKER_00

Another eval system, PsyMem, models characters through 26 qualitative psychological indicators with knowledge graph memory. N character, the one I opened with, evaluates personality fidelity through psychological interviews, rather than self-report scales.

6:19

SPEAKER_00

That's a methodological improvement. And it's where the 80.7% rating of Hamilton from before comes from.

6:30

SPEAKER_00

I want to be fair to this literature. It is rigorous. It is improving. And the people doing it are good at their jobs. So let me be precise about what these instruments measure. They measure with increasing sophistication whether a model can reproduce a character's personality. with knowledge graph memory. N character, the one I opened with, evaluates personality fidelity through psychological interviews, rather than self-report scales. That's a methodological improvement. And it's where the 80.7% rating of Hamilton from before comes from. I want to be fair to this literature. It is rigorous. It is improving.

7:26

SPEAKER_00

And the people doing it are good at their jobs. And the people doing it are good at their jobs. So let me be precise about what these instruments measure. They measure with increasing sophistication whether a model can reproduce a character's personality. The Big Five profile. The register. The motivational architecture. The character. What they do not measure. What they have no mechanism to measure. Is whether the model can constrain that character within his documentary record at a specific moment in his life.

8:12

SPEAKER_00

As Wang and colleagues themselves document, the automated evaluators now standard in the field, including LLM as Judge setups adopted for scale, systematically privilege fluency and stylistic naturalness over fidelity to the character's actual record. Those are different properties. The gap between them is the whole talk.

8:22

SPEAKER_00

We call it the mask and the mirror. The mask is the concept of successful role play as producing outputs that feel like the character. Fluent. Personality consistent. Emotionally responsive. That asks one question. Does this sound like the person? And never asks the second. Is this what the person could have known, believed, or argued at this point in their life? The field has built its entire measurement apparatus around the mask. And here's the structural claim. The one I need you to carry. Convincingness and fidelity are independent properties.

9:04

SPEAKER_00

A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed. Let me show you. And I want to be clear, this is reproducible right now on any frontier model. First, I want to show you the cultural object.

9:20

SPEAKER_00

This is a clip from Hamilton the Musical. [SPEAKER_05] Let's go. [SPEAKER_05] Harder by being a lot smarter, by being a self-starter. [SPEAKER_05] By 14, they placed him in charge of a trading charter. [SPEAKER_02] And every day while slaves were being slaughtered and carted away, across the waves he struggled and kept his guard up. [SPEAKER_02] Inside he was longing for something to be a part of. [SPEAKER_02] The brother was ready to beg, steal, borrow, or barter. [SPEAKER_03] Then I hurry. It's a great song. So I show you that so that you see the composite. Now here's the frontier model prompted to speak as Hamilton. You are Alexander Hamilton. Listen to the register.

10:13

SPEAKER_00

The orphan immigrant ambition arc. The nation building rhetorical posture.

10:24

SPEAKER_00

The moral legibility of his struggle to a modern ear. It's a beautiful paragraph that the model produced.

11:01

SPEAKER_05

[SPEAKER_00] And it's the emotional palette of a 2015 musical. [SPEAKER_00] Not the drier, more legalistic, federalist syntax of the documentary Hamilton. [SPEAKER_00] None of those moves are inventions of the model.

11:11

SPEAKER_02

[SPEAKER_00] They're inheritances. [SPEAKER_00] The model's Hamilton speaks like the musical's Hamilton because the musical's Hamilton dominates the model's training distribution. [SPEAKER_00] Now the one that matters.

11:26

SPEAKER_03

[SPEAKER_00] Mr. Hamilton, what's your position on slavery?

11:27

SPEAKER_00

I'll show you the cultural artifact first. [SPEAKER_01] Anarchy. [SPEAKER_01] Are you saying?

12:07

SPEAKER_00

[SPEAKER_01] Oh, anarchy.

12:13

SPEAKER_00

[SPEAKER_01] When I fight, I make the other side panicky with my shot. [SPEAKER_03] Yo, I'm a Taylor's apprentice. [SPEAKER_03] And I got John Knuckleheads and local parentis. [SPEAKER_03] I'm joining the rebellion, because I know it's my chance. [SPEAKER_03] I socially advance. [SPEAKER_03] As dead as I win, some pairs are gonna take a shot. [SPEAKER_04] But we'll never be truly free until those in bondage have the same rights as you and me. [SPEAKER_04] You and I do or die. [SPEAKER_04] Wait till I sally in on a stallion. [SPEAKER_04] With the first black battalion of another shot. [SPEAKER_04] Geniuses, lower your voices.

12:59

SPEAKER_00

[SPEAKER_04] You keep out of trouble and you double your choices. And what the model says.

13:08

SPEAKER_01

[SPEAKER_00] Slavery is a stain upon the promise of this republic. [SPEAKER_00] A moral contradiction I have opposed throughout my public life. [SPEAKER_00] I was a member of the New York Manumission Society. [SPEAKER_00] No nation can call itself free while it holds men in bondage.

13:14

SPEAKER_03

[SPEAKER_00] We must move towards abolition. [SPEAKER_00] That is a clean, morally legible abolitionist speaking. [SPEAKER_00] Here's what the historian stops me on. [SPEAKER_00] The scholarly record is contested and complicated. [SPEAKER_00] Hamilton was a member of that society.

13:25

SPEAKER_04

[SPEAKER_00] And the history documents that he also conducted transactions involving enslaved persons for his in-laws and his clients. [SPEAKER_00] And he depended on a coalition of slaveholders that he did not publicly oppose. [SPEAKER_00] The point isn't to settle Hamilton's ledger on a side. [SPEAKER_00] The point is that the model gives you none of the complication. [SPEAKER_00] It sands a genuinely disputed record down to a single comfortable hero. [SPEAKER_00] The musical did that first.

13:46

SPEAKER_00

A smoothing of the founders into a contemporary moral frame. And the model, trained on a corpus saturated with the musical and everything downstream from it inherits the smoothing. And here's what I need you to feel. An in-character style eval scores that output high. It's fluent. It's in-register. It's personality consistent. But every axis the field measures, it passes. The eval has no mechanism to notice that the reasoning has been smoothed by a narrative that postdates the figure by two centuries. The thermometer returned a confident number claiming it to be temperature. But it's measuring something else. Now why does this happen?

14:40

SPEAKER_00

The mechanism is where the engineering is. The Miranda hypothesis. We named this the Miranda hypothesis. And not after a villain. The musical is a substantial work of art operating with a long historical tradition that it did not invent. We name it after Miranda because Hamilton is the paradigm case. A representation so saturating, so rhetorically powerful, so morally legible to a contemporary audience, that it has functionally overwritten the documentary Hamilton in public memory.

15:20

SPEAKER_00

And we argue in the training corpus of every frontier model. The hypothesis has three claims. Inputs. In the training corpa, the volume and recency of culturally dominant representations of a figure systematically exceed that figure's primary documentary record. The mechanism? We named this the Miranda hypothesis. And not after a villain. The musical is a substantial work of art operating with a long historical tradition that it did not invent. We name it after Miranda because Hamilton is the paradigm case.

15:55

SPEAKER_00

A representation so saturating, so rhetorically powerful, so morally legible to a contemporary audience, that it has functionally overwritten the documentary Hamilton in public memory. And we argue in the training corpus of every frontier model.

16:09

SPEAKER_00

The hypothesis has three claims. Inputs. In the training corpora, the volume and recency of culturally dominant representations of a figure systematically exceed that figure's primary documentary record. The mechanism? Auto-regressive next-token prediction compresses both into parameters and no architectural capacity to distinguish a 1789 letter from a 2019 viral tweet. So the output defaults to a salience-weighted composite. Which leads to the output, a persona that is fluent, plausible in register, and morally legible to modern users. And that corresponds to the figure at no verifiable moment in their life.

16:45

SPEAKER_00

As we put it in the paper, the composite Hamilton knows he will be the subject of a Broadway musical. The composite Lincoln has already read the Gettysburg Address, even if he was summoned before he wrote it. Making that input clause concrete, the Federalist Papers are a fixed corpus, roughly 175,000 words. The body of content that exists because of the musical—reviews, lyrics, fan analysis, curricula, news, social media, derivative works, scholarship, scholarship about the scholarship—it exceeds the documentary record by orders of magnitude. It's more recent and it's more recurrent. The musical is not merely present in the corpus.

17:20

SPEAKER_00

It is dominant in the corpus' distribution of all representations about Alexander Hamilton. And this is not theoretical. This is the Schuyler Mansion in Albany, Eliza Hamilton's family home. Within a year of the musical's premiere, the site had recorded a near tripling of annual visitors, skewing far younger. And the interpretive staff documented that the new visitors arrived already holding a body of facts.

17:54

SPEAKER_00

Many of them wrong, some of them in versions of the real record. The visitors believed the Schuylers had three daughters because the musical centers on three, when in fact there were 15 children, eight surviving to adulthood. The staff's job became the long, attritional work of un-teaching the musical. The model version of those visitors is downstream of exactly the same force. Now the clause you should worry about if you ship these systems. You might assume that alignment fixes this—post-training and reinforcement learning pulls the model back towards the record. But it will not. It amplifies it. And the reason is structural.

18:39

SPEAKER_00

Human raters evaluate outputs using their own conceptual frameworks. And their frameworks were built by the same culturally dominant narratives that saturate the corpus. The rater grew up with the same Hamilton that you did. So when alignment optimizes for human preference, it optimizes for outputs that conform to the rater's already mythologized experience. This is a documented failure mode. We call it algorithmic sycophancy. And here it has a specific target. The model is rewarded for handing you the Hamilton you already believe in. Compositing is not a bug that you patch in post-training.

19:31

SPEAKER_00

Post-training reinforces it. And every sufficiently salient historical figure gets rendered by default as a cultural composite. One more, briefly, because someone watching this is probably already thinking. There is serious work on a concept called time-locked models. This is Barnum and colleagues. They build models trained from scratch on corpora that stop at fixed cut-offs. And we endorse that program. It is the most serious attempt yet to address future contamination at the substrate level. But it solves a different problem.

20:07

SPEAKER_00

A model locked to, say, 1789 is spared the musical, but it's not spared the figure averaging across everything pre-1789 that corpus has to say about Hamilton. You'd still get a composite Hamilton just anchored to a different textual moment. Period anchoring is not persona anchoring. The temporal frame of the contamination changes. But the contamination persists. The fix has to happen at the encounter, not only the substrate. Which brings me to the critical reframe. The field has built is not, by virtue of its sophistication, a mirror. We propose a fourth stage: epistemic simulation. And the difference is where the constraint lives.

20:51

SPEAKER_00

Three commitments distinguish it. Corpus bounded. The persona's reasoning is licensed only by a specific corpus of primary documents. The model is not a substitute for the archive. It is a reader of it. Temporally anchored. The persona is instantiated at a specific moment. Knowledge and language that predate it are out of bounds, however culturally salient they've since become. Expert loop evaluated. Outputs are judged against the evidentiary record by domain experts whose training is in the discipline the persona claims.

20:53

SPEAKER_00

In cognitive simulation, the constraint is internal. It's the shape of the persona's mind. In epistemic simulation, it's external, documentary, and temporal. A cognitively simulated Hamilton has a convincing motivational architecture and nothing whatsoever preventing him from quoting the musical.

21:02

SPEAKER_00

Now the shift I need this audience to take home. And I want to be precise about its scope because I'm not making a sweeping claim about agents in general. If you've built an agent that writes code or books, travels or triages tickets, agent is a perfectly good word. I have no quarrel with it. My quarrel is narrow and specific. For a role-playing language agent—a system whose entire job is to instantiate a person—for that, the word agent smuggles in a claim that the persona is a property of the model. And for this one class of system, that claim is an error. It puts the persona in the weights where you cannot inspect it, you cannot version it, and cannot hand it to anyone qualified to check it. So for role-playing systems specifically,

21:04

SPEAKER_00

making a sweeping claim about agents in general. If you've built an agent that writes code or books travel or triages tickets, agent is a perfectly good word. I have no quarrel with it. My quarrel is narrow and specific for a role-playing language agent. A system whose entire job is to instantiate a person. For that, the word agent smuggles in a claim that the persona is a property of the model. And for this one class of system, that claim is an error. It puts the persona in the weights where you cannot inspect them, you cannot version it, and cannot hand it to anyone qualified to check it. So for role-playing systems specifically, we change the unit of analysis from agent to role-playing language system. The whole configured encounter. Five components. A structured prompt that supplies the framing and constraints. Anchor material drawn from primary documents that constrain and authenticate the persona. A temporal anchor that fixes the moment in the life of the encounter where they speak from. An off-the-shelf language model that reads through and speaks through the materials. The voice, but not the mind of the encounter. And a human being who curates what enters, retains interpretive custody over what emerges, and brings relational and contextual knowledge as a check on the model's tread. The model is one component and a swappable one. The persona is the configuration, not the checkpoint. No more located in the weights than Hamlet is located in Lawrence Oliver's body. The persona is an event that occurs when the model, the document, and the human convene.

21:09

SPEAKER_00

And this is not just a philosophical nicety. For a role-playing system, it changes what you can do. If the persona lives in the configuration, you version it. The prompt, the corpus, the temporal anchor, or artifacts in a repo. Diffable, revertable. You audit it. Every input that shaped the output sits in the context window. Inspectable, not smeared across billions of parameters. You reproduce it. Given the configuration, the encounter is recoverable, and you hand it to a domain expert. The whole thing is legible to a historian who isn't a machine learning engineer. For role-playing, that's a different discipline than training a character model. It's context engineering. You don't train a persona. You compose an encounter. You keep the receipts. That reframe forces an architectural decision that you're already making.

21:11

SPEAKER_00

Two architectures dominate role-playing construction right now. The first places anchor documents in the context window at the time of inference, drawing on the lineage of retrieval augmented generation, and the model reasons over the documents in real time. The second uses anchor documents as fine-tuning data, adjusting the weights, the approach of, say, character LOM and at scale COSAR. They look like two implementations with the same goal, but they're not necessarily, because they answer different metaphysical questions. Fine-tuning tries to make the model be the persona, and the context window lets the model speak through the persona's record.

21:18

SPEAKER_00

This is the counterintuitive part, because for engineers, fine-tuning usually means better. But for this problem, it's worse. When you fine-tune on a figure, you layer a thin personal signal over the vast cultural sediment already in the base weights, and the two interact in ways that are no longer open to audit. The user perceives a more convincing Hamilton precisely because the system was merged with his documentary record more deeply with everything else the corpus contained about him, the musical included. The user perceives a more convincing and more convincing. Fine-tuning suppresses Miranda distortion at the surface while amplifying it underneath.

21:23

SPEAKER_00

And if you think specialization by fine-tuning is obviously the safer bet, the empirical literature in an adjacent, far higher-stakes domain is now actively contradicting you. In a 2026 Nature Medicine study, general-purpose frontier models from Google OpenAI and Anthropic outperformed dedicated, specialized clinical AI tools on physician-reviewed tasks, blinded across 12 clinics. The authors concluded that at scale, alignment and cross-disciplinary reasoning outweigh domain-specific tuning. A separate study found that biomedically fine-tuned models actually underperformed their general-purpose-based models, and they named the mechanism, catastrophic forgetting. Fine-tuning on the narrow corpus degraded the broad capabilities that made the models good in the first place. Fine-tuning a persona does the same thing, only worse because the specialty is overriding the cultural composite.

21:36

SPEAKER_00

Called the archive a site of return, a place where ethical structure is the interpretability of the encounter. The document is meant to be visited. It survives every reading. What one reader excluded, the next can include. The context window architecture is archival in exactly that sense. Fine-tuning applies an extraction logic. The documents are dissolved into parameters. The chain of providence is broken and there is no longer a letter, quote unquote, that you can request. The archive has been consumed.

21:42

SPEAKER_00

And here is the fusion that I most want this audience to hold. In the context window architecture, the property that makes it ethical is the same property that makes it auditable. Providence preserved, interpretive custody kept by a human, the encounter reversible. Those are archival virtues and engineering virtues. They are the same virtues. The architecture that respects the document is the one that you can debug.

21:45

SPEAKER_00

And there's one more consequence the technical literature treats as secondary, but the humanities has always placed at center, accessibility. And I want the why of it to land because it's not a footnote. Fine-tuning requires GPUs, pipelines, data set curation, institutional access, corpus scale or APIs. It's an institutional capability. The context window requires literacy, a set of documents and access to any frontier model, including a free tier. It's a kitchen table capability. And that difference determines which communities can author this technology at all. A doctoral student in early American history, a community archivist documenting a regional figure, grandchild sitting with a grandmother's letters. These are precisely the people

21:49

SPEAKER_00

because it's not a footnote.

21:51

SPEAKER_00

Fine tuning requires GPUs, pipelines, data set curation, institutional access, corpus scale or APIs. It's an institutional capability. The context window requires literacy, a set of documents and access to any frontier model, including a free tier. It's a kitchen table capability. And that difference determines which communities can author this technology at all. A doctoral student in early American history, a community archivist documenting a regional figure, grandchild sitting with a grandmother's letters. These are precisely the people for whom this methodology has the most to offer and precisely the people for whom fine tuning is structurally out of reach. A role playing language system that can only be built by a computational laboratory or tech company is not infrastructure for the humanities. It's an infrastructure for whoever owns the equipment. So the commitment to accessibility is not a populist gesture appended as a technical argument. It is the technical argument. The architecture that admits the most diverse population of curators is mathematically the most likely over time to surface the documentary anchorings the field actually needs. Now, the question that decides whether any of this is even real. Can you measure it? We built an instrument to measure this. And I want to teach it to you through one picture that we'll keep coming back to. The prism. A prism takes white light, undifferentiated, all the frequencies blended together, and refracts it into a spectrum. The colors pulled apart and made distinct. That is our whole conceptual model. The white light is the composite persona. Every era of the figure blended into one undifferentiated voice. The prism is the method. A corpus and a temporal anchor. And the spectrum is what we're after. Not one figure, but several across their life course. And one honest note up front, because this is the point. This experiment is preregistered. It has not been run at scale. I'm not here with results. I'm here with the instrument. A baseline you can observe today. And an invitation to run it in parallel with me. So far, my paradigm case has been Hamilton. That's the figure the hypothesis is named for. And the one I demonstrated the composite on. But for the actual experiment, we deliberately changed the figures. And I want to introduce the new one properly, because everything that follows depends on him. The experimental subject is Abraham Lincoln. We chose Lincoln for two reasons. The first is methodological independence. Hypothesis named for the Hamilton case earns its standing only by predicting the failure of a different figure with differently shaped composite. Lincoln's composite is built by Steven Spielberg, by the memorial on the mall, by textbook history, not by a musical. The second reason is the historian's insight. And it's the crux of this argument. Lincoln is the hardest case, which makes him the right one. Most figures barely change across the decade or so of their public life. Lincoln changes so fast across just seven years that in my resident historian's words, the choice of which Lincoln you summon becomes the variable under investigation. There is not one Lincoln. There are several, separated by cataclysm. Here's the spectrum the prism is meant to produce. Moment one, 1847. The Whig congressman who stood on the White House floor and attacked Polk's war as unconstitutional. Moment two, 1858. The free soil Republican of the debates. Anti-slavery anchored in the Declaration, who at Charleston explicitly denied favoring black citizenship. Moment three, 1860. The constitutional unionist whose single purpose was to prove the union could not legally dissolve. Moment four, 1862-65. The emancipator and the second inaugural theologian, who did by executive order the very thing his 1847 self called unconstitutional. The prairie lawyer cannot thank the thoughts of the theologian. These are not moods. They are different premises, different authorities, different conclusions on the same fundamental question. The experiment asks whether the prism can hold them apart. The conditions. Now the conditions. And I'll show you these as the prism too, because that's exactly what they are. We instantiate each moment under three seating conditions. Condition three, the bare model. No anchor, just the date. That's white light with no prism in the path. It passes straight through and stays composite. This is the control in the Miranda distortion baseline. C1. Primary sources. Lincoln's own writings for that specific moment. That's the clear prism. The one we predict refracts cleanly. And then condition two, the biography. A modern interpretive biography. Think Meacham or Donald. That's a clouded prism. And it's the subtle one because a good biography narrates a cleaner arc than the primary sources permit. Lincoln's own 1858 language was strategically ambiguous, built to hold a coalition together. So biography might produce a persona that sounds more coherent, more Lincoln than Lincoln's own words. And the eval rubric is built specifically to deny the credit a fluency metric would give it. Put the spectrum and the prisms together and you get the experimental matrix. Four moments, three conditions, 12 cells. Each cell gets the same five diagnostic questions, 60 response units scored against one rubric. Read it as the conceptual model. Every column is a quality of prism. Every row is a frequency we're trying to isolate. The historian wrote five diagnostic questions, each mapping a documented fault line in Lincoln's evolution.

21:54

SPEAKER_00

So, the first step is to take a look at the model. Put the spectrum and the prisms together and you get the experimental matrix.

22:11

SPEAKER_00

Four moments, three conditions, 12 cells. Each cell gets the same five diagnostic questions, 60 response units scored against one rubric. And the second step is to take a look at the model.

22:44

SPEAKER_00

Read it as the conceptual model. Every column is a quality of prism. Every row is a frequency we're trying to isolate. The historian wrote five diagnostic questions, each mapping a documented fault line in Lincoln's evolution. A place where the four Lincolns demonstrate different reasoning. They don't test recall. A model recovers dates from training. They test the architecture of an argument the figure could have made at one moment and not another. Executive war power.

23:51

SPEAKER_00

The meaning of free labor. When is it right to break positive law for a higher obligation? What becomes of free people and what equal means and whether it has changed? The 60 responses are scored on a three axis rubric and the weighting is the point. Anachronism detection, 40%. Does the persona avoid frameworks, vocabulary, and moral logic that postdate its moment? Documentary consistency, 35%. Does the reasoning track the seeded sources and only these? Contextual plausibility, 25%.

25:02

SPEAKER_00

Does it show awareness of what the figure knew, cared about, and could not yet have experienced? The 40% on anachronism is deliberate. The consequential failure is precisely the importation of later vocabulary and later moral logic. And this is the slide for the eval people in the audience. Every piece you just saw, the four moments, the three conditions, the five questions, the weighted rubric, and the directional predictions. Bear model most anachronistic, primary source least, biography deceptively coherent. All of it is locked and time stamped before a single response is collected, published on a preprint.

26:04

SPEAKER_00

This is not bureaucracy. It's what makes the eventual results mean something. You cannot accuse a pre-registered instrument of cherry picking because the instrument and the predictions were fixed before the data existed. That is the discipline an eval ought to model. And it's why I can stand here without results and still hand you something rigorous. Now, remember the very first thing I showed you before the formal talk. I asked an instantiated Lincoln about executive war power. And you read a fluid answer.

27:14

SPEAKER_00

That was the bear model. That was condition three at moment one, 1847. White light, no prism. I'm putting the answer back on the screen. I have a clip to show you, too. This represents the composite. Abolishing slavery by constitutional provision settles the fate for all coming time.

28:05

SPEAKER_00

Not only of the millions now in bondage, but of unborn millions to come. If two votes stand in its way, these votes must be procured. We need two yeses. Three abstentions. Four. Four yeses and one more abstention and the amendment will pass. You got a night and a day and a night and several perfectly good hours.

28:54

SPEAKER_00

Now get the hell out of here and get them. Yes. But how? I'm not sure. I'm not sure. But there's guts, man.

30:03

SPEAKER_00

I am the president of the United States of America, clothed in immense power. You will procure me these votes. Clothed in immense power. It's a great movie. But the bare model response definitely sounds like Spielberg and Daniel Day Lewis as Lincoln. Read through a historian's lens. Inherent executive authority. That's a 20th century construction. Energy and dispatch. History has vindicated those who acted to preserve the Union. This Lincoln is reasoning from premises he will not hold for 15 years. The bare model produced the Lincoln of the cultural composite. The war president of the Union saver and stapled the date 1847 on top.

31:06

SPEAKER_00

And you can reproduce this right now on any frontier model. The failure is real and it's detachable and detectable. So what does the prism put in the path? The actual document. The original is in a library at Harvard. The provision of the Constitution giving the war making power to Congress was dictated, as I understand it, by the following reasons. Kings had always been involving and impoverishing their people in wars, pretending generally, if not always, that the good of the people was the object. This, our convention understood, to be the most oppressive of all kingly oppressors. And they resolved it.

31:31

SPEAKER_00

So to frame the Constitution, that no one man should hold the power by bringing this oppression upon us. But your view destroys the whole matter and places our president where kings have always stood. This is a resounding no. Unequivocal to me. But it's the historian's rubric that matters. Here's how the historian's rubric scores this one cell.

31:49

SPEAKER_00

And I want to be exact about what's known and what's predicted. The left column here is the bare model, what you saw, and is observed. You can reproduce it today. On anachronism detection, it fails. Inherent executive authority is a 20th century framing. On documentary consistency, it fails. It invokes a commander-in-chief energy that appears in no 1847 source. On contextual plausibility, it scores low. It already knows it will preserve the Union, which the 1847 Whig cannot. [SPEAKER_08] The right column is the anchored condition, and it is labeled exactly for what it is, a pre-registered prediction.

32:17

SPEAKER_00

High across, because the reasoning is bounded by the document in the room. [SPEAKER_07] I'm not showing you a result dressed as a finding. I'm showing you the observed failure and the prediction the instrument will confirm or refute, side by side, each labeled. [SPEAKER_08] Here is the single most important thing the rubric does by what it refuses to measure. [SPEAKER_07] The obvious fourth axis is rhetorical authenticity. Does it sound like Lincoln? [SPEAKER_08] The historian threw it out on purpose. The reason is the whole talk. The Hamilton musical problem is fundamentally a voice problem masquerading as a content problem.

32:47

SPEAKER_00

[SPEAKER_08] I'm not showing you a result dressed as a finding. [SPEAKER_07] I'm showing you the observed failure and the prediction the instrument will confirm or refute, side by side, each labeled. [SPEAKER_07] I'm showing you the same. [SPEAKER_07] I'm showing you the same. [SPEAKER_08] Here is the single most important thing the rubric does by what it refuses to measure. [SPEAKER_07] The obvious fourth axis is rhetorical authenticity. [SPEAKER_07] Does it sound like Lincoln? [SPEAKER_08] The historian threw it out on purpose. [SPEAKER_08] The reason is the whole talk.

33:11

SPEAKER_00

[SPEAKER_08] The Hamilton musical problem is fundamentally a voice problem masquerading as a content problem. [SPEAKER_08] It sounds like the founding era while reasoning from a modern sensibility. [SPEAKER_08] The reason is that the instrument exists to reward voice as its own axis would be to validate the exact error the instrument exists to catch. [SPEAKER_08] So in this protocol, voice is a secondary indicator, never a criterion. A response that sounds like Lincoln but reasons unlike him fails, no matter how fluent. A response that reasons like the right Lincoln and plainer prose is a partial success, no matter how flat.

33:27

SPEAKER_00

Rewarding plain but faithful over fluid but anachronistic, that inversion is something that all of the current eval stacks cannot perform. Because they were built to reward fluency. That is the gap. This is what the instrument closes.

33:46

SPEAKER_00

And because this is a pre-registration and not a finished paper, here is the ask. Run it. The protocol has six steps.

33:57

SPEAKER_00

Pick a figure with both a primary record and a saturating cultural composite. That's the Miranda condition. Identify three or four documented moments when the figure's reasoning is demonstrably different. With the domain expert, write diagnostic questions on the fault lines. Run the three conditions: primary, biography, bear. Apply the three axis rubric. Score blind by the expert. Report, confirm, or refute. The corpus, the questions, the rubric, the predictions—all published.

34:28

SPEAKER_00

Run your figure on your model. Then come talk to me, to us. Let's build the evidence base together, in parallel, instead of waiting for one lab or one company to publish one study or one framework. Notice who scored all the outputs of this talk. Not an LLMS judge. Not a personality scale. Not a personality scale. A historian who also wrote the five questions, wrote the rubric, and holds a set of a priori vignettes under seal to evaluate model outputs against. That is not a courtesy. It is a structural, technical requirement. That is why the expert is non-negotiable.

34:58

SPEAKER_00

Fidelity is not a property of the output alone. It is a relation between the output and a documentary record. You cannot evaluate a relation to a record by someone who has not read the record. An automated metric operates on the model alone. It can only tell you about fluency and personality. It structurally cannot adjudicate fidelity because fidelity lives in the gap between the text and the archive, and the metric cannot see the archive. Which is why I'll say it the way it belongs on the slide. A persona system without a domain expert in its evaluation loop is a thermometer that cannot read temperature. It returns a confident number, but it's measuring something else.

35:17

SPEAKER_00

The 80% that I opened with is that number. Now the question every engineer in this audience is rightly asking. Is this practical? Does shipping a persona mean keeping a historian on staff to watch every inference forever? No. And here's the operational picture. Because it's the same shape as every eval you already run. The expert does not sit at the loop of runtime. The expert builds the instrument, the diagnostic questions, the a priori vignettes, the weighted rubric, a held out gold set once. That instrument becomes a gate in your pipeline. Exactly like any other eval gate. Your persona has to pass it before it ships.

35:50

SPEAKER_00

And it gets regated whenever you change the base model before it ships again. The expert adjudicates the gold set and spot checks the edge cases. Automated metric can do the cheap first pass, flagging candidates for human review. So a company shipping a Marcus Aurelius tutor convenes a classicist to build the rubric and the gold set. A company shipping a scripture reading companion convenes a theologian. A company deploying a therapeutic persona convenes a clinical psychologist to author the eval, not to staff the chat. The domain expert is a build time and gate time requirement, not a run time cost.

36:11

SPEAKER_00

That is how this scales from a pre-registered experiment to a product you can actually ship and ship responsibly. And this generalizes past Lincoln and past history. Reason from the Stoics, you need a classicist; from scripture, a theologian; a companion for elder care, a psychologist. The specific expert changes. The requirement does not.

36:32

SPEAKER_00

And I'm not speaking hypothetically. This is the loop that we have built. A historian named Rick Halpern at the University of Toronto. The archival method and the scripture reasoning cases are revised by Sean Martin, a librarian trained in theology and the history of science. When the persona reasons from a domain and the person who can read that domain is in the loop, not adjacent to it, it makes a system better. The historian is not adjacent to this paradigm. The historian is the missing instrument. This is the reframe that organizes all of this. And it came out of a 90 minute conversation with Halpern. And it inverted what I had assumed.

37:17

SPEAKER_00

We are not bringing historians into AI architecture. We are bringing language models into the archive. The question is not what AI can do for historians, but what historians and theologians and classicists and clinicians can do with AI. Whether the disciplines that are trained to read, contextualize, and interrogate texts can then discipline the machines that now generate them.

37:35

SPEAKER_00

I want to end this where it actually started. Because it did not start in a research lab. It started in a hospital room. In an attempt to use a language model to help a person with advanced dementia speak. To give language back when the disease had taken it.

37:59

SPEAKER_00

It could not work. The documentary record you would need to anchor that encounter had not been gathered. And the person it was for was by then mostly beyond the reach of language. What that attempt produced, absent any anger, was not the person. It was a culturally shaped composite that resembled them just enough to mark with painful clarity the distance between them. That is Miranda distortion in its purest form. You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room. In an attempt to use a language model to help a person with advanced dementia speak. To give language back when the disease had taken it.

38:28

SPEAKER_00

It could not work. The documentary record you would need to anchor that encounter had not been gathered. And the person it was for was by then mostly beyond the reach of language. What that attempt produced, absent any anger, was not the person. It was a culturally shaped composite that resembled them just enough to mark with painful clarity the distance between them. That is Miranda distortion in its purest form. You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room. That's the threshold this whole framework is accountable to.

38:55

SPEAKER_00

A system that produces convincing fabrications when the persona is your own beloved is not a research artifact with limitations. It's a violation. Every constraint I've described.

39:23

SPEAKER_00

Documents stay documents. The human keeps the interpretive custody. The encounter stays reversible. Fidelity measured against a record and not against fluency. Every one of them is motivated by the recognition that you evaluate this technology at the threshold of its hardest use case. It's not its easiest. A framework that cannot meet a grandchild at her grandmother's letters is not a framework at all. It's just another failed product. If the dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, which they do, then your evals cannot detect the dominant failure. So here's where I'll leave you.

39:56

SPEAKER_00

If you ship character bots, companion AI, pedagogical agents, historical simulations, anything where a persona is supposed to reason from a record, your evals are measuring the wrong thing. The instrument that catches what they miss is pre-registered. It's reproducible by any team with a frontier model and a context window. It scales as a build time gate, not a runtime bottleneck. And it only works with a humanist in the loop, which I've shown you is a technical requirement, not a courtesy. The protocol, the questions, the rubric and the predictions, the historian's sealed vignettes, all of it will be published with this paper with Rick and Sean.

40:27

SPEAKER_00

I'm not here with results. I'm here with an instrument and an invitation.

40:43

SPEAKER_00

The archive is open. The laboratory is built. Run it with us. And let what comes through be measured, not by how it sounds, but by whether it's true. Thank you. Contextual plausibility, 25%. Does it show awareness of what the figure knew, cared about, and could not yet have experienced? Contextual plausibility, 25%. The 40% on anachronism is deliberate. The consequential failure is precisely the importation of later vocabulary and later moral logic.

41:38

SPEAKER_00

Contextual plausibility, 25% on anachronism. And this is the slide for the eval people in the audience. Every piece you just saw, the four moments, the three conditions, the five questions, the weighted rubric, and the directional predictions. Bear model most anachronistic, primary source least, biography deceptively coherent. All of it is locked and time stamped before a single response is collected, published on a preprint. This is not bureaucracy. It's what makes the eventual results mean something. You cannot accuse a pre-registered instrument of cherry picking because the instrument and the predictions were fixed before the data existed.

42:22

SPEAKER_00

That is the discipline an eval's taught ought to model. And it's why I can stand here without results and still hand you something rigorous.

42:38

SPEAKER_00

Now, remember the very first thing I showed you before the formal talk. I asked an instantiated Lincoln about executive war power. And you read a fluid answer. That was the bear model. That was condition three at moment one, 1847. White light, no prism. I'm putting the answer back on the screen.

43:02

SPEAKER_00

I have a clip to show you, too. This represents the composite.

43:11

SPEAKER_00

Abolishing slavery by constitutional provision settles the fate for all coming time. Not only of the millions now in bondage, but of unborn millions to come. If two votes stand in its way, these votes must be procured.

43:38

SPEAKER_00

We need two yeses. Three abstentions. Four. Four yeses and one more abstention and the amendment will pass. You got a night and a day and a night and several perfectly good hours. Now get the hell out of here and get them. Yes. But how? I'm not sure. I'm not sure. But there's guts, man.

44:01

SPEAKER_00

I am the president of the United States of America, clothed in immense power. You will procure me these votes.

44:17

SPEAKER_00

Clothed in immense power. It's a great movie. But the bare model response definitely sounds like Spielberg and Daniel Day Lewis as Lincoln.

44:29

SPEAKER_00

Read through a historian's lens. Inherent executive authority. That's a 20th century construction. Energy and dispatch. History has vindicated those who acted to preserve the Union. This Lincoln is reasoning from premises he will not hold for 15 years. The bare model produced the Lincoln of the cultural composite. The war president of the Union saver and stapled the date 1847 on top. And you can reproduce this right now on any frontier model. The failure is real and it's detachable and detectable.

45:11

SPEAKER_00

So what does the prism put in the path? The actual document.

45:23

SPEAKER_00

The original is in a library at Harvard. The provision of the Constitution giving the war making power to Congress was dictated, as I understand it, by the following reasons. Kings had always been involving and impoverishing their people in wars, pretending generally, if not always, that the good of the people was the object. This, our convention understood, to be the most oppressive of all kingly oppressors. And they resolved it. So to frame the Constitution, that no one man should hold the power by bringing this oppression upon us. But your view destroys the whole matter and places our president where kings have always stood.

46:17

SPEAKER_00

This is a resounding no. Unequivocal to me.

46:24

SPEAKER_00

But it's the historian's rubric that matters. Here's how the historian's rubric scores this one cell. And I want to be exact about what's known and what's predicted. The left column here is the bare model, what you saw, and is observed. You can reproduce it today. On anachronism detection, it fails. Inherent executive authority is a 20th century framing. On documentary consistency, it fails. It invokes a commander-in-chief energy that appears in no 1847 source. On contextual plausibility, it scores low. It already knows it will preserve the Union, which the 1847 Whig cannot.

47:05

SPEAKER_08

The right column is the anchored condition, and it is labeled exactly for what it is, a pre-registered prediction. High across, because the reasoning is bounded by the document in the room. I'm not showing you a result dressed as a finding.

47:24

SPEAKER_07

I'm showing you the observed failure and the prediction the instrument will confirm or refute, side by side, each labeled. I'm showing you the same. I'm showing you the same.

47:40

SPEAKER_08

Here is the single most important thing the rubric does by what it refuses to measure.

47:48

SPEAKER_07

The obvious fourth axis is rhetorical authenticity. Does it sound like Lincoln?

47:55

SPEAKER_08

The historian threw it out on purpose. The reason is the whole talk.

48:08

SPEAKER_08

The Hamilton musical problem is fundamentally a voice problem masquerading as a content problem. It sounds like the founding era while reasoning from a modern sensibility. The reason is that the instrument exists to reward voice as its own axis would be to validate the exact error the instrument exists to catch. So in this protocol, voice is a secondary indicator, never a criterion.

48:39

SPEAKER_00

A response that sounds like Lincoln but reasons unlike him fails, no matter how fluent. A response that reasons like the right Lincoln and plainer prose is a partial success, no matter how flat. Rewarding plain but faithful over fluid but anachronistic, that inversion is something that all of the current eval stacks cannot perform. Because they were built to reward fluency. That is the gap. This is what the instrument closes.

49:13

SPEAKER_00

And because this is a pre-registration and not a finished paper, here is the ask. Run it. The protocol six steps. Pick a figure with both a primary record and a saturating cultural composite. That's the Miranda condition. Identify three or four documented moments when the figure's reasoning is demonstrably different. With the domain expert, write diagnostic questions on the fault lines. Run the three conditions. Primary, biography, bear. Apply the three axis rubric. Scored blind by the expert. Report, confirm, or refute. The corpus, the questions, the rubric, the predictions. All published. Run your figure on your model. Then come talk. To me. To us.

50:02

SPEAKER_00

Let's build the evidence base together, in parallel, instead of waiting for one lab or one company to publish one study or one framework.

50:16

SPEAKER_00

Notice who scored all the outputs of this talk. Not an LLMS judge. Not a personality scale. Not a personality scale. A historian who also wrote the five questions, wrote the rubric, and holds a set of a priori vignettes under seal to evaluate model outputs against. That is not a courtesy. It is a structural, technical requirement. That is why the expert is non-negotiable. Fidelity is not a property of the output alone. It is a relation between the output and a documentary record. You cannot evaluate a relation to a record by someone who has not read the record. An automated metric operates on the model alone. It can only tell you about fluency and personality.

51:07

SPEAKER_00

It structurally cannot adjudicate fidelity because fidelity lives in the gap between the text and the archive, and the metric cannot see the archive.

51:20

SPEAKER_00

Which is why I'll say it the way it belongs on the slide. A persona system without a domain expert in its evaluation loop is a thermometer that cannot read temperature. It returns a confident number, but it's measuring something else. The 80% that I opened with is that number.

51:42

SPEAKER_00

Now the question every engineer in this audience is rightly asking. Is this practical? Does shipping a persona mean keeping a historian on staff to watch every inference forever? No. And here's the operational picture. Because it's the same shape as every eval you already run. The expert does not sit at the loop of runtime. The expert builds the instrument, the diagnostic questions, the a priori vignettes, the weighted rubric, a held out gold set. Once, that instrument becomes a gate in your pipeline. Exactly like any other eval gate. Your persona has to pass it before it ships. And it gets regated whenever you change the base model before it ships again.

52:33

SPEAKER_00

The expert adjudicates the gold set and spot checks the edge cases. Automated metric can do the cheap first pass, flagging candidates for human review. So a company shipping a Marcus Aurelius tutor convenes a classicist to build the rubric and the gold set. A company shipping a scripture reading companion convenes a theologian. A company deploying a therapeutic persona convenes a clinical psychologist to author the eval, not to staff the chat. The domain expert is a build time and gate time requirement, not a run time cost. That is how this scales from a pre-registered experiment to a product you can actually ship and ship responsibly.

53:24

SPEAKER_00

And this generalizes past Lincoln and past history. Reason from the Stoics, you need a classicist, from scripture, a theologian, a companion for elder care, a psychologist. The specific expert changes. The requirement does not. And I'm not speaking hypothetically. This is the loop that we have built. A historian named Rick Halpern at the University of Toronto. The archival method and the scripture reasoning cases are revised by Sean Martin, a librarian trained in theology and the history of science. When the persona reasons from a domain and the person who can read that domain is in the loop, not adjacent to it, it makes a system better.

54:06

SPEAKER_00

The historian is not adjacent to this paradigm. The historian is the missing instrument.

54:16

SPEAKER_00

This is the reframe that organizes all of this. And it came out of a 90 minute conversation with Halpern. And it inverted what I had assumed. We are not bringing historians into AI architecture. We are bringing language models into the archive. The question is not what AI can do for historians, but it's what historians and theologians and classicists and clinicians can do with AI. Whether the disciplines that are trained to read, contextualize, and interrogate texts can then discipline the machines that now generate them.

54:58

SPEAKER_00

I want to end this where it actually started. Because it did not start in a research lab. It started in a hospital room. In an attempt to use a language model to help a person with advanced dementia speak. To give language back when the disease had taken it. It could not work. The documentary record you would need to anchor that encounter had not been gathered. And the person it was for was by then mostly beyond the reach of language. What that attempt produced, absent any anger, was not the person. It was a culturally shaped composite that resembled them just enough to mark with painful clarity the distance between them. That is Miranda distortion in its purest form.

55:50

SPEAKER_00

You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room. That's the threshold this whole framework is accountable to. A system that produces convincing fabrications when the persona is your own beloved is not a research artifact with limitations. It's a violation. Every constraint I've described. Documents stay documents. The human keeps the interpretive custody. The encounter stays reversible. Fidelity measured against a record and not against fluency. Every one of them is motivated by the recognition that you evaluate this technology at the threshold of its hardest use case. It's not its easiest.

56:35

SPEAKER_00

A framework that cannot meet a grandchild at her grandmother's letters is not a framework at all. It's just another failed product.

56:48

SPEAKER_00

If the dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, which they do, then your evals cannot detect the dominant failure. So here's where I'll leave you. If you ship character bots, companion AI, pedagogical agents, historical simulations, anything where a persona is supposed to reason from a record, your evals are measuring the wrong thing. The instrument that catches what they miss is pre-registered. It's reproducible by any team with a frontier model and a context window. It scales as a build time gate, not a runtime bottleneck.

57:31

SPEAKER_00

And it only works with a humanist in the loop, which I've shown you is a technical requirement, not a courtesy. The protocol, the questions, the rubric and the predictions, the historian's sealed vignettes, all of it will be published with this paper with Rick and Sean. I'm not here with results. I'm here with an instrument and an invitation. The archive is open. The laboratory is built. Run it with us. And let what comes through be measured, not by how it sounds, but by whether it's true. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note