AI Engineer

Build for the Memo, Not the Demo — Shawn Chan, China Resources Holdings

1702 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI systems used for consequential financial or enterprise decisions must be designed to defend every claim under skeptical review, not merely produce fluent outputs that impress in a demo.
  • Why it matters: The speaker frames trustworthiness as an architecture and workflow problem—provenance, contradiction handling, numerical controls, uncertainty labeling, and human accountability—not primarily a model-quality problem.
  • Best use: Use this as a concise product-design and diligence checklist for any agent or AI workflow that produces decision-support artifacts, especially investment, finance, compliance, legal, or board-facing outputs.

Executive Summary

Shawn Chan argues that most AI-finance products are optimized for the "demo" machine: a clean input produces a polished answer that appears convincing for a few minutes. Real investment work runs on the "memo" machine, where committees scrutinize long, conflicting source sets and expect every assertion to survive challenge. The relevant standard is therefore not fluency but defensibility.

Drawing on roughly 200 investment-committee meetings, Chan says deal approval depends on trust rather than apparent intelligence. A single unverifiable, inconsistent, or overstated claim can undermine the credibility of an entire document because reviewers infer that visible errors may signal hidden ones. He uses Google Bard's 2023 factual-error launch demo, the Zillow home-buying write-off, fabricated legal citations, and Air Canada's chatbot case to show that confident output can create material financial and legal consequences.

His implementation prescription is deliberately unglamorous: attach each claim to its exact source paragraph and source-quality level; visibly distinguish facts from estimates; automatically reconcile figures; surface rather than silently resolve source conflicts; and require an auditable human approval gate. He calls these "plumbing and honesty" controls, arguing that stronger models alone do not solve the operational trust problem.

The same logic applies to fundraising. A founder's deck becomes an internal investment memo after the meeting, and every verbal or written metric may be checked against the data room. Product trust design and investor diligence therefore demand the same discipline: show provenance, expose uncertainty, maintain consistency, and make accountability explicit.

Key Takeaways

  • Claim: AI products should be built for adversarial memo review, not for a five-minute demo, because every externally visible statement can become decision-critical once money or liability is involved. | Evidence: Chan contrasts a demo's clean-document-to-fluent-answer flow with a real memo assembled from filings, transcripts, broker notes, spreadsheets, and call notes that may conflict. He cites the February 2023 Bard launch error about the James Webb Space Telescope, after which the company's stock fell about 8% in a day—roughly $100 billion in value. | Implication: Evaluate AI outputs as if a skeptical reviewer will ask for proof of every material sentence; marketing, product, and live-deal workflows should use the same core verification standard. | Caveat: The $100 billion figure is presented as an illustrative market-value estimate tied to a broader market reaction, not proof that one sentence alone mechanically caused the entire loss.
  • Claim: Source retrieval must distinguish evidence quality rather than treating semantically similar text as equally reliable. | Evidence: Chan ranks an audited filing as akin to testimony under oath, an analyst note as less reliable, and an internal email as hearsay. In his example, an expensive AI selected an enthusiastic six-month-old group-chat estimate even though the audited number was nearby in the filings. | Implication: Retrieval and synthesis layers need explicit source-tier metadata, with audited or primary records prioritized over commentary, informal communications, and unsupported estimates.
  • Claim: Small numerical inconsistencies destroy trust because reviewers use them as evidence of weak controls, not because the arithmetic difference itself is material. | Evidence: A hypothetical memo says revenue growth is 18% on page one and 17.4% in a later table; Chan says the committee will ask what else was unchecked. He also references Zillow's algorithmic home-buying business, which wrote off around half a billion dollars, shut the unit, and cut roughly a quarter of staff after pricing assumptions diverged from reality. | Implication: Deploy automated consistency checks across narrative, tables, spreadsheets, and revisions, and block release when material figures cannot be reconciled or explained. | Caveat: Chan's point is not that all numerical variance is an error; it is that unexplained variance in a supposedly controlled deliverable signals that reconciliation and monitoring are missing.
  • Claim: Contradictory evidence should be surfaced to humans rather than silently harmonized by the model. | Evidence: Chan describes a case in which a CEO's earnings-call figure materially differed from the official filing; the system selected the more favorable version without flagging the conflict, and the discrepancy was found only because someone had both documents open. | Implication: An AI workflow should represent disagreement as a review item—with competing claims, sources, dates, and confidence—not collapse it into one smooth narrative.
  • Claim: Facts, estimates, forecasts, and assumptions need persistent visual or structural separation so that probability judgments cannot silently become assertions of fact. | Evidence: Chan traces how "approval expected soon" can evolve across repeated drafts into "approval received" without an intentional lie, because the original estimate loses its label. His proposed low-cost control is a tag, color, or other marker that survives copying into later documents and slides. | Implication: Require typed claims and durable uncertainty labels in generated documents; do not rely on prose wording alone to preserve epistemic status through edits and republishing.
  • Claim: The core product is rapid, exact provenance plus accountable human sign-off, not simply a more capable model. | Evidence: Chan's "30-second test" is whether a reviewer can click once from a claim to the exact source paragraph instead of opening multiple browser tabs. He cites the lawyer sanctioned after submitting six fabricated chatbot-generated cases and Air Canada's unsuccessful argument that its chatbot was a separate legal entity after it invented a bereavement-fare policy. His vendor checklist requires claim-level receipts, source trust levels, visible facts versus guesses, automatic reconciliation, contradiction alerts, and a locked approval audit trail. | Implication: Treat citations, review queues, version history, permissions, and approval logs as first-class product capabilities. Position agents as systems that strengthen a named decision-maker's judgment, not as autonomous substitutes for accountability. | Caveat: Human approval is not a ceremonial final click; it must identify who reviewed the output, what changed, and when the approval occurred.

Detailed Brief

What investment committees infer from an error

  • Claims: A visible mistake is damaging because it changes the committee's assessment of the process that generated the whole document.; Trust is fragile in capital-allocation settings because the reviewer is paid to resist polished persuasion rather than reward presentation quality.
  • Evidence: Chan recalls a polished banker presenting confident but incorrect figures; a senior attendee's quiet question—"Where does this number come from?"—exposed the weakness through the banker's inability to answer.; He says he has participated in approximately 200 investment-committee meetings and reviewed hundreds of founder and banker decks.
  • Caveats: The talk is principally a practitioner perspective from investment and transaction diligence, so its examples emphasize high-consequence document workflows rather than all consumer AI use cases.
  • Implications: For high-stakes deployments, quality measurement should include reviewability and error containment, not only answer accuracy or user satisfaction.; Sales materials should anticipate diligence questions by making the evidence chain and control design demonstrable before a buyer asks.

Fundraising follows the same memo standard

  • Claims: A startup pitch is not evaluated only in the meeting; it is converted into an internal investment memo after the founders leave.; Product claims, spoken metrics, and deck figures are all subject to cross-checking against the data room.
  • Evidence: Chan states that every number said aloud can be checked against diligence materials and explicitly says the product-document standards apply personally to founders raising capital.
  • Implications: Maintain one reconciled source of truth for commercial, usage, financial, and product-security metrics used in decks, data rooms, and customer conversations.; Avoid overstating model autonomy, accuracy, or customer outcomes when the underlying audit evidence cannot support the assertion.

Notable Concepts & Terms

  • Demo machine vs. memo machine: The central distinction: demos optimize for immediate plausibility, while memos must withstand evidence-based disagreement before consequential decisions.
  • Memo test: The standard that every material output claim should survive scrutiny about its source, accuracy, uncertainty, and ownership.
  • Claim-level receipt: A direct link from an output sentence to the exact supporting source passage, ideally accompanied by a source-trust classification.
  • Source trust level: Metadata that differentiates primary, audited, or official materials from analyst commentary, internal notes, and informal communications.
  • Contradiction surfacing: A workflow behavior in which conflicting sources are presented to a reviewer instead of being silently resolved by the model.
  • Facts and guesses in separate boxes: A requirement to preserve the distinction between verified statements and forecasts, estimates, or assumptions throughout document creation and reuse.
  • 30-second test: A practical provenance test: a reviewer should be able to locate the precise support for a claim in one click and roughly 30 seconds.
  • Locked human approval gate: A non-bypassable review step that records the responsible person, reviewed version, changes, and approval time as an audit trail.

Operator Notes / Why Ken Should Care

  • Add source-tier fields to every retrieval corpus and require the generation layer to display both the source link and source tier for material claims.
  • Define a release gate for decision-support outputs: no publication when numeric checks fail, claim provenance is missing, or unresolved source conflicts remain.
  • Implement structured claim types—fact, estimate, forecast, assumption, recommendation—and ensure labels persist in exports, copied text, and downstream slides.
  • Instrument an exception queue for contradictions rather than using model selection logic to choose a single preferred value without disclosure.
  • Require named approvers and immutable version/audit logs for any agent output used in finance, legal, compliance, customer-policy, or board workflows.
  • Before fundraising or enterprise diligence, reconcile every deck metric and product-control claim to evidence that can be produced directly from the data room.

Source/Metadata

  • Title: Build for the Memo, Not the Demo — Shawn Chan, China Resources Holdings
  • Transcript words: 2959
  • Duration seconds: 1462
  • Timestamp note: No timestamps or chapters were provided. The transcript's final section repeats part of the preceding content and appears to end abruptly.
Full transcript 2502 words · 13 min read
0:00

Reza Zang-Tung Good afternoon. Thank you for being here, day four of our conference. A room with no windows, right just after lunch. You are the strongest people in this building. Let me start with a confession. For 15 years, my job has been one thing. I'm sitting in a room, a very smart, confident person hands me a piece of paper. And before my company spends $100 million, I have to decide, do I believe this paper? This year, my job is still exactly that. Except now, some of the time, the very confident person handing me the paper is a chatbot. And honestly, the chatbot is often better written than humans.

0:49

Better grammar, nicer formatting. Never gets defensive when you ask a follow-up question. Very, very sure of itself. The problem is, being sure of yourself and being right are two different scales. Some of the most confident people I have ever met in finance were also the most wrong. AI just learned that trick faster than the rest of us. So, here, my whole talk in two sentences. First, almost every AI finance product is built to impress people for five minutes. And almost none are built to survive a room whose entire job is to not be impressed. Second, and this is the part I promised.

1:26

The organizers had the exact same skills that fix your product as the skills that get investors like me to write you a check. Same muscle. I approve both. And I'm telling this now, this week, because a lot of you are about to get pulled into exactly this. Your CEO saw a demo somewhere. Your biggest customer suddenly has a compliance department. Or you are three months from raising your next round. When any of those days arrive, I'd rather you hear the hard part from someone who sits on the other side of the table. Than discover them live in a room in front of the people who bill by hour. Quick bit about me. I promise this is the boring part, and I will keep it fast.

2:11

15 years of cross-border deals, Hong Kong, mainland China, the UK, the US, mergers, IPOs, big strategic investment. On names, you'd actually recognize the kind of companies that go public, and your LinkedIn feed won't shut up about it for a week. Along the way, I've sat in about 200 investment committee meetings. I have the gray hairs to prove it. I checked. I checked. It's not genetics. And here, the part that matters for the second half of this talk, I've also read hundreds and hundreds of pitch decks, founders' decks, bankers' decks, 14-7 slide seed decks.

2:55

I have seen fonts that should be illegal. 200 committee meetings taught me one thing. No textbook says out loud, the number on the page is not what gets a deal approved. Trust is what gets a deal approved. Money doesn't follow intelligence. Money follows trust, and that trust is fragile, especially when the thing that wrote the page has never once in its entire life said the words, I'm not sure. Let me give you one small taste of what those rooms feel like. Early in my career, a very polished, very expensive banker presented a beautiful slide full of confident, wrong numbers. One senior person in the room asked one quiet question, where does this number come from?

3:36

The banker paused for what felt like an entire fiscal quarter. The pause told me more about finance than three years of exams did. So I'm not here as a builder. I don't build these systems. I sit across the table from them, and from the founders selling them, deciding whether to trust them. Today, I'm going to give you both halves of that. What builds trust in your product and what builds trust in your pitch? Let's define two machines that everyone keeps confusing. Machine one is a demo. One clean document in, one fluent answer out. Its whole job is to make a room go ooh for five minutes. Machine two is a memo.

4:36

The real document a real committee reads before real money moves. Hundreds of pages, filings, transcripts, broker notes, spreadsheets, someone's rushed notes from a call last Tuesday. Half the sources disagree with each other, and the memo's job is not to make you go ooh. Its job is to survive an argument. A family dinner, except somebody's uncle brought a spreadsheet. Here's the everyday version of the gap. Ask your phone to summarize a long email. That's a demo. It just has to sound plausible for 10 seconds. Now imagine that same phone has to stand in front of a bank and defend out loud why you deserve a mortgage. Suddenly, plausible isn't enough.

5:25

Now it has to be right, and it has to prove it. That second situation is what a memo actually is. Now you might think the demo world and the memo world never touch. Let me tell you about the most expensive typo in history.

5:45

February 2023, one of the biggest tech companies on the earth launches its shiny new AI assistant with a promotional demo. In that demo, the assistant answers a simple question about a space telescope, and it gets the wrong answer. One sentence, one wrong fact about a telescope. In a marketing demo, the market noticed. The company's stock dropped around 8% in a day. That's roughly $100 billion of value gone because of an unchecked sentence. $100 billion for one sentence. Nobody in that company asked one question every junior analyst on my team is trained to ask before anything leaves the building, wait. Where does this claim come from? Did anyone check it?

6:31

So here's the punchline. Even the demo failed the memo test. The moment real money is watching, and the real money is always watching, every sentence becomes a memo sentence. There is no safe demo anymore. Keep that story in your head. Because the same trust-breaking moment happens in six smaller, quieter, very predictable ways inside of your product every day. Let's go through them fast. Six ways trust quietly breaks. Which source do you believe? Do the numbers agree? Do you have contradictions or show them? Is that a fact or a guess?

7:18

Can you prove it in 30 seconds? And whose name is actually on the decision? Six, keep count with me. I will keep each one short, and I will bring receipts. The next one is model one. Not every source deserves the same trust, but most AI systems treat them like they do. Think of it like this. A number from an audited filing is your accountant speaking under oath. A number from an analyst note is upfront at a party, confident, probably wrong. Something from someone's internal email is a thing you overheard in an elevator. Most retrieval systems can't tell this apart. They grab whichever text is close to your question and hand it over like gospel.

8:27

True story, I once watched a very expensive AI confidently quote a number from a group chat. Someone's rough guess texted six months earlier. The model loved the confident phrasing. The real audited number was three rows away in the actual filings. The AI just liked the group chat version better because it sounded more enthusiastic. If your system can't tell audited content under oath from a rumor in a group chat, it is not ready for real money. Model 2, the numbers have to agree with each other. Everywhere, every time, here's a memo that already died. Page one says revenue growth 18%. Page 11, a little table nobody reads for fun says 17.4.

9:00

Nobody in the room cares about the missing 0.6. They care about what it means.

9:10

If this person didn't check the easy mathematics, what did they not check on the hard stuff? That memo didn't pass. Not because of the number, but because of what the number implied. And if you want the industrial-strength version of this failure, remember the giant American real estate company that let an algorithm buy houses at scale. The algorithm was extremely confident about the house prices. The houses disagreed. The company ended up writing off around half a billion dollars, shut the whole unit down, and let a quarter of the staff go. The model wasn't stupid. The model was unsupervised.

9:44

Nobody built the boring machinery that forces the numbers to keep agreeing with reality after launch day. Confident was not the same as right. Model 3 surprises people. A contradiction is not a bug. A contradiction is a gift. A contradiction is not a bug. If the CEO says one gross number on the earnings call and the official filing says a different one, the gap is the single most interesting thing in the real story. Real diligence lives for that gap. AI does the opposite. It is trained to sound smooth and helpful. So when it hits a conflict, it quietly picks whichever version reads nicer and moves on. You never even learn there was a disagreement.

10:48

I once sat through exactly this, the CEO's number and the filing number, meaningfully different. The system picked the nicer one. Nobody flagged it. Everyone just used the nicer one. We caught it because one person happened to have both documents open at once.

11:21

Pure luck. Luck is not a control. Your job as a builder isn't to resolve the argument. It's to make sure that the argument happens in front of a human instead of quietly alone inside a box. Model 4. Facts and guesses have to live in separate boxes. And fluent AI loves melting them into one smooth sentence. Example. The company will likely receive approval next quarter. Reads like a fact. Sounds like a fact.

12:30

It is a guess. Somebody's estimate wearing fact-shaped clothing. A committee's entire job is to disagree with the guesses while trusting the facts. If your system melts them together, the committee can't find the seams. And then all they can do is approve or reject the general vibe of the document. You should not spend $100 million on vibes. I watched approval expected soon turn across three drafts of a demo into approval received. Nobody lied.

13:11

The guess just wore its fact costume a little longer each rewrite until nobody remembered it started as a guess.

13:19

The approval didn't arrive on schedule. That was an uncomfortable phone call. This fix is almost embarrassingly cheap. Label your guesses, a tag, a color, anything that survives being copied and pasted into someone else's slide three weeks later. Model 5, if nobody can find where a claim comes from, it doesn't matter how right it is. You've all heard about the New York lawyer, the first one. He filed a legal brief written with a chatbot's help. The brief cited six court cases. Beautiful citations. Proper formatting. Very convincing. One small issue. The cases didn't exist. The AI invented all six. And here, my favorite detail, the part that should be taught in school.

14:27

Before filing, the lawyer got suspicious. So he asked the chatbot, are these cases real? And the chatbot said, yes. That is like asking the guy who sold you the watch whether the watch is real. The judge fined him. The story went around the world. And my second favorite detail, the fake cases even had realistic-sounding names and page numbers.

14:58

The AI didn't just lie. It styled the lie beautifully. Wrong, but beautifully. The lesson is a 30-second test. When someone points at a sentence and says, show me where this comes from, you either click once and land on the exact source paragraph, or you open seven browser tabs and start sweating. I have personally been the guy with seven tabs in a real meeting where a room full of people watched me scroll. 10 out of 10 would not recommend. If you remember only one sentence from this whole talk, the click-through is the product. Everything else is well-written packaging. Model 6, my favorite, because it's the most human. Someone has to sign.

15:42

Here's the story you probably know. An airline website chatbot told a grieving customer he could book a full-price ticket now and claim a bereavement discount afterward. That policy didn't exist.

16:01

The chatbot made it up politely, fluently, confidently. The customer took the airline to a tribunal, and the airline's defense, this is real, was that the chatbot is, quote, a separate legal entity responsible for its own action. That is the corporate version of my dog ate my homework. The tribunal didn't buy it. The airline paid, and every one of us in our boardroom quietly took a note that day. You can't outsource accountability to your own software. At the bottom of every real decision, a human signs. If your architecture doesn't have a fundable human at the end of it, you have not built a product. You have built an excuse generator.

16:38

So build your AI around that accountable person, not instead of them.

16:46

Okay. The fix. Five things. Each one is a direct cure for a story you have just heard. This is almost word for word what I demand from a vendor before letting their system near a live deal. One, every claim comes with a receipt. Each sentence linked straight to its source paragraph with the source trust level attached. Not a citation tab at the end. Two, facts and guesses stay visibly separate. I glance at the page. I instantly see what's proven and what's somebody's best estimate. Three, numbers agree with each other automatically. The system refuses to ship a memo where the figures don't match. No human checking at 2 in the morning.

17:44

Four, contradictions get surfaced, never smoothed over. When sources disagree, the system raises its hand instead of picking the friendlier answer. Five, a real human approval gate, and it's locked. Who reviewed what changed when they signed, that log is audit trail. Now, notice what's not on the list. A smarter model. Not one of these is a bigger-brain problem. All five are plumbing and honesty problems. The winners in this category won't win on benchmark points. They will win because a tired, skeptical finance person can trust their output at 11 at night without opening seven tabs. Now, the part I promised, the money. Many of you are not just building AI products.

18:39

You are raising for them. Or you will be. So let me tell you what actually happens after you leave the pitch meeting. Your deck becomes a memo. Literally, someone like me sits down and writes an internal memo about you.

19:11

Every number you said out loud gets checked against your data room, which means everything I just told you about the documents applies to you personally.

19:20

Okay, my time's up. Thank you. An airline website, chatbot, told a grieving customer he could book a full price ticket now and claim a briefment discount afterwards. That policy didn't exist. The chatbot made it up politely, fluently, confidently. The customer took the airline to a tribunal, and the airline's defense, this is real. Was that the chatbot is, quote, a separate legal entity responsible for its own action. That is the corporate version of my dog ate my homework. The tribunal didn't buy it. The airline paid, and every one of us in our boardroom quietly took a note. That day, you can't outsource accountability to your own software.

20:26

At the bottom of every real decision, a human science, if your architecture doesn't have a fundable human at the end of it, you have not built a product. You have built an excuse generator. So build your AI around that accountable person, not instead of them.

20:55

Okay. The fix. Five things. Each one is a direct cure for a story you have just heard. This is almost a word for word. What I demand from a vendor before lighting their system near a live deal. One, every claim comes with a receipt. Each sentence linked straight to its source paragraph with the source trust level attached. Not a citation tab at the end. Two, fact. And the guesses stay visibility separate. I glance at the page. I insistently see what's proven and what somebody best estimate. Three, numbers agree with each other automatically. The system refuse to ship a memo where the figures don't match. No human checking at 2 in the morning.

22:04

Four, contradictions get surfaced. Never smoothed over. When sources disagree, the system rises its hand instead of picking the front layer answer. Five, a real human approval gate and is locked. Who revealed what changed when they signed that log is audit trial? Now, notice what's not on the list. A smarter model.

22:39

Not one of these is a bigger brain problems. All five plumbing and honesty problems. The winners in this category won't win on benchmark points. They will win because a tired, skeptical finance person can trust their output at 11 at night without opening seven tabs.

23:13

Now, the part I promised, the money. Many of you are not just building AI products. You are rising for them. Or you will be. So let me tell you what actually happens after you leave the pitch meeting. Your deck becomes a memo. Literally, someone like me sits down and writes an internal memo about you. Every number you set out loud gets checked against your data room, which means everything I just told you about the documents apply to you personally for license. Okay, my time's up. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note