AI Engineer

Building a Chess Coach — Anant Dole and Asbjorn Steinskog, Take Take Take

709 summary words 3 min summary Watch video

Start with the signal

3 min read

Summary

Building a Chess Coach: Summary

Main Topics

  • Take Take Take Application: An iOS/Android platform founded by Magnus Carlsen for playing chess and sharing games
  • AI Chess Commentary Generation: How the team built an AI system to provide intelligent game reviews and coaching
  • Historical Context: Evolution of chess AI from brute-force engines to neural network approaches
  • Technical Pipeline: Integration of traditional chess engines with LLMs for explainable analysis
  • Autonomous Agents: Using Claude-powered agents to continuously improve commentary quality
  • Performance Optimization: Balancing latency requirements with output quality for consumer applications

Key Points

The Problem with LLMs and Chess

  • LLMs cannot calculate and tend to hallucinate when playing chess
  • They were trained on language, not chess positions, making pure chess play unreliable
  • However, they excel at explaining moves when given proper context

The Solution: Multi-Layer Pipeline

  • Stockfish Analysis: Classical chess engine identifies the best move and evaluation
  • Context Extraction: Custom detectors identify tactical patterns (forks, pins, skewers) and positional themes (double pawns, structural weaknesses)
  • Maya Engine: Novel neural network that predicts human move probabilities at different skill levels, helping explain difficulty of finding optimal moves
  • LLM Translation: Gemini Flash converts extracted data into natural language explanations
  • User Insights: Track player statistics to identify learning opportunities

Architecture Design Philosophy

  • Separation of concerns: Keep data pipeline separate from language generation
  • LLM's role is strictly to translate structured information into English, not to reason independently
  • Prevents hallucination by grounding all outputs in pre-computed data

Autonomous Agent Loop

  • Users can flag bad commentary, which triggers Claude via Slack integration
  • Claude runs automated "commentary triage" to investigate issues
  • Agent can modify prompts, adjust detectors, regenerate commentary, and verify improvements
  • Agent requests human feedback before submitting changes as pull requests

Performance Metrics

  • Target latency: Sub-3 seconds for coach feedback
  • Gemini 3 Flash benchmark:
  • Time to first token: ~1 second
  • End-to-end latency: ~3 seconds average
  • Accuracy: ~75% on 16 test scenarios
  • Claude with reasoning: Better accuracy (~60%) but unpredictable, longer latency
  • GPT-4 Mini: Lower latency, lower accuracy

Quality Assurance

  • Evaluation methodology: 16 custom chess scenarios testing tactical patterns, blunders, hallucination limits
  • Evaluation sources: Real game extraction + LLM-as-judge + human expert validation
  • Human expertise: Both speakers are skilled chess players who serve as final arbiters on output quality
  • Multi-model comparison: Continuously test new models as they release to identify improvements

Notable Quotes

> "We really don't want it to try to figure out too much on its own because it quickly leads to hallucination."

> "It's really important to separate the data pipeline from the language generation."

> "Really try to close the loop with autonomous agents. The flow...is now very common and very powerful and really allows you to iterate quickly."

> "Always try to build a very clear context extraction model. This, unfortunately, in the beginning is a very slow, painful process."

Takeaways

Technical Implementation

  • Don't use LLMs for core logic when accuracy is critical—use specialized tools (Stockfish) and use LLMs only for explanation
  • Grounding is essential: Feed LLMs pre-computed, structured context rather than asking them to reason from scratch
  • Separate concerns: Keep data extraction completely distinct from language generation for better quality control and faster iteration
  • Latency matters for consumer apps: Sub-3 second responses critical for game analysis use case; may differ for chat-based features

Operations & Iteration

  • Autonomous agents accelerate improvement: Closing the loop with Claude agents enables rapid testing and deployment of fixes without manual intervention
  • Domain experts are irreplaceable: Domain knowledge required for meaningful evaluation; can't rely solely on automated metrics
  • Start broad, prune carefully: Context extraction JSON files should start comprehensive, then be refined iteratively based on quality improvements
  • Model comparison is ongoing: Continuously benchmark new releases (Gemini, Claude, GPT) as they emerge to optimize for your specific use case

Design Principles

  • Clear evaluation frameworks: Create specific test scenarios (16 in this case) covering known challenges before scaling production
  • Human oversight at scale: Even with autonomous agents, domain experts should review outputs, especially for high-stakes domains
Full transcript 2989 words · 19 min read
0:00

[SPEAKER_00] Afternoon everyone.

0:14

SPEAKER_00

Our next talk will be something a little bit different. We're going to dive into the world of chess. Quick show of hands. Who has heard of Magnus Carlsen? Fantastic. No introduction needed, but widely considered the best chess player in the world. He also founded a company called Take Take Take. This is where myself, Anans, and my colleague Osborne currently work at. And we're going to talk to you today about how we built our AI chess coach that now you can use and is in production. So first up, quick agenda. We'll quickly discuss a bit more about Take Take Take, what it is we actually built, what it is we actually launched.

0:45

SPEAKER_00

Osborne will then go into a quick history of chess and AI. A lot of links there. We'll briefly touch on why LLMs are actually bad at chess and how we managed to solve this problem. We're then going to deep dive into actually understanding our game review and closing the loop with our autonomous agent, and you'll get a demo. And then finally, some latency versus quality trade-offs, as this is a consumer-focused AI application. And then lastly, some learnings. So first up, what is Take Take Take? In its simplest form today, it's currently an iOS and Android application. You can go on and play your friends, and you can post about your games.

1:11

SPEAKER_00

What's relevant for our particular talk is that after you play a game, you get presented with our game review. And this is powered by our AI pipeline. So for example, just showing you how it works, in this particular position, it's leading to a checkmate. The last move that White has played has moved the knight from this yellow square over here on F3, captured the pawn on E5. It is a brilliant move, so it automatically gets the brilliant notation, and the commentary below is actually generated by our system. And we're using an LLM, and the pipeline we'll get into in a second.

1:27

SPEAKER_00

But what's quite interesting about it is we're able to give you the nuance of why it is a tactic, what detectors from a positional and tactical sense have fired, what are the threats you're trying to do, and actually explain the why behind the move. So that's the system we're going to be talking about today. Finally, on the last step of our application, we've started revealing insights about your play. And this could be things like how accurate you played in a particular game phase, maybe your current rating, or your current depth in a particular opening.

1:42

SPEAKER_00

And these insights form the next layer of analysis that we present to the coach, who then gives them back to you as opportunities for learning and improving. We hope by using this, you'll be able to improve and become better at the game. [SPEAKER_02] All right. [SPEAKER_02] So first, a brief history of chess and AI since they've been intertwined for so long. [SPEAKER_02] Just to give you a little bit of a back story. [SPEAKER_02] 1949, Claude Shannon, the OG Claude, wrote the paper, Programming a Computer to Play Chess. [SPEAKER_02] And here he envisioned that, or he proposed that there are two types of chess engines, Type A and Type B.

2:08

SPEAKER_00

[SPEAKER_02] Type A were these brute force engines that search through all possible moves and figure out the best move. [SPEAKER_02] While Type B were those who we know from 2017 and onward that can selectively pick out the best moves. [SPEAKER_02] Back then, he assumed that we would need Type B computers to play chess because computers were so weak back then. [SPEAKER_02] You couldn't search through the whole tree of moves. [SPEAKER_02] But computers quickly became better and people just started scaling these Type A computers.

2:27

SPEAKER_00

[SPEAKER_02] They got better and better until they, in 1997, Deep Blue versus Kasparov, the first time a chess engine beat the best chess player at the time. [SPEAKER_02] So people didn't really bother about these Type B computers for a while, these intuitive engines, until DeepMind, shout out to DeepMind, released first AlphaGo, because Go is a much more complex game than chess. [SPEAKER_02] So you can't solve this with these Type A computers. [SPEAKER_02] You would need this intuitive approach, neural network approach, to actually selectively figure out which lines to calculate.

2:42

SPEAKER_00

[SPEAKER_02] But after that, they released AlphaZero, who could play not only Go but also chess and Shogi. [SPEAKER_02] And some years later, LLMs came and people started playing chess against the LLMs and quickly turned out that they can't really play chess. [SPEAKER_02] Sometimes they make some right moves and they can, to an extent, play a nice opening, but they quickly start to hallucinate. [SPEAKER_02] So, let's see if we can show that.

2:59

SPEAKER_02

[SPEAKER_01] Yeah, there's a video of... [SPEAKER_01] Grok went for the Poisoned Pawn line with Qb6 early on and lost pretty badly. Not necessarily because of the opening, but because it doesn't really know how to play chess. That was Magnus Carlsen commentating a LLM chess tournament from our office in Oslo. There was a tournament organized by Kaggle when they launched their game arena, which was a benchmark for benchmarking LLMs when playing different types of games. One of them was chess, and now they've started to add more games.

3:20

SPEAKER_02

Also, I added Werewolf recently, where you can watch LLMs try to deceive each other in social deduction games, which I can recommend watching. But, yeah, we see that LLMs often hallucinate because, obviously, they're trained on language. They're not...they can't calculate. They can't... High reasoning models can, to an extent, calculate through the reasoning steps where they can actually play out moves, but they quickly fall apart. But there's nothing inherently wrong about the architecture of the transformer architectures to play chess.

3:57

SPEAKER_02

DeepMind has trained a transformer to, instead of predicting the next token, they predict the evaluation based on a chess position, where they've trained it on millions of chess positions to Stockfish evaluations pair. And that has actually led the transformer to play at a grandmaster level strength. But these aren't trained on language, so these can't explain chess. So, how do we bridge the gap between these old chess computers that can understand and play really good chess, between the LLMs that can explain chess? So, we're going to go through our pipeline of how our game review explains chess in our app.

4:35

SPEAKER_02

When you play a game, the first thing we do is we run Stockfish through the whole game. Stockfish is the leading chess engine now. That's a classical chess engine that calculates the best move.

5:04

SPEAKER_02

So it's what Stockfish says is considered to be the solution in the chess position.

5:06

SPEAKER_01

[SPEAKER_02] We then extract a lot of context in the position, because we want to explain not only the best move, we want to explain the threats, the plans, the tactics that could arise in the position, what you should have played. [SPEAKER_02] So, we're going to go through our pipeline of how our game review explains chess in our app.

5:16

SPEAKER_02

When you play a game, the first thing we do is we run Stockfish through the whole game. Stockfish is the leading chess engine now. That's a classical chess engine that calculates the best move. So, it's what Stockfish says is considered to be the solution in the chess position. We then extract a lot of context in the position, because we want to explain not only the best move, we want to explain the threats, the plans, the tactics that could arise in the position, what you should have played. A lot of these nuances that are useful if you want to learn how to become better at chess. So, we have a lot of detectors that tries to figure out all of this.

6:05

SPEAKER_02

The forks, pins, skewers, positional structural themes. Double pawns, for example, is a disadvantage. So, we need to be aware of all of those things. And there's also a new novel chess engine called Maya, which is behind a research project by the University of Toronto, where instead of building a chess engine that is trained to play the best, they've trained a chess engine. It's a neural network that predicts the moves that humans would play in certain positions. So, given a rating, for example, an online rating of 1500, it outputs the probability distribution over all the moves in the position. And by doing this, we could actually say that a move is the best move.

6:59

SPEAKER_02

We know that because of Stockfish, but you also know it's really hard to find that move because the probability of playing it at certain levels are so low. And all of this information, we feed that to the LLM and that the LLM's job is only to translate this information into English. Because we really don't want it to try to figure out too much on its own because it quickly leads to hallucination. It still does, but we want everything to be grounded in the information that we give it. And that could result into a comment like this. If you play chess or know about chess, this is a game that I played. My opponent played F5 here, which is a bad move.

7:47

SPEAKER_02

So, by using Stockfish, you could see that you get a bad move indicator. But that's not so useful to just know it's a bad move. So, we are running our detectors to figure out that, okay, F5 is threatening to trap my queen. You can also see it draws a line with Bishop G5. But you can also say, while it threatens to trap the queen, I can just capture the pawn in the middle because that's defense this square so that my queen can get out of the situation. So, that's how we get to that situation. Now, I'm going to explain a bit on how we improve our game review using agents. We have closed the loop from user feedback to the public request, essentially, with humans in the loop.

8:41

SPEAKER_02

But what happens when users download the commentary in your app? Because you can't download it if you think it's bad. It posts it to Slack, but it also sends it to Cloud Code Channel. Channel is a new feature in the research preview that is essentially an MCP server that can inject events into a running Cloud Code session. So, similar to OpenClaw, if you use that. So, you have this continuously running channel, and you can inject events to it. So, and then Cloud Code starts working on the commentary. It gets all the information. It runs a commentary triage skill that we created that outlines its process, how it should go about to investigate what's wrong in the position.

9:29

SPEAKER_02

It has some scripts to actually run the generation, so it can modify, for example, the prompt. It could change some of the detectors, create some new detectors. And then it can generate the commentary again, given this new information, and verify its own work. And then it will also ask questions back to Slack, so that I could be on the bus and I could get a message from Claude, who is working on this problem, where it will ask me if this seems right, and I can guide it. And if it looks right, I'll just tell it to submit the PR, and I open GitHub on my mobile, and it works through it, and merge it. I'm gonna show how this works.

10:07

SPEAKER_02

By, so we have a running Cloud Code channel here. Here is the Slack channel, where the commentary appears. This is just me having tested a bunch of time. I'm gonna open up the app on my phone, go to comment, and report it as bad. Now we see it posts a comment. We can see the position, the commentary that was generated. Now I haven't really looked at the commentary, so it could be, it's probably good. But we can also see that it injects it to the Claude channel, who invokes the commentary triage skill, and starts working. Now, this is now running on high effort, so this could take a while.

10:50

SPEAKER_02

So I'm thinking we should just go to the next slide, and then we could get back to it to see if it is something that's happening. [SPEAKER_00] Fantastic. So we'll come back to that in a few seconds. [SPEAKER_00] So as we built this for end users, we had to really consider this trade-off between latency versus quality. [SPEAKER_00] So typically, when you finish a chess game, you want to get the analysis and the results pretty quick. [SPEAKER_00] You want to cycle through the moves one by one. [SPEAKER_00] So we really couldn't show you a coach's thinking screen indefinitely, while reasoning tokens are running in the background, as an example.

11:27

SPEAKER_02

[SPEAKER_00] So we had to get this done, which felt almost instant. In AI world, that's a few seconds at best. [SPEAKER_00] So we were aiming for sub three seconds to generate our coach feedback. [SPEAKER_00] How do we do this? We use Gemini 3 Flash. [SPEAKER_00] Time to first token is typically being about a second, end-to-end latency on average is about three seconds, which meets our criteria. [SPEAKER_00] We have experimented with other reasoning models, and we'll get into that on the next slide. [SPEAKER_00] The analysis is not incorrect. So the quality is definitely good.

11:53

SPEAKER_02

[SPEAKER_00] But the challenge is it's unpredictable as to how long it's going to take to finish. [SPEAKER_00] So we have a new set of features planned for a more chat with your coach type experience, where we can expect the user to be more patient and wait for a response rather than in this sort of phase where it needs to be more instantaneous. [SPEAKER_00] The last thing about quality is Osborne and I are both good chess players. [SPEAKER_00] So we ultimately have the final say when we look at a position to actually use how we would calculate and how we would play and compare it to the LLM's response.

12:19

SPEAKER_02

[SPEAKER_00] This allows us to actually evaluate whether it's doing the right thing or not. [SPEAKER_00] So if we talk about evals in more detail, Gemini Flash is our benchmark, but we have multiple chess scenarios. [SPEAKER_00] Currently, we have 16 different scenarios that we created.

12:39

SPEAKER_00

These are around themes like tactical patterns, blunders and limiting hallucination. The last thing about quality is Osborne and I are both good chess players. So we ultimately have the final say when we look at a position to actually use how we would calculate and how we would play and compare it to the LLM's response. This allows us to actually evaluate whether it's doing the right thing or not.

12:49

SPEAKER_00

So if we talk about evals in more detail, Gemini Flash is our benchmark, but we have multiple chess scenarios. Currently, we have 16 different scenarios that we created. These are around themes like tactical patterns, blunders and limiting hallucination. So as an example, there might be a knight fork on the particular chess position. And we're trying to assert that the LLM can actually understand and mention this when we run it through with our context engine.

12:53

SPEAKER_00

And how do we do this? We extract scenarios from real games. We use LLM as a judge, a very powerful technique to test. We then run the model in Open Router. Open Router comes in handy because new models are being released so fast and frequently. We just want to be able to quickly swap in and swap out a new version of, say, Gemini. We want to check out the latest GPT-5 model or one of the Claude models. So we'll then compare and see the quality, ultimately relying on our own skill to detect whether this is good or bad.

12:55

SPEAKER_00

And as a final point on this, we ran our three models: Gemini, Claude, and GPT-5. Typically, Gemini Flash is about 75%. It still doesn't pass all the scenarios we've set it. So we're always seeing if a new model will actually exceed some of the tricky cases we've set up. Claude with more thinking gets us to about just under 60%, but the latency is much longer. GPT-5 mini gives us a smaller model with lower latency, but lower accuracy as well. So we continuously run through these to update.

13:03

SPEAKER_00

Last thing on our learnings and how this can apply to your world sitting in front of us. Number one, it's really important to separate the data pipeline from the language generation. LLMs can do a lot of different things, but if you need high latency or quick latency, it's a good technique. Really try to close the loop with autonomous agents. The flow that Osborne showed is now very common and very powerful and really allows you to iterate quickly. Always try to build a very clear context extraction model. This, unfortunately, in the beginning is a very slow, painful process. It's a large JSON file that you keep starting big and you start to prune step by step and you see how quality improves over time. Automated evals really do help. And I hope in your domains you also have a set of SMEs that you can rely on to evaluate the output. And sometimes that's not necessarily the person actually building it. It could be someone else who is a domain expert. So remember to partner if needed.

13:09

SPEAKER_00

The last thing on a fun note before we go back to the output of the coding agent. We do have some chess sets on the third floor at the entrance. You may have seen. We're going to host a chess simultaneous display today in the afternoon around 3:45 p.m. A chess simultaneous display for those who are unfamiliar is when one person, in this case me or Osborne, plays multiple people at the same time. So we have four chess sets. We will play four people at the same time. We'll have slightly more time for us because we have to walk around to play multiple boards. If you happen to play and you happen to win, you will get one of the wooden chess boards at the end of the event. They're very nice, high quality chess sets. If no one wins, we will still determine who the two best players are and we will still give you a set. And if everyone wins, we need to buy more boards. Hopefully not everyone wins. There's a QR code if you want to sign up or just stop by 3:45. You're welcome to do that.

13:15

SPEAKER_00

And then, yeah, just close it off. Let's go back to see what has been happening here. Oh, is it still thinking? It has actually added a comment. Looking into this now, investing in position. [SPEAKER_02] Quick question. What specifically feels wrong about the commentary? Yeah, it got me there. There's nothing wrong. You're absolutely right. Nothing wrong. So, yeah, now it's going to close this off because it worked well. Fantastic. Well, thank you so much and happy to take any questions. The analysis is not incorrect. So the quality is definitely good. But the challenge is it's unpredictable as to how long it's going to take to finish.

13:38

SPEAKER_00

So we have a new set of features kind of planned for a more, you know, chat with your coach type experience, where we can kind of expect the user to be more patient and wait for a response rather than in this sort of phase where it needs to be more instantaneous. The last thing about quality is Osborne and I are both good chess players. So we ultimately kind of have the final say is when we look at a position to actually use how we would calculate and how we would play and compare it to the LLM's response. This allows us to actually evaluate whether it's doing the right thing or not.

14:10

SPEAKER_00

So if we talk about evals in more detail, like I said, Gemini Flash is kind of our benchmark, but we have multiple chess scenarios. Currently, we have 16 different scenarios that we created. These are around themes like tactical patterns, blunders and sort of limiting hallucination. So as an example, you know, there might be a night fork on the particular chess position. And we're trying to assert that can the LLM actually understand and mention this when we run it through with our sort of context engine. And how do we do this? We extract scenarios from real games. We use LLM as a judge, very powerful sort of technique to test. We then run the model in open router.

14:47

SPEAKER_00

Open routers come in handy because new models are being released so fast, so frequently. We just want to be able to quickly swap in and swap out maybe a new version of say Gemini. We want to check out the latest GPT-5 model or one of the Claude models. So we'll then compare and see sort of the quality, ultimately relying on our own skill to detect whether this is good or bad. And as a sort of final point on this, we ran our three models, Gemini, Claude and GPT-5. And, you know, typically Gemini Flash is about 75%. It still doesn't pass all the scenarios we've set it. So we're always kind of seeing if a new model will actually exceed some of the tricky cases we've set up.

15:27

SPEAKER_00

Claude on more thinking gets us to about just under 60%, but the latency is much longer. GPT-5 mini given us a smaller sort of model, lower latency or the slower latency, but lower accuracy as well. So we kind of continuously run through these to update. Last thing on sort of our learnings and how this can sort of apply to your world sitting in front of us. Number one, you know, really important to separate that sort of data pipeline from the language generation. LLMs can do a lot of different things, but if you need, you know, high latency or quick latency, it's a good sort of technique. Really try to close the loop with autonomous agents.

16:01

SPEAKER_00

Kind of the flow that Osborne showed is now very common and very powerful and really allows you to iterate quickly. Always try to build a very clear sort of context extraction model. This, unfortunately, in the beginning is a very slow sort of painful process. It's a large, you know, ultimately JSON file that you keep sort of starting big and you start to prune step by step and you see how quality improves over time. Automated evals really do help. And I hope in your domains you also have a set of, you know, SMEs that you can help rely on to evaluate the output. And sometimes that's not necessarily the person actually building it.

16:33

SPEAKER_00

It could be someone else who is a domain expert. So remember to sort of partner if needed. The last thing just on the fun sort of note before we go back to the output of the sort of coding agent. We do have some chess sets on the third floor at the entrance. You may have seen. We're going to host a chess symbol today in the afternoon around 3.45 p.m. A chess symbol for those who are unfamiliar is when one person, in this case me or Osborne, plays multiple people at the same time. So we have four chess sets. We will play four people at the same time. We'll have a slightly more time for us because we have to walk around to play multiple boards.

17:09

SPEAKER_00

If you happen to play and you happen to win, you will get one of the wooden chess boards at the end of the event. They're very nice high quality chess sets. If no one wins, we will still determine who the two best players are and we will still give you a set. And if everyone wins, we need to buy more boards. Hopefully not everyone wins. There's a QR code if you want to sign up or just stop by 3.45. You're welcome to do that. And then, yeah, just close it off. Let's go back to see what has been happening here. Oh, is it still thinking? It has actually added a comment. Looking into this now, investing into position.

17:43

SPEAKER_02

Quick question. What specifically feels wrong about the commentary? Yeah, it got me there. There's nothing wrong. You're absolutely right. Nothing wrong. So, yeah, now it's going to close this off because, yeah, it worked well. Fantastic. Well, thank you so much and happy to take any questions.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note