Running a Chess YouTube Channel entirely by AI — Stephan Steinfurt, TNG
Description
Daily chess puzzle explanations on YouTube: Our agent analyzes and describes chess puzzles in an accessible way - arrows included!
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: An AI agent system built on Gemini 3.1 Pro + chess engines + structured tools can autonomously produce high-quality chess educational videos from Lichess games, analyzing positions, generating narration, and uploading nightly without human oversight—at ~20-30 cents per video.
- Why it matters: This is a working example of agentic AI generating monetizable, domain-specific content at scale; it demonstrates tool-use patterns, cost structures, error rates, and trade-offs in human-in-the-loop vs. full automation that apply to Ken's agent systems, content automation, and AI-native product thinking.
- Best use: Study the architecture (LLM + chess engine + CCT tool + conflicting info strategy) as a template for multi-tool agent design; note cost/error trade-offs and the shift from Python scripts to reasoning-model-driven tool orchestration; consider parallels for AI-generated analysis/content in other verticals.
Executive Summary
Stephan Steinfurt and his team at TNG have built an end-to-end AI pipeline that downloads chess games from Lichess nightly, analyzes them using a custom agent, and uploads fully narrated, visually annotated chess tutorial videos to YouTube—all without human intervention. The system uses Gemini 3.1 Pro (chosen for its superior chess understanding after post-training) as the reasoning core, supplemented by a suite of tools: legal move validation, chess engine evaluation, checks-captures-threats (CCT) analysis, Maia engine (for human-like move prediction), board state manipulation, and optional web search for historical context. The agent decides which moves to highlight, which arrows to draw, which positions to call 'brilliant' or 'blunder,' and generates the script that ElevenLabs v3 TTS converts to audio. The result is videos like the two-minute tactical breakdown shown in the talk, which explain complex chess positions with narrative arc, emphasis cues (e.g., 'excited'), and multi-variation analysis.
The key architectural insight: the team initially used Python scripts to assemble chess data and feed structured prompts to an LLM, but with reasoning models (O1, Gemini Flash Thinking, Grok 4—Grok 4 was the best before Gemini 3.1), they shifted to letting the agent think and call tools itself. This delegation increased quality and flexibility. The agent is intentionally given conflicting information (best engine move vs. human-likely moves, checks vs. captures) so it can reason about what a human viewer might consider and explain why certain intuitive moves fail. The CCT tool, for example, lists all checks, captures, and threats in a position; some are obviously bad, but the agent uses that diversity to explain common misconceptions—e.g., in the example shown, the best move is a non-obvious rook sacrifice, but the agent can discuss why checking moves look tempting yet fail.
The channel has grown to 4,000+ subscribers and ~500k views, mostly in the past month. Each video costs 20-30 cents in API calls; longer, more in-depth videos can cost a few euros. The error rate is ~1 in 20 videos (e.g., missing a checkmate or a bad tool call), which Steinfurt considers acceptable for automated upload. The team has not yet monetized (below YouTube's threshold) and is still deciding whether to review every video or trust full automation. Trade-offs include cost optimization (redundant tool calls, agent re-analyzing games) vs. quality, and audience targeting (checkmate-in-one is trivial for strong players but valuable for beginners). The system is positioned not to compete with top streamers like GothamChess, but to scale personalized analysis—users could submit their own games and receive custom videos. The German newspaper Süddeutsche Zeitung called the approach 'the holy grail of chess programming,' quoting chess researcher William Weber, because the system explains chess as well as a human trainer—a milestone traditional engines never achieved.
Key Takeaways
- Claim: Gemini 3.1 Pro is the best LLM for chess understanding due to post-training, outperforming previous models and enabling agent-driven tool orchestration. | Evidence: Steinfurt: 'Gemini 3.1 Pro, which recently came out, is actually the best model I've seen so far on chess. I'm pretty sure they've done some in-depth post training on the model. And you can really see in the reasoning traces that it really understands chess a lot better than the previous models.' Before Gemini 3.1, Grok 4 was the best, used in autumn 2023. | Caveat: No public confirmation of post-training from Google; Steinfurt's assessment is based on observed reasoning quality and tool-use behavior, not documented training details. | Implication: For domain-specific agent systems, model selection should prioritize domain understanding (via fine-tuning or post-training) over general benchmarks; reasoning traces are a useful diagnostic for model fit. | Timestamp: timestamp unavailable
- Claim: Shifting from script-driven analysis (Python assembles data, LLM narrates) to agent-driven tool use (LLM reasons and calls tools) dramatically improved quality and flexibility. | Evidence: Steinfurt: 'When we started this whole project, we initially had some Python scripts which would analyze a chess position and would then assemble information... But what changed last year when reasoning models came out was actually that the agents could rather think themselves about the positions. And they were already pretty good.' The agent now decides which tools to call, which variations to explore, and how to explain. | Caveat: Reasoning models are more expensive and can make redundant tool calls (e.g., analyzing the same game twice), which the team has not yet optimized for cost. | Implication: For Ken: reasoning-model-driven agent design can replace hand-coded orchestration logic, but cost/efficiency trade-offs remain; the pattern is to provide tools and let the model decide, rather than pre-compute everything. | Timestamp: timestamp unavailable
- Claim: Providing conflicting or diverse information (best move, human moves, checks, captures, threats) helps the agent explain common human misconceptions and create richer narratives. | Evidence: Steinfurt: 'The checks, captures, and threats tool... we are giving the LLM now in this situation is not only the best move, but we are giving it also access to the checks, captures, and threats... this kind of conflicting information which we provide via the tools is actually very beneficial. So we also have other kind of tools which are more geared towards getting out the more human moves... that also helps to balance the description of not being too much focused on the best moves but also on the most human moves.' Example: in the complex endgame shown, the best move is a non-obvious rook sacrifice, but the agent can discuss why obvious check moves fail. | Caveat: The Maia engine (used for human-like moves) is 'not perfect'; it approximates moves for a given rating but doesn't guarantee human-like reasoning. | Implication: For agent design: deliberately including suboptimal or alternative data can improve explanatory quality and user engagement; don't just feed the agent the 'correct' answer. | Timestamp: timestamp unavailable
- Claim: The system runs nightly, fully automated, with an error rate of ~1 in 20 videos; Steinfurt is moving toward full trust and post-hoc review rather than pre-publication screening. | Evidence: Steinfurt: 'We are automatically creating these videos every night and uploading them to YouTube... I would say every 20th video maybe has a very weird description in which there's, I don't know, a checkmate and that's missed or something like that... I've actually much more now switched to okay, I don't care anymore. And even if there's a bad video, then I'll take it down afterwards maybe.' He initially wanted to review every video but now accepts occasional errors as part of scaling. | Caveat: Errors include missing checkmates or bad tool calls; error diagnosis (e.g., missing tool call at the end) is valuable for iteration but not yet eliminated. No formal error-handling or quality gates are described. | Implication: For AI content systems: a 95% success rate may be acceptable for scale if errors are non-catastrophic and post-hoc fixable; the shift from 100% human review to spot-checking is a key scaling threshold. | Timestamp: timestamp unavailable
- Claim: Each video costs 20-30 cents in API calls; longer, more complex videos cost a few euros; the channel is not yet monetized but has 4,000+ subscribers and 500k views, mostly in the past month. | Evidence: Audience question on cost; Steinfurt: 'It's something on the order of 20, 30 cents... we also have created a couple of other videos which are much, much longer. And there it can also get to euros and something like that.' On monetization: 'Currently we don't make any money yet because we haven't reached the monetization stage yet. And so it's net minus at the moment.' Growth: '500k views... more than 4000 subscribers. Most of them actually were subscribed within the last month.' | Caveat: Cost estimates exclude infrastructure, Lichess data access, hosting, and human review time; only API calls (Gemini, ElevenLabs) are counted. YouTube monetization requires 1,000 subscribers + 4,000 watch hours (or 10M Shorts views in 90 days); they recently crossed the subscriber threshold but watch hours are unclear. | Implication: Unit economics are viable for scale if monetization kicks in; cost per video is low enough to support personalized content (e.g., user-submitted games) as a product; rapid subscriber growth suggests product-market fit. | Timestamp: timestamp unavailable
Detailed Brief
System Architecture: LLM + Chess Engine + Tool Suite
- Claims: The agent uses Gemini 3.1 Pro for reasoning, supplemented by a chess engine, legal move validator, checks-captures-threats (CCT) tool, Maia engine for human-like moves, board state manipulation, and optional web search.; The CCT tool lists all checks, captures, and threats in a position, including obviously bad moves, to give the agent a diverse set of candidate moves to reason about.; The agent decides which moves to highlight, arrows to draw, and which moves to label 'brilliant' or 'blunder'; it also generates the script with emphasis cues (e.g., 'excited') for TTS.; ElevenLabs v3 TTS converts the script to audio; visual annotations (arrows, highlights) are generated programmatically from the agent's structured output.
- Evidence: Steinfurt lists tools: 'legal moves, to just prevent it from ever thinking about something completely illegal... complete chess board... run a chess engine... checks, captures, and threats... web search.'; CCT tool example: in a complex endgame, the tool shows several check moves, some obviously bad (e.g., Qxa4), and the best move (rook sacrifice on e3); the agent uses this diversity to explain why intuitive moves fail.; Agent autonomy: 'the agent also decides by itself which squares it wants to highlight, which arrows it wants to draw and if something should be considered a brilliant move or not.'; TTS: '11labs v3 for text to speech... audio texts where you can put this excited in there and it would then sound excited.'
- Caveats: The team has not yet optimized for redundant tool calls (e.g., agent analyzing the same game twice), which inflates costs.; The Maia engine for human-like moves is 'not perfect' and only approximates play at a given rating.; No detail on video rendering pipeline (how arrows/highlights are composited, frame rate, resolution).
- Implications: The tool suite is a blueprint for multi-domain agent design: domain-specific oracle (chess engine), structured search (CCT), human-behavior model (Maia), and general reasoning (LLM).; The trade-off between cost and quality is managed by accepting some redundancy and not micro-optimizing agent behavior yet.; The system is modular enough to extend to other games, puzzles, or decision-analysis domains (e.g., poker, Go, military strategy).
Reasoning Model Shift: From Script Assembly to Agent Orchestration
- Claims: Initially, Python scripts analyzed positions, assembled data, and passed structured prompts to an LLM for narration; the LLM was a writer, not a thinker.; With reasoning models (O1, Gemini Flash Thinking, Grok 4), the team shifted to letting the agent think and call tools itself, improving quality and flexibility.; Grok 4 was the best model in autumn 2023; Gemini 3.1 Pro is now the best due to post-training on chess.; The agent now reasons about which positions to analyze, which variations to explore, and how to explain common mistakes.
- Evidence: Steinfurt: 'When we started this whole project, we initially had some Python scripts which would analyze a chess position and would then assemble information from various positions... And then we would pass it on to the language model to then have a whole description.'; Shift: 'What changed last year when reasoning models came out was actually that the agents could rather think themselves about the positions. And they were already pretty good. So the best model which we've been using in autumn last year was GROC 4.'; Current state: 'the agents could rather think themselves about the positions... the other models from OpenAI and so on, they also have quite some decent base chess knowledge to be able to then call the tools at the right position and then assemble all the knowledge.'
- Caveats: Reasoning models are more expensive and slower than prompt-based generation.; The team has not published reasoning traces or ablation studies comparing script-driven vs. agent-driven approaches.; No discussion of when script-driven is sufficient (e.g., for simple positions or tactical puzzles).
- Implications: For Ken: the shift from 'assemble data, prompt LLM' to 'give LLM tools, let it reason' is a general pattern applicable to research, analysis, and content generation agents.; Reasoning models unlock richer, more context-sensitive outputs but require accepting higher cost and some inefficiency.; The pattern suggests a two-tier architecture: reasoning model for high-value, complex tasks; cheaper models for repetitive or simple tasks.
Cost, Error Rate, and Scaling Trade-offs
- Claims: Each video costs 20-30 cents; longer videos cost a few euros; the team has not optimized for cost yet.; Error rate is ~1 in 20 videos (missing checkmates, bad tool calls); errors are usually diagnosable and inform tool/prompt improvements.; Steinfurt is moving from pre-publication review to post-hoc takedown of bad videos, accepting occasional errors for scale.; The channel is not yet monetized (below YouTube's threshold) but is growing rapidly (4,000+ subs, 500k views, mostly in the past month).
- Evidence: Audience Q&A: 'What is the cost of the video? It's something on the order of 20, 30 cents... we also have created a couple of other videos which are much, much longer. And there it can also get to euros and something like that.'; Error rate: 'I would say every 20th video maybe has a very weird description in which there's, I don't know, a checkmate and that's missed or something like that... that's usually also then valuable information.'; Trust shift: 'I've actually much more now switched to okay, I don't care anymore. And even if there's a bad video, then I'll take it down afterwards maybe.'; Monetization: 'Currently we don't make any money yet because we haven't reached the monetization stage yet. And so it's net minus at the moment.'
- Caveats: Cost estimate excludes infrastructure, data ingestion, hosting, and human review time.; No detail on whether the 20-30 cent cost includes Gemini API, ElevenLabs TTS, and rendering, or just LLM calls.; YouTube monetization threshold is 1,000 subscribers + 4,000 watch hours (or 10M Shorts views in 90 days); watch hours not disclosed.; No discussion of content policy risk (e.g., AI-generated content labeling, copyright on Lichess games, fair use).
- Implications: Unit economics are viable for scale if monetization is achieved; at 20-30 cents per video, even low CPMs (e.g., $2-5 per 1k views) could be profitable.; The 95% success rate is sufficient for an automated content pipeline; post-hoc review is a pragmatic scaling strategy.; Rapid subscriber growth suggests the content quality is competitive with human-created chess content, validating the AI-native approach.; For Ken: cost per unit (video, report, analysis) is a key metric for AI content systems; accepting a small error rate enables scale, but error diagnosis must be systematic.
Use Case and Market Positioning
- Claims: The system is not competing with top streamers (e.g., GothamChess) but enables personalized, scalable chess analysis for non-elite players.; The input is human games from Lichess; users could submit their own games for custom video analysis (though this is not yet a product).; The team avoids 'slop' aesthetics (e.g., exploding kings on checkmate) and focuses on chess quality and educational value.; Audience targeting is unclear: checkmate-in-one puzzles are trivial for strong players but valuable for beginners; the team is still deciding what to publish.
- Evidence: Steinfurt: 'There are a lot of interesting and great streamers. So for example, Gotham Chess is one, maybe the most well known streamer. But he would probably not describe one of your games or my games in his videos. But we now have a way to also scale for other people who are not the best players in the world.'; Input: 'We basically download chess games from Lichess every night and analyze them... we could be creating videos of your games and you could send the videos then to your friends and family.'; Aesthetics: 'We are not trying to put in these artificial things like exploding positions or exploding kings or something like that on checkmate which might be beneficial for the view count... we are really trying to get the most out of the chess quality.'; Targeting challenge: 'Sometimes we also have videos with a checkmate in one, which for good players is ridiculous... But on the other hand there are real beginners who really don't see that and need an explanation.'
- Caveats: No product or pricing for personalized videos yet; the system is not yet user-facing.; No data on which video types (tactical puzzles, endgames, opening traps) perform best on YouTube.; Audience segmentation (beginner vs. advanced) is manual and not yet data-driven.; No discussion of competition (e.g., chess.com's auto-analysis, Lichess studies, other AI chess channels).
- Implications: The product opportunity is personalized chess education at scale, enabled by AI; this is a template for AI-native SaaS in education, coaching, and entertainment.; The decision to avoid 'slop' aesthetics is a quality/brand trade-off vs. virality; it positions the channel as premium/educational, not clickbait.; For Ken: the challenge of audience segmentation (beginner vs. advanced) is common to AI content systems; data-driven targeting and A/B testing are missing but high-leverage.; The system demonstrates that AI can compete with human creators on quality, not just cost, in niche educational content.
Notable Concepts & Terms
- Checks, Captures, and Threats (CCT): A beginner chess heuristic for candidate move selection; the team built a tool that lists all checks, captures, and threats in a position, including bad moves, to give the agent diverse options to reason about and explain why intuitive moves fail.
- Maia Engine: A chess engine trained by the University of Toronto to predict human-like moves at a given rating (e.g., 1200, 1800); used as a tool to surface moves a human might consider, even if suboptimal, for educational explanations.
- Gemini 3.1 Pro post-training on chess: Steinfurt's hypothesis that Google post-trained Gemini 3.1 Pro on chess data, making it the best LLM for chess understanding (better than O1, Grok 4); evidenced by reasoning traces showing superior chess logic.
- Reasoning model shift: The architectural transition from script-driven (Python assembles data, LLM narrates) to agent-driven (LLM reasons and calls tools itself); enabled by models like O1, Grok 4, Gemini 3.1, and improves quality/flexibility at higher cost.
- Conflicting information strategy: Intentionally providing the agent with diverse or conflicting data (best move, human moves, checks, captures) so it can reason about common misconceptions and create richer explanations, rather than just narrating the optimal line.
- Holy grail of chess programming: A quote from chess researcher William Weber (via German newspaper Süddeutsche Zeitung) describing AI systems that explain chess as well as a human trainer, not just calculate best moves; the team's system is presented as achieving this milestone.
Operator Notes / Why Ken Should Care
- This is a working example of agentic AI generating monetizable, domain-specific content at scale; the architecture (LLM + domain oracle + structured tools + conflicting info) is a template for agent design in research, analysis, and content verticals.
- The shift from script-driven to reasoning-model-driven orchestration is a general pattern for AI systems; it trades cost/efficiency for quality/flexibility and is a key design decision for Ken's agent stack.
- Cost per unit (20-30 cents per video) and error rate (~5%) are critical metrics for scaling AI content; the team's move from pre-publication review to post-hoc takedown is a pragmatic scaling threshold.
- The rapid subscriber growth (4,000+ in the past month) validates that AI-generated content can compete on quality with human creators in niche educational verticals; this has implications for AI-native media, education, and SaaS.
- The product opportunity (personalized chess analysis for non-elite players) is a template for AI-native services: scale what human experts can't economically deliver (e.g., personalized coaching, game analysis, tutoring).
- The decision to avoid 'slop' aesthetics (no exploding kings, focus on chess quality) is a quality/brand trade-off; it suggests that AI content can be premium, not just cheap/viral.
- The team has not yet solved audience segmentation (beginner vs. advanced) or optimized for cost; these are high-leverage next steps and common challenges for AI content systems.
- The system demonstrates that reasoning models (O1, Gemini 3.1) can replace hand-coded orchestration logic, but require accepting higher cost and some inefficiency; the trade-off is context-sensitive.
Watch Map
- 00:00: Intro: title change to 'holy grail of chess programming' based on German newspaper (Süddeutsche Zeitung) quote from William Weber; presentation of AI-generated chess video example.
- 02:00: Two-minute demo video: AI-narrated tactical breakdown of a chess position (rook sacrifice, knight fork, queen trap); shows voice emphasis, visual annotations, variation analysis.
- 04:00: System overview: nightly automated pipeline from Lichess games to YouTube upload; uses Gemini 3.1 Pro + chess engine + tools (legal moves, CCT, Maia, board state, web search).
- 06:00: Checks-Captures-Threats (CCT) tool example: complex endgame position, tool lists all checks/captures/threats (including bad moves), agent uses diversity to explain why intuitive moves fail; best move is non-obvious rook sacrifice.
- 08:00: Reasoning model shift: from Python scripts assembling data + LLM narration to agent-driven tool use; Grok 4 was best in autumn 2023, Gemini 3.1 Pro now best due to post-training.
- 10:00: Conflicting information strategy: providing best move + human moves + checks/captures helps agent explain misconceptions; Maia engine used for human-like moves.
- 12:00: Video creation workflow: agent generates structured format, ElevenLabs v3 TTS for audio (with emphasis cues), agent decides highlights/arrows/labels.
- 14:00: Channel stats: 4,000+ subs, 500k views (mostly past month); not yet monetized; error rate ~1 in 20 videos; Steinfurt moving from pre-review to post-hoc takedown.
- 15:00: Q&A: cost per video 20-30 cents (longer videos a few euros); audience targeting (beginner vs. advanced) still unclear; no user-facing product yet but could enable personalized game analysis.
Source/Metadata
- Title: Running a Chess YouTube Channel entirely by AI — Stephan Steinfurt, TNG
- Transcript words: 2782
- Duration seconds: 991
- Timestamp note: No timestamps in transcript; watch_map entries are approximations based on segment order and 991-second duration.
Transcript
Hello everyone. Yes, Swix wrote a blog post and said, okay, don't write boring titles. So I changed my title again actually and said, okay, we're working on the holy grail of chess programming. And if you knew me, I'm usually not the kind of guy who oversells stuff. So it's actually a quote from someone else about this. Because roughly a week ago there was an article on one of the biggest newspapers in Germany, which was discussed a couple of approaches to, new approaches to doing, combining AI with chess. And they had said it could easily take another five years until AI explains chess as well as a human trainer. William Weber calls it the holy grail of chess programming. And work is already underway in Munich and that's where we are from, right? And we showed them a couple of videos that my bosses mentioned and myself. And they were completely created by our AI engine. And I want to quickly tell you a little bit how that's working. So maybe you don't want to only show it to the newspaper, but also show it to you. There's a two minute video of what the outcome is. So I'm going to start it now. [SPEAKER_02] In this position, Black's queen is currently under fire from the white rook on the H file. [SPEAKER_02] Trying to find safety, Black slides the queen over to G4. [SPEAKER_02] This is a crushing blunder. [SPEAKER_02] It looks like a completely safe square, but moving off the H file allows white to unleash a spectacular mind-bending sacrifice. White plays rook takes H5. White is offering up a full exchange, but this is a masterful trap built on incredible knight geometry. Let's look at what happens if Black takes the bait. First, if Black simply recaptures with the G pawn taking on H5, White springs the trap. The knight jumps into the action with knight to F6 with check. Look at this beautiful octopus knight on F6. It hits the king on G8 while simultaneously skewering that newly placed queen on G4. A lethal fork. The king is forced to step aside to F8, and the knight simply scoops up the queen on G4. But wait. Let's back up. What if Black tries to be clever and avoids the pawn capture? Black could capture the rook with the queen, playing queen takes H5. But it is the exact same trap. The queen is attracted right into the danger zone. By pulling the queen to H5, White set up the exact same trick and plays knight to F6 check anyway. Once again, the king is attacked, and the knight reaches across the board to attack the queen on H5. No matter how Black captures the sacrificed rook, the queen gets forked, because this knight magically controls both G4 and H5. The king must step aside to F8, and the knight captures the queen on H5. This is a gorgeous double duty fork, demonstrating the terrifying hidden power of the knight. Always watch out for these tricky jumping pieces when your king is exposed. [SPEAKER_00] So actually when I prepared the talk yesterday, I wanted to show a different video. [SPEAKER_00] But we are automatically creating these videos every night and uploading them to YouTube. [SPEAKER_00] And when I this morning had a quick look if there's anything embarrassing there, which I could hide from all of you, I thought, okay, actually there's this video, and it's actually even better than the other one. [SPEAKER_00] So I might go for that one. [SPEAKER_00] And as I said, it's automatically created. We basically download chess games from Lichess every night and analyze them in the background, and then let our agent run and analyze it in more depth. From that analysis we then create some special format from which we can later on then create a video. And the video in the end shows variations being explained. It shows brilliant moves, blunders, and it's all automated. So how does it work? Maybe backing up a little bit, what's the general problem? The problem is we've had really good chess engines for multiple decades actually, but they can't really explain chess. On the other hand, we have now LLMs. They can say, have words and describe things, but they can't play chess well. So we have to somehow combine them. That's the main challenge. And we now have built an agent which has a lot of tools, which is basically the main important ingredient here. But what is also important is the LLM which we're using under the hood. So Gemini 3.1 Pro, which recently came out, is actually the best model I've seen so far on chess. I'm pretty sure they've done some in-depth post training on the model. And you can really see in the reasoning traces that it really understands chess a lot better than the previous models. But what we have now put on top of that particular model is a list of tools. So what have you got? We have got a tool for legal moves, to just prevent it from ever thinking about something completely illegal on a chess board, which might otherwise happen. Then we basically give the agent a complete chess board and let it play moves and take them back and go to various variations itself. It can then always run a chess engine and also get some other kind of chess data, for example, looking at checks, captures, and threats. And for some kind of videos we also include web search because then it might be interesting to also describe the historic context of a game or something like that. So looking at one particular tool in detail, there's the checks, captures, and threats tool. That's actually quite well known in chess as a beginner's explanation of what you should be focusing on if you've got a position. So this is a relatively complicated position in which I'm guessing not that many people in the audience would immediately know which one is the best move. I mean, does anyone want to have a guess at it? It's pretty complicated to be honest. It's an endgame. But what we are giving the LLM now in this situation is not only the best move, but we are giving it also access to the checks, captures, and threats. Because otherwise it might miss that. And you can see that there are a couple of check moves on the left board, some of which are obviously wrong. So for example, with a queen taking the bishop on a4 is obviously a bad move. Some other queen moves to give a check, they are quite reasonable. But actually the best move in this whole situation is to sacrifice the rook on e3, which is not completely obvious here, but that's actually the best move. In other situations it might be, however, much better to look at the check, at the capture moves. But what we are giving the LLM now in this situation is not only the best move, but we are giving it also access to the checks, captures, and threats. Because otherwise it might miss that. And you can see that there are a couple of check moves on the left board, some of which are obviously wrong. So for example, with a queen taking the bishop on a4 is obviously a bad move. Some other queen moves to give a check, they are quite reasonable. But actually the best move in this whole situation is to sacrifice the rook on e3, which is not completely obvious here, but that's actually the best move. In other situations it might be, however, much better to look at the check, at the capture moves. And by providing all this via one tool, we give quite some diversity to the agent to then maybe later on explore other kinds of variations. So maybe it wants to check all of these moves and describe which ones are actually bad, because a human might think about them. And so it's not always about the very, very best move, which we have to describe. In general, the big question is who should do the thinking. So when we started this whole project, we initially had some Python scripts which would analyze a chess position and would then assemble information from various positions and would say, okay, here are the check moves and that's the engine evaluation. And then we would pass it on to the language model to then have a whole description. But what changed last year when reasoning models came out was actually that the agents could rather think themselves about the positions. And they were already pretty good. So the best model which we've been using in autumn last year was GROC 4. Surprisingly, that was the best one. But yeah, also the other models from OpenAI and so on, they also have quite some decent base chess knowledge to be able to then call the tools at the right position and then assemble all the knowledge. And yeah, as I already mentioned, this kind of conflicting information which we provide via the tools is actually very beneficial. So we also have other kind of tools which are more geared towards getting out the more human moves and positions. And yeah, that also helps to balance the description of not being too much focused on the best moves but also on the most human moves. Also in the historical context, there might be some valuable information which moves have actually been played or someone might have described something in detail which we might want to analyze. So all these kind of things get into the mix in the context. Yeah, after we did the whole analysis, we would then create it into a special format which we could then easily transfer into a video. We then use 11labs v3 for text to speech. Yeah, it's actually pretty nice that there are now also these audio texts where you can put this excited in there and it would then sound excited. And yeah, and also the agent also decides by itself which squares it wants to highlight, which arrows it wants to draw and if something should be considered a brilliant move or not. Yeah, and now maybe also some other questions. So is that all a slotwatch we are creating? I think not obviously. But yeah, I mean we are able to create many, many, many such videos. So we have to really balance how much do we want to put out there on YouTube. So what is the input? I mean the input is actually human games. So in a sense what we are now positioned at is we could be creating videos of your games and you could send the videos then to your friends and family. And we are not trying to put in these artificial things like exploding positions or exploding kings or something like that on checkmate which might be beneficial for the view count and so on. But we are really trying to get the most out of the chess quality. In general, why are we doing this? I mean there are a lot of interesting and great streamers. So for example, Gotham Chess is one, maybe the most well known streamer. But he would probably not describe one of your games or my games in his videos. But we now have a way to also scale for other people who are not the best players in the world and who might want to have a video. And yeah, we've got a YouTube channel and yeah, currently it's something like 500k views. And yeah, more than 4000 subscribers. Most of them actually were subscribed within the last month. So it's actually going up quite a bit there. Yeah, and that's it. If you've got any more questions, I mean ask them or send me a message via these platforms here, whatever. And yeah, that's it. Yeah. How would it cost? [SPEAKER_01] Are you monetizing YouTube? [SPEAKER_01] Are you covering the element in the world's cost? [SPEAKER_01] Yeah. [SPEAKER_00] Well, currently we don't make any money yet because we haven't reached the monetization stage yet. [SPEAKER_00] And so it's net minus at the moment. [SPEAKER_00] But okay, it might change. [SPEAKER_00] We'll see. [SPEAKER_00] Yeah, I mean, so this whole automation we only enabled a couple of weeks ago. And so currently I'm still a little bit skeptic. Should I really just leave it upload videos? And yeah, I'm mostly still leaning on okay, I still want to watch them first once. But yeah, the error rate is actually pretty low. So I would say every 20th video maybe has a very weird description in which there's, I don't know, a checkmate and that's missed or something like that. But I mean, that's usually also then valuable information, right? Because maybe some tool call was not done at the very end and okay, we can learn something from that. But I've actually much more now switched to okay, I don't care anymore. And even if there's a bad video, then I'll take it down afterwards maybe. Anything else? Yeah. What is the cost of the video? It's something on the order of 20, 30 cents, something like that. So, I mean, we also have created a couple of other videos which are much, much longer. And there it can also get to euros and something like that. Currently, we are not trying to optimize the costs too much because we rather want to err on having a too good description. [SPEAKER_00] And, but there are also a couple of optimization possibilities in there, I don't know, redundant tool calls. Sometimes the agent goes through a game twice or something like that. And yeah, that's obviously stupid in a sense. [SPEAKER_00] Yeah? [SPEAKER_01] You were talking about human move. Is it Grandmaster level move or the human move that you might put in the mix? Yeah, basically all. I mean, so there's this Maia engine which was also mentioned yesterday in the talk. So by the University of Toronto which, yeah, they trained a model in which you can basically put a rating of a player. And then it would roughly give you a move which that player might want to play. But it's not perfect in that sense, right? Sometimes the agent goes through a game twice or something like that. [SPEAKER_00] And yeah, that's obviously stupid in a sense. [SPEAKER_00] Yeah? You were talking about human move. Is it like Grandmaster United move or like the human move that you might put in the mix? Yeah, all. I mean, so there's this Maya engine which was also mentioned yesterday in the talk. So by the University of Toronto which, yeah, they trained a model in which you can put a rating of a player. And then it would roughly give you a move which that player might want to play. But it's not perfect in that sense, right? But it still might be valuable information that maybe some move like this deserves a description. So we don't necessarily need to have what a human would be playing of that strength, but rather a mix of different things to consider. [SPEAKER_00] Yeah. To explain a sub-one learning. Yes, exactly. One in 16, hello, or something like a public page for one year. Yeah. One to explain to you. [SPEAKER_01] Yes, I mean these things could also be used then for targeting a little bit more like for the audience. Rather have videos for better players or videos for worse players. And I mean currently we are also not really sure exactly which videos we should be putting up there or not. I mean sometimes we also have videos with a checkmate in one, which for good players is ridiculous. I mean you see it immediately. But on the other hand there are real beginners who really don't see that and need an explanation why that's a checkmate and so on. [SPEAKER_00] So yeah, it's all a bit of a balancing question which we are not really sure yet. Have you tried on the case? No, not yet. But yeah, I mean it's possible to extend in some, yeah. Okay, yeah, then that's it. [SPEAKER_00] Thanks. Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks name