SPEAKER_00
Hello everyone. Yes, Swix wrote a blog post and said, okay, don't write boring titles. So I changed my title again actually and said, okay, we're working on the holy grail of chess programming. And if you knew me, I'm usually not the kind of guy who oversells stuff. So it's actually a quote from someone else about this. Because roughly a week ago there was an article on one of the biggest newspapers in Germany, which was discussed a couple of approaches to, new approaches to doing, combining AI with chess. And they had said it could easily take another five years until AI explains chess as well as a human trainer.
SPEAKER_00
William Weber calls it the holy grail of chess programming. And work is already underway in Munich and that's where we are from, right? And we showed them a couple of videos that my bosses mentioned and myself. And they were completely created by our AI engine. And I want to quickly tell you a little bit how that's working. So maybe you don't want to only show it to the newspaper, but also show it to you. There's a two minute video of what the outcome is. So I'm going to start it now. [SPEAKER_02] In this position, Black's queen is currently under fire from the white rook on the H file. [SPEAKER_02] Trying to find safety, Black slides the queen over to G4.
SPEAKER_02
[SPEAKER_02] This is a crushing blunder. [SPEAKER_02] It looks like a completely safe square, but moving off the H file allows white to unleash a spectacular mind-bending sacrifice. White plays rook takes H5. White is offering up a full exchange, but this is a masterful trap built on incredible knight geometry. Let's look at what happens if Black takes the bait. First, if Black simply recaptures with the G pawn taking on H5, White springs the trap. The knight jumps into the action with knight to F6 with check. Look at this beautiful octopus knight on F6. It hits the king on G8 while simultaneously skewering that newly placed queen on G4. A lethal fork.
SPEAKER_02
The king is forced to step aside to F8, and the knight simply scoops up the queen on G4. But wait. Let's back up. What if Black tries to be clever and avoids the pawn capture? Black could capture the rook with the queen, playing queen takes H5. But it is the exact same trap. The queen is attracted right into the danger zone. By pulling the queen to H5, White set up the exact same trick and plays knight to F6 check anyway. Once again, the king is attacked, and the knight reaches across the board to attack the queen on H5. No matter how Black captures the sacrificed rook, the queen gets forked, because this knight magically controls both G4 and H5.
SPEAKER_02
The king must step aside to F8, and the knight captures the queen on H5. This is a gorgeous double duty fork, demonstrating the terrifying hidden power of the knight. Always watch out for these tricky jumping pieces when your king is exposed. [SPEAKER_00] So actually when I prepared the talk yesterday, I wanted to show a different video. [SPEAKER_00] But we are automatically creating these videos every night and uploading them to YouTube. [SPEAKER_00] And when I this morning had a quick look if there's anything embarrassing there, which I could hide from all of you, I thought, okay, actually there's this video, and it's actually even better than the other one.
SPEAKER_02
[SPEAKER_00] So I might go for that one. [SPEAKER_00] And as I said, it's automatically created.
SPEAKER_00
We basically download chess games from Lichess every night and analyze them in the background, and then let our agent run and analyze it in more depth. From that analysis we then create some special format from which we can later on then create a video. And the video in the end shows variations being explained. It shows brilliant moves, blunders, and it's all automated. So how does it work? Maybe backing up a little bit, what's the general problem? The problem is we've had really good chess engines for multiple decades actually, but they can't really explain chess. On the other hand, we have now LLMs.
SPEAKER_00
They can say, have words and describe things, but they can't play chess well. So we have to somehow combine them. That's the main challenge. And we now have built an agent which has a lot of tools, which is basically the main important ingredient here. But what is also important is the LLM which we're using under the hood. So Gemini 3.1 Pro, which recently came out, is actually the best model I've seen so far on chess. I'm pretty sure they've done some in-depth post training on the model. And you can really see in the reasoning traces that it really understands chess a lot better than the previous models.
SPEAKER_00
But what we have now put on top of that particular model is a list of tools. So what have you got? We have got a tool for legal moves, to just prevent it from ever thinking about something completely illegal on a chess board, which might otherwise happen. Then we basically give the agent a complete chess board and let it play moves and take them back and go to various variations itself. It can then always run a chess engine and also get some other kind of chess data, for example, looking at checks, captures, and threats.
SPEAKER_00
And for some kind of videos we also include web search because then it might be interesting to also describe the historic context of a game or something like that. So looking at one particular tool in detail, there's the checks, captures, and threats tool. That's actually quite well known in chess as a beginner's explanation of what you should be focusing on if you've got a position. So this is a relatively complicated position in which I'm guessing not that many people in the audience would immediately know which one is the best move. I mean, does anyone want to have a guess at it? It's pretty complicated to be honest. It's an endgame.
SPEAKER_00
But what we are giving the LLM now in this situation is not only the best move, but we are giving it also access to the checks, captures, and threats. Because otherwise it might miss that. And you can see that there are a couple of check moves on the left board, some of which are obviously wrong. So for example, with a queen taking the bishop on a4 is obviously a bad move. Some other queen moves to give a check, they are quite reasonable. But actually the best move in this whole situation is to sacrifice the rook on e3, which is not completely obvious here, but that's actually the best move.
SPEAKER_00
In other situations it might be, however, much better to look at the check, at the capture moves. But what we are giving the LLM now in this situation is not only the best move, but we are giving it also access to the checks, captures, and threats. Because otherwise it might miss that. And you can see that there are a couple of check moves on the left board, some of which are obviously wrong. So for example, with a queen taking the bishop on a4 is obviously a bad move. Some other queen moves to give a check, they are quite reasonable.
SPEAKER_00
But actually the best move in this whole situation is to sacrifice the rook on e3, which is not completely obvious here, but that's actually the best move. In other situations it might be, however, much better to look at the check, at the capture moves. And by providing all this via one tool, we give quite some diversity to the agent to then maybe later on explore other kinds of variations. So maybe it wants to check all of these moves and describe which ones are actually bad, because a human might think about them. And so it's not always about the very, very best move, which we have to describe. In general, the big question is who should do the thinking.
SPEAKER_00
So when we started this whole project, we initially had some Python scripts which would analyze a chess position and would then assemble information from various positions and would say, okay, here are the check moves and that's the engine evaluation. And then we would pass it on to the language model to then have a whole description.
SPEAKER_00
But what changed last year when reasoning models came out was actually that the agents could rather think themselves about the positions. And they were already pretty good. So the best model which we've been using in autumn last year was GROC 4. Surprisingly, that was the best one. But yeah, also the other models from OpenAI and so on, they also have quite some decent base chess knowledge to be able to then call the tools at the right position and then assemble all the knowledge. And yeah, as I already mentioned, this kind of conflicting information which we provide via the tools is actually very beneficial.
SPEAKER_00
So we also have other kind of tools which are more geared towards getting out the more human moves and positions. And yeah, that also helps to balance the description of not being too much focused on the best moves but also on the most human moves. Also in the historical context, there might be some valuable information which moves have actually been played or someone might have described something in detail which we might want to analyze. So all these kind of things get into the mix in the context. Yeah, after we did the whole analysis, we would then create it into a special format which we could then easily transfer into a video.
SPEAKER_00
We then use 11labs v3 for text to speech. Yeah, it's actually pretty nice that there are now also these audio texts where you can put this excited in there and it would then sound excited. And yeah, and also the agent also decides by itself which squares it wants to highlight, which arrows it wants to draw and if something should be considered a brilliant move or not. Yeah, and now maybe also some other questions. So is that all a slotwatch we are creating? I think not obviously. But yeah, I mean we are able to create many, many, many such videos. So we have to really balance how much do we want to put out there on YouTube. So what is the input?
SPEAKER_00
I mean the input is actually human games. So in a sense what we are now positioned at is we could be creating videos of your games and you could send the videos then to your friends and family. And we are not trying to put in these artificial things like exploding positions or exploding kings or something like that on checkmate which might be beneficial for the view count and so on. But we are really trying to get the most out of the chess quality. In general, why are we doing this? I mean there are a lot of interesting and great streamers. So for example, Gotham Chess is one, maybe the most well known streamer.
SPEAKER_00
But he would probably not describe one of your games or my games in his videos. But we now have a way to also scale for other people who are not the best players in the world and who might want to have a video. And yeah, we've got a YouTube channel and yeah, currently it's something like 500k views. And yeah, more than 4000 subscribers. Most of them actually were subscribed within the last month. So it's actually going up quite a bit there. Yeah, and that's it. If you've got any more questions, I mean ask them or send me a message via these platforms here, whatever. And yeah, that's it. Yeah. How would it cost? [SPEAKER_01] Are you monetizing YouTube?
SPEAKER_00
[SPEAKER_01] Are you covering the element in the world's cost?
[SPEAKER_01] Yeah. [SPEAKER_00] Well, currently we don't make any money yet because we haven't reached the monetization stage yet. [SPEAKER_00] And so it's net minus at the moment. [SPEAKER_00] But okay, it might change. [SPEAKER_00] We'll see. [SPEAKER_00] Yeah, I mean, so this whole automation we only enabled a couple of weeks ago. And so currently I'm still a little bit skeptic. Should I really just leave it upload videos? And yeah, I'm mostly still leaning on okay, I still want to watch them first once.
SPEAKER_00
But yeah, the error rate is actually pretty low. So I would say every 20th video maybe has a very weird description in which there's, I don't know, a checkmate and that's missed or something like that. But I mean, that's usually also then valuable information, right? Because maybe some tool call was not done at the very end and okay, we can learn something from that. But I've actually much more now switched to okay, I don't care anymore. And even if there's a bad video, then I'll take it down afterwards maybe. Anything else? Yeah. What is the cost of the video? It's something on the order of 20, 30 cents, something like that.
SPEAKER_00
So, I mean, we also have created a couple of other videos which are much, much longer. And there it can also get to euros and something like that. Currently, we are not trying to optimize the costs too much because we rather want to err on having a too good description.
SPEAKER_00
[SPEAKER_00] And, but there are also a couple of optimization possibilities in there, I don't know, redundant tool calls. Sometimes the agent goes through a game twice or something like that. And yeah, that's obviously stupid in a sense. [SPEAKER_00] Yeah? [SPEAKER_01] You were talking about human move. Is it Grandmaster level move or the human move that you might put in the mix? Yeah, basically all. I mean, so there's this Maia engine which was also mentioned yesterday in the talk. So by the University of Toronto which, yeah, they trained a model in which you can basically put a rating of a player.
SPEAKER_00
And then it would roughly give you a move which that player might want to play. But it's not perfect in that sense, right? Sometimes the agent goes through a game twice or something like that. [SPEAKER_00] And yeah, that's obviously stupid in a sense. [SPEAKER_00] Yeah? You were talking about human move. Is it like Grandmaster United move or like the human move that you might put in the mix? Yeah, all. I mean, so there's this Maya engine which was also mentioned yesterday in the talk. So by the University of Toronto which, yeah, they trained a model in which you can put a rating of a player. And then it would roughly give you a move which that player might want to play.
SPEAKER_00
But it's not perfect in that sense, right? But it still might be valuable information that maybe some move like this deserves a description. So we don't necessarily need to have what a human would be playing of that strength, but rather a mix of different things to consider.
SPEAKER_01
[SPEAKER_00] Yeah.
SPEAKER_00
To explain a sub-one learning.
SPEAKER_01
Yes, exactly.
SPEAKER_00
One in 16, hello, or something like a public page for one year.
SPEAKER_01
Yeah.
SPEAKER_00
One to explain to you. [SPEAKER_01] Yes, I mean these things could also be used then for targeting a little bit more like for the audience. Rather have videos for better players or videos for worse players. And I mean currently we are also not really sure exactly which videos we should be putting up there or not. I mean sometimes we also have videos with a checkmate in one, which for good players is ridiculous. I mean you see it immediately. But on the other hand there are real beginners who really don't see that and need an explanation why that's a checkmate and so on. [SPEAKER_00] So yeah, it's all a bit of a balancing question which we are not really sure yet.
SPEAKER_00
Have you tried on the case?
SPEAKER_00
No, not yet. But yeah, I mean it's possible to extend in some, yeah. Okay, yeah, then that's it. [SPEAKER_00] Thanks. Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks Thanks name