Chess is a game that has fascinated AI researchers. Alan Turing wrote a chess algorithm on paper in 1951, though the computers of the time were too slow to run it. In 1957 an IBM engineer called Alex Bernstein wrote a chess program for an IBM 704 computer, which ran on vacuum tubes. As silicon chips became faster and cheaper, chess programs developed in strength, until in 1997 IBM’s Deep Blue outplayed world champion Garry Kasparov. However, at this point all these computers used “brute force” algorithms to calculate moves based on explicit rules that were programmed in by humans. Humans encapsulated chess opening and endgame theory into programs and the ability to evaluate middlegame positions. As processors became faster the sheer hardware power meant that even the best human players were unable to compete.
Things got interesting with the development of reinforcement learning, a branch of AI. The company DeepMind (acquired by Google in 2014) was a pioneer in this area, and created a program called AlphaZero in 2017. This was taught only the basic moves of chess, and was left to play itself time after time to teach itself which approaches and strategies worked. At first, its moves were almost random, but as it played more and more it discovered what moves and strategies worked and which did not. After a few million games against itself its performance reached a plateau, and it was pitted against the strongest specialist program at that time, called Stockfish. By that time, chess computers were far stronger than any humans, and competed against other chess programs in tournaments. This was the first time that a specialist chess program had been pitted against a program that learned entirely by self-play. The result? Across 100 games, Stockfish drew 28 games and lost 72 to AlphaZero, which didn’t lose even a single game against the strongest specialist chess program ever written.
At this point you will notice that we have not talked about LLMs, but just about reinforcement learning as an AI approach to chess (and Go). When ChatGPT was released to the public in November 2022, some people were curious about its chess abilities. It turned out that large language models (LLMs) like GPT are excellent at pattern recognition, and have seen millions of chess games in their training data. However, they are language models, predicting text one token (a short word) at a time. They are not calculating chess positions and evaluating them as a human would do. Consequently, it turned out that LLMs were embarrassingly bad at chess beyond the first few moves, where they are able to play plausible moves because they have seen many real chess games in their training data. Unlike LLMs, engines such as Stockfish repeatedly search millions of legal continuations before selecting a move, maintaining an exact representation of the board throughout .
In chess, there is considerable theory about the opening moves, so many professional-level games have similar initial moves. Depending on the openings chosen by the players, this theory stage of the game, which humans memorise with varying degrees of success, can go on for a few moves, or in some cases many moves. Some grandmaster opening lines extend as far as move 30 or more in certain contested positions, though that is rare. Most common chess openings get to around moves 10-20 before the opening theory runs out and players have to think for themselves instead of relying on memory.
LLMs can mimic the opening moves of chess well, but quickly end up in lines that they have not seen in their training data . This is a problem as they do not maintain an internal representation of how to play chess at all: they are merely good at mimicking the patterns found in a game. Once in a chess position that they have not seen, they start to hallucinate moves, inventing pieces that are not on the board or making illegal moves with those pieces that actually are on the board .
In August 2025 an organisation called Kaggle set up an LLM chess championship, pitting different models against each other . Aware of the issue of hallucination, each model was allowed up to three retries to play a legal move at each turn (for a total of four attempts) . If it hallucinated and failed to submit a legal move, it lost the game . Note that this pitifully low bar would easily be met by almost any young child that had been taught the rules of chess. A human who knows the rules of chess would virtually never generate an illegal move, even if they overlooked the best continuation. In the tournament, eight LLMs played, including OpenAI’s o4-mini, Google’s Gemini 2.5 Pro, Anthropic’s Claude 4 Opus, Grok 4, DeepSeek R1 and Moonshot’s Kimi K2 .
The resulting games saw Kimi K2 regularly disqualified due to too many illegal moves, though other models managed rather better. The final was a match between Grok 4 and OpenAI o3, which o3 won by the margin of 4 games to 0 . The standard of play was, er, not good, with some quite basic tactical errors. The final saw an amused Magnus Carlsen (the strongest human player) trying to politely commentate on the games, with Grok blundering a piece on move 7 followed by further blunders . The models showed their textual “reasoning”, which appeared to be rationalisations rather than faithful descriptions of how the moves were generated.
So, how have LLMs moved on since then? There have been numerous new LLM models and so in July 2026 I spent some time playing against the four best known models, namely Grok, ChatGPT, Gemini and Claude. For context, I am a decent club level chess player, but far below the level of a chess international master or grandmaster. How did I get on?
It turned out that LLMs still hallucinate moves, though probably less than they did a year ago when I first tried playing with them. Claude played very weakly and lost quickly, as did Grok. ChatGPT played reasonably for a while until it started hallucinating and Gemini did best of all. It actually managed 37 moves with hardly any hallucinations until it resigned in a hopeless position. All of them managed to play respectable moves for at least the first 7-10 moves, until they started to run out of opening theory from their training data . At this point the play deteriorated, in some cases quite radically. At present all four models that I tried played at a level of a weak (in some cases very weak) club player. There is a huge gulf between their chess ability and that of a strong human player, never mind a top human player.
It is interesting to see how LLMs have indeed become better at chess than they were a year ago, but in a way that parallels their development in other areas. Hallucination rates of LLMs have, in general, improved for simple factual checks over the years but have actually worsened for more complex reasoning. Their hallucinations are, though, harder to spot as their invented answers are increasingly plausible. In the case of chess, they have also become better at mimicking how to play, but without actually reasoning through positions in the way that a human player would. This is a reminder that LLMs are very plausible and lucid, but are not thinking in the way that a human would. It is also perhaps an indication that we should be wary of putting too much faith in the logical reasoning ability of models that cannot actually master a simple board game. If they struggle with chess, we should be wary of putting too much faith in their reasoning abilities in more complex situations. They are good at pattern recognition and are plausible liars in terms of making up answers. However, they need rigorous checking if placed in situations where complex judgement is required.
Chess demonstrates that fluent language should not be confused with maintaining a consistent internal model of a problem. Modern LLMs can produce convincing explanations of why a move is good while simultaneously proposing an illegal move. That disconnect is a useful reminder that eloquence is not the same as reliable reasoning. As models improve, they may hallucinate less often, but they will also become increasingly persuasive when they are wrong, making independent verification more, not less, important.
Appendix: The Games
To understand this section fully, you need to know how to follow a chess game, which is written using a standard notation. You can easily enter these moves onto a chess board if you wish to follow the games. In all cases I played white, the games happening on 3rd July 2026.
In summary:
| Model | Legal Moves | Illegal Moves | Playing Strength |
|---|---|---|---|
| Gemini | 36 | few | weak club player |
| ChatGPT | 16 | several | weak club player |
| Grok | 13 | many | novice |
| Claude | 11 | none | Beginner |
Here are the actual games. In each case I had the white pieces.
Grok
1. e4 e5
2. Nf3 Nc6
3. d4 exd4
4. Nxd4 Bc5
5. Nb3 Bb6
6. Qe2 Nf6
7. Nc3 O-O
8. Be3 Re8
9. f3 d5
10. O-O-O dxe4? (losing the queen)
11. Rxd8 Rxd8
12. Bxb6 axb6
13. fxe4 Nxe4?
14. Qxe4 Bf5?
15. Qxf5 …
Now Grok tried an illegal move 15. … Nxc3. After the illegal move was pointed out it corrected to 14. … Nd4 to which I replied 15. Nxd4 … Grok replied 15. … Rxd4 and I played 16. Bd3 … threatening h7. Instead of defending against this directly it played 16. … Rxd3 17. Qxd3 Re8. 18. Rd1 Re4? Instead of capturing the rook with either queen or knight I played 19. Qd8+ .. which is mate in two after the forced 19. … Re8 20. Qxe8 … mate. Grok tried another illegal move here with 20. …Kf8
ChatGPT
1. e4 c5
2. Nc3 Nc6
3. f4 g6
4. Nf3 Bg7
5. Bc4 e6
6. f5 Nge7
7. fxe6 dxe6
8. O-O O-O
9. d3 b6
10. Qe1 Bb7
11. Qh4 Nd4
12. Bg5 f6
13. Bh6 Bxh6
14. Qxh6 Nxc2
15. Bxe6+ Kh8
16. Rad1 …
All fine so far and I was quite impressed until its next move “16. … Nxd1”. This is illegal as its knights are on c2 and e7. After some “good catch” comments from the model, it played 16. … Qc7 to which I replied 17. Ng5 … exploiting a pin on the f file. Black is best to play 17. … Ng8 but instead blundered with 17. … gxf5, which loses immediately to 18. Rxf8 …. Black is best to play 18. … Ng8 with a losing position, as recapturing on f8 with the queen is mate. However, ChatGPT tried the very surprising “18. … Qxf8” which is impossible as its queen is on c7. After some dialog it eventually tried “18. … Kh8”, which makes no sense as its king is already there.
So ChatGPT managed 16 perfectly decent moves but immediately fell apart. The first ten moves are in the Chess Opening Explorer database so are “book” lines. As soon as it is out of established chess opening theory, ChatGPT could only manage six more moves before hallucinating, and being unable to recover for more than one move.
Gemini
1. e4 e5
2. Nf3 Nc6
3. d4 exd4
4. Nxd4 Bc5
5. Nb3 Bb6
6. Qe2 d6
7. Be3 Nf6
8. Nc3 O-O
9. O-O-O Re8
10. f3 Bxe3
11. Qxe3 a5
12. a3 a4
13. Nd4 Nxd4
14. Rxd4 Bd7
15. g4 c5? (losing a pawn)
16. Rxd6 Qe7
17. Qxc5 Rac8
18. Qb4 Bc6
19. Be2 Red8
20. Rhd1 Rxd6
21. Qxd6 Qxd6
22. Rxd6 Ne8
23. Rd4 Nc7
24. Nxa4 Bxa4
25. Rxa4 Ne6
26. Ra7 Rb8
27. c3 Nc5
28. Bc4 Kf8
29. Kc2 Ke7
30. Bd5 Kd6
31. b4 f6?
32. bxc5+ Kxc5
33. Rxb7 Rxb7
34. Bxb7 g5
35. Kd3 h5
36. gxh5 …
At this point Gemini hallucinated with recapturing on h5, but eventually settled on 36. … f5 37. exf5 Kd6 (after a couple more illegal move attempts). Then 38. h6 Ke7 39. h7 Kf7 40. h8=Q … and Gemini resigned.
Gemini actually played a perfectly reasonable game right up until move 36, when it started hallucinating. Sure, it did not play particularly well, being worse by move 15 and effectively lost by move 22, but at least it played a game with no illegal moves until into the endgame.
Claude
1. e4 e5
2. Nf3 Nc6
3. d4 exd4
4. Nxd4 Nf6
5. Nxc6 bxc6
6. e5 Qe7
7. Qe2 Nd5
8. c4 Nf4? (losing a piece)
9. Bxf4 Qxe5? (losing the queen)
10. Qxe5+ Bg7
11. Qxg7 Bf6? (losing the bishop)
12. Qxf6 … and Claude resigned
This was not an advert for LLM chess ability.







