Prompting An LLM For Chess: What Works And What Hallucinates
Ask ChatGPT for a move and you will get a move. It will arrive with confidence, a plausible-sounding justification, and roughly a coin flip’s chance of being legal in the position you described. Ask the same model to explain a Stockfish line you paste in, and something different happens: you get a coaching session that is often better than what a 2200 club player could give you in the same five minutes.
That gap is the whole story. The useful chatgpt chess prompts are not the ones that ask for chess ability. They are the ones that ask for language ability applied to chess data you supply.
The failure mode, with actual output
Here is a position from a Lichess rapid game, Black to move, White having just played 15.Ne5:
r2q1rk1/pp2bppp/2n1pn2/3pN3/3P4/2N1P3/PP3PPP/R1BQ1RK1 b - - 3 15
I asked GPT-5 and Claude Opus 4.5 the same thing: “Black to move. What’s the best move and why?” No engine, no extra scaffolding.
GPT-5 suggested …Nxe5, called it “the cleanest way to relieve the pressure,” then explained that after 16.dxe5 the knight on f6 is “attacked twice.” It is attacked zero times. Claude offered …Bd6, describing it as “activating the bishop toward the kingside while defending e5.” The bishop on e7 cannot reach d6 in one move and there is nothing of Black’s on e5 to defend.
Both answers read beautifully. Neither is wrong in a way a 1400 spots instantly, because the prose is fluent and the reasoning template is correct. That is what makes it dangerous. Run the same prompt eight or ten times on middlegame positions with six or more pieces per side and you should expect two or three answers containing an illegal move or a flatly false tactical claim.
Why? The model has no board. It has a string of characters that it is pattern-matching against millions of PGNs it has read. Openings work fine, because the first ten moves of the Najdorf appear in the training data verbatim thousands of times. By move 25 of your particular game, your position has quite possibly never existed before, and the model is improvising notation that looks like chess.
The prompt that works
Swap the model’s job. Stockfish generates, the LLM translates.
Open your Lichess analysis board, let Stockfish 17 run to depth 22 or so, and copy the top three lines. Then:
You are a chess coach. I am 1450 Lichess rapid. Do NOT suggest
moves of your own, do NOT evaluate the position yourself, and do
NOT extend any line beyond what I give you.
POSITION (FEN): r2q1rk1/pp2bppp/2n1pn2/3pN3/3P4/2N1P3/PP3PPP/R1BQ1RK1 b - - 3 15
I played: 15...Nd7
STOCKFISH 17, depth 24:
1. -0.3 15...Nxe5 16.dxe5 Nd7 17.f4 c5
2. -0.5 15...Nxd4 16.exd4 Bd6 17.Bf4
3. -1.4 15...Nd7 16.Nxd7 Qxd7 17.Bd2 Rfd8
For line 1 only, answer these in order:
1. What does 16.dxe5 change about the pawn structure?
2. Why does Stockfish want the knight on d7 rather than e8?
3. What is 17.f4 preparing, in one sentence?
4. My move scored 1.1 pawns worse. Name the single concrete
thing my move gives away. Do not list more than one.
5. What pattern should I drill so I find this next time?
The model cannot hallucinate a move because it has been forbidden from producing moves. Every square it names is a square you handed it. The answers to questions 1 through 3 are explanations of structure and plan, which is exactly what language models are good at and what an engine’s numeric output cannot give you.
The answer I got to question 4 on this position: “You left the e5 knight alive. Stockfish’s line trades it off; yours invites 16.Nxd7 on White’s terms, and after 16…Qxd7 your d5 pawn is a fixed target with no counterplay on the c-file.” That is coaching. It is correct, it is specific, and I could act on it that afternoon.
Constraints that do the heavy lifting
Six rules, ranked by how much hallucination each one kills.
Forbid move generation explicitly. “Do not suggest moves of your own” is worth more than every other instruction combined. Models comply with negative constraints on output format far more reliably than they comply with requests to be accurate.
Paste the FEN, not a move list. “In my game I played 1.e4 e5 2.Nf3…” forces the model to replay 30 moves in its head, and errors compound. A FEN is one string describing one position. Lichess gives you it under the board, or press f in the analysis view.
Cap the ply depth. “Do not extend any line beyond what I give you” stops the single most common leak, which is the model helpfully continuing 17.f4 c5 18.Nf3?? into fiction.
Give your rating and ask for one thing. A model told “I’m 1450” and “name the single concrete thing” produces one actionable sentence. A model given no rating produces a paragraph mentioning Nimzowitsch, prophylaxis, and four candidate plans you will remember none of.
Number your questions. Unnumbered multi-part asks get answered in a blur. Numbered ones get answered in order, and you notice when one is skipped.
Include the eval numbers. The model uses -0.3 versus -1.4 as an anchor for how emphatic to be. Strip the numbers and it treats all lines as equally interesting.
Reading the eval so the prompt is honest
Your prompt is only as good as the engine output you feed it, and club players routinely feed the model garbage.
Depth matters more than most people assume. Lichess’s browser Stockfish at depth 18 will sometimes flip its top move entirely by depth 26. Before you paste anything, let it settle: on a modern laptop, depth 24 on a middlegame position takes maybe 8 to 20 seconds. If the top move is still changing, you are asking the LLM to explain noise.
Centipawn scale, in terms you can act on:
| Eval | What it means at club level |
|---|---|
| 0.00 to ±0.30 | Nothing. Do not build a lesson on this. |
| ±0.30 to ±0.80 | A real but survivable edge. Worth understanding. |
| ±0.80 to ±1.50 | One clear concession. This is the sweet spot for prompting. |
| ±1.50 to ±3.00 | A serious error, usually a tactic you missed. |
| beyond ±3.00 | You dropped a piece. You do not need an LLM for this. |
Blunders over 3.00 are self-explanatory and waste your prompts. The 0.8 to 1.5 band is where you actually lose rating points, and it is precisely the band where you cannot tell what went wrong without help. That is the band to mine.
One more thing: an engine’s “best move” at 3000+ Elo is frequently a move you should not play. Stockfish will happily recommend sitting in a defensive crouch for eleven moves. When the top line requires precision you do not have, ask the model directly: “Of these three lines, which is most playable for a 1450 who will be on a 10-minute clock?” It will usually point at line 2, and it will usually be right, because playability is a linguistic judgement about complexity rather than a calculation.
Turning this into a week of training
Three prompts, run on real games, produce a plan.
Mine your own losses first. Export your last 20 Lichess rapid games, run each through the built-in analysis, and list every position where you dropped more than 0.8 centipawns. Paste that list of FENs with this: “Here are 14 positions where I lost more than 0.8 eval. Do not analyse any position. Group them by what they have in common. Give me at most 4 groups, named.”
I ran this on my own games in April and got back four buckets: bishop trades that left me with a bad knight (5 positions), premature pawn pushes in front of my own king (4), missed intermediate checks (3), and rook placement on closed files (2). That is a curriculum, and I did not have to notice the pattern myself.
Then convert the biggest bucket into drills: “Bucket 1 is bishop trades leaving me a bad knight. Name three specific Lichess Study or Chessable patterns that drill this, and write me one sentence I can say to myself at the board before trading a bishop.” The tool names it returns are worth checking, since models invent course titles, but the one-sentence heuristic is reliably good and costs nothing to verify.
Finally, use the model as a rubber duck before you look at the engine at all. Paste the FEN, say “I am considering …Nd7 and …Nxe5. Ask me four questions about this position. Do not answer them and do not tell me which move is better.” Four sharp questions about a position you are staring at is genuinely close to what a coach does, and it is the one use where the model’s chess weakness does not matter, because you are the one calculating.
For a broader comparison of which tools are worth paying for and how they stack up against a bare LLM plus an engine, the pillar piece at AI coaching tools and LLM chess prompts, tested runs the numbers.
What to do tomorrow
Pick your three worst rapid losses from the past fortnight. Find the moment in each where the eval swung between 0.8 and 1.5. Run the five-question prompt above on each of them, with the move-generation ban intact.
You will finish in about twenty minutes with three named weaknesses, in your own words, from your own games. Nothing the model produced along the way required it to know how a knight moves, which is exactly why you can trust it.