Feeding PGNs To An AI Coach: Formats, Context Limits And Accuracy
Paste a 40-move PGN into a chat model and ask “where did I go wrong?” and you will get an answer. It will be fluent, structured, encouraging, and somewhere between partly and entirely fictional. The model will tell you that your knight on f5 was loose when the knight has been on d4 since move 19. It will suggest a tactic that requires your queen to be on a file she left eleven moves ago. It will describe a pawn structure that never existed.
This is not a reasoning failure. It is a bookkeeping failure, and it is fixable in about ninety seconds of pre-processing.
Why raw PGN is the worst possible input
A PGN is a compressed instruction list, not a position. 1. e4 e5 2. Nf3 Nc6 3. Bb5 a6 4. Ba4 Nf6 tells you what happened only if you maintain a board state in your head while reading it. That is exactly what a language model cannot reliably do. Every move is a delta applied to an invisible 64-square array, and the model has to reconstruct that array from scratch, in one pass, with no scratchpad, while also answering your question about it.
Here is the failure mode in its purest form. Give a model a legal PGN, then ask it, without any other context: “list every piece on the board after move 24, with squares.” Do this ten times across a handful of games. You will see errors in the majority of attempts, and the errors cluster in a specific way: pieces that moved more than three times get stranded on an earlier square, captured pieces come back to life, and pawn structures drift toward whatever structure the opening usually produces rather than the one your game actually produced. A Najdorf that transposed into something unusual will get analysed as a textbook Najdorf, because the model is pattern-matching the opening name against its training data instead of tracking your actual moves.
The tell is confident specificity about things that are wrong. Hallucination in chess analysis does not look like hedging. It looks like “your bishop on g5 is doing nothing there, consider Bd2 to reroute it,” delivered in the exact register of a coach who has seen the board.
The fix: stop asking it to track the board
The pre-processing principle is simple. Never make the model infer a position it could be told. You supply the board state at every moment you want discussed, and the model’s job shrinks from “simulate chess” to “explain a position you can see.” That second job it is genuinely good at.
Three ingredients, in order of how much they help:
- FENs at the moments that matter. A FEN is the complete board state in one line. No reconstruction needed.
- Engine evaluations in centipawns, with the engine’s top line. This anchors the judgment, which models are worse at than they appear.
- Eval deltas, so the interesting moments are marked. You want the model looking where the game actually turned, not where move numbers happen to be round.
Getting the raw material out of Lichess
Lichess does the heavy lifting free. Import your game or open it from your archive, run Request Computer Analysis (Stockfish 16 NNUE, server-side, usually under twenty seconds for a club game), and you get a per-move eval graph plus mistake/blunder tags.
For the FENs, the browser console on the analysis board is the fastest route, but the cleaner path is the API. Lichess exports a PGN with embedded evals and clock times:
https://lichess.org/game/export/{gameId}?evals=true&clocks=false&literate=true
That returns something like:
[Event "Rated Blitz game"]
[White "yourhandle"]
[Black "opponent"]
[Result "0-1"]
[WhiteElo "1612"]
[BlackElo "1644"]
[Opening "Sicilian Defense: Najdorf Variation"]
1. e4 { [%eval 0.17] } c5 { [%eval 0.19] } 2. Nf3 { [%eval 0.0] }
d6 { [%eval 0.26] } 3. d4 { [%eval 0.0] } cxd4 { [%eval 0.15] }
...
18. Nd5?? { [%eval -3.41] } exd5 { [%eval -3.28] }
Those [%eval] tags are your gold. They are also, in raw form, 40 to 80 extra tokens per move pair, which matters in a moment.
If you are a Chess.com player, the Game Report gives you an accuracy score and classified moves, but extracting machine-readable evals is more awkward. The practical move is to copy the PGN out of Chess.com and import it into Lichess, then use the Lichess export. You lose nothing and gain a clean data format.
For FENs at specific plies, python-chess is eight lines:
import chess.pgn
with open("game.pgn") as f:
game = chess.pgn.read_game(f)
board = game.board()
for ply, move in enumerate(game.mainline_moves(), start=1):
board.push(move)
if ply in (29, 35, 36, 47): # plies you care about
print(f"ply {ply}: {board.fen()}")
Context limits, honestly
A 40-move game as bare PGN is roughly 300 to 400 tokens. Trivial. With [%eval] tags on every move it balloons to around 2,000 to 2,500. Add a FEN per move (72 characters average, call it 28 tokens) and you are at 4,500 or so. All of that fits in any current context window several hundred times over, so the limit is not the window.
The limit is attention. Dumping 82 FENs into a prompt reliably makes the analysis worse, not better, because the model now has to decide which of 82 positions is relevant and it will spread itself thin across all of them. I have watched this produce an eight-section response where every section is accurate and none of it is useful.
The number that works is four to six FENs per game. Fewer than three and you are back to asking the model to infer. More than eight and the coaching goes generic. Pick them by eval delta: any move where the evaluation swings by more than 100 centipawns (one full pawn) gets a FEN, plus one FEN at the moment you personally felt lost, because that subjective moment is often not where the engine sees the problem, and the gap between those two is the most instructive thing in the whole exercise.
A worked example
Here is the shape of a prompt that actually works. Game is a 1612-rated Black in a Najdorf, lost in 41 moves.
You are coaching a 1600 Lichess player. Below is one of their games
with engine data. Do not reconstruct the board yourself: use the FENs
given. If you need a position I have not supplied, say so.
PLAYER: Black (1612). Result: 0-1. Time control: 5+3.
PGN:
[full PGN with %eval tags]
CRITICAL POSITIONS (Stockfish 16, depth 22):
Ply 29 (after 15.f4), eval +0.42, Black to move
FEN: r1bq1rk1/1p2bppp/p1nppn2/8/3NPP2/2N1B3/PPPQ2PP/2KR1B1R b - - 0 15
Black played 15...b5. Engine best: 15...Nxd4 (eval +0.31)
Ply 33 (after 17.e5), eval +0.38, Black to move
FEN: r1bq1rk1/4bppp/p1nppn2/1p2P3/3N1P2/2N1B3/PPPQ2PP/2KR1B1R b - - 0 17
Black played 17...dxe5. Engine best: 17...Nd7 (eval +0.44)
Ply 36 (after 18...Nd5), eval -3.41 -> after 19.Nxd5 eval +2.10
FEN: r1bq1rk1/4bppp/p1np4/1p1nP3/5P2/2N1B3/PPPQ2PP/2KR1B1R w - - 0 19
Black played 18...Nd5?? Engine best: 18...Qc7 (eval +0.51)
Eval swing: 291 centipawns
Ply 47 (after 24.Rxd6), eval +4.80, Black to move
FEN: [...]
TASK:
1. For the 291-centipawn swing at ply 36, explain in plain terms what
Black missed. Name the tactical pattern.
2. Identify the decision at ply 29 or 33 that made ply 36 likely.
Treat this as the real error.
3. Give me two specific drills. No general advice.
The structural moves here are worth naming. The explicit “do not reconstruct the board yourself” instruction cuts a surprising amount of drift, because without it the model will sometimes helpfully re-derive a position it was handed. The “if you need a position I have not supplied, say so” line gives it a legal exit from guessing, and it uses that exit. And task 2 is the one that earns its keep: engines flag the blunder, but blunders are usually the bill for a decision made two or three moves earlier, and the preceding-decision question is where an LLM adds something Stockfish does not.
What this actually buys you, measured
Run the same game through four prompt formats and count factual errors about board state. Across a batch of twenty club games, the pattern is consistent:
| Input format | Board-state errors per game | Eval judgment errors |
|---|---|---|
| Bare PGN, “analyse this” | 4 to 7 | frequent |
PGN + all [%eval] tags | 3 to 5 | rare |
| PGN + 5 FENs, no evals | 0 to 1 | frequent |
| PGN + 5 FENs + evals + best lines | 0 to 1 | rare |
Two separate things are being fixed by two separate ingredients. FENs kill positional hallucination. Evals kill judgment hallucination. Supply one without the other and you get a model that describes the right board with confidently wrong assessments, or the right assessments about an imaginary board. Both failure modes read as authoritative.
The second column deserves a note, because it is the one people underestimate. Without engine numbers, models are drawn to moves that look instructive rather than moves that are good. Sacrifices get praised. Quiet consolidating moves get called passive. A model with no eval anchor told me a rook sac was “thematically justified” in a position where Stockfish had it at -3.8 and the refutation was a two-move king walk.
Where the engine stops and the model starts
This is the division of labour worth internalising. Stockfish tells you that 18…Nd5 cost 291 centipawns and that 18…Qc7 held. It will not tell you that you played Nd5 because you had spent 40 seconds hunting for activity in a position that wanted patience, that this is your third game this month where a knight jump into a pin followed a long think, or what to drill on Tuesday.
That translation layer is the whole value of putting an LLM in the loop, and it only works when the model is not simultaneously trying to remember where your pieces are. The broader taxonomy of which tools handle which part of this, and which prompt patterns hold up under testing, is laid out in AI coaching tools and LLM chess prompts, tested.
The three-game pattern pass
Once the single-game format works, the genuinely useful application is batching. Take ten games where you lost, extract the blunder ply from each with its FEN and eval delta, and feed all ten blunder positions in a single prompt with no surrounding PGN at all. About 1,500 tokens total.
Ten positions. Each is a moment where a 1650 player lost 200+ cp.
For each, state the move played, the engine's move, and classify the
error type. Then tell me which error types repeat, ranked by frequency.
1. FEN: r1bq1rk1/4bppp/p1np4/1p1nP3/5P2/2N1B3/PPPQ2PP/2KR1B1R w - - 0 19
Played: 18...Nd5 (-291 cp). Best: 18...Qc7
2. FEN: [...]
The output from this is categorically different from single-game analysis. Four of my ten positions turned out to be the same error: a piece move that broke the defence of a square I had already spent a tempo defending. Nobody notices that reviewing one game at a time, including Stockfish, which analyses games in isolation and has no concept of your recurring habits.
Build the extraction script once and this becomes a twenty-minute Sunday routine instead of a project.
A short checklist you can run tonight
Pick your last three losses. For each, request Lichess computer analysis, export the PGN with evals=true, find every move with a swing above 100 centipawns, and grab the FEN at each of those plies plus one at the moment you remember feeling uncomfortable. Cap it at six positions. Write the prompt with the explicit no-reconstruction instruction and the “tell me if you need a position” escape hatch. Ask for the preceding decision, not just the blunder.
Then do the thing that separates this from entertainment: when the model names a pattern, go find three more examples of that pattern in your own archive before you accept it. A model that has been handed accurate FENs and honest evals is a good diagnostician. It is still describing your games from the outside, and the confirmation is yours to do.