Analysing Your Own Games With An Engine
Stockfish will tell you that you were losing by 3.4 pawns on move 22. It will not tell you why you played the move, what you were afraid of, or which of your habits produced the position in the first place. That gap is where most engine analysis dies: people run the automatic review, collect a list of red exclamation marks, nod at them, and play the same game again next week.
This page is about closing that gap. You’ll get a repeatable process for interrogating an engine rather than just reading it, the specific settings that change what the engine tells you, and a method for turning output into a small number of things you actually practise.
Before The Engine: Write Down What You Thought
The single biggest upgrade to engine analysis costs nothing and happens before you open the engine. You annotate the game yourself, from memory and from the position, recording your thinking at each decision point.
Do it because the engine’s evaluation is only useful in contrast to yours. If Stockfish says a move drops 1.2 pawns and you had no idea, that’s a blind spot. If it drops 1.2 pawns and you’d already written “not sure about this, felt loose”, that’s a calculation failure, and those get fixed differently. Same red mark on the screen, two completely different training responses.
Practically: open the game in Lichess’s study feature or a local board, and add a text comment at every move where you spent more than 30 seconds, plus every capture, check, and pawn move you made without much thought. Write what you were calculating and what you decided against. Three sentences maximum per move. If you can’t remember, write “no recollection”, which is itself a finding.
The full pre-engine routine, including how to sort your own mistakes into categories before any silicon gets involved, lives in The Post-Loss Review Checklist. Read that first if you’re building the habit from scratch; the rest of this page assumes you’ve done some version of it.
Getting An Engine That Tells You Something Useful
Any Stockfish is better than no Stockfish, but the default settings on most interfaces are tuned for speed over insight.
Lichess’s server-side analysis runs Stockfish 16 NNUE at a fixed depth for free, on any game, including ones played elsewhere if you import the PGN. It gives you an evaluation graph, move classifications, and the “Learn from your mistakes” drill mode. Good for triage, weak for interrogation: you can’t ask it questions.
Lichess’s local engine (the “cloud/local evaluation” toggle in the analysis board) runs Stockfish in your browser via WebAssembly. It’s genuinely strong, and crucially it’s interactive, so you can play moves out and watch the evaluation change. Set Multiple lines to 3 and Memory as high as it lets you.
Chess.com’s Game Review is slicker and worse for learning. It hands you an accuracy percentage and a coach-voiced explanation of each mistake, which feels great and teaches you to be a passive consumer. The underlying engine is fine. Use the Analysis board tab rather than Game Review if you want to drive.
Desktop Stockfish in a GUI is where serious work happens. Get the current Stockfish binary from the official releases (pick the build matching your CPU: the bmi2 build if your processor supports it, avx2 otherwise), and load it into a free GUI. On Windows, Arena or the free tier of Chessbase’s Fritz interface; cross-platform, Nibbler is purpose-built for engine analysis and shows you the search tree in a way browser interfaces don’t. En Croissant is a newer free option that handles PGN databases and multi-engine setups well.
Three settings matter more than the rest:
| Setting | Default | Set it to | Why |
|---|---|---|---|
| MultiPV | 1 | 3 or 4 | You need to see the second-best move, not just the best. Most of your improvement lives in the gap between them. |
| Threads | 1 | cores minus 1 | Speed, so you can afford depth 30+ on demand. |
| Hash | 16 MB | 1024–4096 MB | Stops the engine re-searching positions it already understands. |
MultiPV is the one people skip and it’s the one that changes everything. With MultiPV 1 the engine is an oracle. With MultiPV 4 it’s a conversation partner: you can see that your move was third-best and only 0.15 worse, which means the position wasn’t your problem at all.
Reading The Numbers Without Fooling Yourself
An evaluation of +0.60 means Stockfish thinks White’s position is worth about six-tenths of a pawn. It does not mean White is winning, and at your level it barely means anything at all.
Here’s the practical scale for club players:
0.00 to ±0.30 Equal. Ignore any "mistake" that lands in here.
±0.30 to ±0.80 Small edge. Real, but nobody converts this by force.
±0.80 to ±1.50 Clear advantage. A good opponent is uncomfortable.
±1.50 to ±3.00 Winning with accurate play. Still losable at 1400.
±3.00 and up Decided, absent a blunder.
#5 Mate in 5, forced.
The trap is treating evaluation drops as uniformly bad. A move that takes you from +0.20 to -0.10 is noise. Lichess may not flag it; a stricter tool might call it an inaccuracy. Either way it belongs nowhere in your training plan. Meanwhile a move that takes you from -0.40 to -1.30 is a genuine error that no interface will shout about because it doesn’t cross a dramatic threshold.
So set your own thresholds and apply them consistently. A reasonable set for 1000–1900:
- Drop of 0.5 to 1.2: worth a look if it’s in a position type you meet often.
- Drop of 1.2 to 2.5: analyse properly. This is where your recurring errors live.
- Drop over 2.5: usually a tactic you missed or allowed. Extract it as a puzzle.
- Any drop from a position that was already ±3.0 or worse: skip. Losing by 4 instead of 3 taught you nothing.
That last rule deletes roughly a third of the flagged moves in a typical loss, all of them from the phase where you were already lost and thrashing. Your engine report looks worse than your play actually was, because it keeps scoring after the game has been decided.
Also: watch depth. An evaluation at depth 18 and the same position at depth 32 can differ by a full pawn, and the shallow number is the one that shows up first. Never write down an evaluation you saw for less than a few seconds on a position that matters. In Nibrose or Nibbler the depth ticks up in the corner; wait for it to reach 28–30 before you trust a sharp tactical position, and be aware that some positions (fortress-like endings, closed structures with locked pawns) fool Stockfish even at depth 40.
The Interrogation Loop
This is the core technique. Once you’ve found a moment that matters, you stop reading and start asking.
Take a real example. Suppose you’re Black in a Queen’s Gambit Declined and the engine flags move 14:
14... Nbd7? (0.35 -> 1.28)
Stockfish 17, depth 31, MultiPV 4:
1. (+0.31) 14... c5 15. dxc5 Bxc5 16. Ne4 Be7
2. (+0.42) 14... Ne4 15. Bxe7 Qxe7 16. Nxe4 dxe4
3. (+0.55) 14... h6 15. Bf4 c5 16. dxc5 Nxc5
4. (+1.28) 14... Nbd7 15. Bxf6 Nxf6 16. Ne5 Qc7 17. f4
The report gives you the number. The interrogation gives you the lesson. Work through five questions in order:
One: what does the engine’s move actually do? Play 14…c5 on the board and look. It breaks in the centre, trades off White’s space advantage, and activates the bishop that was sitting behind the pawn chain. That’s the idea: freeing the position before White’s pieces get to the good squares.
Two: what does my move fail to do? Play 14…Nbd7 and continue with the engine’s line. After 15.Bxf6 Nxf6 16.Ne5 you can see it: your knight got in the way of the c-pawn, so the freeing break never happens, and White gets e5 for free. Your move wasn’t tactically refuted. It was positionally passive, which is an entirely different failure and points at a different fix.
Three: does the engine still like its move if I take it three moves further? Play out 15.dxc5 Bxc5 16.Ne4 Be7 and check the evaluation holds at around +0.30. It does. Now you know the line isn’t a computer illusion that collapses on the next move.
Four: what happens if my opponent doesn’t cooperate? This is the step almost nobody does, and it’s the most valuable. Instead of following the engine’s main line, play the move your opponent would actually pick at your level. After 14…c5, try 15.Bxf6 instead of 15.dxc5 and see what the engine says. If the evaluation stays around equal, you’ve learned that the break works against multiple replies, which makes it a rule you can use. If one specific reply is annoying, that’s the line to remember.
Five: can I now state the lesson in one sentence without naming any moves? For this example: “In the QGD with White’s pieces aimed at my kingside, I need the c5 break before I develop the b8 knight, because Nbd7 blocks it.” That sentence is portable. “14…c5 was better” is not.
If you can’t produce the sentence, you haven’t finished analysing. Go back to question two.
Three Things Engines Say That Beginners Misread
“Blunder” on a move you had to play. Sometimes every move loses and the engine labels the least-bad one a blunder because the alternative was a slightly slower loss. Check the second-best line’s evaluation. If the whole MultiPV list is between -3.1 and -3.6, the blunder happened earlier and you should be analysing that move instead. Engine reports attribute the loss to the move where the number moves, not the move that caused it.
Best-move sequences that require inhuman precision. Stockfish will happily recommend a queen sacrifice that wins in 14 moves with only moves. Noted, and irrelevant. When the engine’s recommendation depends on a long forced sequence you’d never find and never hold, look at what MultiPV 2 or 3 suggests. A move that’s 0.4 worse and comprehensible is a better addition to your repertoire than a move that’s objectively best and unplayable.
Opening evaluations of roughly +0.4 for basically everything. Almost any sane opening move in the first ten moves evaluates between 0.00 and +0.50, and the engine has no idea which one suits you. Do not use raw evaluation to pick an opening. Use the Lichess opening explorer’s win-rate data at your rating band instead: click Openings in the analysis board, set the database to Lichess games, filter to 1400–1800 and your time control, and look at actual results plus how often each move gets played. A line that scores 54% for White in 12,000 games at 1600 is more useful information than an engine’s +0.28.
Turning Output Into A Training Plan
Analysing one game well is satisfying. It also changes almost nothing, because a single game is one sample and your errors are patterns. The payoff comes from aggregating.
Keep a plain text file or spreadsheet. One row per analysed error, with five columns: date, phase (opening / middlegame / endgame), category, one-sentence lesson, position FEN. Categories should be specific enough to count:
- Missed opponent’s tactic (what kind: fork, pin, back-rank, discovered attack, trapped piece)
- Missed my own tactic
- Went passive when a break or trade was available
- Wrong plan (played on the side of the board where I had nothing)
- Time trouble: move played in under 5 seconds in a critical position
- Endgame technique: specific ending, named
- Opening: left theory before move 12 and got a worse structure
After 15 or 20 analysed games you’ll have 40 to 70 rows, and the pattern will be embarrassingly obvious. A typical 1500 finds something like: 18 entries of “missed opponent’s tactic, mostly discovered attacks and knight forks on f7/f2”, 11 of “went passive”, 9 of “time trouble in the first 15 moves of a rapid game”, and a long tail of singletons. That distribution is your training plan, and it took you ten minutes to produce.
Now act on it. The tactical cluster means Lichess puzzle themes, specifically, not random puzzles: go to the puzzle dashboard, pick Fork and Discovered Attack, do 20 a day in that theme, and check your theme rating weekly. The passivity cluster means playing out your logged FENs against the engine, with a rule that you must find a forcing move or a structural change in every position before looking. The time-trouble cluster is not a chess problem at all and no amount of analysis fixes it; that’s a change to how you allocate clock, and you test it by playing five games where you deliberately spend under 90 seconds on the first 12 moves.
Sparring Against The Engine
Reading lines is passive. Playing them is not.
Take a position from your log, load it into an analysis board, and play both sides against Stockfish at a reduced strength. On Lichess, “Play with the machine” from a position lets you set Stockfish level 1–8; level 6 is roughly 1900-strength in practice, level 8 considerably more. In a desktop GUI you can set Skill Level directly (0–20) or use UCI_LimitStrength with UCI_Elo, which lets you specify something like 1700 and get an opponent that makes human-scale errors rather than an engine that plays perfectly and then randomly throws.
Two exercises worth the time:
Convert your own winning positions. Pull three positions from your log where you were at +2.0 or better and drew or lost. Play each one out against Stockfish at roughly your rating plus 200. If you can’t convert +2.0 three times out of three, technique is your bottleneck, not tactics, and that reframes the whole plan.
Defend your own losing positions. Same idea, other direction. Take a position where you were at -1.5 and collapsed, and defend it. You’ll discover that -1.5 is extremely holdable and that your actual loss came from panic rather than the position. That’s worth knowing about yourself.
For both exercises, turn the evaluation display off while you play. Looking at the number destroys the exercise. Nibbler has a hide-eval option; on Lichess, you’re playing the machine rather than analysing, so the number isn’t shown anyway.
A Worked Session, Start To Finish
Here’s what 40 minutes on one game looks like, so you can copy the shape of it.
You lose a 15+10 rapid game as White, rated 1550 against 1610. First, before anything else: five minutes writing your own notes on the game, flagging move 11 (felt uncomfortable), move 19 (spent four minutes, still unsure), and move 27 (played fast, immediately regretted it).
Then run the Lichess computer analysis. It reports 2 inaccuracies, 3 mistakes, 1 blunder, accuracy 82%. Six flagged moves, plus your three self-flagged ones, and there’s overlap at move 27.
Apply the thresholds. Two of the flagged moves are sub-0.5 drops in a roughly equal position: gone. One is a mistake on move 34 when you were already at -4.2: gone. That leaves move 19, move 27, and the blunder on move 31, plus move 11 which the engine liked fine but which felt bad to you.
Move 11 goes first, because a move that the engine approves of and you dislike is pure gold. Set MultiPV 4, depth 30. The engine gives +0.35 for your move and +0.30, +0.28, +0.25 for three others, all within noise. Your discomfort wasn’t about the move, it was about not having a plan in that structure. Log it as “wrong plan, isolated queen’s pawn structures: no idea what to aim at” and note the FEN. That one line is worth more than the blunder analysis.
Move 19 is where you spent four minutes. The engine has your move at +0.10 and its top choice at +0.75, with the difference being a knight manoeuvre to a square you never considered. Interrogate: play the engine’s line out four moves, watch the knight land on c5 and attack two things, then check what happens if your opponent prevents it. Lesson sentence: “When my opponent’s pieces are tangled on the queenside, look for a knight route to c5 before I look for a pawn move.” Logged.
Move 27, the fast one, is a 1.9-pawn drop to a knight fork you allowed. This is a tactic entry, and specifically a pattern (fork on a square you’d stopped watching after a trade). Log the FEN, and note the trigger: the fork became available the moment a defender left the board two moves earlier. That trigger is the actual lesson: after every trade, re-scan for what the departing piece was defending.
Move 31’s blunder, on inspection, has the whole MultiPV list between -2.8 and -3.3. Nothing to learn. Skip it and say so in the log, because “no lesson available” is a legitimate result and stops you from re-analysing it later.
Total: four log entries, three portable sentences, two FENs saved for engine sparring. That’s a good return on 40 minutes, and it’s a very different output from a list of six moves with red icons next to them.
What To Do When The Engine Won’t Explain Itself
Sometimes you run the whole loop and the answer is still opaque. The engine wants a quiet rook shuffle, the evaluation swings 0.6 pawns, and you cannot see why. This happens most often in closed positions and in endings.
Force the engine to show its reasoning by pushing the position further. Play 10 or 12 moves of the engine’s line against itself, then look at the resulting position and compare it to what you’d have had. Usually the difference is visible at the end of the line even when it’s invisible at the start: your rook ended up passive, or a pawn that was flexible got fixed on a bad square, or the king never found shelter. The engine’s evaluation is a compressed statement about a position 12 moves away, so go and look at that position.
If it’s an endgame with seven pieces or fewer, stop guessing and check the Lichess tablebase, which gives you the exact result and the distance to mate or conversion. Tablebases are truth, not evaluation: “win in 31” is a fact. They’re also brutally instructive about how much you don’t know, since you’ll find plenty of positions you’d have called drawn that are wins in 20.
And if the answer is genuinely beyond you right now, write “don’t understand, revisit” in the log and move on. A position you’ll understand in six months is not a good use of tonight.
The Habits That Make This Stick
Analyse losses, but analyse at least one win a month, because wins hide errors that your opponent failed to punish and those are the most dangerous ones you own. A game you won 1-0 while spending 20 moves at -1.8 is a loss that got rescued, and the engine is the only witness.
Cap your analysis time. Twenty to forty minutes per game, hard stop. People who spend two hours on one game analyse four games a year.
Re-read your log every fourth week. Not the whole thing, just the category counts. If “missed opponent’s tactic, discovered attack” has gone from 8 entries to 1 over two months of themed puzzles, the method worked and you can redirect the puzzle time. If it hasn’t moved, the puzzles weren’t the fix and you need a different intervention (usually: a slower blunder-check routine, not more puzzles).
Keep the engine off during your own first pass, every single time. The moment you know the answer you stop being able to reconstruct how you’d have found it, and reconstructing that is the entire point.
In this section
The supporting pages under this subject.