AI Chess
026 Analysing Your Own Games With An Engine 1,848 words · 8 min

Turning Engine Findings Into A Weekly Training Plan

You ran the analysis. Stockfish flagged eleven inaccuracies, four mistakes and two blunders across your last six rapid games. You nodded along, muttered “yeah, that was bad”, closed the tab, and then opened a puzzle rush because it was there.

That is not training. That is watching a replay of your own defeat with a commentary track. The engine gave you a diagnosis and you treated it as a highlight reel.

The conversion step is the whole game. An error category on its own tells you nothing about what to do on Tuesday night. A chess training plan from game analysis has to answer three questions per finding: which drill, how much of it, and when you check whether it stuck. Miss any one of those and you are back to studying by mood, which in practice means studying whatever YouTube put in front of you at 9pm.

Here is the machinery. It takes about forty minutes on a Sunday and it fills your week.

Step one: harvest, don’t admire

Open your last eight to ten serious games. Rapid and classical only; blitz findings are mostly clock findings and they will pollute your data. If you play on Lichess, the URL lichess.org/@/yourname/search lets you filter by Perf: Rapid, Result: Loss, Rated: Yes and sort by date. Chess.com users: Archive, then filter Time Class to Rapid.

Run each game through the analysis board. On Lichess that is “Request a computer analysis” (Stockfish 16 NNUE, usually depth 22 to 26 on the cloud eval). On Chess.com, Game Review gives you the same categories with different labels. Either way you now want a spreadsheet, not vibes. Four columns:

move_no | your_move | engine_top | eval_swing_cp | category
------- | --------- | ---------- | ------------- | --------
18       Bxf6?       Rfd1         -180            traded_my_good_piece
24       h6?!        Kh8          -90             pawn_move_no_purpose
31       Qd7??       Rxc3         -640            missed_opponent_tactic
9        Nbd7?!      c5           -70             opening_move_order

The eval_swing_cp column is the difference in centipawns between the engine’s top line and what you played, measured at the same depth. A 70cp swing at 1300 is noise you can ignore for now. Anything over 150cp is a real leak. Anything over 500cp is a bleeding wound and gets priority.

That last column is the one that matters, and it is the one the engine will never write for you. Stockfish says “-640”. It does not say why. You have to name the mistake in your own words, and you have to name it in a way that implies a fix. “Bad move” implies nothing. “I stopped checking my opponent’s captures once I had a plan” implies a drill.

If you are shaky on getting clean numbers out of the engine in the first place (depth, multipv, when the eval bar is lying to you), read analysing your own games with an engine first and come back. This piece assumes you can already get a trustworthy reading.

Step two: the category-to-drill map

Now the conversion. Each named category gets exactly one drill, one dose and one review date. No category gets “read more about it”.

Error categoryDrillDoseReview
Missed opponent’s tactic (you had a plan, they had a fork)Lichess Puzzle Storm, 3 runs, but announce every enemy capture aloud before moving15 min × 4 daysDay 7: 20 mixed puzzles, log any miss where the threat was the opponent’s
Missed your own tacticChess Tempo “Standard” tactics, rated, no hints, 10 problems20 min × 3 daysDay 10: Puzzle Rush Survival, compare score to baseline
Opening move-order slip (before move 12)Build the line in a Lichess Study, 4 branches deep, then quiz yourself from the diagram25 min × 2 daysDay 5: blank-board recital of the first 10 moves plus one deviation
Traded a good piece for a bad one12 positions from Simple Chess or Chessable’s Common Sense; state the worst piece on the board before choosing20 min × 3 daysDay 8: review your 2 newest games for the same swap
Pawn move with no purposePlay 3 games where every pawn move gets a one-line written justification in the notes field3 gamesDay 9: reread the justifications, count how many were honest
Converted winning position into a draw or lossLichess “Practice” endgames: rook+pawn vs rook, then 6 positions from your own games set up on the board vs Stockfish level 630 min × 3 daysDay 12: the same 6 positions, cold
Time trouble (over 60% of clock spent before move 20)4 games with a hard rule: 30 seconds max on any move where the position is quiet4 gamesDay 7: check move-time graph on your last 4 games

Dose matters more than people think. “Work on tactics” spreads into nothing. “15 minutes, four days, and you write down every miss where the threat was your opponent’s” is a thing you can either do or not do, and you will know which.

Step three: a worked week

Take a real profile. Say you are 1420 rapid on Chess.com, eight games analysed, and the harvest produced this:

  • 6 findings tagged missed_opponent_tactic, average swing 410cp
  • 4 findings tagged pawn_move_no_purpose, average swing 95cp
  • 3 findings tagged opening_move_order, all in the same Caro-Kann Advance line, average swing 120cp
  • 2 findings tagged bad_trade, average swing 260cp
  • 1 finding tagged botched_conversion, swing 780cp

Total leaked evaluation, roughly: 6×410 + 4×95 + 3×120 + 2×260 + 780 = 4,530cp across eight games. Call it 566cp per game. Now look at where it comes from. The opponent-tactic bucket alone is 2,460cp, over half your total bleed, from a single named habit.

So the week writes itself, and it is not balanced. It shouldn’t be.

Monday. 15 min opponent-threat drill. Puzzle Storm, but before each move you say out loud (or type in a scratch file) what the opponent’s most annoying reply is. Then 10 min setting up the botched_conversion position on the analysis board and playing it against Stockfish at level 6, three times.

Tuesday. 15 min opponent-threat drill again. Then 25 min building the Caro-Kann Advance line in a Lichess Study: main line to move 12, plus the three deviations your opponents actually played. Not twenty branches. Three, because three is what beat you.

Wednesday. Two rated rapid games (15+10). The rule for the session: every pawn move gets typed into the notes field with a reason. “b5 to stop Nc4” counts. “b5 because it looked active” does not count and you write that down honestly, because the count of dishonest ones is your actual measurement.

Thursday. 15 min opponent-threat drill. 20 min on the bad-trade category: twelve positions where you name the worst piece on the board before you pick a move.

Friday. Off. Genuinely off. A plan you can’t sustain is a plan you’ll abandon by week three.

Saturday. 15 min opponent-threat drill (fourth and final session). Two rated games, no special rules, played normally. These are your fresh data.

Sunday. Review day. Analyse Saturday’s two games. Then the only question that matters: did the missed_opponent_tactic count per game drop? You had 0.75 per game. If Saturday’s two games show one or zero, the drill is working and you keep it for another week at half dose. If they show two or more, the drill is wrong for the leak and you change the drill, not the diagnosis.

That last distinction gets skipped constantly. When the numbers don’t move, people conclude the weakness is deeper than they thought and add more volume. Usually the drill was just mismatched. Puzzle Storm at 1400 rewards pattern speed, so if your real problem is that you stop looking once you’ve found a plan you like, speed puzzles may actively reinforce it. Swap to slow, untimed, one-position-at-a-time work with a written threat list. Different drill, same diagnosis.

Why the review date is non-negotiable

A drill without a review date is a hobby. You’ll do it for four days, feel vaguely more competent, and never find out if the feeling was real.

Put the dates in a calendar with the metric attached. Not “review chess”. Write: “Oct 7: count missed_opponent_tactic in games 9-10. Baseline 0.75/game.” Ten words, and it converts a feeling into a number you either beat or didn’t.

Two practical notes on measurement. First, don’t compare raw accuracy percentages between games. Chess.com’s accuracy score swings wildly with position sharpness; a quiet 40-move draw will show 94% and a messy tactical win 78%, and neither tells you about the specific habit you’re training. Count category instances per game instead. Second, give any drill three weeks before you judge it. One week of data at club level is four to six games, which is well inside noise. Three weeks is fifteen-odd games and the trend becomes readable.

The failure modes, named

Harvesting everything. Twenty-two findings across eight games, all of them logged, all of them turned into drills, none of them practised properly. Cap it. Three categories a week, weighted by centipawn bleed. The rest waits, and most of it will turn out to be the same three categories wearing different hats.

Studying the engine’s move instead of your process. Stockfish plays 31.Rxc3 and you memorise 31.Rxc3. Useless: that position will never occur again. The transferable thing is the question you failed to ask before playing 31.Qd7, which was “what does the rook on c8 do if I move my queen off the file”. Write down the question, not the move.

Trusting the eval bar during noise. At depth 18 in a locked middlegame, +0.4 and +0.9 are the same position. Don’t log a 50cp finding as a mistake; you’ll fill your week with drills for problems you don’t have. Set your floor at 150cp and hold it.

Rebuilding the plan from scratch every Sunday. The whole point of the review date is continuity. A category you’re mid-way through doesn’t get replaced because this week’s games threw up something shinier. Finish the three weeks, measure, then retire it.

What this costs you

Forty minutes of harvesting on Sunday, plus roughly three and a half hours of actual work spread over five days. That is less than most club players already spend on chess in a week. The difference is that none of it is chosen by whatever you feel like doing, and all of it has a number attached that will tell you in seven days whether it worked.

Start with the spreadsheet. Eight games, four columns, and the honest naming of what went wrong in the fifth. The plan falls out of the data almost mechanically once the data exists, which is why the people who skip this step keep studying openings while losing rook endings.