Accuracy Percentage And Game Review Scores: What They Hide
You finish a game, click Game Review, and get a number. 94.2%. Or 71.8%. It feels like a grade. It is not a grade. It is a compression of your entire game into a single figure that depends more on what kind of position you played than on how well you played it, and if you are using that figure to track your improvement, you are tracking the wrong thing with real consequence.
Let me be blunt about the claim, because this article commits to it: accuracy percentage rewards quiet positions and punishes sharp ones. Two players of identical strength, one steering into a Carlsbad structure and one entering a Najdorf Poisoned Pawn, will post wildly different accuracy scores while playing equally well relative to the difficulty they faced. The number does not know this. It cannot know this. And every time you screenshot a 96% to a friend, or feel deflated by a 74%, the metric is teaching you something false about your own chess.
What the number actually is
Chess.com’s accuracy is derived from what they call Expected Points Loss, or CAPS2. It looks at your evaluation before and after each move, converts centipawns into a win-probability figure, and measures how much win probability you shed. Lichess uses a similar approach built on a formula from Lichess’s own research, mapping centipawn evaluations to a winning-percentage curve and averaging the per-move accuracy across the game.
The conversion is the whole story. Neither site treats a 30-centipawn loss as equally costly everywhere. They use a sigmoid: win probability changes fast when the evaluation is near zero, and barely moves when the evaluation is already ±400. That is reasonable in principle. It is also exactly what creates the distortion.
Consider the arithmetic. On Lichess, the winning-percentage function is roughly:
Win% = 50 + 50 * (2 / (1 + exp(-0.00368208 * centipawns)) - 1)
Run a few numbers through it:
| Eval (cp) | Win% | Eval after dropping 50cp | Win% | Win% lost |
|---|---|---|---|---|
| 0 | 50.0 | -50 | 45.4 | 4.6 |
| +200 | 67.6 | +150 | 63.2 | 4.4 |
| +600 | 89.9 | +550 | 88.6 | 1.3 |
| +1000 | 96.9 | +950 | 96.3 | 0.6 |
Same 50-centipawn error. Four costs, ranging from 4.6 points of win probability down to 0.6. If you are already winning by a rook, you can play sloppily for fifteen moves and your accuracy barely flinches. If the game is balanced and every move matters, small imprecisions are charged at full price.
The worked example: 96% and 78% are the same player
Here is the scenario that made me stop trusting the number entirely.
Game A. Exchange Slav, 1.d4 d5 2.c4 c6 3.cxd5 cxd5 4.Nc3 Nf6 5.Nf3 Nc6 6.Bf4. Symmetrical pawn structure, one open file, no imbalance worth the name. Both sides develop, trade on c-file, reach a queenless middlegame where the evaluation sits between -0.2 and +0.3 for twenty consecutive moves. There are typically two or three moves that are all within 0.15 of each other. You play one of them. Stockfish shrugs. You draw on move 41.
Run it through Game Review: 96.4%.
Game B. Same player, a week later, Sicilian Najdorf, 6.Bg5 e6 7.f4 Qb6. You are Black. From move 12 to move 26 the engine evaluation swings between -1.4 and +0.9, because that is what the Poisoned Pawn does. On move 18 you play a move Stockfish rates as the third choice, 0.55 worse than best. On move 22 you miss a tactic and drop 0.7. You find a saving resource on move 29, hold the balance, and grind out a win on move 58.
Game Review: 78.1%.
Same player. Same week. Same objective strength. In Game A, the player never once had to calculate a forcing line four moves deep. In Game B, they navigated fourteen consecutive moves where every candidate carried real tactical content, missed two things, and still won. The 78% game is unambiguously the harder and more instructive performance. The metric scored it as a disaster.
Now flip it. Take a genuinely bad game where you hang a piece on move 11 and your opponent, also bad, hangs it back on move 14, and then the position becomes +6 and you shuffle to victory. Because the eval is off the sigmoid’s steep section for the last 40 moves, those 40 moves are nearly free. That game can score 91%.
Why this matters more than it sounds
The failure mode is not “the number is imprecise.” The failure mode is that the number creates an incentive, and the incentive is backwards.
If you optimise for accuracy percentage, you learn to avoid complications. You start preferring the London System to the Anti-Moscow Gambit, not because it suits your style but because your score goes up. You trade queens earlier. You avoid the messy but objectively fine continuation because messy positions cost you accuracy points even when you play them well. Improvers at 1200 who do this for six months end up with a higher average accuracy and a flat rating graph, because they have been training the one skill that does not transfer: playing positions where nothing happens.
There is a related trap. Accuracy punishes long games. A 90-move endgame where you are up a piece and converting will contain dozens of moves that are “inaccurate” by a tenth of a pawn, and every one of them dilutes your score. A 25-move miniature where you won with a known trap can score 97%. Which game taught you more?
And accuracy is asymmetric in a way nobody mentions: it depends heavily on your opponent. If your opponent blunders on move 9 and the position becomes +4, your remaining moves become cheap and your score inflates. If your opponent plays well and keeps the game balanced to move 50, every move you make is priced at maximum. Your own accuracy percentage is substantially a function of how good your opponent was. Track your progress with a metric that moves when other people play differently, and you are measuring noise.
What the classification labels hide too
The Brilliant / Great / Best / Excellent / Good / Inaccuracy / Mistake / Blunder ladder has its own problems.
“Brilliant” (Chess.com’s !!) fires on a sacrifice that is the best move and not immediately obvious. Fine as entertainment. It is not a measure of insight, it is a measure of whether your move matched a pattern rule. Many “brilliant” moves are the only move, played because the alternatives were obviously losing. Players collect them like achievements.
More importantly, the thresholds are absolute centipawn bands applied to a relative situation. A “Mistake” is defined by a win-probability drop, so in a sharp position with three plausible moves separated by 0.4 each, you can play a perfectly human, defensible move and get tagged. In a dead-drawn rook endgame, you can play the second-best move twenty times and get twenty greens.
The labels tell you where the engine disagreed. They do not tell you whether the disagreement was findable, whether the move was a reasonable practical choice, or whether you understood the position. That last one is the only thing worth training.
Use the engine properly instead
The fix is not a better single number. It is asking the engine questions instead of accepting its verdict. This is what the pillar piece on reading engine output, evals, depth and lies goes into properly, and it is worth your time if you have only ever looked at the number in the top-left corner.
For now, three habits that replace accuracy percentage entirely:
Count your critical moments, not your errors. Go through the game and mark every move where the evaluation could have swung by more than 0.8 depending on your choice. In a quiet Slav there might be three. In a Najdorf there might be eighteen. Now score yourself only on those. A player who gets 12 of 18 critical moments right in a sharp game is playing better chess than one who gets 3 of 3 in a dead position, and nothing on the review screen will tell you that.
Read the multi-PV spread, not the top line. Open the position in a local Stockfish (the Lichess analysis board does this too: set MultiPV to 3 or 4 in the engine settings). If the top three moves are +0.31, +0.28, +0.24, the position had no single right answer and your “inaccuracy” was noise. If they are +0.9, -0.2, -0.6, there was one move and you needed to find it. Same eval loss, completely different lesson. Here is what that looks like in raw Stockfish output:
info depth 24 multipv 1 score cp 31 pv d4d5 c6d5 c3d5
info depth 24 multipv 2 score cp 28 pv c1f4 e7e6 e2e3
info depth 24 multipv 3 score cp 24 pv e2e3 e7e6 f1d3
Three moves inside 7 centipawns. Any “accuracy” penalty for picking the third is meaningless.
Log the reason, not the score. Keep a text file. After each loss, one line: what kind of error was it? Calculation, evaluation, time, opening knowledge, missed opponent resource, wrong plan? After thirty games you will have a distribution, and that distribution is an actual training plan. “I miss opponent counterplay in 40% of my losses” is a thing you can work on. “My average accuracy is 81.3%” is not.
One more thing about depth
Game Review runs at a fixed, shallow depth on Chess.com’s servers, and Lichess’s server analysis runs Stockfish at around depth 18 to 22 depending on load. Your local engine at depth 30 will sometimes disagree with the classification you were given. A “blunder” at depth 18 is occasionally a fine move at depth 30. This happens most often in closed positions and fortress-adjacent endgames, exactly the positions where shallow search misleads.
If a classification surprises you, check it yourself before you internalise it. Open the position, let Stockfish 17 run to depth 30 on your own machine for thirty seconds, and see whether the verdict holds. It usually does. When it doesn’t, you have just learned something the review screen was actively hiding from you.
Next time you finish a sharp game and see a 76%, the right response is not disappointment. It is to open the position at move 18, set MultiPV to 4, and find out how many of those fourteen critical moments you actually got right.