Reading Engine Output: Evals, Depth And Lies
You have Stockfish one click away. That is not the same as having an analyst. Most club players use an engine the way a nervous patient uses a thermometer: glance at the number, feel good or bad, close the tab. The number goes up, you played well. The number goes down, you blundered. Nothing gets learned, and six months later the same mistakes are still in the games.
An engine is a witness, not a judge. It will tell you exactly what it saw, at what depth, under what settings, if you know how to ask. It will also mislead you in about eight predictable ways, and every one of those ways is worth knowing because club-level analysis sessions run into them constantly.
This page is about interrogation. What the eval number is measuring, what depth does and does not promise, which classes of position the engine reports badly, and how to convert a session of analysis into something you actually train.
What Does Chess Engine Evaluation Mean, Precisely
A chess engine evaluation is a score in centipawns, where 100 centipawns is nominally one pawn, reported from the perspective of one side. Your GUI shows +0.85; the engine internally reported score cp 85. That much is common knowledge and it is also where most people stop, which is a shame, because the interesting part is what the scale is anchored to.
In modern Stockfish the scale is not “material plus positional bonuses” in any human sense. Since Stockfish 14 the output is deliberately normalized so that +1.00 means roughly a 50% chance of winning the game from that position in self-play, at the time controls used on Fishtest. Read that again, because two words in it are doing a lot of work. Self-play: the win probability assumes both sides are Stockfish. Winning, not “not losing”: the other 50% is split between draws and losses.
So +1.00 does not mean “a pawn up.” It means “a position where a 3600-rated machine converts about half the time against another 3600-rated machine.” When you are 1450 and your opponent is 1480, that number tells you very little about your actual practical chances and a great deal about whether the position is objectively sound.
Lichess converts the centipawn score into a win percentage with a published logistic curve:
Win% = 50 + 50 * (2 / (1 + exp(-0.00368208 * centipawns)) - 1)
Plug in a few values and the shape becomes obvious:
| Eval | Lichess Win% for the better side |
|---|---|
| +0.20 | 53.7 |
| +0.50 | 59.1 |
| +1.00 | 67.6 |
| +2.00 | 82.4 |
| +3.00 | 90.9 |
| +5.00 | 98.0 |
| +8.00 | 99.8 |
The curve is steep in the middle and flat at both ends. That is the single most useful thing to internalise about the scale. Going from +0.20 to +0.70 costs your opponent about 10 points of win expectancy. Going from +5.00 to +8.00 costs them under 2. Which is why a player who drops 0.4 six times in a game has lost more than a player who dropped 3.0 once, and why raw centipawn averages mislead. That whole argument gets unpacked at /centipawn-loss-explained/, and it is the companion piece to this one.
Reading A Raw UCI Info Line
Talk to Stockfish directly at least once. Download the binary, open a terminal, type uci, and then feed it a position. The GUI is hiding information you want.
setoption name Threads value 6
setoption name Hash value 2048
setoption name MultiPV value 4
setoption name UCI_ShowWDL value true
position startpos moves e2e4 c7c5 g1f3 d7d6 d2d4 c5d4 f3d4 g8f6 b1c3 a7a6
go depth 26
What comes back, once per improvement, looks like this:
info depth 26 seldepth 36 multipv 1 score cp 31 wdl 121 823 56
nodes 58412233 nps 4102884 hashfull 486 tbhits 0 time 14238
pv f1e2 e7e5 d4b3 f8e7 e1g1 ...
Every field earns its keep:
depth 26is the nominal full-width depth. Thirteen moves for each side, before extensions and reductions.seldepth 36is how deep the deepest line went. Forcing sequences get extended well past nominal depth. A seldepth far above depth means the position is full of checks and captures.score cp 31is centipawns from the side to move, not from White. The UCI protocol is side-relative; your GUI flips it. If you are reading raw output for a Black-to-move position,score cp 31means Black is better by 0.31.wdl 121 823 56is win, draw and loss in parts per thousand, again for the side to move. Here: 12.1% win, 82.3% draw, 5.6% loss.hashfull 486means the transposition table is 48.6% full. Once this pins at 999 for a long search, your hash is too small and results get noisier.tbhits 0counts endgame tablebase probes. Zero here because there are 30 pieces on the board.
Two more score forms matter. score mate 7 means mate in seven moves for the side to move, score mate -4 means being mated in four. And you will sometimes see score cp 140 upperbound or lowerbound. Those are aspiration-window failures: the engine only knows the true score is above or below that figure. Never quote a bounded score as an eval. Wait for the clean line.
Also try eval at a position. Stockfish prints a per-bucket NNUE breakdown and a final number labelled from White’s point of view, which is a static assessment with no search at all. Comparing the static eval to the searched eval is a fast way to see how much of an assessment rests on tactics.
Depth Is Not A Quality Score
Depth 30 sounds twice as trustworthy as depth 15. It is not twice as trustworthy, and it is not comparable across engines, positions or sessions.
Here is why. Modern engines prune aggressively: late move reductions, null-move pruning, futility pruning. “Depth 30” means the search frame was 30 plies, while most branches were cut far shallower and a handful of forcing lines ran to 45. Two engines reporting depth 30 have searched wildly different trees. Leela Chess Zero does not report depth in a comparable sense at all, because it reports node counts and an average depth over a Monte Carlo tree.
Depth also inherits history. If you have been analysing a position for two minutes and then play one move forward, the new position reaches depth 28 in three seconds because the transposition table is warm. That depth 28 is not a fresh depth 28. Clear hash (setoption name Clear Hash) when you want a clean read.
What depth does buy you is roughly this: each extra ply costs somewhere between 1.5x and 2x the nodes, and each doubling of thinking time is worth on the order of 40 to 50 Elo for a top engine now, down from about 70 in the era of Rybka. Diminishing, but not zero.
Practical settings for a club-level analysis machine. On an 8-core laptop, set Threads to 6 and leave two cores for your browser. Set Hash to 2048 MB for interactive work, 4096 or more for long overnight runs. Be aware that multithreaded search is non-deterministic: run the same position twice with 6 threads and you can get different best moves and evals differing by 0.15. That is normal and it is also a hint about how much to trust small differences.
The Eval Belongs To The End Of The Line, Not The Position In Front Of You
This is the misreading that does the most damage.
When Stockfish says +1.40 with a principal variation twelve moves long, the +1.40 is the evaluation of the position at the end of those twelve moves. It is not a property of the board you are looking at. It is a claim: “if both sides play this exact sequence, we arrive somewhere I score at +1.40.”
So the first thing to do with any eval you care about is play out the PV on the board, at least six moves deep, and look at the resulting position with your own eyes. Do that ten times and you will start noticing things: the eval was resting on a queen sortie you would never find, or on a bishop retreat that only makes sense if you already saw move nine, or on a piece sacrifice whose point is a bind three moves later.
You also start noticing the horizon effect. Give an engine a position where it is losing a piece in four moves and a shallow search will happily throw two pawns at the problem to shove the loss past the edge of its vision. Run the same position to depth 28 and the pawn-throwing evaporates, replaced by an honest assessment. If a suggested move looks like a panicky concession, check it deeper before you write it down as “the computer move.”
Turn On MultiPV Or Stay Fooled
By default an engine gives you one line. One line answers the question “what is best?” and hides the question you usually need answered: “how much worse is what I actually played, and were there three reasonable moves or one?”
Set MultiPV to 4. On Lichess the control is the number of arrows in the analysis settings. In Nibbler, En Croissant, ChessBase or Banksia GUI it is a spinner. Now you get output like this from a middlegame in a Queen’s Gambit structure:
multipv 1 score cp 46 pv c4c5 b7b6 b2b4 ...
multipv 2 score cp 41 pv d1c2 h7h6 f1d1 ...
multipv 3 score cp 12 pv f3e5 c6e5 d4e5 ...
multipv 4 score cp -38 pv c1g5 f6e4 ...
Read that as a landscape, not a ranking. Moves one and two are within 5 centipawns, which is noise: pick whichever you understand. Move three costs you a third of a pawn, which is a real but survivable concession. Move four loses nearly a pawn of equity and is the kind of move worth understanding, because Bg5 looks natural and the refutation Ne4 is the sort of resource club players miss.
Contrast with a position where the spread is:
multipv 1 score cp 210
multipv 2 score cp -145
multipv 3 score cp -160
multipv 4 score cp -190
That is a one-move position. Only one move works and everything else collapses. These are the positions worth putting in a training file, because they test calculation rather than judgement.
One caveat that catches people: raising MultiPV weakens the search. The engine cannot prune the alternative lines as brutally, so each individual line gets less effective depth. If you want the single most reliable best move, drop back to MultiPV 1 for the final check.
Where Engines Genuinely Lie
Not lies exactly. Reports that a naive reader will misinterpret, which amounts to the same thing at the board.
Fortresses. A position where one side has extra material that can never be converted because the defender’s structure is impenetrable. Search-based engines are bad at these because “no progress possible ever” is not a thing you can see in 40 plies. Opposite-coloured bishops with a blocked pawn chain is the classic family. You will see +1.80 on a dead draw. Leela, with its policy network and native win-draw-loss head, usually reports these better than Stockfish does, which is one good reason to keep lc0 installed as a second opinion.
Wrong-bishop and other theoretically drawn endings. A light-squared bishop plus an h-pawn against a bare king, with the defending king able to reach h8, is a draw regardless of how it looks. Without endgame tablebases loaded, an engine can score it well above +5 for a long time. With Syzygy tablebases loaded it returns 0.00 instantly. Get them: the 5-piece WDL and DTZ set is about 1 GB, the 6-piece set is around 150 GB. Point your GUI at the folder via SyzygyPath. If you would rather not host them, Lichess’s tablebase covers up to 7 pieces and is free to query from the analysis board. Endgame analysis without tablebases is analysis of the wrong problem.
The fifty-move rule. Stockfish tracks the halfmove clock, so a winning position can drift toward 0.00 as the counter climbs, and then jump back up after a capture resets it. If you see an eval mysteriously decay in a technical endgame, look at the move counter before you look for a defensive resource. Syzygy handles this differently again: DTZ is distance-to-zeroing, and a tablebase “win” can still be a 50-move-rule draw in practice.
Opening lines that need twenty-five moves of memory. The engine’s near-equal assessment of the Botvinnik Semi-Slav after 1.d4 d5 2.c4 e6 3.Nc3 c6 4.Nf3 Nf6 5.Bg5 dxc4 6.e4 b5 7.e5 h6 8.Bh4 g5 9.Nxg5 hxg5 10.Bxg5 Nbd7 is correct and completely useless to you. Same with the Najdorf Poisoned Pawn after 6.Bg5 e6 7.f4 Qb6 8.Qd2 Qxb2 9.Rb1 Qa3. The eval is 0.00 and the practical result at 1500 is a loss for whoever knows less. An engine has no concept of “how likely am I to remember this at move 14 with eight minutes left.”
Best move versus best move for you. The engine is optimising against a perfect defender. You are playing someone who will, statistically, miss a knight fork on move 25. This is where Maia is genuinely useful: it is a set of neural nets trained to predict what a human at a given rating actually plays, with separate models around 1100, 1500 and 1900. Running Maia 1500 alongside Stockfish on your own games shows you the gap between “objectively best” and “what a player like me is likely to do here,” and the second column is where your practical decisions live.
Eval instability. Watch the score across depths rather than reading the final figure:
depth 14 score cp +22
depth 16 score cp +48
depth 18 score cp -15
depth 20 score cp +118
depth 22 score cp +96
depth 24 score cp +103
depth 26 score cp +101
The settled answer is about +1.00. The story is in the middle: an 80-centipawn swing appearing at depth 20 and a sign flip at 18. A position that unstable under machine search is a position where a human will get it wrong, and the eval alone does not tell you that. Instability is a complexity metre and you should treat it as a first-class signal.
A Six-Step Interrogation You Can Run In Ninety Seconds
Whenever you hit a position that matters, in your own game or in someone else’s:
- Guess first, in writing. Your move and your eval, before the engine speaks. No exceptions. If you look first you learn nothing because you will agree with whatever you see.
- MultiPV 4, depth 24 minimum. Look at the spread, not the top line. Is this a one-move position or a five-move position?
- Check the WDL split.
wdl 121 823 56andwdl 480 460 60can both round to similar centipawn scores, and they are completely different positions to play. The first is a grind, the second is a fight. - Play out the PV six to eight moves and look at the final position. Ask what changed. If you cannot articulate why the engine likes it, the eval is not yet yours.
- Check your own move by force. Make your actual move on the board and read the eval from the other side. The difference between the best move and your move, in win-percentage points, is the only measure of the error that matters.
- Write one sentence. “I played Bg5 because it looked developing; Ne4 wins a tempo on the bishop and I never check knight jumps to e4.” That sentence is the whole point of the exercise.
What The Web Interfaces Are Actually Doing
Both major sites hide the machinery, and the machinery differs.
Lichess runs Stockfish compiled to WebAssembly inside your browser, using your own CPU. It is the real engine, with real NNUE, but it is single or limited-thread unless your browser enables shared memory, so it is slower than a native binary. Lichess also serves cloud evals for positions it has already analysed, often at depth 40 or more. When you see depth jump instantly to 40 on a common opening position, you are reading a cached cloud eval, not your browser’s work. Its Inaccuracy, Mistake and Blunder tags are computed from win-percentage drops, not centipawns, with bands in the region of 5, 10 and 20 percentage points. Its Accuracy figure uses a published formula:
Accuracy% = 103.1668 * exp(-0.04354 * (Win%_before - Win%_after)) - 3.1669
Chess.com’s Game Review uses Stockfish server-side at a depth that varies with your membership and the position, plus a proprietary accuracy model and a classification set including Brilliant, Great, Miss and Blunder. The classifications are opinionated: a “Brilliant” label is triggered by a good sacrifice, which means you can earn one for finding a move that any 2000 would play. The Miss tag is genuinely useful because it flags a failure to punish, which centipawn accounting tends to smear.
For serious work, get off the web. En Croissant is free, open source, and gives you engine management, MultiPV, accuracy reports and puzzle generation from your own games. Nibbler is excellent for Leela and shows policy percentages and WDL clearly. Both let you set threads, hash and tablebase paths, which the browser will not.
From Output To Training Plan
Here is the process that turns a Saturday morning of analysis into a month of training.
Pull your last twenty serious games. Not blitz. Run each through full analysis, then list every position where your move cost more than 15 percentage points of win expectancy. In a typical 1400-level game that is three to six positions. Twenty games gives you somewhere around 80 moments.
Now classify each one by cause, not by size. The categories that actually predict improvement look like this:
| Cause | Example | What it implies |
|---|---|---|
| Missed opponent’s one-move threat | Hung a piece to a fork you could see | Blundercheck routine, not more theory |
| Missed own tactic | Failed to punish a loose piece | Pattern volume, tactics trainer |
| Bad plan in a quiet position | Drifted, no target, pieces on poor squares | Annotated master games in that structure |
| Traded into a worse endgame | Went into a bad rook ending “to simplify” | Endgame study, specific to the ending |
| Opening confusion by move 12 | Out of book and out of ideas | Ten lines deep, not thirty |
| Clock | Error under 60 seconds | Time management, not chess |
Count the buckets. One of them will hold 30% or more of your equity loss, and that bucket is your training plan for the next four weeks. Everything else waits. The most common mistake at club level is not misreading an eval; it is doing generalised “improvement” when the analysis has already pointed at one specific hole.
The measurement side of this, including why your average centipawn loss can go up in a month where you genuinely improved, sits at /centipawn-loss-explained/. Read it before you start tracking numbers, or you will draw the wrong conclusions from your own data.
Drills That Use The Engine As An Examiner
Build a position bank. Take the 10 most instructive moments from those 20 games, put each into a Lichess study as its own chapter, and in the comment field record three things: your move, the engine’s move, and the one-sentence reason. Come back to the study a week later and solve them cold. If you get one right for the wrong reason, note that too.
Sparring is better than staring. Set Stockfish’s Skill Level to a value that beats you narrowly, or use UCI_LimitStrength with UCI_Elo around 200 points above your rating, and play out the positions you lost from. A position you have lost once and then played five more times against resistance is a position you own.
Try the reverse read as well. Take a position, form your own assessment in words and a number, then check the WDL split rather than the centipawn score. “I said slightly better for White; the engine says 34% win, 61% draw, 5% loss” teaches you something about your own calibration that a centipawn figure will not. Most improvers are systematically overconfident in positions with a space advantage and systematically underconfident when a pawn down with activity, and you will only find out which way you lean by making the prediction first.
Then go and find a position from your own games where you and the engine disagree and you still think you are right after twenty minutes of looking. Those are worth more than a hundred positions where you nodded along.
In this section
The supporting pages under this subject.