AI Chess

Depth Is Not Truth: Why Stockfish Changes Its Mind At Depth 30

You paste a position into Stockfish, wait for depth 30, write down +0.42, and close the tab. Fifteen seconds later, if you’d left it running, the same engine would have said -0.80. Nothing about the position changed. What changed is which handful of branches the search decided were worth extending, and at depth 31 it found one it had been throwing away since depth 14.

This is the single most misunderstood thing about engine output at club level. The chess engine depth meaning that most players carry around, “the engine looked 30 moves ahead”, is wrong in a way that costs you real analysis. Depth 30 does not mean every legal line was examined to 30 plies. It means the iterative deepening loop reached its 30th iteration, and in that iteration the main line got roughly 30 plies while the overwhelming majority of sibling moves got 8, or 5, or were cut off before a single node was spent on them.

The number in the corner is an advertised depth, not a guarantee

Stockfish searches by iterative deepening: it completes a full search to depth 1, then 2, then 3, and so on, using each pass to order moves better for the next. That much is honest. The dishonest-feeling part is everything layered on top.

Late move reductions cut the depth of any move that isn’t near the front of the move-ordering list. Search the 15th-ordered move at a nominal remaining depth of 12, and Stockfish will actually give it 7 or 8, sometimes less. Null move pruning assumes that if you can give your opponent a free move and still stand well, the branch is winning and needs no further work. Futility pruning kills quiet moves near the leaves when the static eval is already hopeless. Aspiration windows start each iteration with a narrow score window around the previous result and only re-search wider if the score escapes it.

Each of those is a bet. Every bet is usually right, which is why Stockfish 17 at depth 30 beats Stockfish 8 at depth 30 into pulp. But “usually right” is exactly the condition under which a tactical position produces a late eval flip: the bet was wrong, and it took fifteen extra iterations of improved move ordering before the search stopped discarding the refutation.

Look at seldepth, because that’s where the honesty lives

Fire up a GUI that shows raw UCI output. Nibbler is the best one for this at club level (it’s a Stockfish-specific front end and it puts nodes-per-move and seldepth right in front of you). En Croissant works too. So does plain Arena, and so does running stockfish.exe in a terminal and typing UCI commands by hand, which takes about ninety seconds to learn:

uci
setoption name Threads value 4
setoption name Hash value 1024
setoption name MultiPV value 3
position fen r1b1kb1r/1p1n1ppp/p1n1p3/q7/3NP3/2N1B3/PPPQ1PPP/2KR1B1R w kq - 0 1
go infinite

What comes back looks like this:

info depth 22 seldepth 34 multipv 1 score cp 41 nodes 12845201 nps 2481000 hashfull 421 time 5177 pv f1e2 ...
info depth 22 seldepth 29 multipv 2 score cp 18 nodes 12845201 ...
info depth 22 seldepth 31 multipv 3 score cp -12 nodes 12845201 ...

depth 22, seldepth 34. The selective depth is the deepest single line the search reached anywhere in the tree, usually a forcing capture sequence chased down by quiescence search and check extensions. The gap between 22 and 34 is the whole point: the tree is not a wall, it’s a spike. Twelve plies of that spike exist along one narrow path. Everywhere else, the engine is being far more frugal than the headline suggests.

When you see seldepth sitting only two or three above depth, you’re usually in a position with nothing forcing in it. When it’s fifteen above, there’s a tactical rabbit hole, and the eval you’re reading is resting on the engine having resolved that rabbit hole correctly.

Three kinds of eval change, and only one of them matters

Not every wobble is meaningful, and lumping them together is how people end up distrusting engines entirely.

Window noise. The score moves by 0.05 to 0.15 between depths and settles back. Aspiration windows and hash table effects. Ignore it completely. If you’re using it to choose between two moves, you’re doing astrology.

Horizon resolution. The score walks steadily in one direction over several iterations: +0.30, +0.45, +0.60, +0.75. The engine is gradually cashing in a positional feature it can now see the consequences of, like a passed pawn queening or a bad bishop never getting out. This is informative but calm. Your judgment of the position was directionally fine.

A flip. +0.40 at depth 26 becomes -1.10 at depth 31 and stays there. A move that was pruned or reduced turned out to work. This is the signal. It does not mean the engine was lying at depth 26; it means the position contains a line that requires deep, specific calculation to evaluate, and no human sitting at the board is going to find it by feel.

The upperbound and lowerbound tags in the UCI output are worth knowing here. When you see score cp -150 upperbound, the search failed low: the true score is at most -1.50, and Stockfish is in the middle of re-searching with a wider window. You’re watching the flip happen in real time. Don’t screenshot that number.

The depth ladder: a protocol you can actually run

Stop reading a single figure. Take a ladder of them. For any position you care about, record the eval at four checkpoints and write down the spread.

DepthEvalBest move
15+0.38Bd3
20+0.30Bd3
25+0.44Bd3
30+0.51Bd3

Spread: 0.21, best move stable throughout. That’s a quiet position. Your positional understanding is the thing being tested, and you can trust yourself to evaluate it over the board. Study the plan, not the line.

Now the other kind:

DepthEvalBest move
15+0.62Qxb2
20-0.15Rb1
25+0.10Rb1
30-0.94Nd5

Spread: 1.56, and the best move changed twice. That’s a position where the evaluation is a function of a specific line, not of structure. The Najdorf Poisoned Pawn (1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 6.Bg5 e6 7.f4 Qb6 8.Qd2 Qxb2 9.Rb1 Qa3) is the classic laboratory for this; run it at MultiPV 4 and watch the ordering of the top four moves shuffle for thirty seconds. Engines disagreed about that variation for two decades and the disagreement lived exactly in this kind of instability.

Your numbers will not match mine. They shouldn’t. Stockfish 17.1 on eight threads with 2 GB of hash reaches different depths than the Lichess server-side eval, which reaches different depths than Chess.com’s Game Review. Multi-threaded search is nondeterministic by design, so two runs on your own machine will differ slightly. The shape of the ladder is what transfers, not the digits.

What to do with an unstable position

Here’s the conversion step nobody teaches. Sort every critical position from your games into two buckets based on the spread.

Spread under about 0.30 with a stable best move: this is a understanding position. Write a sentence of prose about why the move is good. “White’s knight is going to f5 and Black can’t stop it without weakening g6.” Add it to your notes, move on. No memorisation required, because you can rederive it.

Spread over about 0.80, or a best move that changes above depth 20: this is a calculation position. Set the clock for ten minutes, cover the engine line, and calculate it yourself on a board. Write down your tree. Then compare against the engine’s PV at depth 35 and find the exact ply where you and Stockfish parted company. That ply is your actual weakness, and it’s almost always the same three or four motifs repeating across your whole game collection: you stop at a check, you don’t consider the quiet retreat, you assume a piece is hanging when it’s defended twice.

Positions in the middle need a judgment call, and they’re the ones where running MultiPV 3 pays off most. If the top three moves are within 0.15 of each other at depth 30, the position has multiple good plans and your move choice is a style question, not an accuracy question. Go with the one you understand.

There’s a lot more to unpack about what centipawns actually represent and why a +1.00 in a locked position is a different animal from +1.00 with queens on. The pillar piece on reading engine output covers the eval scale itself in the detail it deserves.

Practical settings that stop depth from lying to you harder

Hash size matters more than most club players realise. Stockfish’s default is 16 MB, which is absurd for analysis. The hashfull field in the output runs from 0 to 1000; if it’s sitting at 950+ after twenty seconds, the engine is discarding useful positions and your depth numbers are inflated relative to the actual work done. Set Hash to 1024 or 2048 on a machine with 8 GB of RAM.

Threads help throughput but hurt reproducibility. If you’re trying to compare two candidate moves precisely, run at Threads 1 and compare node counts instead of times. You’ll get deterministic, repeatable numbers you can actually put in a spreadsheet.

Don’t compare depth across engines or versions. Leela Chess Zero doesn’t even use depth in the same sense (it reports a nodes figure that corresponds to MCTS playouts, and a Leela “depth 20” is not remotely a Stockfish depth 20). Stockfish 11 at depth 30 searched far more nodes than Stockfish 17 at depth 30, because seven years of more aggressive pruning went in between. Depth is an internal bookkeeping counter that each program defines for itself.

Go pull your last ten losses. For each one, find the move where the eval swung, run the ladder at depths 15, 20, 25 and 30, and write the spread next to the move number. By the tenth game you’ll have a list, and the list will tell you whether you’re losing games to calculation or to understanding. Those need completely different training, and right now you’re probably guessing which one you need.