Endgame Lies: Tablebases Versus Engine Evaluation
You load a rook ending into Stockfish and it says +0.00. Or +4.3. Or, worse, -0.3 for a position you’ve just drawn in a tournament. Which of those numbers can you bank, and which are an educated shrug?
The answer depends on one thing: how many pieces are on the board. With seven or fewer (kings count), the evaluation is not an opinion. It’s a lookup in a database that was computed by exhaustive search. With eight or more, you’re back to a heuristic guess, and in certain kinds of positions that guess is badly wrong. The chess tablebase vs engine evaluation question is really a question about which regime you’re standing in.
What a tablebase actually is
A tablebase stores the exact game-theoretic result of every legal position with a given material set, plus the distance to the end under perfect play. Nobody searched it with clever pruning. It was built backwards: start from checkmates, mark every position that reaches mate in one, then mate in two, and so on until nothing is left unlabelled.
Two families matter to you:
- Syzygy (Ronald de Man): stores win/draw/loss (WDL) and distance-to-zeroing (DTZ, moves until a capture or pawn move resets the fifty-move counter). All 6-piece files are about 150 GB. The 7-piece set is roughly 16 TB, which is why you won’t download it.
- Lomonosov (Moscow State University): the first complete 7-piece set, with distance-to-mate. Online-only for most people.
Free access: Lichess has a tablebase page and API for up to 7 pieces (analysis board, click the book icon or just play into a 7-piece position and the panel appears). Stockfish itself can probe local Syzygy files through the SyzygyPath UCI option.
When a tablebase says “win”, it means a win against any defence, including the most stubborn possible. When it says “draw”, no human or engine on earth can squeeze out more. That is a different kind of claim from “+1.8 at depth 38”.
Regime one: seven pieces or fewer
Here’s a position every improver should know. White: Kb6, Pa5. Black: Kd8… fine, that’s too thin. Take a famous one instead: KRP vs KR, Philidor.
White: Ke5, Pe4… let’s use something concrete.
White: Kd5, Re1, Pe5 (hypothetical Lucena-style setups aside)
Rather than invent squares, use the classic Lucena shape, which you can paste into Lichess’s tablebase page yourself:
FEN: 1K1k4/1P6/8/8/8/8/r7/2R5 w - - 0 1
Tablebase: White wins, DTZ 16 (mate in 25 with best play)
Run the same FEN in Stockfish 17 at depth 30 and you’ll see something like +7 or higher, and the engine finds the bridge-building idea (Rc4 and Kc7 style). Both agree. That’s the normal case: the engine is right and the tablebase tells you exactly how right.
The interesting case is where they differ in what they tell you, even when they agree on the result. The tablebase tells you the whole move tree: which moves keep the win and which throw it away to a draw. Engines show you one best line. Tablebases show you that, say, in a particular KQ vs KR position 11 of 14 legal moves still win and 3 throw the win. That matters for training, and we’ll return to it.
The fifty-move rule trap
Here is a lie the tablebase exposes in a fully legal position. Syzygy DTZ values above 50 plies mean “win in theory, but the fifty-move rule will rescue the defender”. These are called cursed wins (WDL = +1 in Syzygy, reported as “cursed win”), and the mirror image is a blessed loss. Famous example: KNN vs KP positions and some KQ vs KRBN-type endings take well over 100 moves to convert.
A plain engine without tablebases will show +9 or +12 on a cursed win because it thinks it’s mating. The tablebase says “draw in practice”. If you are going to play on in a rapid game on that basis, you have been lied to by the number.
Regime two: eight pieces or more
Now switch the lights off. Stockfish still probes tablebases once material drops to seven, and that’s why its endgame play is superb. But your analysis often starts earlier: a rook ending with five pawns each is 12 pieces, and the engine has to search to reach the table. It will usually get there, but not always, and not on the right branch.
Three kinds of positions where the number degrades.
1. Fortresses. Engines count material and activity. A fortress is a position where the side down material cannot be broken. The classic: queen against rook and bishop-pawn setups, or a bishop of the wrong colour with a rook pawn.
FEN: 8/8/8/8/8/2k5/1p6/1K1B4 b - - 0 1 (not a fortress; wins)
FEN: 7k/8/8/8/8/8/6BP/7K w - - 0 1 (wrong bishop, rook pawn)
The second is a draw: the bishop does not control h8, so Black’s king sits in the corner. A classical engine with a shallow search sees White up a bishop and a pawn, roughly +3.5 to +4, and will keep that for many plies. In an 8-piece version, say with a pair of extra blocked pawns, depth 35 may still say +3. The truth is 0.00. This is the single most common way engine output fools club players: a big plus number in a position that’s dead drawn.
2. Pawn races and mutual zugzwang. Here the engine’s horizon bites. In a position where both sides promote, the evaluation depends on a tempo count that sits at the edge of the search. At depth 24 you may see +0.4 swing to -2.6 at depth 31, because one queen check turned out to win a queen. The number is jumping because the engine hasn’t finished the calculation. A stable evaluation across several depths is a clue; a number that keeps lurching tells you the engine is still guessing.
3. Opposition and triangulation in king and pawn endings. Pure K+P endings reach tablebase territory fast, but with 8+ pieces (say 3 pawns each) the engine can mis-order moves in the key zugzwang. Humans solve these by counting; engines solve them by brute force and sometimes miss that the wrong move merely drops half a point.
How to interrogate the engine
Practical routine, using Lichess analysis or a local Stockfish with a GUI such as Arena, Nibbler or ChessBase’s engine panel:
- Count the pieces. Seven or fewer? Trust the tablebase panel over any Stockfish number. Eight or more? Treat the eval as a guess whose quality depends on step 2.
- Check depth stability. Let it run to depth 30+, then look at whether the top move and score have held for at least five depth increments. In Nibbler the score graph per depth shows this directly.
- Look for fortress warning signs. Material edge but a big eval that doesn’t climb as depth increases. A rising number is the engine finding progress; a flat +3 that never moves while the engine shuffles is a fortress smell. Set MultiPV to 3 and see if the top three moves are all just king shuffles.
- Walk it into the table. Force the simplifying line (trade the last pieces) and see if the panel flips to tablebase. If the tablebase says draw while the engine said +3, you’ve found the lie.
- Check the fifty-move counter. In the FEN, the last number matters. A DTZ of 78 with a counter at 30 is a draw.
If you want the general theory of what depth, multi-PV and centipawn scores mean, the pillar piece Reading Engine Output: Evals, Depth And Lies covers it. This article is about the one place the numbers become facts.
Turning it into a training plan
Do not go through endings by memorising tablebase output. Seven-piece lookup is for checking, not for learning. Use it this way:
- Pick one 5-piece family a week: KRP vs KR, KQ vs KRP, KBP vs KB. Set up positions on Lichess’s “Practice” or the board editor, play against the tablebase (Lichess “Play against tablebase” via analysis) with the winning side, and see how often you let the win slip.
- Log your errors by type: “gave up the win by wrong king move”, “missed a draw”. Fifteen positions will show a pattern; most 1400s lose Philidor positions by pushing the pawn too early.
- Review your own games’ endings. Where the game reached 7 pieces, check the result against the table. If you resigned or agreed a draw in a won position, or fought on in a lost one, write the position down. That’s your most reliable source of real weaknesses.
- Where the game had 8+ pieces, don’t let the engine’s number be the lesson. Ask: could I have drawn this by building a fortress? Could I have misjudged a race? Then verify by simplifying into the table.
A last concrete exercise: take one of your own rook endings from the last month where Stockfish said you were worse by about -1.5. Trade down the final pieces in analysis until the tablebase takes over. If the verdict is “draw”, you now know the eval overstated the problem and your actual error was probably psychological: you surrendered a half point. That’s a far more useful finding than “I was -1.5”.