AI Chess
§3 Section 3 of 6 2,555 words · 12 min

Engines Compared: Stockfish, Leela, Maia, NNUE

Most club players use an engine the way you’d use a spellchecker: paste the game in, watch for red marks, nod at the squiggly lines, close the tab. That gets you a list of moves you already suspected were bad and almost no information about why you played them. The gap between “I have Stockfish” and “I can interrogate an engine” is where most of the available rating points are sitting.

This page compares the four things people usually lump together when they ask about the best chess engine for analysis: Stockfish, Leela Chess Zero, Maia, and NNUE. One of those isn’t an engine at all. Two of them answer completely different questions. And the one you already have installed is probably running on default settings that make it worse at the specific job you’re using it for.

The four things, and what each one actually is

Stockfish is an alpha-beta searcher. It generates moves, prunes hard, and evaluates leaf positions with a neural network. Version 17.1 sits somewhere north of 3600 on CCRL’s 40/15 list, which is several hundred Elo above any human who has ever played. It runs on a CPU, ships as a single executable, and is what Lichess and Chess.com both use under the hood.

Leela Chess Zero (the binary is called lc0) is a Monte Carlo tree search engine with a much bigger network. It needs a separate weights file, and it wants a GPU. Leela evaluates far fewer positions per second than Stockfish but each one more expensively, which gives it a different personality: better at long-term positional compensation, historically weaker at deep forcing tactics. TCEC superfinals in recent seasons have gone Stockfish’s way, usually by a handful of decisive games out of a hundred.

Maia is a Leela network trained to predict human moves rather than good moves. The CSSLab team trained nine models on roughly 12 million Lichess games each, bucketed by rating: maia-1100.pb.gz through maia-1900.pb.gz in 100-point steps. Maia-1900 predicts the actual move played by a 1900 more than half the time. It is not trying to find the best move and you should never read its evaluation as a verdict.

NNUE is not an engine. It stands for Efficiently Updatable Neural Network, an architecture that came out of the Japanese shogi community (Yu Nasu, 2018) and landed in Stockfish 12 in August 2020, worth about 80 Elo overnight. The trick is incremental updates: when you play a move, you don’t re-run the whole network, you patch the input layer for the two or three features that changed. That’s why Stockfish can run a neural net at millions of nodes per second on a laptop CPU while Leela needs a graphics card. When someone says “NNUE engine,” they mean an alpha-beta engine using this kind of evaluation. Stockfish, Koivisto, Ethereal, and Berserk all qualify.

Evaluations are win probabilities wearing a pawn costume

Here’s the single most useful thing to internalise. Stockfish’s +1.00 does not mean “a pawn up.” Since the evaluation was normalised, +1.00 is calibrated to mean roughly a 50% chance of winning the game, with the remainder being draws. Material and win probability drifted apart years ago and most club players never got the memo.

Lichess converts centipawns to win percentage with this formula:

winPct = 50 + 50 * (2 / (1 + exp(-0.00368208 * cp)) - 1)

Run the numbers and the shape of the function becomes obvious:

Stockfish evalWin %What it means in practice
+0.2051.8%Noise. Stop looking.
+0.3553.2%A pleasant opening. Nothing more.
+0.5054.6%Real but not worth tension.
+1.0059.1%You should be winning this against a peer.
+2.0067.6%Clearly better. Still losable.
+3.0075.1%Convert it.
+5.0086.3%Technique only.

Notice the compression at the top and the flatness at the bottom. Going from +0.20 to +0.50 buys you 2.8 percentage points. Going from +2.00 to +3.00 buys you 7.5. That asymmetry is why chasing a 0.15 edge in the opening is a waste of study time and why botching a +3.00 is a catastrophe in a way that botching a +0.50 is not.

Lichess’s blunder classification works directly in this currency, not in centipawns: a win-percentage drop of 10 points is an inaccuracy, 20 is a mistake, 30 is a blunder. And the per-move accuracy score is 103.1668 * exp(-0.04354 * winDiff) - 3.1669, so a single 10-point drop scores about 65% on that move and a 30-point drop scores about 26%.

Turn on UCI_ShowWDL in Stockfish and you skip the conversion entirely. The engine reports per-mille win/draw/loss directly:

info depth 28 score cp 78 wdl 312 651 37 ...

That’s 31.2% win, 65.1% draw, 3.7% loss. Far more honest than “+0.78” for a position where the practical question is whether you can create a second weakness.

A ten-second interrogation, start to finish

Take the Elephant Trap, which claims a steady supply of 1400s every day. After 1.d4 d5 2.c4 e6 3.Nc3 Nf6 4.Bg5 Nbd7 5.cxd5 exd5, White to move:

position fen r1bqkb1r/pppn1ppp/5n2/3p2B1/3P4/2N5/PP2PPPP/R2QKBNR w KQkq - 0 6
setoption name Threads value 7
setoption name Hash value 2048
setoption name MultiPV value 3
go movetime 10000

Three things about that input matter more than the position. Threads defaults to 1, so if you’ve never touched it you have been analysing with one eighth of your CPU. Hash defaults to 16 MB, which for a ten-second search is like taking notes on a Post-it; 1024 to 4096 MB is the sane range for interactive work. And MultiPV 1, the default, hides the thing you actually need: whether the second-best move is 0.05 worse or 1.40 worse.

The output will look roughly like this:

info depth 26 multipv 1 score cp 22 nodes 21884103 nps 2188410 pv e2e3 c7c6 g1f3 f8e7
info depth 26 multipv 2 score cp 17 nodes 21884103 nps 2188410 pv g1f3 f8e7 e2e3 h7h6
info depth 26 multipv 3 score cp 9  nodes 21884103 nps 2188410 pv d1c2 c7c6 e2e3 f8e7
bestmove e2e3

Three moves inside 13 centipawns of each other. That spread is the answer to “which move should I play”: all of them, pick by taste, this is not a critical position. Now notice what is missing. 6.Nxd5 doesn’t appear, and if you force it with searchmoves c3d5 the evaluation lands around -2.2 by depth 12, because 6...Nxd5 7.Bxd8 Bb4+ 8.Qd2 Bxd2+ 9.Kxd2 Kxd8 nets Black a knight for a pawn. You did not need depth 26 for that. You needed depth 12 and the discipline to ask.

This is the pattern worth building into a habit: use MultiPV to learn how wide a position is, and use searchmoves to make the engine defend the move you actually wanted to play. An engine that only ever tells you its favourite move teaches you nothing about your own candidate list.

Where Stockfish will lie to you

Fortresses and drawn endgames are the classic failure. Set up a Philidor-position rook-and-bishop versus rook and Stockfish will sit at around +1.0 for as long as you let it, because the win probability model doesn’t know the position is dead. Load tablebases and the same position resolves instantly to 0.00:

setoption name SyzygyPath value C:\syzygy\3-4-5
setoption name Syzygy50MoveRule value true

The 5-piece Syzygy set is about 939 MB and settles nearly every endgame a club player will ever reach. The 6-piece set is roughly 150 GB, which is a real decision rather than an obvious one. Lichess’s analysis board queries a 7-piece tablebase online for free, so for one-off endgame questions just paste the FEN there.

Long-term compensation is the second blind spot, and it’s where Leela earns its keep. Positions with a pawn sacrificed for a bind, or a piece sacrificed for two connected passers eight moves from promotion, tend to get a cooler read from Stockfish and a warmer one from Leela. When the two engines disagree by more than half a point in a quiet position, that disagreement is itself information: it usually marks a position whose merits are structural rather than calculable.

Then there’s the problem no engine can solve, which is that “objectively fine” and “playable by you” are unrelated properties. The Najdorf Poisoned Pawn after 1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 6.Bg5 e6 7.f4 Qb6 8.Qd2 Qxb2 evaluates within a few centipawns of level. It is also thirty moves of memorised forcing play where one slip loses on the spot. Stockfish will cheerfully hand a 1300 a repertoire that requires a 2600’s memory, and it will never mention the problem.

Leela, and the question “what does this position look like?”

The reason to install lc0 isn’t strength. It’s the policy head. Before searching a single node, Leela’s network assigns every legal move a prior probability based on pattern recognition alone, and you can read it:

lc0 --weights=BT4-1740.pb.gz --backend=cuda-fp16 --verbose-move-stats

With verbose move stats on, each move reports N (visits), P (policy prior), Q (average value), and D (draw probability). A move with P: 34.12% is one the network considers obvious. A move with P: 0.41% that nonetheless wins after search is, almost by definition, the kind of move you will never find at the board without being told it exists. Filtering your analysis for high-value, low-policy moves is a very efficient way to build a list of your blind spots rather than a list of chess’s blind spots.

For a CPU-only machine, --backend=eigen works and is slow. Leela on a GPU at 30,000 nodes per second is a different engine from Leela on a laptop CPU at 300.

Maia: the engine that plays your mistakes back to you

Stockfish answers “what is the best move.” Maia answers “what would a 1500 actually play here,” which is the question that decides your results. The distinction matters enough that it has its own page: Maia vs Stockfish: Human Mistakes vs Perfect Moves goes through the move-matching data and what the nine rating-bucketed models get right and wrong.

Running Maia is a one-liner, because it’s just an lc0 network and you use a single node so that you’re reading the policy head rather than a search:

lc0 --weights=maia-1500.pb.gz --backend=eigen
position fen <your position>
go nodes 1

The practical application, in two directions. First, threat prediction: when you’ve got a winning position and you want to know what to prepare for, ask Maia at your opponent’s rating what they’ll try. Its top move at the right bucket is the move you’ll face maybe half the time, which is a far better use of clock than calculating Stockfish’s best defence that nobody plays.

Second, trap auditing. Take every line in your repertoire where you go slightly worse and check what Maia-1300 through Maia-1900 play there. If your dubious gambit scores because five of the nine Maia models walk into the refutation, you have quantified something real about why it works for you at 1450 and will stop working at 1800.

Nibbler is the GUI that makes this pleasant: point it at lc0 with a Maia weights file, switch the display to policy, and you get every move annotated with its human-likelihood as a percentage on the board.

Which engine for which question

Your questionToolSettings that matter
Did I blunder, and by how much?Stockfishdepth 20-24, MultiPV 1, Hash 2048
How many playable moves were there?StockfishMultiPV 4-5, 15s per position
Is this endgame actually winnable?Stockfish + SyzygySyzygyPath set, or Lichess tablebase
Why is this sacrifice good?lc0big net, watch D and Q separately
Would a human find this move?lc0--verbose-move-stats, read P
What will my opponent play?Maia at their ratinggo nodes 1
Why do I keep losing this structure?Maia at your ratingcompare its move to yours across 20 games
Sparring at a target levelStockfishUCI_LimitStrength true, UCI_Elo 1600

On that last row: Stockfish’s UCI_Elo runs from 1320 to 3190, and the handicap is implemented by making the engine pick deliberately worse moves. It plays strong chess punctuated by inhuman gifts. Maia at the equivalent bucket plays plausibly weak chess, which is much better sparring, because the mistakes it makes are the mistakes you need practice punishing.

A review protocol you can actually sustain

Skip the “run game review, read the numbers” ritual. It produces the illusion of work. Try this instead, four passes, roughly 25 minutes per game:

Pass one, no engine. Play through your game and mark every position where you remember being uncertain. Write your candidate moves and a one-line reason. Five to ten marks is typical.

Pass two, Stockfish at your marks only. MultiPV 3, ten seconds each. For each, note the win-percentage spread between move one and move three. If it’s under 5 points, you agonised over nothing and the lesson is about time management. If it’s over 20, you found a genuinely critical moment, and that is the position to study.

Pass three, the full scan. Now let the engine sweep the whole game at depth 18 and list every move whose win-percentage drop exceeded 15 points. Cross-reference against your marks. The blunders you didn’t mark are the important ones: you weren’t uncertain, which means you didn’t see the problem at all. Those go in a spreadsheet with a tag.

Pass four, Maia. At the critical positions from pass two, ask Maia at your rating what it plays. If Maia plays your move, you made a normal mistake for your level and the fix is a pattern. If Maia plays something better than your move, the fix is probably attention, not knowledge.

Automating pass three takes about fifteen lines of python-chess:

import math, chess.engine, chess.pgn

def win_pct(cp):
    return 50 + 50 * (2 / (1 + math.exp(-0.00368208 * cp)) - 1)

eng = chess.engine.SimpleEngine.popen_uci(r"C:\engines\stockfish.exe")
eng.configure({"Threads": 7, "Hash": 2048})

game = chess.pgn.read_game(open("mygames.pgn"))
board = game.board()
prev = 50.0
for move in game.mainline_moves():
    board.push(move)
    info = eng.analyse(board, chess.engine.Limit(depth=18))
    cur = win_pct(info["score"].white().score(mate_score=10000))
    if abs(cur - prev) >= 15:
        print(board.fullmove_number, board.peek(), round(prev, 1), "->", round(cur, 1))
    prev = cur

Run that over fifty games and you get something no game-review button will give you: a frequency table of your own failure modes. The output is usually embarrassing and specific. Eleven of thirty blunders allowing a knight fork within two moves. Six from moving a piece that was holding a back rank. Four from capturing on a square defended twice.

That table is your training plan. A tactics set aimed at the fork count is worth forty times a general “do 100 puzzles a day,” and you can only build it by turning the engine from an oracle into an instrument.

Pick your last twenty rated games and run the script tonight. Whatever tops the frequency table is what you study this month.

In this section

The supporting pages under this subject.