AI Chess

NNUE In Plain English: Why Stockfish Grew A Neural Network

Somewhere in your chess education, you absorbed a rule that sounded permanent. Doubled pawns are a weakness. A knight on the rim is dim. Two bishops are worth roughly a quarter pawn on an open board. A rook on the seventh is worth a pawn of pressure. Those rules didn’t come from nowhere: they came from a century of master annotation, and for a long stretch they were also literally what chess engines believed, because someone had typed them in.

That’s the part most club players don’t know. Until August 2020, Stockfish evaluated positions using a hand-written scoring function. A human programmer decided that an isolated pawn on a half-open file costs you some number of centipawns, and a different human argued on a forum about whether that number should be 12 or 17. The engine was strong because the search was brilliant and the rules were tuned to death against millions of games, but the rules themselves were still statements somebody wrote down in C++.

Then Stockfish 12 shipped with NNUE, and the rules went away.

So what is NNUE in Stockfish, concretely

NNUE stands for Efficiently Updatable Neural Network, written backwards, because it came out of Japanese shogi programming (Yu Nasu’s 2018 paper) before Hisayori Noda ported the idea to Stockfish. Two things about it matter for you as a player.

First, it’s a learned evaluation. Nobody wrote “bad bishop” into it. The network was trained on tens of millions of positions, each labelled with a score, and it worked out its own internal features. Those features do not map onto words. There is no line of code inside Stockfish 17 that says anything about doubled pawns, and yet it plays doubled-pawn structures better than any human.

Second, and this is the “efficiently updatable” bit, it runs fast on an ordinary CPU. Leela Chess Zero uses a big convolutional network and really wants a GPU. NNUE is deliberately small, and the trick in the name is that when you make a move on the board, the network doesn’t recompute from scratch. Only a couple of inputs changed (a piece left one square, arrived at another), so the first layer just gets patched. That means Stockfish can still search tens of millions of nodes per second on your laptop while calling a neural net at every leaf. On a modern eight-core machine you’ll see something like 8–20 Mnodes/s. Leela on the same machine without a GPU does maybe 1500 nodes/s, and compensates with much better positional judgement per node.

If you want the fuller picture of how these engines differ in temperament and what each is good for, the engines compared breakdown covers Stockfish, Leela, Maia and the NNUE family side by side.

The strength jump, in numbers

Stockfish 11, the last classical-eval version, sat around 3450 Elo on the CCRL 40/15 list. Stockfish 12 with NNUE came in roughly 130 Elo higher on identical hardware. That is an enormous single-release gain: for context, the whole of 2016–2019 development added about 130 Elo combined. Stockfish 17 is now near 3650. The classical evaluation was kept as a fallback for a while, then removed entirely in Stockfish 16 (2023). There is no hand-written eval left to fall back to.

Here’s what that means for you in practice: any chess advice written before late 2020 that cites engine evaluations was quoting a different engine with different opinions.

Where the old advice and the new engine disagree

The disagreement isn’t uniform. NNUE didn’t shift every evaluation by a bit. It shifted specific kinds of position a lot, and left others nearly untouched.

Closed positions. The classical eval was built on mobility counts and pawn-structure penalties, and it was persistently too generous to the side with the “better” structure in a blocked position. NNUE learned that a blocked position is about who can generate play on the flanks, which pieces have routes, and whether the king can be opened at all. Load a King’s Indian Petrosian structure into Stockfish 17 and you’ll often get an evaluation near 0.00 where Stockfish 11 would confidently say +0.60 for White.

Sacrifices. This is the biggest and most useful change. Old Stockfish was material-anchored: a piece for two pawns and some initiative read as roughly -0.80 until the search could actually see the payoff, which for a long-term positional sacrifice is well beyond any practical horizon. NNUE evaluates compensation as a pattern. Exchange sacs on c3, the Greek gift, a rook for bishop-plus-attack, piece-for-three-pawns in a locked centre: the network has seen tens of thousands of these and recognises the shape.

Fortresses. Still a weakness, honestly. NNUE improved fortress detection but doesn’t solve it. You’ll still see +2.5 on positions that are dead drawn by human standards. Knowing this stops you trusting a number where it shouldn’t be trusted.

A worked example you can run right now

Take the Marshall Attack exchange sacrifice line. Or better, take something shorter you can set up in thirty seconds. Open Lichess’s analysis board, go to the FEN input, and paste this Sicilian Dragon Yugoslav position after White has played the standard sacrificial build-up:

r2q1rk1/1b1nbppp/p2ppn2/1p4B1/3NPP2/2N2Q2/PPP3PP/2KR1B1R w - - 0 12

Now run the engine and watch the top line for ten seconds. Then do something most club players never do: read the whole output rather than just the number. In the Lichess analysis board, turn on multiple lines (the “lines” slider, set it to 3) and note what happens to the second and third choice as depth increases. That volatility is the actual signal.

The raw engine output in a terminal looks like this, and the fields are worth learning:

info depth 28 seldepth 38 multipv 1 score cp 47 nodes 41830912
     nps 12938... pv e5 dxe5 fxe5 Nh5 Qh3 Nxg5 ...
info depth 28 seldepth 41 multipv 2 score cp 12 nodes ...
     pv Bxf6 Nxf6 e5 dxe5 fxe5 Nd7 ...

score cp 47 means 0.47 pawns in the side-to-move’s favour. depth 28 is the nominal full-width depth; seldepth 38 is how far the most interesting line was actually chased, and the gap between them tells you the position has forcing content. nodes divided by nps gives you the time spent. When multipv 1 and multipv 2 are within 0.35 of each other, you are not looking at a best move, you’re looking at a position with several playable moves and an engine breaking a near-tie on tiny margins.

The thing that makes engine numbers usable

Stop reading the evaluation as a truth claim and start reading it as a stability claim. Three questions, every time:

  1. Does the number move as depth increases? Let it run from depth 20 to depth 35. An eval that sits at +0.4 the whole way is a positional assessment you can learn from. An eval that swings +0.3 → +1.8 → +0.6 is a tactical position where the engine is discovering things, and the final number is about a concrete line, not a feature of the position.

  2. How far apart are the top three moves? A 1.5-pawn gap between best and second-best means the position has one move and you should know why. A 0.1 gap means the engine is expressing a preference, and copying it teaches you nothing.

  3. Is the eval coming from something a human can execute? If the top line is a 14-move forced sequence, the +2.0 is real but it is not yours to have. NNUE’s happy acceptance of sacrifices makes this trap worse than it used to be, because the network will cheerfully tell you a piece sac is +0.9 based on a pattern it can’t articulate and you can’t reconstruct over the board.

Turning this into a training plan

Pick ten of your own games from the last three months, the ones where you lost and weren’t sure why. Run each through Stockfish at depth 30 or so, using Lichess’s server analysis or the local engine in the desktop app. Don’t look at the blunder list.

Instead, find every position where your move and the engine’s top move differ by more than 0.6 and the evaluation was stable across depth. Those are the positional mistakes. Ignore the rest for now. You’ll probably get four or five per game, and a striking number of them will cluster: the same structure, the same piece left on the same bad square, the same premature pawn push.

Now the NNUE-specific part. For each of those positions, ask the engine what happens if you play the move your old rulebook told you to play. Play out the “correct” recapture that fixes your pawn structure, the developing move that avoids the doubled pawn, the retreat that saves the bishop pair. Quite often modern Stockfish will rate the “wrong” move higher, and the gap between your rule and the network’s opinion is the single most valuable piece of feedback available to a 1400. You are watching a 3600-rated player disagree with a book written in 1953.

One caution worth taking seriously: an engine at 3650 Elo is not a coach. Its evaluations are correct and its explanations are nonexistent, because there are no explanations inside it, just weights. The gap between “Stockfish says +0.8” and “I understand why this is better” is yours to close, and depth 40 won’t close it for you. What NNUE gives you is a far more honest sparring partner for positions the old rules were bad at judging: the murky, blocked, unbalanced, pawn-sac-flavoured positions that actually decide club games.

Use it there. That’s where the 130 Elo went.