AI Chess

Maia vs Stockfish: Human Mistakes vs Perfect Moves

Stockfish will tell you the best move in any position you give it. That is not usually your problem. Your problem is that you played 24.Rfd1 instead, and you want to know why that felt right at the time, whether your opponent was ever going to punish it, and which of your bad habits it belongs to. Stockfish has no opinion about any of that. Maia does.

The two engines are built to answer genuinely different questions, and once you see the split, the maia chess engine vs stockfish comparison stops being about strength ratings and starts being about what you can actually do with the output. If you want the wider landscape first, including how NNUE and Leela fit in, the engines compared overview covers the family. This page goes narrow: two engines, one analysis board, and a method for turning their disagreement into a study list.

Two engines, two different questions

Stockfish answers: what is the objectively strongest continuation, and how good is the resulting position?

Maia answers: what move would a human rated 1500 actually play here, and with what probability?

That second question has a different shape. Its output is not a single best move plus a number, it is a probability distribution over every legal move. Maia is a prediction model that happens to play chess, trained by the CSSLab group at Toronto and Microsoft Research on roughly twelve million Lichess games per rating band. There are nine separate networks, maia-1100 through maia-1900 at 100-point intervals, each one fitted to the moves of players in that band. Peak move-matching accuracy is around 52%, and crucially each model matches its own band better than the neighbouring ones. That is the whole trick: Maia is not a weak engine, it is a calibrated one.

Stockfish 17.1Maia (lc0 + maia weights)
QuestionBest move, position valueP(human of rating R plays move m)
MethodAlpha-beta search + NNUE evalOne neural net forward pass, no search
Outputcentipawns, WDL, PV linespolicy % per legal move
Typical depth25 to 40 ply0 ply
Strength~3600 CCRL~1100 to 1900, by design
Good forVerdicts, tactics, endingsPractical risk, blunder prediction, prep

Under the hood: search versus a single forward pass

Stockfish’s power is search. It evaluates leaf nodes with a small NNUE network (fast, quantised integer arithmetic) and then searches millions of nodes per second, pruning aggressively. Its evaluation of a position is really the evaluation of a position twenty-five moves deeper, backed up to the root. That is why its verdicts are trustworthy and why they tell you nothing about difficulty. A move that requires a nine-move forced sequence and a move that is obvious both print as score cp 120.

Maia deliberately throws search away. It is a Leela-architecture network run at exactly one node, so you only ever see the policy head’s raw prior. No lookahead, no tree, no rollouts. This matters if you try to “improve” it: crank lc0 up to 800 nodes with Maia weights and you get a stronger engine that is no longer human-like, because the errors humans make are precisely the errors that search removes. Keep it at one node or you have broken the instrument.

Getting both running on your own machine

You need the lc0 binary (from the LeelaChessZero releases) and the Maia weight files from github.com/CSSLab/maia-chess, in the maia_weights directory. Both are small, a few megabytes each, and run fine on CPU with the eigen backend. No GPU required.

$ lc0
setoption name WeightsFile value maia-1500.pb.gz
setoption name Backend value eigen
setoption name VerboseMoveStats value true
position startpos moves e2e4 e7e5 g1f3 b8c6 f1c4 g8f6 f3g5 d7d5 e4d5 f6d5 g5f7 e8f7 d1f3
go nodes 1

With VerboseMoveStats on, lc0 prints one line per legal move before its bestmove. The field you care about is P:, the policy prior. Everything else (N, Q, V) stays empty or zero because there was no search:

info string f7e6  (322) N: 0 (+ 0) (P: 19.4%) (Q: 0.00000) (D: 0.000) (V: -.----)
info string f7g8  (401) N: 0 (+ 0) (P: 33.8%) (Q: 0.00000) (D: 0.000) (V: -.----)

For Stockfish, three UCI options change everything about how usable the output is:

$ stockfish
setoption name UCI_ShowWDL value true
setoption name MultiPV value 4
go depth 30

UCI_ShowWDL adds wdl 412 574 14 (win/draw/loss in permille) to every info line, which is far more honest than centipawns. MultiPV 4 gives you the top four moves with separate evaluations, which is what you need for the gap metric below. Ignore UCI_LimitStrength and UCI_Elo (valid range 1320 to 3190) for training purposes: a throttled Stockfish plays five accurate moves and then drops a rook for no reason, which is not how anyone at your club loses.

If you would rather not touch a terminal, Lichess runs Maia as three bot accounts you can challenge directly: maia1 (1100), maia5 (1500) and maia9 (1900). Scripting it is easiest through python-chess, where engine.analyse(board, chess.engine.Limit(nodes=1), multipv=250) gets you the whole distribution in one call.

Worked example 1: the Fried Liver, where “only move” meets “nobody plays it”

Take the position after 1.e4 e5 2.Nf3 Nc6 3.Bc4 Nf6 4.Ng5 d5 5.exd5 Nxd5 6.Nxf7 Kxf7 7.Qf3+. Black has one move. 7…Ke6 holds the knight and, per Stockfish at depth 30, leaves White around +0.95: uncomfortable, playable, not lost. Anything else concedes material immediately (7…Kg8 8.Qxd5+, 7…Ke8 8.Bxd5).

Here is what the two engines report side by side. Treat the exact percentages as representative rather than gospel, since they shift with the weights version, but the shape is stable and you can reproduce it in ten minutes:

Black’s moveStockfish (d30)maia-1100maia-1900
Ke6 (only move)+0.9519%58%
Kg8+3.434%14%
Ke8+3.922%9%
everything else~+2.525%19%

Now compute the policy-weighted expected loss, which is what nobody does and everybody should. Multiply each move’s probability by how much worse it is than Ke6, then sum:

  • maia-1100: (0.34 × 2.45) + (0.22 × 2.95) + (0.25 × 1.55) = 1.87 pawns
  • maia-1900: (0.14 × 2.45) + (0.09 × 2.95) + (0.19 × 1.55) = 0.90 pawns

Stockfish’s +0.95 is the truth about the position. The 1.87 is the truth about the position as played by an 1100, and it is the number that justifies playing 4.Ng5 against a beginner and abandoning it against a 1900. Two engines, one arithmetic step, a real practical decision.

Worked example 2: flat eval, split policy

Flip the situation around. Queen’s Gambit Declined, Exchange Variation: 1.d4 d5 2.c4 e6 3.Nc3 Nf6 4.cxd5 exd5 5.Bg5 Be7 6.e3 0-0 7.Bd3 Nbd7 8.Qc2 c6 9.Nf3 Re8 10.0-0. White to move, and Stockfish at MultiPV 4 clusters four candidates inside 0.15 of each other: Rab1, Ne5, Rae1, h3. As a verdict, this is useless. Four moves, one eval, no guidance.

Run maia-1300 on it and the distribution is lopsided in a very specific way. The minority attack starter (Rab1, preparing b4-b5) collects almost nothing, because club players do not volunteer pawn moves on the side where they have fewer pawns. Mass concentrates on Ne5 and Rae1, plus a fair chunk on moves Stockfish rates 0.4 worse, like an early Ne2 or a2-a3. The engine agreement tells you the position is balanced; the policy split tells you which balanced move your opponent will be least prepared for, and which plan you have quietly been avoiding for years.

Reading Stockfish’s numbers without fooling yourself

Centipawns are not linear in anything you care about. Modern Stockfish normalises its output so that +1.00 corresponds roughly to a 50% win expectancy between equal engines, which means a +0.35 you have been agonising over is close to noise at 1500. Lichess makes the conversion explicit, and you can do it with a calculator:

winPct  = 50 + 50 * (2 / (1 + exp(-0.00368208 * cp)) - 1)
accuracy = 103.1668 * exp(-0.04354 * (winPctBefore - winPctAfter)) - 3.1669

Plug in the Fried Liver’s +95: winPct = 58.7%. Not “White is better,” just 59-41. Now suppose you then blunder into -300, which is 24.9%. The drop of 33.8 percentage points feeds into the accuracy formula and scores that single move at 20.5%. This is exactly how the accuracy percentage on your Lichess game report is built, and knowing the formula stops you treating a 91% game and an 88% game as meaningfully different.

Three more habits worth forming. Watch the eval stabilise rather than reading the first number that appears: depth 12 and depth 30 frequently disagree by a pawn in sharp positions, and the shallow number is the one that flatters your intuition. Treat score cp 0 with suspicion, because a dead-drawn evaluation that depends on a unique 30-move perpetual is a practical loss for you. And when Stockfish reports score mate 7, check whether the mate needs a quiet move, since those are the ones humans miss.

Using Maia on your opponent, not just yourself

The most underused application is prep. Pick a line you play, walk it to the branch point, and ask maia-1500 what the opponent’s distribution looks like. You are hunting for a specific shape: a move with high policy mass and a poor Stockfish evaluation. That is a trap that works on the population rather than on one person.

A second application is calibration. Export your own games (https://lichess.org/api/games/user/<you>?max=200&perfType=blitz,rapid), replay every position where you were to move, and count how often each Maia model’s top move equals yours. If maia-1500 matches 47% of your choices and maia-1900 matches 38%, your move-selection habits look like a 1500’s regardless of what your rating says this week. Rerun it in three months. The number moving is a better progress signal than rating, which is polluted by time trouble and opponent variance.

Where Maia falls over

No search means no tactics beyond pattern recall, so Maia’s policy on a position requiring a five-move combination is close to uninformative about whether a strong club player finds it. The weights are frozen on Lichess blitz and rapid games from a particular era, which bakes in that era’s opening fashions: it has no idea what happened to the London System since. Clock state is invisible to it, so it cannot tell you that the 34% blunder becomes 60% with eleven seconds left. And the classic Maia models are not you, they are a rating band, which is why Maia-2 exists: one model conditioned on both players’ ratings, installable with pip install maia2, that gives per-move probabilities plus a win estimate for a specified Elo pair.

Endgames are the clearest no-go zone. For anything with seven pieces or fewer, skip both engines and query the Lichess tablebase, which returns provably correct results with distance-to-mate. Maia’s opinion about a rook endgame is an opinion about how humans butcher rook endgames, which is interesting sociology and terrible technique.

A four-week loop you can actually run

Week one, play twenty rapid games against maia9 on Lichess and resist the urge to analyse as you go. Week two, run each game through Stockfish at MultiPV 3 and flag every position where your move cost more than 100cp of win expectancy, which usually lands between six and fifteen positions across twenty games. Week three, take each flagged position and ask maia-1500 for its distribution, then split your list in two: positions where Maia also puts most mass on a bad move (a population-level blind spot, worth drilling as a pattern) and positions where Maia’s top move was fine and you found something worse (a personal habit, worth writing down in words). Week four, drill only the first list, and reread the second.

The split matters because the two piles need different medicine. Population blind spots respond to repetition. Personal habits respond to a sentence you can say to yourself at the board, of the form “before I move a rook, check whether the knight is loose.”

One position, three questions

Next time you open an analysis board, run the position through all three: what does Stockfish say is best, what does maia-your-rating say you would play, and what does maia-opponent’s-rating say they will play. The first is a verdict, the second is a diagnosis, the third is a weapon. Most players collect the first, have never seen the second, and never realise the third was available for the price of a 5MB download.