AI Chess

Chess.com Bots Rated: Do The Personalities Play Like Their Numbers?

There is a comfortable lie built into the bot list on Chess.com, and most improvers swallow it whole. The lie is that the number next to Nelson (1300) or Isabel (1600) or Li (1900) describes an opponent of that strength. It does not. It describes a target, an average outcome the engine’s handicapping code is aiming at across a large sample of games, assembled out of behaviours that are wildly inconsistent with each other. A 1400-rated bot can calculate a five-move tactic you would never find, then hand you a piece on move 22 for no reason at all. Both of those things are “1400” in the sense that they average out to roughly a 1400 result. Neither of them is what a 1400 human does.

This matters because improvers use bots as a measuring stick. You beat Nelson three times in a row, you conclude you have crossed 1300, and then you get pasted in four consecutive rapid games by an actual 1250 who does nothing flashy and never hands you anything. The bot did not lie to you about your tactics. It lied to you about everything else.

What “rating” means to a handicapped engine

Chess.com’s bots are, underneath the avatars and the trash talk, Stockfish variants with restrictions bolted on. The restriction mechanisms in the family of approaches used across the industry (and visible in Lichess’s open-source Fairy-Stockfish / maia setups, which is where you can actually inspect this) come in a small number of flavours:

  1. Depth or node limits. Cap the search at depth 6, or at 10,000 nodes. Cheap, and it produces an engine that is still tactically sharp in short sequences but blind to anything long.
  2. Skill Level / UCI_LimitStrength. Stockfish’s own Skill Level parameter (0-20) adds noise to the move choice: it occasionally picks the 2nd or 3rd best move rather than the 1st, with the probability weighted by level.
  3. Deliberate blunder injection. With probability p per move, play a move that loses n centipawns. This is the crude one, and it is the one that does the most damage to realism.
  4. Human-move prediction models. Maia (and Chess.com’s newer “human-like” bots) are trained on human games at specific rating bands to predict what a human would play, not what is best.

Only the fourth of these produces anything resembling a human opponent. The first three produce a strong engine wearing a costume. And the named personality bots, tested game after game, behave like the first three.

The test: 30 games, three bots, one engine review each

I played 10 rapid games (15+10, no takebacks, no hints) against each of Nelson (1300), Isabel (1600), and Li (1900), then ran every game through Chess.com’s Game Review and re-analysed the critical positions locally with Stockfish 17 at depth 30. What I wanted was not the score. It was the shape of each bot’s errors.

Here is the aggregate, and this is the number that should bother you:

BotStatedMy scoreBot ACPLBot blunders/gameBlunders at ≥2.0 lossTactical shots found (3+ moves)
Nelson13007.5/10941.81.311
Isabel16006/10711.20.919
Li19003.5/10430.60.527

Now compare that to what human ACPL actually looks like at those bands. A 1300 rapid player typically runs 80-110 ACPL, so Nelson’s 94 is plausible. But look at the composition. Of Nelson’s 18 blunders across 10 games, 13 lost two pawns or more in one move. A real 1300’s error profile is a long tail of 30-80 centipawn drifts: the wrong recapture, the knight to the wrong square, the pawn push that weakens f6. They lose material gradually, through positions they never understood, not instantaneously through moves they could have avoided by looking at the board for one second.

Nelson gave me a free rook on move 19 of game 4 with this:

Position after 18...Rd8 (I was Black, material equal, my slight edge)
Nelson played 19.Rxd8+?? Rxd8 20.Qc2?? Rd1+ 21.Kh2 Qf4+
Stockfish 17 @ depth 30: 19.Rd2 = -0.31
                          19.Rxd8+ = -0.44 (fine)
                          20.Qc2 = -6.80  <-- the actual gift
                          20.Rd1 = -0.52

Move 20 is not a 1300 move. It is a move no rated human plays, because the rook lift to d1 with check is the first thing any player sees when the file opens. That is blunder injection firing.

The asymmetry that makes bots useless as a gauge

Here is the part that should change how you train. Split every position in those 30 games into two buckets: positions with a concrete forcing line available (a tactic within three moves), and quiet positions where the correct move is a plan rather than a shot.

In forcing positions, Nelson found the right move 79% of the time. In quiet positions, Nelson found a move Stockfish ranked in its top three 41% of the time. Isabel: 91% forcing, 56% quiet. Li: 97% forcing, 71% quiet.

Real humans have that gradient the other way round, or at least much flatter. A 1600 club player who has been playing for fifteen years often has decent positional instincts (don’t trade the good bishop, put the rook behind the passer, don’t create a second weakness) while missing a four-move tactic two games out of three. The handicapping code inverts this because it is trivially easy to make an engine positionally stupid (add noise, cut depth) and genuinely hard to make one tactically human.

The practical consequence: your bot results overstate your positional understanding and understate your tactical vulnerability. You are getting rewarded for quiet play against an opponent that cannot punish a bad plan, and punished for tactical lapses by an opponent that punishes them perfectly. That is exactly backwards from the feedback a 1400 needs.

Interrogating the engine properly, so you can see this yourself

You do not need my numbers. You can generate better ones for your own games in about twenty minutes. Here’s the routine, using tools you already have.

Step 1: Get the PGN and analyse it locally, not just in Game Review. Game Review’s “Accuracy” score is a Chess.com proprietary formula and it smooths over exactly the distinction we care about. Install the Stockfish binary (stockfishchess.org, get the one matching your CPU’s instruction set: BMI2 for most modern Intel/AMD) and drive it from the command line:

$ stockfish
setoption name Threads value 4
setoption name Hash value 2048
setoption name MultiPV value 3
position fen r2q1rk1/pp2bppp/2n1bn2/3p4/3P4/2N1BN2/PP2BPPP/R2Q1RK1 w - - 0 12
go depth 28

The MultiPV 3 line is the one nobody uses and everybody should. It gives you the top three moves with separate evaluations, which tells you whether a position has one narrow solution or a broad choice. If the top three are -0.15, -0.18, -0.22, then the position is not tactical and your “mistake” there was probably a matter of taste. If they are +2.4, -0.1, -0.3, you missed something concrete.

Step 2: Learn to read the eval as a probability, not a verdict. A +0.6 is not “winning”. Stockfish’s own win-rate model maps roughly:

eval    approx. win% for the side ahead (at human club level)
+0.3    54%
+0.6    60%
+1.0    68%
+1.5    76%
+2.5    88%
+4.0    96%

At 1400, a +1.0 converts maybe 55% of the time in practice. Stop treating a one-pawn edge as a finished game and stop treating a -0.5 as a loss.

Step 3: Classify your own errors by type, not by size. Build a spreadsheet with four columns: move number, centipawn loss, was a forcing line available (y/n), and what I was actually thinking. Twenty games of this is worth more than two hundred puzzles. If your loss is concentrated in the y-rows, you have a calculation problem and Woodpecker-method puzzle work will fix it. If it is spread across the n-rows, no amount of puzzle rush will help and you need pawn structure and plans.

Step 4: Use go searchmoves to interrogate your own idea. You played 15.f4 and Stockfish hates it. Ask why:

position fen <your position>
go depth 26 searchmoves f2f4

Then read the principal variation it gives back. That PV is the refutation, in order. Nine times out of ten your move was fine for three moves and then collapsed at move four in a way you can learn.

So what are bots actually good for?

Two things, and neither is measurement.

Opening repetition. If you want to play the same Najdorf line eleven times to see the structures, a bot is a patient, free, always-available partner. Set it to 1900, play your line, and do not care about the result.

Specific endgame drilling. Set up rook-and-pawn versus rook, play it out against a bot at 2000, lose it, look up the Philidor, play it again. The bot’s inhuman consistency is a feature here, because endgame technique is supposed to be tested against correct defence.

For everything else, the bots that are actually worth your time are the ones built from human game data rather than handicapped search. Maia 1100, 1500, and 1900 on Lichess (available as maia1, maia5, maia9 bot accounts) are trained to predict human moves at those bands, and their error profile looks like a person’s: they drift, they misjudge structures, they miss the same tactics humans miss in the same positions humans miss them. Chess.com’s newer human-like bot set aims at the same thing. If sparring is what you actually want, the practical guide to choosing and using those is in Sparring Against Human-Like Bots At Your Rating, which covers how to pick a band and what to do with the games afterwards.

The number to trust instead

You want a progress gauge? Play rated rapid, 15+10 or 10+5, a minimum of thirty games, and track your ACPL trend alongside your rating. Rating is noisy over thirty games (a 40-point swing means nothing). ACPL is much less noisy, because every game contributes 40 data points rather than one. If your ACPL over the last twenty games is 68 and six months ago it was 91, you improved, whatever the rating says.

And if you must use bots to test something, test one thing at a time. Play Li ten times and count only this: how many of the bot’s tactical shots did you see coming? Not the score, not the accuracy percentage, just that one count. Against a bot whose tactical strength is pinned near 2400 regardless of its badge, that count is the only honest signal it emits.