AI Chess
§6.1 Engine Setup: GUIs, Settings And Cloud Depth 2,165 words · 10 min

Stockfish Hash And Threads Settings That Actually Change Your Analysis

Almost every club player who installs a local engine leaves it on defaults, and the defaults are deliberately tiny. Stockfish 17 ships with Hash at 16 MB and Threads at 1, because the developers cannot know whether you are running it on a Raspberry Pi or a 64-core server. On the machine you actually own, those two numbers are the difference between an engine that reaches depth 24 in a minute and one that reaches depth 33 in the same minute. That gap matters most in exactly the positions you are trying to learn from: closed middlegames, fortress-ish endings, and long tactical sequences where the refutation is eleven moves deep.

This page is about getting the stockfish hash and threads settings right for your hardware and your task, with numbers you can check yourself. It is not a general engine tour. If you need the wider view of GUIs, cloud evaluation and when local hardware stops being the bottleneck, the parent guide on engine setup, GUIs, settings and cloud depth covers that ground.

What the two options are really doing

Hash is the transposition table. Every position Stockfish evaluates gets a compact record stored in memory: a partial hash key of the position, the best move found, the score, the depth that score was proved to, and a flag saying whether the score is exact or just a bound. When the search reaches the same position by a different move order (1.e4 e5 2.Nf3 Nc6 3.Bb5 and 1.e4 e5 2.Bb5 is illegal, but 2.Nf3 Nc6 3.Bc4 versus 2.Bc4 Nc6 3.Nf3 transposes perfectly), it looks the answer up instead of re-searching it. In Stockfish the table is laid out as clusters of three entries in 32 bytes, so roughly 100,000 positions fit per megabyte.

Threads is the number of parallel search workers. Stockfish uses Lazy SMP: every thread runs its own near-identical search with slightly different depth and move-ordering offsets, and they cooperate only by hammering the same shared transposition table. There is no clever work-splitting tree. Threads find something useful, dump it in the hash, and other threads pick it up. Which is why the two options are coupled: more threads pouring into a table that is too small just means more collisions and more overwritten entries.

The 16 MB default is the biggest single waste at club level

Do the arithmetic on a typical modern laptop. A 2023-vintage machine runs Stockfish 17 with NNUE at roughly 2 million nodes per second on one thread. Ask it for 30 seconds of thinking and you have generated about 60 million nodes. The default 16 MB table holds about 1.6 million entries. You are asking a filing cabinet with 1.6 million slots to hold 60 million documents, and it deals with that by throwing away almost everything.

Symptoms are specific and recognisable. The evaluation jumps around between depths for no apparent reason. Depth climbs quickly to 20 and then crawls. You stop the analysis, restart it on the same position, and get a visibly different principal variation. Re-examining a position you looked at two moves ago takes just as long as the first time, because nothing survived.

Bump the same run to 512 MB and the engine will typically reach two to four plies deeper in the same wall-clock time. Two extra plies does not sound dramatic until you notice that a two-ply-deeper search is what separates “+0.4, White is a bit better” from “+2.1, White has a forced win of a piece in six”.

A hash-sizing rule you can do in your head

The rule that actually works: hash in megabytes ≈ nodes per second, divided by 100,000, multiplied by your search time in seconds.

Take a 6-thread search at 9 Mnps that you plan to leave running for 60 seconds:

9,000,000 / 100,000 = 90
90 × 60 = 5,400 MB

Round to 4096 MB if you have the RAM free, and you will be in good shape. The formula overstates the need slightly because quiescence nodes are not all stored and real games have heavy transposition, so landing at half the calculated figure is generally fine. Landing at a twentieth of it, which is what defaults do, is not.

Practical ceilings by machine, assuming you want the operating system and your browser to keep working:

RAM installedReasonable HashNotes
8 GB1024 MBClose Chrome first; 2048 will swap
16 GB2048–4096 MB4096 is safe with a light desktop
32 GB8192 MBDiminishing returns past this for 1-minute searches
64 GB+16384 MBOnly pays off for multi-hour correspondence analysis

The hard rule underneath all of this: never set Hash larger than your genuinely free physical memory. If the table spills into the page file, every hash probe becomes a disk read and your effective nps collapses by an order of magnitude. A 4 GB hash on a machine with 3 GB free is far worse than a 256 MB hash. Check Task Manager’s Performance tab for “Available” memory before you commit, not “Total”.

One piece of folklore worth retiring: you no longer need powers of two. Stockfish rounded hash size down to the nearest power of two in versions before 12; current releases index the table with a multiply-shift and use the full allocation, so 3000 MB really is 3000 MB. Powers of two remain a tidy habit, not a requirement.

Reading hashfull, because the engine tells you when it is full

Every UCI engine reports table saturation, and almost no GUI shows it to you by default. Run Stockfish from a terminal instead:

> uci
> setoption name Hash value 256
> setoption name Threads value 4
> position fen r1bqkb1r/pp1n1ppp/2n1p3/3pP3/3P4/2N2N2/PP3PPP/R1BQKB1R w kq - 0 8
> go movetime 60000

info depth 28 seldepth 38 multipv 1 score cp 31 nodes 412883091
  nps 6881384 hashfull 1000 tbhits 0 time 60000 pv f3g5 f8e7 g5h7 ...

hashfull 1000 is per-mille. A thousand means completely saturated. Anything above about 700 and Stockfish is discarding useful work to make room. Raise Hash to 2048 on that same position and you would expect hashfull around 300 to 500 at 60 seconds, with depth landing two or three plies higher.

Nibbler exposes this in its info panel, En Croissant and Banksia GUI both surface it, and raw stockfish.exe in PowerShell always will. It is the only feedback loop that tells you whether your hash setting is right for your actual time controls rather than for a rule of thumb.

Threads: why 8 threads is not 8 times faster

Lazy SMP scales, but not linearly. Measured in time-to-depth on typical middlegame positions, the rough picture is:

ThreadsSpeedup vs 1 threadElo gain (approx)
21.7×+45
42.9×+80
84.8×+110
167.5×+135

Two things follow. First, going from 1 to 4 threads is a huge win and always worth doing. Second, going from 8 to 16 buys you about a quarter of a ply at fixed time, which will not change a single verdict in a game review.

Set Threads to your physical core count minus one, not your logical thread count. An Intel Core i7-12700H reports 20 logical processors, but that is 6 performance cores with hyperthreading, 8 efficiency cores, and a scheduler that will happily park Stockfish workers on the slow cores. Setting Threads 20 on that chip is frequently slower than Threads 10. The efficiency cores run at a lower clock, and Lazy SMP is only as fast as the threads that finish their iterations. Start at physical performance cores minus one and test upward.

On anything with two CPU sockets or an AMD Threadripper, Stockfish 17’s NumaPolicy option matters more than the thread count. Leave it on auto and it will bind workers sensibly across NUMA nodes. Setting it to none on a dual-socket box can cost 30% of your nps to cross-node memory traffic.

The nondeterminism tax, which nobody warns you about

Run Stockfish with Threads 4 on the same position twice and you will get different node counts, sometimes different depths, and occasionally a different move. This is intrinsic to Lazy SMP: thread timing is not reproducible, so what lands in the hash table is not reproducible either.

For game review this does not matter. For two specific jobs it matters a great deal. If you are building an opening file and want to be able to re-verify an evaluation you wrote down six months ago, set Threads 1 and pin your search with go depth 30 rather than a time limit. Same engine version, same depth, one thread, and you get the same answer every time. Likewise if you are comparing two candidate moves and the gap is 0.15 pawns, a multi-threaded search cannot tell you whether the difference is real or scheduling noise. Drop to one thread, raise the depth, and the comparison becomes meaningful.

Worked example: reviewing one game on a 16 GB laptop

Say you have an 8-core machine (8 physical, 16 logical), 16 GB RAM, and a 40-move rapid game to go through in Nibbler.

Set Threads 6, leaving two cores for the GUI and the OS. Set Hash 2048. Set MultiPV 3, because for training purposes you want to see the second and third best moves, not just the engine’s pick. The single-best-move display tells you what to play; the top three tell you why your move was bad, which is the thing you can actually learn from.

Those settings on that hardware give roughly 8 to 10 Mnps. At 20 seconds per position that is 180 million nodes and hashfull around 550, which is comfortable. Forty moves at 20 seconds is about fourteen minutes of compute, and you will be reading and thinking for far longer than that anyway.

Now the same machine, same evening, but you have found one critical position and want a real answer. Change to Threads 1, Hash 1024, MultiPV 1, and go depth 35. It will take several minutes. It will be reproducible, it will be deeper than any multi-threaded 20-second glance, and you can write the resulting evaluation in your notes with a depth number next to it and trust it later.

Settings by task

TaskHashThreadsMultiPVSearch limit
Blunder-check a batch of games512 MBcores − 115s/move
Study one game properly2048 MBcores − 1320s/move
Verify a critical position1024 MB11depth 32–35
Build an opening file1024 MB14depth 28 fixed
Endgame with Syzygy tablebases512 MBcores − 1110s/move

The endgame row has a smaller hash on purpose. Once SyzygyPath is set and you are inside 6 or 7 pieces, tablebase hits replace search entirely, and tbhits in the info line will climb into the millions while the transposition table does comparatively little.

Three traps that cost real depth

Changing Hash mid-analysis wipes the table. Stockfish reallocates and zeroes it, so if you have been thinking for three minutes and then raise the hash, you start from nothing. Set it before you begin. The same applies to the Clear Hash button, which is useful when you want a clean measurement and destructive when you hit it by accident.

Windows users can get 5 to 10% extra nps by enabling large pages. Stockfish attempts this automatically for the hash allocation, but it requires the “Lock pages in memory” privilege, granted through secpol.msc under Local Policies then User Rights Assignment, followed by a sign-out. The engine’s startup line tells you whether it worked. If your hash is 4 GB or larger, this is the easiest free speedup available.

Cloud evaluation beats your hash settings for common positions and loses badly for your own. Lichess serves cloud analysis at depth 40+ with MultiPV 5 for positions someone has already computed, which is most opening theory and nothing from move 20 of your own game. Knowing which side of that line you are on saves you from tuning a local engine for work the server has already done.

Pick one position from your last loss where the engine’s verdict surprised you. Run it at Hash 16, Threads 1, 30 seconds. Then Hash 2048, Threads 6, 30 seconds. Write down both principal variations and both hashfull figures. The size of the gap on your specific hardware is the only number in this piece you should actually trust.