For the puzzle-savvy
How Guess Grid ranks a guess
A good guess isn't the one most likely to be right. It's the one that tells you the most, whatever the answer turns out to be. Here is exactly how Guess Grid measures that, with real numbers from its five-letter word lists.
- Every guess sorts the words into groups
- Scoring a split
- The best openings
- Why not "fewest words left"?
- A small bonus for words that could win
- The opening shortcut
- Double letters
- How practice grades you
- The word lists
1. Every guess sorts the words into groups
Take a puzzle down to six possible answers that differ only in their first letter:
- BATCHCATCHHATCHLATCHMATCHPATCH
Guess WATCH and every one of the six shows the same colours:
Whatever the answer is, you're left with all six. One group, nothing learned, a wasted turn.
Now guess BLIMP, which can't be the answer but tests four of the first letters at once. Each answer shows a different pattern, except for two:
Five groups. Four of them hold a single word, so four times out of six you know the answer for certain. This is the whole idea: the colours you'll see sort the possible answers into groups, you end up holding exactly one group, and a guess is good when those groups are many and small.
2. Scoring a split
Guess Grid scores a split by how much information it gives you on average, measured in bits. If a guess divides N possible answers into groups of sizes n1, n2, …, the chance of landing in group i is ni / N, and the score is:
bits = − Σ (ni / N) × log2(ni / N)
One bit halves the field. WATCH, with a single group, scores 0 bits. BLIMP scores 2.25 bits, close to the most possible for six words, which is 2.585, when every word lands in its own group.
The score rewards two things at once: more groups, and more even ones. A split into many groups where one holds most of the words scores worse than fewer groups of similar size.
3. The best openings
At the start, a guess faces every possible answer. Here is how three openings split them. Each segment is one group, largest first; groups too small to draw are shown as a striped tail.
| Opening | Groups | Largest group | Bits | Avg words left | Rank by bits | Rank by avg |
|---|---|---|---|---|---|---|
| SLATE | 148 | 209 | 5.881 | 66.5 | 1st | 12th |
| RAISE | 131 | 161 | 5.874 | 57.8 | 2nd | 1st |
| CRIME | 111 | 324 | 5.211 | 111.9 | — | — |
Ranks are among every word the solver can play as a guess. "Avg words left" is how many words you'd have left on average after the guess. Rank by bits includes the small bonus for words that could win, described in section 5; all three of these could.
SLATE and RAISE are 0.007 bits apart, a statistical tie. SLATE makes more groups; RAISE keeps its largest group smaller. CRIME is a reasonable guess, better than 93% of words, but its largest group holds 324 words, and that is where it loses.
4. Why not "fewest words left"?
The obvious alternative is to rank guesses by how many words they leave on average. It's easy to explain, and by that measure RAISE is the best opening and SLATE only 12th. So which measure is right?
Rather than argue, Guess Grid's solver was played against every answer, both ways, always taking its own top suggestion:
| Ranking by | Opening | Avg guesses | Most guesses |
|---|---|---|---|
| Five letters, every answer in the list | |||
| Bits (as shipped) | SLATE | 3.437 | 6 |
| Fewest words left | RAISE | 3.497 | 5 |
| Six letters, every answer in the list | |||
| Bits (as shipped) | SALTER | 3.114 | 5 |
| Bits, with its own top opening | SATIRE | 3.124 | 5 |
| Fewest words left | SATIRE | 3.129 | 5 |
Bits solved faster on average at both lengths. It isn't a clean sweep: at five letters, a single answer took the shipped solver six guesses, while ranking by words left never needed more than five. Guess Grid optimises the average, the way you'd judge a solver over many puzzles, and says so.
Why the difference? Averaging words left squares each group's size, so the largest groups dominate it. Bits weigh every group by its share of the words, which rewards breaking up all of them. Neither measure looks past the next guess, so which one plays better over a whole game can only be found by playing. On these lists, bits won.
5. A small bonus for words that could win
Information alone can't see winning. A guess that could itself be the answer might end the game on the spot; a probe that splits the field just as well can't. So a word that is still possible gets a small bonus of 0.15 bits on top of its score. It's enough to break near-ties in favour of a word that could win, and too small to prefer a possible answer that splits the field badly.
That's why the suggestions are labelled. Possible answer could win now. Narrows it down can't be the answer, but it earns its place by splitting the field better than any possible answer would. And once only one or two words are left, the solver stops probing and plays one of them.
6. The opening shortcut
Scoring every allowed guess against every possible answer takes millions of pattern comparisons, every time the screen updates. That's fine for a single calculation, and it's what produced the tables above, but too slow to repeat on a phone every time you tap.
So while a great many words are still possible, which in practice means the opening turn, Guess Grid ranks the possible answers by how common their letters are among the words left, counting a letter in the right position double. From the second guess the full calculation takes over. The app says as much under the opening suggestions: "Sharper ranking kicks in once fewer words remain."
The shortcut does well. At five letters it picks SLATE, the same word as the full calculation. At six it picks SALTER over the full calculation's SATIRE, and SALTER played better, 3.114 guesses on average against 3.124. No one-turn measure predicts a whole game perfectly; this is where that shows.
7. Double letters
The tempting rule is "a red letter isn't in the word". It's wrong whenever a guess repeats a letter. Guess SPEED when the answer is ABIDE:
The second E is red, yet the answer contains an E. Each colour is handed out the way puzzles score it: greens first, then ambers left to right, one per matching letter in the answer. ABIDE has one E, the first E in SPEED claims it, and the second has nothing left to claim.
Guess Grid doesn't filter with rules like "red means absent". For every possible answer it works out the exact pattern your guess would have shown, and keeps the words whose pattern matches what you saw. That makes double letters correct by construction rather than by special case.
8. How practice grades you
In practice, when the game ends, each of your guesses is set beside the move the solver would have made in the same position, scored the same way as above.
- Best available: you played the solver's pick, or a word within 0.02 bits of it. SLATE and RAISE are both best available; the numbers can't separate them.
- Better than N% of words: how your guess ranked against every word you could have played, rounded down, so nothing ever claims 100%.
- A wasted turn: one word was left and you played something else.
Where the solver's pick ranked higher, Why SLATE ranked higher shows both splits side by side, the same bars as above. The standard is the solver's actual move, not an abstract ideal: if you follow the app's own top suggestion, the review always calls it best available.
9. The word lists
The lists are built from SCOWL, an openly licensed collection of English word lists graded by how common each word is. Nothing is taken from any particular puzzle's list. There are three, at each length:
- Answers: common words. What the solver suggests, and what practice picks from.
- Guesses: a spell checker's vocabulary, with British and other spellings. The words ranked as guesses.
- Accepted: a much wider list, for what practice lets you type. Never suggested.
The answer lists leave out plurals, past tenses and archaic words, which puzzles rarely use as answers, along with anything offensive. The six-letter answer list was also read end to end by hand. If no common word fits your clues, perhaps because the answer is unusual, the solver falls back to the full guess list rather than giving up.