Originally posted on LinkedIn.
The NYT WordleBot's explanations of my Wordle games are often completely foreign to me. The bot knows every possible answer and treats each turn as a search problem, the kind usually modeled with Shannon entropy. It picks the guess that minimizes the expected number of turns left, which often means playing a word that cannot be the answer because it splits the remaining pool more cleanly. Even its creators note that the best guess for a computer is not necessarily the best guess for you.
I play by pushing letters around until something jogs my memory. The skeleton is the same, prune first and then commit, but my version is shaped around the tools I have. I do better once a couple of consonants are pinned, because consonants are what trigger my recall. The bot and I are not disagreeing. We are solving the same problem with different technology.
My only real shot at beating the bot on any given day is a high-variance move: commit to a real guess early and hope it lands. That gets the occasional win and is even worse in the long run, though I do not track my scores, so I will never find out how much worse. Over any real number of games the bot wins, because Wordle is shaped exactly like the problems deterministic computing is built for: a closed search space with a clear way to score every guess. When a problem fits the machine that well, it is no contest.
Same lesson, new machine
An LLM has different technology at its disposal. It reads about three orders of magnitude faster than I do and writes about two orders of magnitude faster. It writes code and operates a computer fluently, so anything that can be done through a terminal it can do at machine speed. It does not get bored, so it will grind through repetitive work I would quietly cut corners on. It is also sycophantic and bad at checking its own work.
So my job is to reshape problems around that profile: point the model at work that plays to its strengths, and build a harness that checks its output where it is weak.
Do that and the execution capacity is effectively unlimited. Two shapes keep coming up. If a problem can be turned into code, a model is great at it. If it can be turned into search, with a clear way to test candidates, a model can grind its way to an answer. In September, OpenAI announced a resolution of the Navier-Stokes Millennium Prize problem produced by roughly 10,000 agents running for 88 hours. Benn Stancil calls this brute intelligence: not monkeys at typewriters, but a million researchers taking swings at the same problem. No human mathematician would approach it that way, the same way no human plays Wordle like the bot.
