Jev plays chess
I wanted to see how well Jev plays chess.
So I had it play 470 games against chess engines and bots that imitate human players, and then gave it 1,000 puzzles from Lichess to solve.
The short version is that it’s pretty bad. It never won a game against an opponent that was trying to win. It solved 244 of the 1,000 puzzles, which works out to three out of four of the easiest ones and almost none of the hardest. If you turn those games into a rating, you get somewhere between 600 and 800 Elo, which is beginner territory.
But the parts that didn’t need any chess skill went great. Jev never played an illegal move, because the only way it can answer is by picking one of the moves I give it. Each move took about a tenth of a second, a whole game cost about a tenth of a cent, and the entire project, testing and all, cost $0.96.
So how do you lose almost every game when you can’t even make an illegal move? In Jev’s case, mostly by giving check every chance it gets.
Why I thought this would work
Most people who try to get a language model to play chess run into the same problem first. The model writes its move as text, and the text is often wrong. It names a move that isn’t legal, or writes something that isn’t a move at all, and now you have to ask again or pick a random move for it. Your results get messy fast.
Jev doesn’t have this problem. It only ever picks one of the options you give it.
A choice question takes up to 255 options, and a chess position never has more than 218 legal moves. This means that for every move in a game, I could hand Jev the whole list of legal moves and let it choose. There’s nothing to parse and nothing to retry. Whatever it picks is legal.
I also had a reason to be hopeful. Right before this, I had Jev play the Wikipedia Game, where you get from one Wikipedia article to another by clicking links. Jev picked the link on every page and reached the target in 95% of 1,015 games, for $1.74 in total.
Picking a move from a list of legal moves didn’t seem that different from picking a link from a list of links.
Spoiler: it’s different.
How I set it up
For every move, my code sends Jev one request.
The request describes the position in plain text: where each piece is, whose turn it is, and the moves played so far. Then it lists every legal move, each with a short plain-English description, and asks one question: “Which of these legal moves is the strongest move for the side to move?”
That’s it. One request, one pick.
Here’s what that looks like. This is move 8 of a game where Jev played Black against a bot that plays like an 1100-rated human. There were 38 legal moves to choose from. Jev put 37% on taking the pawn with its queen and 28% on taking it with a knight.
It took with the queen.
White’s knight captured the queen on the next move.
Now, I gave myself one rule here, and it matters for everything that follows. My code could describe the position and each move, but it couldn’t think ahead for Jev.
In other words, a move’s description says which piece goes where and what it captures. It never says whether the piece will be in danger, what the opponent can do next, or what a chess engine thinks of the move. The code also never hides or reorders moves. Every move is one request to Jev and nothing else.
Well, almost.
There is one small exception. When a move puts the opponent in check, the description says “check”, and when a move ends the game, it says “checkmate”. Working that out takes one move of lookahead. I left it in because it’s part of how chess moves are normally written down, and it seemed harmless.
It was not harmless. We’ll come back to that.
The first test
I started small. Two games against Stockfish, the strongest open-source chess engine, turned down to its weakest rating of 1320, and 100 puzzles from Lichess.
Jev lost both games. It solved 24 of the 100 puzzles, mostly easy ones, and none of the hard ones.
Hm.
That was worse than I expected, but it was also my first prompt. So before running the real test, I tried to improve it.
I tried 20 prompts and most did nothing
To avoid fooling myself, I tested every prompt on a separate set of 300 practice puzzles. That way I wouldn’t tune the prompt on the same puzzles I’d use for the final test.
Here’s what each change did.
| Change to the prompt | Practice puzzles solved |
|---|---|
| The starting prompt | 21% |
| Tell Jev it’s a chess grandmaster | 16% |
| Give step-by-step instructions on what to look for | 22% |
| Remove the list of moves played so far | 22% |
| Describe the position in a structured data format | 20% |
| List the moves in a random order | 21% |
| Ask the same question three times and average the answers | 21% |
| Replace the drawing of the board with a list of where each piece is | 24% |
| Add a plain-English description to every move | 24% |
| Both of the last two together | 26% |
Most changes did nothing.
My favorite is the grandmaster one. Telling Jev it was a chess grandmaster made it worse, mostly because it stopped taking checkmates it would otherwise have found. Apparently grandmasters are above checkmate.
What did help was describing the pieces and the moves in plain words, so that became the final prompt. On the final puzzle set, which neither prompt had seen, it solved 3.2 more puzzles out of every 100 than the starting prompt.
So, a real improvement. Not a big one.
The real test
With the prompt locked in, I ran the full test. Jev played 50 games against each of eight opponents, another 50 against Maia 1100 with the starting prompt, and 20 games against itself. Then it solved 1,000 Lichess puzzles it had never seen.
But first, a quick primer on Elo, since the rest of this depends on it. If you already know how Elo works, skip ahead to the results.
Elo is the usual way to measure a chess player’s strength. A higher Elo means a stronger player, and the gap between two ratings predicts how often each one wins. So to give Jev an Elo, you play it against opponents whose ratings you already know.
I used two groups of those. The first was Maia, a set of bots trained on millions of online games to play like humans of a given rating. The second was Stockfish at four weak settings.
Here's how Jev did against each one.
| Opponent | Opponent’s Elo | Jev’s points out of 50 |
|---|---|---|
| Maia 1100 | 1100 | 0.5 |
| Maia 1500 | 1500 | 1.5 |
| Maia 1900 | 1900 | 0.5 |
| Stockfish skill 0 | about 1310 | 1.5 |
| Stockfish, rating limit 1320 | 1320 | 2.5 |
| Stockfish skill 2 | about 1620 | 0 |
| Stockfish skill 4 | about 1850 | 0.5 |
A win is one point and a draw is half a point.
So, yeah. Jev never won. Its few points all came from draws, and every draw against Maia happened because Maia stalemated it or repeated moves.
From these games, I estimated Jev’s Elo two ways:
- Against Maia, Jev comes out at about 570, with a 95% confidence interval of 283 to 697. Maia’s ratings come from the human players it learned from, so this is the closest thing to a human rating. Maia’s authors say the bots play a bit stronger than their labels, so Jev’s real number is probably somewhat higher.
- Against Stockfish, Jev comes out at about 770, with a 95% confidence interval of 603 to 867. This uses Stockfish’s own rating scale, which isn’t the same as a human scale.
Both numbers are rough, because Jev lost almost every game, and you need some wins and draws to get a precise rating. Put them together and Jev plays somewhere around 600 to 800 Elo.
Puzzles tell a slightly different story. Lichess rates every puzzle by how hard it is, and from the puzzles Jev solved, its puzzle rating comes out around 1085. That’s higher than its game rating, which makes sense when you think about it. A puzzle tells you there’s a good move to find. A game doesn’t.
Jev against a player who moves at random
The clearest sign of where Jev is at came from somewhere I didn’t expect: games against a player making completely random moves.
Let’s think about what that means for a second. The random player isn’t trying to win. It doesn’t know what a queen is worth. Every turn, it picks any legal move it likes. A beginner who knows the rules should beat it almost every time.
Jev won 23 of 50 games. It drew 26 and lost one.
How do you draw 26 games against random moves? You get a winning position and then shuffle your pieces back and forth until the game is drawn by repetition. Jev did this over and over.
And the one game it lost? It gave away its queen for a pawn, and ten moves later the random player checkmated it.
Why it loses: it loves checks
Remember that exception to my rule? This is where it comes back.
When I went back through the games, one pattern stood out. Jev loves to give check.
When a move that gives check was among Jev’s top five choices, Jev played it 70% of the time. And those checks were mostly bad moves. On average, a check cost Jev 1.7 pawns’ worth of position, and one in five was a blunder that lost three pawns or more. Its other moves cost 0.7 pawns on average, and about one in fourteen was a blunder.
My guess is that the word “check” in a move’s description pulls Jev toward that move, whether or not the check is any good.
In puzzles, this helps. More than half of the moves in puzzle solutions are checks or checkmates. Jev solved every single mate-in-one puzzle, because the answer was the only move labeled “checkmate”. Take the mate-in-one puzzles out, and its puzzle rating drops from about 1085 to about 1000.
In real games, the same habit is how Jev gives away its pieces. That queen it lost to the random player? It went with a check.
The other pattern is that Jev makes its big mistakes early. In 346 of its 350 games against Stockfish and Maia, it made a move that lost at least three pawns’ worth of position by move 10. After that, it was usually playing from far behind, which isn’t a place you come back from against Stockfish.
Speed and cost
Okay, now for the good news.
Across 10,376 moves in the test games, Jev took 121 milliseconds per move at the median, and 99% of moves came back within 294 milliseconds. That’s about the same time I gave Stockfish to think about each move.
It also didn’t slow down when there were more moves to choose from. Positions with fewer than 10 legal moves took 117 milliseconds, and positions with 40 to 58 took 125 milliseconds. And that’s including the trip over the internet from my laptop.
A whole game used about 2.3 seconds of Jev time and cost about a tenth of a cent. The entire project, all the prompt testing and all 470 games included, cost $0.96.
That’s less than a pack of gum.
How this compares
None of this is new, by the way. Other people have tried Jev at chess too, and they found the same thing:
- Maxim Saplin put Jev on his LLM Chess leaderboard, where it rated about 243 on that leaderboard’s own scale, close to
o4-mini-medium, at $0.0015 a game. - Marcos Pazzarelli’s Jev Chess Lab tested Jev on 104 positions and found that Maia 1100 beat it on every measure.
- sliday’s jev-chess-algo gives Jev every legal move too, but adds code that checks for pieces left undefended.
For comparison, the language model people usually point to as good at chess, gpt-3.5-turbo-instruct, is estimated at around 1,750 Elo. Jev is a long way from that.
So what did I learn?
Two things.
First, picking from a list is not the same as understanding the list.
Jev did well on the Wikipedia Game because a good link usually looks like a good link. The link text tells you a lot. In chess, a great move and a terrible move can look almost the same written down, and the only way to tell them apart is to think about what happens next.
Jev doesn’t think ahead. It reads the list once and picks the move whose description sounds best. Which is exactly why the word “check” fools it.
Second, everything that didn’t depend on chess skill held up. Every move was legal, each one came back in about a tenth of a second, and the whole project cost $0.96.
So if you have a problem where the right option can be recognized from its description alone, Jev is fast and cheap enough to be worth trying.
Chess isn’t that kind of problem.
Details
- Model: Jev, reported by the API as
jev-1.13.0. - Engines: Stockfish 19 and Lc0 0.32.1 running Maia 1100, 1500 and 1900, all on my laptop. Stockfish had 100 milliseconds per move. Maia played its first instinct, with no search.
- Games: each started from one of 34 balanced openings, two to four moves deep, with colors alternating. Ratings were fitted with the standard Elo formula and a bootstrap for the confidence intervals. The Stockfish skill levels were placed on Stockfish’s rating scale using 500 extra games between the engines.
- Puzzles: from the free Lichess puzzle database, 50 for every 100 rating points from 600 to 2599. The practice and test sets don’t overlap.
- Move quality: measured by Stockfish at depth 14. A blunder is a move that loses three pawns or more.
The code, every game and every puzzle result are on GitHub at byrencheema/jev-chess, if you want to run it yourself.
For more on language models playing chess, read dynomight’s Something weird is happening with LLMs and chess and the follow-up.