0

Aevion Chess Mate

Fresh

Forced-mate chess positions with deterministic, dependency-free verification

Type
RL Env
Publisher
Aevion
License
apache-2.0
Size
v0.3.6
Published
Oct 2026
Updated
Oct 2026

Cite

Notes

Only stored in your browser.

AEVION Chess Mate — verifiable reasoning tasks

Forced-mate chess positions with deterministic, dependency-free verification: every accepted move is precomputed by exhaustive search, so checking an answer is a set lookup, not an engine call.

Why this environment is unusual

Most environments verify with a judge model or a fuzzy rubric. Here the outcome is binary by the nature of the subject: the move either delivers mate or it does not. There is nothing to tune and nothing to disagree about.

There is no chess engine and no chess library inside this package. All legal moves that reach the goal were enumerated in advance on AEVION's side using chess.js (BSD-2-Clause). That keeps this package free of copyleft dependencies — relevant because python-chess is GPL-3.0+, and importing it would make any environment that uses it a derivative work.

What is in the bank

Measured on the shipped puzzles.json, 2026-10-05:

positions2,491
themesmate in 1 — 1,702, mate in 2 — 688, mate in 3 — 101
difficulty (Glicko rating)400 … 2522, median 1007
positions with a single solution2,450 of 2,491 (98.4 %)
positions with alternative solutions41 — every one of them is accepted, which a one-string comparison would reject

Difficulty is a number, not a label, so you can build a curriculum (rating_min, rating_max) and see exactly where a model breaks.

The traceable set — 200 positions, added 2026-10-07

load_traceable() returns a second, smaller set that carries the two things the bank above cannot:

traceable setbank above
positions200 (mate-in-1: 100, mate-in-2: 100)1,237
difficulty615 … 1,982400 … 2,488
exactly one accepted move196 of 2001,213 of 1,237
source id on every positionyes — 200 of 200, li_<PuzzleId>yes — 2,491 of 2,491
accepted-move set proven completeyes, exhaustive searchyes — all three depths, 101 of 101 mate-in-3 sets re-derived twice by independent implementations

Why it exists. An incomplete accepted-move set scores a genuine alternative mate as 0 — the signal punishes a correct move, invisibly. Completeness can only be proven where the search finishes: one ply for mate-in-1, three plies for mate-in-2. Mate-in-3 is five plies, tens of millions of positions per task, which we do not complete — so the traceable set contains none, and proven completeness is never mixed with assumed completeness.

Verification of this set (2026-10-07): reference line legal and mating with the declared length — 200 of 200; complete mate-in-1 sets re-derived by a second, different predicate (mate marked in move notation rather than play-then-test) — 100 of 100, exact set equality; declared mate length confirmed by Stockfish 18 Lite at depth 14 — 200 of 200; four deliberate data corruptions — 4 of 4 rejected.

from aevion_chess_env import load_traceable
task = load_traceable()[0]
task.source_id          # "li_0Rrzp" → row of the open Lichess database

Grading has three outcomes, not two

grade() separates "the model was wrong" from "we could not parse what the model wrote":

from aevion_chess_env import grade
grade(task.solutions[0], task.solutions)  # verdict "correct",      reward 1.0
grade("h1h2",            task.solutions)  # verdict "wrong",        reward 0.0
grade("I think Ra8",     task.solutions)  # verdict "unverifiable", reward None

Collapsing those two makes every formatting failure count as a reasoning failure: comparing a model that answers a1a8 against one that answers "I think the answer is Ra8 because…" then measures format compliance, not chess. What to do with unverifiable — score zero, drop, or re-ask — is your decision, and the environment will not make it silently.

Input is UCI, with unambiguous decoration removed (a1a8#, 1. a1a8, 12...a1a8). SAN is not accepted: SAN→UCI without a board is guesswork. A leading move number is stripped only when a separator follows, so 1a1a8 stays unverifiable — a false correct inflates a score silently and is worse than a refusal. reward_mate() keeps its original binary behaviour for existing callers.

Install and run

From the Hub, which is what you most likely want:

prime env install aevion/aevion-chess-mate

From a checkout:

pip install -e .
python test_env.py                 # 46 self-checks, exit 0 on success
python test_letter_claims.py       # every number we publish, re-measured

Check us instead of trusting us. test_installed_package.py walks the path you just took — it imports the installed package, counts the bank, re-derives the composition, scores a correct move and a fabricated one, and refuses to run (exit 2) if it finds itself importing from a source tree instead of site-packages:

python -m venv /tmp/buyer
/tmp/buyer/bin/pip install aevion_chess_mate   --extra-index-url https://hub.primeintellect.ai/aevion/aevion-chess-mate/install/simple/
/tmp/buyer/bin/python test_installed_package.py    # 13 checks, exit 0

That refusal is not decoration. Every defect listed in this file as "fixed in 0.3.3" was invisible to checks that ran against the source tree and visible within seconds to checks that ran against the installed wheel.

from aevion_chess_env import ChessMateEnv

env = ChessMateEnv(seed=42, theme="Мат в 1", rating_max=1200)
obs  = env.reset()        # {"fen": ..., "goal": "mate1", "rating": 491, "side": "black", ...}
_, reward, done, info = env.step("g6g3")
# reward 1.0, info["accepted"] lists every move that would have been correct

Episodes are reproducible: the same seed yields the same order of positions, and sampling is without replacement within an epoch — otherwise an agent sees far fewer distinct positions than episodes, and part of the bank never participates in training.

Using it on the Environments Hub

The core package has no dependencies. The Hub wrapper lives in aevion_chess_env/taskset.py and is installed separately, so the core stays dependency-free:

pip install -e ".[hub]"

It exposes ChessMateTaskset (verifiers v1 API: vf.Taskset / vf.Task / @vf.reward). Config options: theme, rating_min, rating_max, limit.

Ignore the usage snippet the CLI prints after installing. prime env install ends with

from verifiers import load_environment      # v0 API — not how this package loads
env = load_environment('aevion-chess-mate')

which the CLI prints unconditionally for every environment regardless of the runtime it declares (checked in the CLI source: the two lines have no branch on runtime). This package targets the v1 API, so load the taskset class directly:

from aevion_chess_env.taskset import ChessMateTaskset, ChessMateConfig
taskset = ChessMateTaskset(ChessMateConfig(theme="Мат в 2", limit=256))

Two measured notes on the extra, both of which were our own defects before 0.3.3:

  • the extra pinned verifiers>=1.0, which cannot be satisfied — that package has never reached 1.0 (latest on PyPI is 0.3.1). pip install "aevion-chess-mate[hub]" failed with "No matching distribution found". The v1 in verifiers.v1 is a submodule name, not a version number. Now verifiers>=0.3, tested against 0.3.1.
  • up to and including 0.3.2 the Hub listed this package as Unclassified, because it infers the runtime from a required dependency and ours is deliberately optional. Declared explicitly from 0.3.3 on; the dependency is still optional, so the core stays at zero.

Unlike most environments, verification needs no isolated runtime: the reward is a pure function of the trace, because every accepted move is already precomputed. validate() checks each shipped position against its own solution set, with no model involved.

Answer format

Moves are expected in UCI (g6g3, e7e8q). Decorations are stripped (G6G3# works). SAN (Qg3#) is not accepted: converting SAN to UCI requires a board, and this package deliberately has no board. Say so in your system prompt.

How this differs from the other chess environments on the Hub

There are already several, and one is close to this one — so here is the honest comparison rather than a claim of novelty.

albertklorer/chess-rlvr is single-turn, takes a FEN plus the list of legal SAN moves, and rewards the chosen move with a precomputed Stockfish regret (0.0 for the best move, negative for worse ones). It covers arbitrary positions, not just mates.

This environment is narrower on purpose, and differs in three ways that matter for training reasoning rather than move selection:

  1. The reward is a proven fact, not an evaluation. A forced mate was established by exhaustive search; it does not depend on an engine or its depth. A regret score does — change the depth and the numbers change, and so does what the model learns.
  2. No list of legal moves is given. The model has to find the move itself, which is a harder task than picking from a supplied list.
  3. All correct moves are accepted. 41 of the 2,491 positions have more than one solution, and every one of them scores 1.0; a single-answer comparison would punish a correct alternative.

What the other environment does better, plainly: it covers the whole board of positions rather than mates only, and it accepts SAN, which is closer to how a model naturally writes moves.

On "proven complete" for mate-in-3

Two different claims were made about this inside the team, and both were wrong in opposite directions, so here is the measured answer.

The claim "completeness cannot be proven for mate-in-3 — five plies, tens of millions of positions" was an upper bound, not a measurement: in mate positions the tree collapses as soon as the defender has a reply with no mate, so the worst case never materialises. The opposite claim, "proven by exhaustive search", was equally unsafe as a blanket statement, because the search was not run on every position.

Measured on real positions from this bank, recomputing each accepted-move set from scratch:

mate-in-3 positions re-verified from scratch101 of 101 — all of them
differences found0
cost45 min total (~27 s per position)
a deliberately easy control (starting position, five plies)1.2 s — which is why an easy control misleads here

Every mate-in-3 accepted-move set has now been recomputed from scratch, and every one matched the shipped set. The check is two-sided, which is the part that matters: not only does the listed move force mate, but no other legal move does — that second half is what completeness means, and it is the half that protects a model from being scored 0 for a correct alternative.

A second, independent implementation — different code, same data, written by another team member — has now finished on every position:

verified by two independent implementations101 of 101, zero differences
accepted-move sets of size 1101 of 101 — each position has exactly one forcing move, and that is a search result, not an assumption
cost of the second passmedian 14.9 s per position, max 157 s, 47.7 min for 99 positions

The two implementations were kept apart on purpose. "My search agrees with itself" and "two separately written searches agree" are different claims, and only the second rules out a shared bug in a single implementation. Both halves of the two-sided check were repeated by the second pass, and its controls were run before the work: a mate-in-1 is found by a depth-3 search, a position with no mate yields an empty set, and the starting position has no mate in 3.

One measurement detail, because it would otherwise travel as a false number: two of the 101 timings (54 min and 781 min) were taken by wall clock across an interrupted run and a machine sleep, and one of them accounted for 95% of the naive total. They are artefacts of measurement, not hard positions, so the cost above is stated over the 99 clean timings. Both implementations land in the same range — the first pass took 45 min — and an earlier claim that one was ~6× slower was itself an artefact of measuring under a concurrent heavy run.

Why this matters beyond pedantry: an incomplete accepted-move set scores a genuine alternative mate as 0. The signal then punishes a correct move, and it does so invisibly.

Provenance — every position traceable to its source

⚠️ Before 0.3.3 this was true of the data and false of the API. The two data files carry the identifier under different keys — source_id in the main bank, id in the traceable set — and the loader read only id, so load_puzzles()[0].source_id came back empty for all 2,491 positions while the file held an identifier for every one of them. The one test that checked source_id ran against the traceable set, where the key happened to match, so it stayed green. It was caught by installing the published wheel into a clean environment and walking through what a reader of this file would actually type. Fixed in 0.3.3, and the bank now has four checks of its own that print the denominator, plus a control that goes red on an empty identifier.

All 2,491 positions carry a source_id of the form li_<PuzzleId>, matched against the source bank by full FEN — halfmove counters included, so each match is an exact record, not merely "the same position". Matching ran against the complete mate slice of the bank: 158,045 records.

The match is reproducible, and that was verified rather than assumed. Two independent passes over the bank, run back to back, produced identical mappings for all 2,491 positions: same positions, same ids, zero differences. The comparison itself was control-tested — a deliberately altered id is detected, so "identical" does not mean "unable to tell apart".

This section previously said the opposite, and the history is worth stating because it explains what to trust. The first pass reported 97 positions with no origin. Those 97 were an artefact: the paginated API ordered rows by a non-unique field, so pages silently skipped records, and a skipped record looked like "not in the bank". After the ordering was fixed at the source, two consecutive passes agreed exactly and every position matched. Nothing about the positions changed — only our ability to count them did.

Multi-step episodes — the agent plays the mate out

A one-move answer tests one decision. Training reasoning needs a horizon: does the model hold a plan when the opponent replies in a way it did not expect? So the forcing tree is precomputed, and the agent plays the line to the end — up to three of its own moves, with the defender answering in between.

from aevion_chess_env.lines import ChessMateLineEnv

env = ChessMateLineEnv(seed=7)
obs = env.reset()                     # {"fen": ..., "moves_left": 2, "plies_left": 3, ...}
obs, reward, done, info = env.step("g4h3")   # reward 1.0, done False
                                             # info["defender_reply"] = "h2g1"
obs, reward, done, info = env.step("h3g2")   # reward 1.0, done True, info["mate"] = True

Every node is re-checked by code that does not share the search which built it, and that took two predicates rather than one, because a single one would have left a gap we found only by counting nodes:

  • at a leaf the mate is one move away, so "which moves mate" is answered by applying each legal move and asking isCheckmate — no search at all;
  • at an intermediate node (mate in two from there) the answer is composed of those same isCheckmate calls — a move qualifies when, after it, every defender reply leaves a mate in one — with no shared recursion;
  • a root set is compared against the single-step data shipped in this package, which was itself re-derived by two independently written implementations.

The gap is worth naming because it was ours. After the first pass only the leaves had been checked independently, and that sounded like the whole tree. Counting nodes showed the mate-in-3 trees hold 481 nodes — 101 roots, 218 intermediate and 162 leaves — so 218 of them still rested on the very search that produced them, which is not evidence. They are now checked separately: 218 of 218 matched, with zero misses in either direction.

Both directions are counted apart, because they fail differently: a move that forces mate but is not accepted punishes a correct answer, and an accepted move that does not force mate credits a wrong one.

Verifiability is unchanged, and that is the whole point. Every node of the tree carries the complete set of moves that force mate in the remaining number of moves, enumerated in advance by exhaustive search. Grading stays a set lookup: no engine, no judge model, no dependencies.

Three design decisions worth stating, because each could reasonably have gone the other way:

  • The defender's reply is chosen by seed, not by strength. After a forcing move every reply loses, so "best reply" is not a meaningful notion here, while reproducibility is: the same seed yields the same episode, or two runs of one model are not comparable.
  • A wrong move ends the episode with 0.0. A wrong move can destroy the forced mate, and from there the task is no longer verifiable — continuing would mean scoring something we never proved. info["reason"] says so explicitly rather than leaving it to be inferred.
  • Reward is 1.0 per correct move, not only at mate. Otherwise a long line gives the same signal as a short one, and the training signal cannot tell at which step the model lost the plan.

Measured, on the data actually shipped:

positions with a precomputed tree789 — every mate-in-2 (688) and mate-in-3 (101) position
nodes in those trees1,990: 789 roots, 218 intermediate, 983 leaves
root accepted-move sets matching the single-step data789 of 789, zero differences
nodes with a missing or empty accepted set0 across the whole tree
leaf nodes re-checked by a search-free predicate983 of 983 matched exactly
intermediate nodes re-checked by an independent depth-2 predicate218 of 218 matched exactly
moves that force mate but were not accepted (would punish a correct answer)0
moves accepted that do not force mate (would credit a wrong answer)0
checkspython test_lines.py — 16, including a denominator per depth and a control that the tree walker detects a broken node

Honest limits

  • Only forced-mate tasks. Positions whose goal is "find the best move" need an engine evaluation, which would reintroduce a GPL dependency; they are excluded.
  • Chains are short. The underlying bank holds ~500k positions but only about 0.4 % have a 3+ move forced line. This package ships a sample, not the whole bank.
  • Multi-step episodes cover mate-in-2 and mate-in-3 (789 positions). Mate-in-1 is single-step by nature, so 1,702 of the 2,491 positions stay one move long.
  • No multi-turn Hub wrapper. The multi-step environment is a plain Python API, verified by its own 13 checks. A verifiers rollout wrapper for it is not shipped, because verifiers.v1 needs fcntl and cannot run on the machine this was built on — writing a wrapper we cannot execute would mean shipping untested code and calling it a feature.
  • No containerisation. Plain Python package.

Data provenance and licence

Positions come from the open Lichess puzzle database, released under CC0:

"Database exports are released under the Creative Commons CC0 license. Use them for research, commercial purpose, publication, anything you like." — https://database.lichess.org/ (checked 2026-10-05)

Precomputed solution sets are derived from those positions by exhaustive search.

Package code: Apache-2.0 (see LICENSE).

Verification of the bank itself

Every shipped position was replayed against its own solution set: 2,491 of 2,491 reference solutions are accepted, and all 2,491 pass the environment's own validate(). For the 41 positions with alternatives, every alternative is accepted too.

Self-checks include negative controls (wrong move, empty answer, SAN input, empty pool) and were validated by mutation. Mutating each branch separately — not just one — is what found the gaps: an earlier validate() was a tautology (it asked a set whether its own members belonged to it) and survived until the mutation return True was tried; a missing length check survived because both malformed-move samples were caught by the square-range check instead. Both are fixed and the mutations now fail the suite.