```
Full recipes (scoring, generation, batch, convert, benchmark) are in the kit's
[README](https://github.com/raphael-anjou/eternity2/tree/main/research/starter-kit#readme).
Boards to test a solver against live in the CC0
[dataset](/research/build/dataset); the approaches to try are surveyed in the
[map of every known approach](/research/build/approaches-map).
## Generate boards right here
The kit's generator, in the browser: the same WASM engine, no install. Pick a
shape and a seed range and it makes a reproducible batch; click any board to open
it in the [viewer](/viewer) and score it edge by edge.
> **[Interactive: BoardGenerator]** Rendered on the canonical page (link above); not shown in this markdown export.
Want the same thing on the command line, or 10,000 boards with clues pinned? Use
the kit's `generate_batch` example: it writes each board as JSON you can feed
straight back to a solver.
## Related
- [Run it yourself](https://eternity2.dev/research/build/run-it-yourself) — The whole site, the engine, and every result in this section run from one repository. Here is how to get it going, rebuild the WebAssembly engine, and reproduce the numbers.
- [The dataset](https://eternity2.dev/research/build/dataset) — A public, CC0 dataset for Eternity II in two parts: fourteen benchmark instances to solve, and a corpus of 7,658 distinct strong boards to learn from. Every score is recomputed from the board itself, and the corpus is checked to be genuinely diverse rather than a thousand copies of one board.
- [A map of every known approach](https://eternity2.dev/research/build/approaches-map) — The survey the community keeps asking for and never finds: every family of attack tried on Eternity II, what each one actually reached, where it walls out, and a link to the deep page. One organising insight runs through all of them.
- [Board & puzzle formats, written down](https://eternity2.dev/research/build/formats) — Every format an Eternity II board or puzzle travels in on this site and in the community: the board_edges letter string and the hints clue list, e2pieces.txt, the puzzle CSV, the site's Puzzle JSON, and the viewer URL, with the exact byte-level rules (how the grey border is encoded in each) and, most of all, what each format can and cannot recover.
---
# The variants and the claims that live there
> Every "Eternity II" that is not the real puzzle: the TopCoder Marathon variant and the unresolved Takahashi 468, McGavin's unframed 480/480, the mixed-set boards, the no-starter challenge, and the claim quarantine, from the 2007 phantom-sets warning to the community's zero-knowledge verification ethic. Ends with a checklist for stating a score properly.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/build/variants/
- Updated: 2026-07-02
- Source: The phantom-sets warning (Sergio Demian Lerner), eternity2@groups.io, 2007-07 — https://groups.io/g/eternity2/message/1210
- Source: The unsaved E2Lab "471" popup, eternity2@groups.io, 2009-12 — https://groups.io/g/eternity2/message/7284
- Source: A hand-solve worth "at least 472", never evidenced, eternity2@groups.io, 2010-01 — https://groups.io/g/eternity2/message/7383
- Source: Brendan Owen asks the 476/480 claimant to publish the grid, eternity2@groups.io, 2011-02 — https://groups.io/g/eternity2/message/8676
- Source: The $1000 no-starter challenge, quantified by Peter McGavin, eternity2@groups.io, 2011-06 — https://groups.io/g/eternity2/message/8924
- Source: The Vautrin "resolved" claim falls to duplicate-piece checks, eternity2@groups.io, 2014-10 — https://groups.io/g/eternity2/message/9300
- Source: The Takahashi 468 claim surfaces (Dima), eternity2@groups.io, 2017-09 — https://groups.io/g/eternity2/message/9694
- Source: The TopCoder rules check: randomly generated pieces (Peter McGavin), eternity2@groups.io, 2017-10 — https://groups.io/g/eternity2/message/9697
- Source: Unframed E2 solved: thousands of 480/480 boards (Peter McGavin), eternity2@groups.io, 2017-11 — https://groups.io/g/eternity2/message/9750
- Source: Why the unframed 480 is disqualified: grey edges not on the rim, eternity2@groups.io, 2017-11 — https://groups.io/g/eternity2/message/9766
- Source: The mixed-set "480": one E2 + Clue-1 + Clue-2 set (Peter McGavin), eternity2@groups.io, 2023-10 — https://groups.io/g/eternity2/message/11169
---
Over twenty years, the archive has accumulated a small zoo of puzzles that
are *almost* Eternity II: same pieces with a rule dropped, same rules with
different pieces, same board with donor pieces mixed in. Each variant is a
legitimate object of study (some produced genuinely beautiful results),
but each also became, sooner or later, the home address of a headline
score that is not a record on the real puzzle. This page is the field
guide: what each variant is, what was actually achieved on it, and how the
community learned to keep variant scores out of the
[records table](/research/records).
## Why variants matter: scores only compare within a game
There is one lesson, and everything else on this page is a case study of
it: **a score is only comparable to another score if both were achieved on
the same piece set under the same rules regime.** The real puzzle is the
board Tomy sold: 256 specific pieces, a 16×16 frame with the grey edges
on the rim, and the starter piece pinned at I8 (the four other clue
placements were optional aids under the contest rules,
[msg 11046](https://groups.io/g/eternity2/message/11046)). The record line
467 → 468 → 469 → 470 all lives in that one regime, verified board by
board ([msg 10554](https://groups.io/g/eternity2/message/10554)); the
strict five-clue line (best known: 464) is tracked separately because it
is a *harder* regime, not a different puzzle
([Records & solvers](/research/records)).
Change the piece set and "480" can become easy. Drop the frame rule and
480/480 has actually been *achieved*. Drop the starter constraint and the
expected number of solutions multiplies by 784. None of those numbers says
anything about the official 480, which remains unfound. Hence the
catalogue below, and the quarantine that follows it.
## The catalogue
### The TopCoder Marathon variant and the Takahashi 468 claim
In 2009, TopCoder ran a Marathon Match on an Eternity-II-*style* problem.
When Peter McGavin checked the contest's problem statement years later, he
found that it used **randomly generated pieces**, apparently with
different colour counts than E2 and possibly no distinction between border
and interior edge types
([msg 9697](https://groups.io/g/eternity2/message/9697)). Each test was
its own puzzle. That makes any score from the contest a statement about
random instances of the puzzle *class*, not about Monckton's fixed 256
pieces: a different game, however similar the rules sheet looks.
The claim attached to it surfaced in September 2017, in the afterglow of
McGavin's [10×10 benchmark](/research/build/benchmarks) solve: Dima
reported that "467 has been broken in 2009". Naohiro Takahashi (chokudai),
a top competitive programmer, had tweeted a 468. Dima stated plainly that
he had never seen the solution and it was unverified
([msg 9694](https://groups.io/g/eternity2/message/9694),
[9696](https://groups.io/g/eternity2/message/9696)). The scrutiny followed
the standard protocol. McGavin, surprised that a "simple" beam-search-style
method would beat everything in
[Verhaard's arsenal](/research/lab/experiments/louis-verhaard/eii)
([msg 9695](https://groups.io/g/eternity2/message/9695)), did the rules
check above and asked for confirmation that the 468 used the original
pieces ([msg 9697](https://groups.io/g/eternity2/message/9697)). David
Barr argued the other side: the tweet says "Eternity 2", and Takahashi's
own profile distinguishes the TopCoder contest from the real puzzle
([msg 9698](https://groups.io/g/eternity2/message/9698)). Dima undertook
to contact him and ask for the solution: "This is the only way to know
for sure" ([msg 9700](https://groups.io/g/eternity2/message/9700)).
No resolution ever appeared in the archive, so the claim's status remains
unresolved. Told in full: *if* the 468 was on the TopCoder variant, it is
rule-incomparable with Verhaard's 467; *if* it was on the real pieces, no
board was ever produced. Either way it stays out of the records table as
a claim, not a record.
### Unframed E2: McGavin's 480/480 that isn't a solution
The most beautiful non-result in the archive. In late 2017 Peter McGavin
set himself a variant: place **all 256 real pieces** on the real 16×16
board, but treat the grey edges as ordinary colours (so the border pieces
migrate into an interior band) and maximize internal joins. On 2 November
he posted 479/480
([msg 9736](https://groups.io/g/eternity2/message/9736)); ten days later
one of his backtrackers, working from around its 4,400th 14×15
sub-solution, found not one but **thousands of full 480/480
configurations**
([msg 9750](https://groups.io/g/eternity2/message/9750), picture at
[9755](https://groups.io/g/eternity2/message/9755)). Henk van der Griendt
laid the pieces out physically and confirmed they fit
([msg 9754](https://groups.io/g/eternity2/message/9754)).
The inevitable question came: so you have solved Eternity 2? McGavin
was precise. The board scores 480/480 internal joins with original pieces
on the original board, but "the competition rules say the grey edges must
be around the border, which clearly they are not. From that point of view,
the entry is disqualified"
([msg 9757](https://groups.io/g/eternity2/message/9757),
[9766](https://groups.io/g/eternity2/message/9766)). His own summary is
the model of how to report a variant result: "I set my own E2-like
challenge with my own rules and then solved it"
([msg 9767](https://groups.io/g/eternity2/message/9767)). He estimated the
variant's difficulty as on par with building an unframed 14×14 from the
196 middle pieces
([msg 9719](https://groups.io/g/eternity2/message/9719)). Hard, but
demonstrably not E2-hard: the frame constraint the variant drops is
precisely where the real puzzle's difficulty concentrates.
### Mixed piece sets: the donor-piece "480"s
The third family keeps the frame rule and changes the pieces. In 2014
McGavin showed complete 16×16 boards built from *two* E2 sets, exploiting
duplicate pieces ([msg 9305](https://groups.io/g/eternity2/message/9305)),
and in October 2023 he built a full 480/480 board from the pieces of one
E2 set, one Clue-1 set and one Clue-2 set
([msg 11169](https://groups.io/g/eternity2/message/11169)), in Vlastislav
W.'s max-distinct-tiles challenge, a deliberate game whose metric is how
many *distinct* official tiles such a board can carry (record: 238,
[msg 11167](https://groups.io/g/eternity2/message/11167),
[11170](https://groups.io/g/eternity2/message/11170)). These
constructions hide nothing: clearly labelled by their makers, they are
the reason the records table carries a "variant" badge on its 2023 row.
The full story of the donor sets (what the clue puzzles were, and how
their pieces ended up in these boards) is on
[the clue puzzles page](/research/build/clue-puzzles).
### The no-starter variant and the $1000 challenge
In June 2011, list member ebeternity put up $1000 for a full E2 solution
**without** the piece-139 starter constraint, valid to the end of
September ([msg 8922](https://groups.io/g/eternity2/message/8922)). It was
the first post-contest cash on the table. McGavin immediately quantified
what the relaxation buys: an estimated **1.1527 × 10⁷ solutions without
the starter constraint versus 14,702 with it**, 784 times more needles,
yet the expected backtracker cost per solution barely moves (9.2766 × 10⁴²
nodes versus 9.2751 × 10⁴², best-known search order)
([msg 8924](https://groups.io/g/eternity2/message/8924)). Fewer
constraints, proportionally bigger haystack: the variant is 784× "easier"
in solution count and essentially unchanged in work.
Nobody claimed the money. The only "solve" that reached the thread was a
second-hand 480 claim promptly labelled bogus, occasioning Johannes
Lindé's summary of a decade of such claims: claimants become "quite
indignant (or silent) upon requests for proof"
([msg 8951](https://groups.io/g/eternity2/message/8951),
[8952](https://groups.io/g/eternity2/message/8952)).
### A different objective: most pieces placed, zero conflicts
Not every variant changes the pieces or the frame. Some change the *scoring
rule*, and the cleanest example is the max-conflict-free-placement objective:
instead of counting matched edges on a full board, count the most pieces you
can place so that **every** placed piece matches all its neighbours, leaving
holes rather than mismatches. The score is "pieces placed", not "edges matched",
and the two are not comparable: a board with a handful of holes says nothing
about how a full board built from the same search would score.
The frontier here is old and small, and it still belongs to Louis Verhaard. His
"Only seven holes" board from around 2008 places **249 of 256** pieces
conflict-free, with the seven holes clustered in the top band (rows 1 to 4). In
June 2026 Laurent Zamofing reached **248**, one short of it, by recombining the
community's published 467-edge boards as crossover donors
([msg 11890](https://groups.io/g/eternity2/message/11890),
[11901](https://groups.io/g/eternity2/message/11901)). What is striking is that
both boards hit a *proven local dead-end* at their score: Verhaard's seven holes
cannot be filled by the seven remaining pieces, and freeing a radius-6
neighbourhood around them still never reaches 250. That the residual holes always land in the same top band, under an
objective that has nothing to do with edge count, is its own small finding, and
it is [the mismatch geometry again](/research/why/mismatch-geometry): the damage
collects against whichever edge the search finishes on, whatever it is
optimising.
## The claim quarantine
Variants are one way a false "record" is born; the other is a claim with
no board attached. The community's response to both matured into a
protocol, and its history is worth telling because the quarantined claims
still circulate.
**The earliest form: the phantom-sets warning (2007).** Within days of the
launch, Sergio Demian Lerner warned the list about "phantom sets":
accepting a stranger's "random" E2-like benchmark could mean unknowingly
solving an isomorphic re-encoding of the real puzzle, and handing the
stranger a $2M solution
([msg 1210](https://groups.io/g/eternity2/message/1210)). The community
adopted the practice of publishing benchmarks with provenance and known
solutions ([msg 1213](https://groups.io/g/eternity2/message/1213)). The
insight generalizes: *a piece set you cannot verify is a claim you cannot
score*. Set-scrutiny was part of the culture before any famous fake
arrived.
**The 471 that was never saved (2009).** One morning Dominique Schaltz
reported that E2Lab had popped up a message box overnight saying 471, but
the score did not appear on his grid and the configuration was apparently
never saved ([msg 7284](https://groups.io/g/eternity2/message/7284)).
Within weeks it echoed through the list as "somebody touched 471"
([msg 7317](https://groups.io/g/eternity2/message/7317),
[7353](https://groups.io/g/eternity2/message/7353)). No board ever
existed. The episode is the cleanest illustration of why the rule is
*board or it didn't happen*: a faithfully reported software popup, once
detached from its caveats, circulated as a record.
**The ≥472 hand-solve (2010).** Faruk Barber reported completing the whole
board by hand with none of the clue pieces on their positions; Henk van
der Griendt computed the claim would be worth at least 472 after swapping
the starter into compliance
([msg 7383](https://groups.io/g/eternity2/message/7383),
[7385](https://groups.io/g/eternity2/message/7385)). Al Hopfer called it
"just silly": a genuine 472 would dwarf the known 467
([msg 7427](https://groups.io/g/eternity2/message/7427)). No evidence was
ever posted; the claimant was simultaneously offering his clue puzzles for
sale ([msg 7395](https://groups.io/g/eternity2/message/7395)).
**The 476-then-480 bloom (2011).** In January 2011 Alain Bidon claimed a
476/480 achieved without the starter piece in place; the original message
is missing from the archive export and survives only through replies
([msg 8247](https://groups.io/g/eternity2/message/8247)). The list
immediately re-scored it under the rules (at best 472 without the starter,
468 guaranteed after compliance). By February the claim had grown to
multiple hintless **480s**
([msg 8636](https://groups.io/g/eternity2/message/8636)) and hit what
one member called "the good old zero knowledge test". Brendan Owen's two
demands in that thread are still the model of the genre: "there is a bug
in your program. Actually place the real pieces"
([msg 8670](https://groups.io/g/eternity2/message/8670)), and "could you
please publish the grid of piece ids"
([msg 8676](https://groups.io/g/eternity2/message/8676)); Bruce made it
concrete: just turn each piece over and post the 16×16 list of the
numbers on the backs
([msg 8681](https://groups.io/g/eternity2/message/8681)). No grid ever
appeared. In the same season an anonymous "new algorithm" announcement was
met with the other standard challenge: prove it by solving
[Brendan's hint-free 10×10 benchmark](/research/build/benchmarks)
([msg 8262](https://groups.io/g/eternity2/message/8262),
[8264](https://groups.io/g/eternity2/message/8264)).
**The Vautrin affair (2014): the protocol working end to end.** Nicolas
Vautrin first claimed his program could prove E2 has *no* solutions,
retracted (a bug), then days later posted "eternity2 resolved ;)" while
withholding the solution because of a prize
([msg 9280](https://groups.io/g/eternity2/message/9280),
[9286](https://groups.io/g/eternity2/message/9286),
[9291](https://groups.io/g/eternity2/message/9291)). This time a checkable
artifact existed, a claimed solution to one of Brendan's benchmark
puzzles, and verification killed it in days: Conny Öström's editor found
**115 duplicate pieces** in it
([msg 9300](https://groups.io/g/eternity2/message/9300)), McGavin
exhibited one piece used five times
([msg 9301](https://groups.io/g/eternity2/message/9301)), Arnaud Carré
identified the exact bookkeeping bug
([msg 9308](https://groups.io/g/eternity2/message/9308)), and Vautrin
conceded ([msg 9303](https://groups.io/g/eternity2/message/9303)). Nobody
had to trust anybody: the grid was posted, so the grid could be checked.
The pattern across all five stories is the community's real institutional
knowledge. Claims die or survive on artifacts (a grid of piece ids, a
solved public benchmark, a checksum), never on reputation, enthusiasm or
plausibility. The 471, the ≥472 and the 476/480s must never enter the
records table; the Takahashi 468 waits outside it for a board; and the
Vautrin episode shows the same machinery clearing a genuine mistake in
four days.
## How to state a score properly
Distilled from the twenty years above, the checklist the archive
effectively enforces. A score means nothing without all four:
1. **Name the piece set.** Official 256-piece E2? A Brendan benchmark set
(which one)? Random pieces? Mixed donor sets? If the set is private or
unverifiable, it scores nothing; the
[phantom-sets warning](https://groups.io/g/eternity2/message/1210) is
about exactly this.
2. **Name the rules regime.** Framed with grey edges on the rim?
Starter-only (the 467→470 record line), strict five-clue (the 464
line), no-starter (the $1000 challenge), or unframed? Scores do not
transfer between regimes: 784× more solutions is still not the same
game.
3. **Post the board.** The 16×16 grid of piece ids with rotations:
Brendan's demand, Bruce's "turn each piece over". A board can be
re-scored by anyone; a number cannot. Every board on
[the notable boards page](/research/community/boards) can be loaded and
re-scored edge by edge in the [viewer](/viewer).
4. **Say who and when, with a link.** The records table only carries
entries with a primary source; claims that live only in a popup, a
tweet or a memory stay quarantined
([Records & solvers](/research/records)).
Set + regime + board, or it didn't happen.
## Related
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [The four clue puzzles](https://eternity2.dev/research/build/clue-puzzles) — Tomy sold four small companion puzzles for Eternity II: solve one, submit the solution, and the official site revealed one piece's placement on the main board. What each puzzle was, the broken online checker, the eBay grey market, why puzzles 5 and 6 never came, and the clue puzzles' second life as complex-theory test cases and donor piece sets.
- [The notable boards](https://eternity2.dev/research/community/boards) — The famous Eternity II boards, treated as first-class citizens: Verhaard's prize-winning 467, the Blackwood record line to 470, the strict five-clue record to 464, and this project's own boards: who found each one, when, and what makes it structurally interesting. Every bundled board opens in the viewer.
---
# History & community
> The record and the people behind it: two decades of Eternity II told as a story, the record boards themselves, the record-holders and theorists, the academic literature, and how to add your own work to the wiki.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/community/
- Updated: 2026-07-13
---
Eternity II is a story before it is an algorithm. This section keeps the
community's memory: how the record climbed, the boards that carry the line,
the people who found them, the papers written along the way, and how to add
what you know.
- [History: the big steps](/research/history) — The whole story at a glance: a scannable timeline of the turning points, from the mailing list founded in 2000 to the 470 record that still stands.
- [The hunt, in full](/research/community/hunt) — The long-form account in two parts: the founding era (2000–2009) and the modern record line (2009–2026), with the posts to prove each turn.
- [Records & solvers](/research/records) — The record timeline and the engines behind it: who reached 470/480, and what those boards reveal about the wall.
- [The notable boards](/research/community/boards) — The record line told through the boards themselves: Verhaard's 467, the line to 470, the strict five-clue record, each with its non-matching edges marked and openable in the viewer.
- [Who's who](/research/people) — The people behind two decades of research (founders, record-holders, theorists), with what each contributed and the posts to prove it.
- [Papers](/research/papers) — The academic literature on Eternity II and edge-matching, ranked by how useful it actually is to a solver-builder.
- [Contribute](/research/contribute) — How to add to the wiki: the sources it trusts, how findings are credited, and where to send what you know.
## The community archive
Much of what this wiki draws on lives, in raw form, in the
[groups.io Files area](https://groups.io/g/eternity2/files): about 300 MB
built up over two decades. It is worth knowing what is in there.
- [The paper archive](https://groups.io/g/eternity2/files/00_Eternity2_articles_papers_presentations_books) — Fifteen dated PDFs, 2007 to 2020: the complexity proofs, the SAT/CSP and MILP papers, theses, and a maintained bibliography. The backbone of our [Papers](/research/papers) page.
- [Puzzle theory & proofs](https://groups.io/g/eternity2/files/Puzzle%20Theory) — Sixteen more: phase-transition bounds, the shared-edge feasibility proof, French seminar notes, and cross-domain work from the SAT and physics literature.
- [Decades of solvers](https://groups.io/g/eternity2/files) — Dozens of contributors' folders (Brendan, Guenter, Peter McGavin, and many more): backtrackers, SAT and exact-cover encoders, GA and GPU attempts, with source, most from before this wiki existed.
- [Databases](https://groups.io/g/eternity2/databases) — Brendan Owen's ["Backtracker estimates"](/research/why/complex-theory) tabulation (the tree sizes and solution counts you'll find worked into the theory pages) and the record leaderboards, all queryable.
Sample puzzles, benchmark sets, and board images are in there too. If you
find something worth a page, [say so](/research/contribute) and it can become
one.
## Pages in this section
- [The hunt, a history (part I: 2000–2009)](https://eternity2.dev/research/community/hunt) — The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
- [The hunt, a history, part II: 2009–2026](https://eternity2.dev/research/community/hunt-part-2) — Seventeen years after the prize: the contest dies with its solution locked in a safe, 467 stands for a decade, the archive survives Yahoo's shutdown by days, and then an outsider from Reddit rewrites the record book. Every event sourced to its original message.
- [The notable boards](https://eternity2.dev/research/community/boards) — The famous Eternity II boards, treated as first-class citizens: Verhaard's prize-winning 467, the Blackwood record line to 470, the strict five-clue record to 464, and this project's own boards: who found each one, when, and what makes it structurally interesting. Every bundled board opens in the viewer.
---
# The notable boards
> The famous Eternity II boards, treated as first-class citizens: Verhaard's prize-winning 467, the Blackwood record line to 470, the strict five-clue record to 464, and this project's own boards: who found each one, when, and what makes it structurally interesting. Every bundled board opens in the viewer.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/community/boards/
- Updated: 2026-07-02
- Source: msg 6337: the January 2009 scrutiny result, $10,000 to Anna Karlsson of Lund for 467/480 — https://groups.io/g/eternity2/message/6337
- Source: msg 10032: Blackwood's 468 reaches the list from Reddit — https://groups.io/g/eternity2/message/10032
- Source: msg 10045: McGavin announces the 469 — https://groups.io/g/eternity2/message/10045
- Source: msg 10074: Fernandez's single-piece-swap 469 — https://groups.io/g/eternity2/message/10074
- Source: msg 10117: Blackwood's 470 — https://groups.io/g/eternity2/message/10117
- Source: msg 11401: Bucas ties the 470 — https://groups.io/g/eternity2/message/11401
- Source: msg 11074: Gauthier's strict five-clue 460 (the record 2023–2026) — https://groups.io/g/eternity2/message/11074
- Source: Riotte's 464, the new strict five-clue record (groups.io, July 2026) — https://groups.io/g/eternity2/message/11919
- Source: e2.bucas.name (Jef Bucas), the community viewer these boards were first shared in — https://e2.bucas.name
---
The history of Eternity II is usually told through the people and the
solvers. This page tells it through the boards themselves: the handful of
artifacts that carry the whole record line, whoever found them. Each entry
says who announced the board and where, and what is publicly known about its
structure. Every board shown here is bundled with this site: click any
preview to load it in the [viewer](/viewer) and re-score it edge by edge. The
full timeline and method history live on the
[records page](/research/records); this is the gallery.
## The record line: 467 → 468 → 469 → 470
### Verhaard's 467 (2008): the prize board
The board that won the only money Eternity II ever paid out. Louis Verhaard's
[distributed *eii* solver](/research/lab/experiments/louis-verhaard/eii), released to
volunteers in September 2008 with the words "I am stuck"
([msg 5940](https://groups.io/g/eternity2/message/5940)), had found 467/480
more than 40 times by different users by the time he
documented it ([msg 6275](https://groups.io/g/eternity2/message/6275)). At
the 31 December 2008 scrutiny the entry, submitted under the name of Anna
Karlsson of Lund (Verhaard's household), took the $10,000 best-partial prize
([msg 6337](https://groups.io/g/eternity2/message/6337), with the group
recognizing the winner immediately at
[msg 6349](https://groups.io/g/eternity2/message/6349)).
Two later footnotes change how the board reads. Re-reading Verhaard's own
account, Jef Bucas noted the 467 was achieved under a self-imposed handicap:
Verhaard had misinterpreted the rules and restricted the kind of board he
allowed himself to submit
([msg 10582](https://groups.io/g/eternity2/message/10582)). And by 2023 the
score had become a calibration target: "last week I found about 30 x 467. I
use this target to calibrate the code I'm running"
([msg 11001](https://groups.io/g/eternity2/message/11001)). It stood as the
record for twelve years anyway.
> **[Interactive: CommunityBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
The viewer also bundles three sibling 467s (Verhaard 467a, 467b, 467c) from
the same campaign, in its board menu.
### Blackwood's 468 (2020): the outsider board
The first advance past 467 in twelve years arrived from outside the
community: a Reddit post by Joshua Blackwood, then unknown to the mailing
list, relayed at [msg 10032](https://groups.io/g/eternity2/message/10032). A
board with only 12 breaks, 468/480. Bucas verified it, added it to the
viewer, and passed on the startling claim that the author could "find a 468
every 4 days" ([msg 10033](https://groups.io/g/eternity2/message/10033)).
Days later Blackwood open-sourced
[the solver that found it](/research/lab/experiments/joshua-blackwood/solver)
([msg 10037](https://groups.io/g/eternity2/message/10037)). Every board below
this point traces back to that code.
> **[Interactive: CommunityBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
### McGavin's 469 (2020): the ceiling board
"Your solver is fantastic!! I ran it for a few days on about a couple of
hundred cores and hit the jackpot. New record score of 469! Only 11 breaks!"
So wrote Peter McGavin, running Blackwood's freshly published solver
([msg 10045](https://groups.io/g/eternity2/message/10045)). Blackwood
explained the solver's break discipline in the same thread: breaks are never
allowed to touch, so this 469 is equivalently a 249-piece partial with 7
holes ([msg 10051](https://groups.io/g/eternity2/message/10051)).
Structurally this board has a signature you can check yourself: all eleven of
its unmatched edges sit in the top five rows, leaving an eleven-row slab that
is locally flawless: the damage is swept up against the edge the search
finished at. That geometry, and why it flips on boards built in the other
direction, is analyzed in
[Where the mismatches live](/research/why/mismatch-geometry).
> **[Interactive: CommunityBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
### The 469 wave (October–November 2020)
Bucas posted his first two 469s in October
([msg 10052](https://groups.io/g/eternity2/message/10052), "Here are 2
other 469"), ported Blackwood's algorithm to C, roughly doubling its speed
([msg 10065](https://groups.io/g/eternity2/message/10065)), and the stream
kept coming: boards c through g in November, from
[msg 10067](https://groups.io/g/eternity2/message/10067) ("Another one yay
\o/") through [msg 10078](https://groups.io/g/eternity2/message/10078),
which shipped together with the open-sourced port, *libblackwood*. All seven
are in the viewer's board menu as Bucas 469a–g.
The odd one out is Carlos Fernandez's. He fed McGavin's 469 to a
piece-exchange program of his own and found another 469 differing by a
single piece ([msg 10074](https://groups.io/g/eternity2/message/10074)),
proof, in Bucas's words when he added it to the viewer, that "even at 469,
there is still a little bit of flexibility left"
([msg 10075](https://groups.io/g/eternity2/message/10075)). That flexibility
is thin: it is exactly the kind of isolated tie the rigidity analysis below
predicts, not a path upward.
### The 470s (2021 and 2024): the current ceiling
Blackwood closed his own arc with a nearly wordless message: a viewer URL, a
piece list (470/480, 10 breaks), and the single line "I haven't physically
put it together yet, but I found one!"
([msg 10117](https://groups.io/g/eternity2/message/10117)). He later
confirmed it was found with exactly the public repo's code
([msg 10161](https://groups.io/g/eternity2/message/10161)), the retuned
schedule he had published in advance, and then stopped: "I have not written
a line of code or run any algorithms since I found a 470"
([msg 10185](https://groups.io/g/eternity2/message/10185)).
> **[Interactive: CommunityBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
The score has been tied once, not beaten. In December 2024 Bucas restarted "a
few threads of Joshua's code" and another 470 appeared
([msg 11401](https://groups.io/g/eternity2/message/11401)); Fernandez noticed
its three unmatched border figures allowed a rearrangement and posted a 470b
([msg 11403](https://groups.io/g/eternity2/message/11403)). The 470 class
remains a Blackwood-solver monopoly, and 470 remains the ceiling: "No one is
even close… The best result to date, is the partial solution from Joshua
Blackwood, 470/480 edges matching"
([msg 11343](https://groups.io/g/eternity2/message/11343)).
> **[Interactive: CommunityBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## What the record boards share
Three publicly documented facts hold across the 468/469/470 line.
**Same clue regime.** All of them place the mandatory starter piece at its
official spot and none of the four optional clue pieces at theirs. The
community checked this for the 470
([msg 10554](https://groups.io/g/eternity2/message/10554)), and Blackwood
confirmed he had the clue tiles and deliberately did not use them
([msg 10185](https://groups.io/g/eternity2/message/10185)). Under the
original contest rules only the starter was mandatory, so these boards were
prize-eligible as submitted.
**Banded mismatches.** The few unmatched edges are not scattered; they pack
into one band of rows against the edge the search finished at.
[Where the mismatches live](/research/why/mismatch-geometry) scores the real
boards live and shows the band; it also shows this project's boards
exhibiting the same band, mirrored.
**Local rigidity.** None of these boards is an almost-solution waiting for a
polish. [The rigidity wall](/research/why/rigidity-wall) uses exact
integer-programming solves to prove that on record boards, freeing whole
regions and refilling them optimally returns the same arrangement: the best
boards sit at the bottom of their own valleys. Fernandez's single-piece-swap
469 is the exception that proves the rule: a tie next door, never a step up.
## The strict five-clue line: from Gauthier's 460 to Riotte's 464
Most published record boards keep only the mandatory starter piece. In March
2023 a thread asked the stricter question: what is the best board that
respects **all five** official clues
([msg 11037](https://groups.io/g/eternity2/message/11037))? A ladder climbed
in days: 412, then 452 and 457 (Carlos Fernandez), then 453/455/456/458
(David Barr), stopping at **Bruno Gauthier's 460**, "made with my own
program and Eternity II Editor"
([msg 11074](https://groups.io/g/eternity2/message/11074)); McGavin converted
it to the shared viewer format
([msg 11081](https://groups.io/g/eternity2/message/11081)).
That 460 stood for more than three years. In late June 2026 a new thread,
"Record of Eternity2 with 5 hints?", reopened the line: **Benjamin Riotte**
posted a 461, then a wave of results climbing to **464/480 (only 16 broken
edges)**, the first advance on the strict record since 2023
([groups.io msg 11919](https://groups.io/g/eternity2/message/11919)).
**Igor Pejic** reached the same 463–464 range independently in the same
thread, using what he described as a "modified Blackwood's DFS algorithm". All
of the boards respect the five clues at their official cells (starter #139,
hints #208, #255, #181, #249).
The strict line matters because of what the five clues imply: the community's
solution-count estimate with all five clues placed is about 0.00000004
(effectively a unique target), against roughly 14,702 expected "solutions"
with the starter alone
([msg 11193](https://groups.io/g/eternity2/message/11193)). A five-clue board
is playing the exact game the designers set. The 464 stands sixteen edges below
the full solution, and that gap is itself a datum.
> **[Interactive: BoardsFoundGallery]** Rendered on the canonical page (link above); not shown in this markdown export.
## This project's boards
This site's own experiments produced four boards worth keeping, labelled for
what they are: experiments several steps below the community line, and
documented end to end. Each links
to the experiment that found it; every score is recomputed from the board's
own edges.
### PALIMPSEST (463)
The project's best, found by reading the whole corpus of strong boards to
separate genuine shared structure from shared bad habits, then attacking the
traps. The full story is in
[the PALIMPSEST experiment](/research/lab/experiments/raphael-anjou/learning/palimpsest).
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
### KEYRING (460)
Built from scratch by letting three learned signals vote on every placement,
reaching 460 in a board family no earlier search here had cracked;
see [KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring). Its mismatches band in the
bottom rows, the mirror image of McGavin's 469. That scan-order fingerprint
is explained in [Where the mismatches live](/research/why/mismatch-geometry).
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
### PRIOR (460)
From scratch plus destroy-and-repair, breaking construction ties by where
pieces tend to sit in strong boards (see
[PRIOR](/research/lab/experiments/raphael-anjou/learning/prior)).
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
### GAUNTLET (458)
The same [beam search](/research/build/construct/beam-search) run across nine
scan orders so it lands in different regions of board space; the zigzag order
found this 458. The full run is written up in
[GAUNTLET](/research/lab/experiments/raphael-anjou/pipelines/gauntlet).
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## Related
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [Where the mismatches live](https://eternity2.dev/research/why/mismatch-geometry) — A near-perfect board doesn't scatter its few errors evenly. It packs them into one band of five rows and leaves all the rest flawless. Which band is decided by the direction the search filled the board, and you can see the mirror on the real record boards.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
---
# The hunt, a history (part I: 2000–2009)
> The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/community/hunt/
- Updated: 2026-07-02
- Source: Monckton's copyright ultimatum and Brendan Owen's disqualification (groups.io message 1342) — https://groups.io/g/eternity2/message/1342
- Source: Brendan Owen derives the 17+5 design as the hardest possible puzzle (groups.io message 1947) — https://groups.io/g/eternity2/message/1947
- Source: eternity2.net shuts down: 1.6 TFlops, 10^19 operations, no solution (groups.io message 3511) — https://groups.io/g/eternity2/message/3511
- Source: Louis Verhaard releases the eii solver: “I am stuck” (groups.io message 5940) — https://groups.io/g/eternity2/message/5940
- Source: First scrutiny result: $10,000 to Anna Karlsson of Lund, score 467 (groups.io message 6337) — https://groups.io/g/eternity2/message/6337
- Source: Wikipedia, "Eternity II puzzle" — https://en.wikipedia.org/wiki/Eternity_II_puzzle
---
Eternity II has an official history (a press release, a prize, a deadline)
and it has a real one, which happened on a mailing list. The
[eternity2 group](https://groups.io/g/eternity2) (originally a Yahoo group)
is where the puzzle was analysed before it existed, digitized the day it
launched, declared unsolvable within a fortnight, and where the only prize
money ever paid out was quietly traced to one of the list's own regulars.
This page tells that story from the archive itself: every event links to the
message where it happened. Part I runs from the group's founding to the first
scrutiny date, 2000 → February 2009.
## The group that predates its puzzle (2000–2006)
The strangest fact about the Eternity II community is its birthday. Brendan
Owen created the eternity_two group in October 2000, six and a half years
before the puzzle was announced, "because we anticipated a follow up
puzzle" to Eternity I
([msg 404](https://groups.io/g/eternity2/message/404)). The earliest
surviving messages guess at what the sequel's pieces might be, betting
wrongly on 3D tetrahedra
([msg 8](https://groups.io/g/eternity2/message/8)).
By 2001 the group was already rehearsing the designer's problem rather than
the solver's: Günter Stertenbrink asked how one would build a puzzle with a
huge prize and only a ~1% chance of being solved in ten years
([msg 15](https://groups.io/g/eternity2/message/15)), almost exactly the
brief Christopher Monckton's team would later execute. Eternity I veterans
drifted in and waited: Dave Clark, author of the distributed E1 solver
ESolve, rejoined in April 2001
([msg 21](https://groups.io/g/eternity2/message/21)). When the first press
reports arrived in December 2005, with Monckton promising the sequel would
need "the lifetime of the Universe" while co-designer Oliver Riordan said
"several years", Stertenbrink drew the obvious conclusion: "So we can
conclude, the end of the universe is in several years."
([msg 34](https://groups.io/g/eternity2/message/34))
## Launch summer (2007)
The announcement came on 22 January 2007: 256 pieces, US$2,000,000 for the
first correct solution, launch on 28 July, and a London Toy Fair stunt in
which a strongman destroyed "the computer containing a solution"
([msg 36](https://groups.io/g/eternity2/message/36)). That computer, a
distributor insider later admitted, contained nothing at all
([msg 148](https://groups.io/g/eternity2/message/148)).
The group did not wait for the pieces. Within two days it was counting edge
types on promotional photos, and Owen had derived the expected-solutions
formula that made difficulty a tunable design parameter
([msg 38](https://groups.io/g/eternity2/message/38),
[msg 44](https://groups.io/g/eternity2/message/44)). He published a
generator for Eternity-II-like puzzles
([msg 41](https://groups.io/g/eternity2/message/41)); Alan O'Donnell had the
first working solver a week later
([msg 64](https://groups.io/g/eternity2/message/64)). By 8 July,
three weeks before anyone had touched a real piece, anr_56 had estimated
the edge-type counts from launch photos and Owen had run the numbers on that
configuration: about 10^30 times harder than the benchmarks the group was
solving in minutes. "No one will solve this puzzle," he concluded
([msg 665](https://groups.io/g/eternity2/message/665)).
Then came the famous 48 hours. Australia got the puzzle first, and Owen was
at Kmart on launch eve, "first one to grab it off their stack"
([msg 1051](https://groups.io/g/eternity2/message/1051)). He fed the piece
panels through a computer-vision program he had written in advance
([msg 789](https://groups.io/g/eternity2/message/789)), reported the edge
distribution "as flat as could be"
([msg 1054](https://groups.io/g/eternity2/message/1054)), and confirmed the
real parameters: 5 border colours, 17 interior colours, the group's
worst-case scenario ([msg 977](https://groups.io/g/eternity2/message/977)).
Then, after Dave Clark's "nice chat to Mr Monckton", he published the first
solution-count estimate for the real piece set: roughly 5,930 solutions
with the mandatory clue, against Monckton's own figure of about 5 million
([msg 987](https://groups.io/g/eternity2/message/987)). The puzzle was two
days old and its difficulty was already measured.
The same week set the community's norms. To verify piece transcriptions
without sharing copyrighted data, members converged on published CRC
checksums ([msg 1063](https://groups.io/g/eternity2/message/1063)); and
newcomer Sergio Demian Lerner warned about "phantom sets": a stranger's
benchmark could be the real puzzle re-encoded, tricking you into handing
over a $2M solution
([msg 1210](https://groups.io/g/eternity2/message/1210)).
## The theory race (2007–2008)
A week into August, the puzzle's inventor arrived in person; he came to
disqualify someone. Posting from an unverifiable Yahoo address, Christopher
Monckton declared he held copyright on the piece designs and that his
"web-trawling software" had flagged Brendan Owen's website: Owen, who had
been aiming for the lesser prize, was out
([msg 1342](https://groups.io/g/eternity2/message/1342),
[msg 1352](https://groups.io/g/eternity2/message/1352),
[msg 1358](https://groups.io/g/eternity2/message/1358)). It later emerged
the flagged file did not even correspond to the real pieces
([msg 1358](https://groups.io/g/eternity2/message/1358)). Members protested
that an anonymous forum account could hardly carry official weight; Monckton
doubled down ([msg 1374](https://groups.io/g/eternity2/message/1374),
[msg 1377](https://groups.io/g/eternity2/message/1377)). The chill was real
and lasting: a year later Owen, as moderator, was still deleting derived
data tables "to be safe"
([msg 5651](https://groups.io/g/eternity2/message/5651)). It was in this
climate that Owen, showing that E2's flat piece statistics killed the
strategy that had cracked Eternity I, drew his line: "If we cannot talk
about something as fundamental as piece frequencies, then I will give up on
this puzzle now" ([msg 1667](https://groups.io/g/eternity2/message/1667),
[msg 1694](https://groups.io/g/eternity2/message/1694)).
The theory came fast after that. kubzpa proved by a
[parity argument](/research/build/analysis/parity-arguments) that a score
of exactly 479 (one mismatched edge) is impossible
([msg 1640](https://groups.io/g/eternity2/message/1640)); hold that thought,
because seventeen months later the argument would meet a counterexample.
Owen then produced the period's signature result: assuming the designers
wanted the hardest possible 16×16, with a flat distribution and about one
expected solution, the interior colour count should be (196! · 4^196)^(1/392) ≈
17.14, hence 17 interior colours and 5 border colours. Exactly the real
puzzle ([msg 1947](https://groups.io/g/eternity2/message/1947)). The 17+5
split was not bad luck; it was [a bullseye on the hardness
peak](/research/why/phase-transition), and the group had reverse-engineered
the targeting.
Theoretical claims were checked against node counts measured on shared
[benchmarks](/research/build/benchmarks). Angel de Vicente, posting as
Txibilis, built the standard suite of E2-like test boards
([msg 1610](https://groups.io/g/eternity2/message/1610)), and a duel began:
Txibilis's hand-designed [fill orders](/research/build/backtracking/fill-order)
against the automated strategy-optimizer of doc_s_smith, driving
full-search node counts down by orders of magnitude; one benchmark stood at
89,794 nodes, answered within two days by 85,729
([msg 2896](https://groups.io/g/eternity2/message/2896),
[msg 2928](https://groups.io/g/eternity2/message/2928)). Mid-duel, the list
discovered who doc_s_smith was: Dietmar Wolz, finder of most of the known
Eternity I solutions ([msg 2972](https://groups.io/g/eternity2/message/2972)).
The estimates converged too. kubzpa's paper put the solution count near 15
million ([msg 3497](https://groups.io/g/eternity2/message/3497)); Owen
measured the search tree's branching depth by depth, finding the node count
peaks at 161 pieces placed
([msg 3147](https://groups.io/g/eternity2/message/3147)); and in April 2008
he announced he had "nailed the theory": an exact model of the search tree
whose predictions sat on top of the empirical curves, independently
cross-checked by Louis Verhaard: on the order of 10^27 CPU-years per
solution ([msg 5197](https://groups.io/g/eternity2/message/5197),
[msg 5193](https://groups.io/g/eternity2/message/5193),
[msg 5209](https://groups.io/g/eternity2/message/5209)). The group's chief
theorist acted on his own numbers: he was now hunting the highest score, not
480 ([msg 4996](https://groups.io/g/eternity2/message/4996)).
## Big iron and syndicates
If one computer was hopeless, perhaps thousands were not. Dave Clark's
eternity2.net, a BOINC-based
[distributed solver](/research/build/faster/distributed-solving), launched
with the puzzle in July 2007
([msg 756](https://groups.io/g/eternity2/message/756)) and had
1,300 members within a month, 160 of them in the USA, where the puzzle had
not even been released
([msg 2122](https://groups.io/g/eternity2/message/2122),
[msg 2132](https://groups.io/g/eternity2/message/2132)). The project
submitted a 462-edge partial to Tomy, then 463
([msg 2663](https://groups.io/g/eternity2/message/2663)), a number that
would function as the community's public score ceiling for over a year.
A prize-sharing "Eternity 2 Syndicate" followed, paying members in
proportion to placements contributed
([msg 3021](https://groups.io/g/eternity2/message/3021)).
It lasted five months. In December 2007 Clark shut eternity2.net down,
publishing the final accounting: over 1.6 TFlops of aggregate computing,
over 10^19 CPU operations, best scores in the mid-460s. His verdict: a
brute-force solution "was always clearly going to be impossible"
([msg 3511](https://groups.io/g/eternity2/message/3511)). The list
immediately spawned threads titled "Eternity2 Must Have Been Solved"; it had
not ([msg 3554](https://groups.io/g/eternity2/message/3554)). The project's
files were preserved in the group archive
([msg 3633](https://groups.io/g/eternity2/message/3633)), and Clark
open-sourced his research solver
([msg 3716](https://groups.io/g/eternity2/message/3716)). He also left the
archive its best primary source on the puzzle's creation: a phone call with
Monckton, who described judges typing entropy into a generator built by the
Eternity I winners Alex Selby and Oliver Riordan, the solution printed once
and vaulted ([msg 4177](https://groups.io/g/eternity2/message/4177)),
"whilst all parties were out of the room", as the Tomy distributor leaflet
he had posted at launch put it
([msg 901](https://groups.io/g/eternity2/message/901)).
The exact-methods flank fared no better. A published
[SAT encoding](/research/build/exact/sat-csp-encodings) of the
full puzzle came with a call for anyone owning a machine with more than
16 GB of RAM ([msg 4084](https://groups.io/g/eternity2/message/4084));
integer programming died at 8×8 boards
([msg 5602](https://groups.io/g/eternity2/message/5602)). Raw speed kept
climbing: istarinz broke 100 million placements per second on multiple
cores ([msg 5804](https://groups.io/g/eternity2/message/5804)), then
measured 558 million on a brand-new Core i7
([msg 6212](https://groups.io/g/eternity2/message/6212)). But against
search spaces measured in powers of forty, throughput was a rounding error.
## The road to 467 (2008)
The scores had crept up all along: 410 from Pierre Schaus in the first weeks
([msg 1568](https://groups.io/g/eternity2/message/1568)), 453 from
philippe.dupond ([msg 3998](https://groups.io/g/eternity2/message/3998)),
461 from e2dude ([msg 4440](https://groups.io/g/eternity2/message/4440)),
with 463 as the "current known high"
([msg 5688](https://groups.io/g/eternity2/message/5688)). The methods
changed character in mid-2008: Schaus posted his constraint-programming
paper, whose key move (remove a set of non-adjacent pieces and re-place
them *optimally* by solving an assignment problem) became the engine of
antminder's hybrid, averaging a 462 per day
([msg 5589](https://groups.io/g/eternity2/message/5589),
[msg 5601](https://groups.io/g/eternity2/message/5601)).
Meanwhile two people had quietly gone past the ceiling. In a long August
exchange, Max reported getting "beyond the 'don't-talk-about-limit' quite
easily"; Louis Verhaard ("Max, you are really a dangerous man!") matched
methods with him and found they had converged on the same heuristic
signature, promising full disclosure "after new-year", that is, after the 31
December scrutiny date
([msg 5767](https://groups.io/g/eternity2/message/5767),
[msg 5780](https://groups.io/g/eternity2/message/5780)). Verhaard would
reveal only the geometry: his best high-score fill orders looked like a
"comb": most rows scanned horizontally, the rest vertically
([msg 6112](https://groups.io/g/eternity2/message/6112)).
Then, on 22 September 2008, Verhaard published
[his solver](/research/lab/experiments/louis-verhaard/eii) for anyone to run
at fingerboys.se: "This because I am stuck and my only hope to improve my
best score is by using brute force"
([msg 5940](https://groups.io/g/eternity2/message/5940)). Any prize would be
split 50-50 with the best-scoring user. antminder, whose own program needed
a week to reach 463, was blunt: eii "completely blows it away"
([msg 5950](https://groups.io/g/eternity2/message/5950)). The prize itself
was still folklore (the rules only promised a discretionary "lesser prize")
until Max traced the $10,000 figure to a Monckton interview on the French
official site ([msg 6073](https://groups.io/g/eternity2/message/6073)).
Verhaard, three months before winning it, claimed not to care: he was in it
for the honour, and only for this year
([msg 6072](https://groups.io/g/eternity2/message/6072)).
The same autumn the hand-solvers surfaced: Christine Raisin down to 21
pieces left in the box lid
([msg 5935](https://groups.io/g/eternity2/message/5935)), Verhaard sorting
pieces by pattern with his seven-year-old daughter and running a
solve-by-hand mini-competition, the threads coining nicknames like "pink
swords" and "kipper ties" for the motifs
([msg 5933](https://groups.io/g/eternity2/message/5933),
[msg 5929](https://groups.io/g/eternity2/message/5929),
[msg 5958](https://groups.io/g/eternity2/message/5958),
[msg 5959](https://groups.io/g/eternity2/message/5959)). The list was never
only a solver bench.
## The scrutiny-date farce (December 2008 → January 2009)
Going into the first scrutiny date, the community was in a holding pattern:
several members refused to commit serious effort until the puzzle proved it
could survive its first deadline
([msg 6216](https://groups.io/g/eternity2/message/6216)); NickB had £20 on
there being no winner and considered his money safe
([msg 5979](https://groups.io/g/eternity2/message/5979)).
31 December 2008 came and went. Nothing. Members emailed Tomy's UK and US
addresses and left voicemail, and got no response
([msg 6245](https://groups.io/g/eternity2/message/6245)); the consumer
Careline knew of "no details as of yet, regarding whether or not there is a
winner" ([msg 6295](https://groups.io/g/eternity2/message/6295)). The rules
said winners would be notified within 14 days and results published in the
London Times and New York Times
([msg 6296](https://groups.io/g/eternity2/message/6296),
[msg 6336](https://groups.io/g/eternity2/message/6336)); no such publication
ever visibly appeared. Meanwhile, on 6 January, Verhaard released his
solver's documentation and disclosed its ceiling: users had found 467 more
than forty times ([msg 6275](https://groups.io/g/eternity2/message/6275)).
On 15 January, roughly the last day of the 14-day window, Henk van der
Griendt found a statement behind an unobtrusive link on the UK site only:
hundreds of entries, none complete, the $2M prize still open, and a $10,000
runner-up prize to **Anna Karlsson of Lund, Sweden, for 467 of 480**
([msg 6337](https://groups.io/g/eternity2/message/6337)). No press release,
no Times, and not a word from Monckton at any point. The list needed about
an hour to decode "Lund + 467": the entry was from Louis Verhaard's
household, submitted in the name of his wife, who had received the
congratulation email days earlier
([msg 6349](https://groups.io/g/eternity2/message/6349)). Congratulations
poured in from every regular; Max revealed his own best had been 465, each
extra edge costing his program a factor of ~30 in time
([msg 6348](https://groups.io/g/eternity2/message/6348)). The Swedish press
ran the family photo: over a year of work, 13 "seams" from a full solution,
part of the money going to the World Wildlife Fund
([msg 6374](https://groups.io/g/eternity2/message/6374),
[msg 6375](https://groups.io/g/eternity2/message/6375)).
The aftermath had three barbs. A researcher from the Lleida SAT group posted
that reaching 470 was "rather easy" with their methods, a claim never
backed by any submitted entry, as Max pointedly noted
([msg 6306](https://groups.io/g/eternity2/message/6306),
[msg 6360](https://groups.io/g/eternity2/message/6360)). Tomy told a member
there would be no further [clue puzzles](/research/build/clue-puzzles),
which most of the list welcomed as keeping the challenge pure
([msg 6381](https://groups.io/g/eternity2/message/6381)). And Verhaard
casually settled a seventeen-month-old theorem: 479 *is* reachable, by
flipping a border piece whose two border edges share a colour in a 480
solution. The 2007 parity proof had missed the outward-facing border edges,
which the score never counts
([msg 6317](https://groups.io/g/eternity2/message/6317)).
Even the community's cleanest impossibility result had a loophole; the
puzzle, unsolved, kept all of its.
## The story continues
> **Part II: 2009–2026**
>
> The story continues in [part II](/research/community/hunt-part-2): the Hungarian "we solved it" rumor resolved, the contest's death with its solution locked in a safe, the decade of 467, the migration that saved this archive by days, and the record wave that took the board to 470.
## Related
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
---
# The hunt, a history, part II: 2009–2026
> Seventeen years after the prize: the contest dies with its solution locked in a safe, 467 stands for a decade, the archive survives Yahoo's shutdown by days, and then an outsider from Reddit rewrites the record book. Every event sourced to its original message.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/community/hunt-part-2/
- Updated: 2026-07-02
- Source: Monckton closes the contest: “Extension of time is not allowed under the rules” (groups.io message 8477) — https://groups.io/g/eternity2/message/8477
- Source: Peter McGavin solves Brendan Owen's 10×10 in ~180 core-years, inside complex theory's error bars (groups.io message 9688) — https://groups.io/g/eternity2/message/9688
- Source: McGavin's 469 with Blackwood's solver: “New record score of 469! Only 11 breaks!” (groups.io message 10045) — https://groups.io/g/eternity2/message/10045
- Source: Joshua Blackwood's 470, the standing record (groups.io message 10117) — https://groups.io/g/eternity2/message/10117
- Source: Bruno Gauthier's 460, the best five-clue board 2023–2026 (groups.io message 11074) — https://groups.io/g/eternity2/message/11074
- Source: Benjamin Riotte's 464, the new five-clue record (groups.io, July 2026) — https://groups.io/g/eternity2/message/11919
- Source: Jef Bucas ties the 470 with Blackwood's code (groups.io message 11401) — https://groups.io/g/eternity2/message/11401
---
[Part I](/research/community/hunt) ended in January 2009 with a $10,000
cheque decoded in an hour and a puzzle that had survived its first deadline
untouched. Part II covers everything since: the contest's slow death, the
decade in which one number, 467, refused to move, the day the archive
itself nearly vanished, and the record wave that finally moved it. As
before, this is the community's story told from the
[eternity2 archive](https://groups.io/g/eternity2), and every event links to
the message where it happened.
## The long silence and the diehards (2009–2010)
The Hungarian "we solved it" rumor that part I left hanging died the way
such claims always die on this list. Once properly translated (a team at
ELTE university claiming a full solve via an algorithm salad that included
alpha-beta pruning, for a one-player puzzle), the native speakers called the
posts confused and the engineers ran the arithmetic
([msg 6575](https://groups.io/g/eternity2/message/6575)). By June the
claimant's account and every comment had been deleted from the Hungarian
forum ([msg 6764](https://groups.io/g/eternity2/message/6764)).
The verified story of 2009 was quieter and better. JSA ran Louis Verhaard's
public [eii solver](/research/lab/experiments/louis-verhaard/eii) on a single PC
and logged the climb: four million 463s, 625 466s, and finally, at day ~82,
a 467, independently reproducing the prize score with the public binary and
quantifying just how thin the air gets above 466
([msg 6687](https://groups.io/g/eternity2/message/6687)). In
August, Verhaard himself gave the definitive first-person account of the
winning entry: "Anna is my wife"; she submitted it, "it was my program that
did the job", and the 467 board contained 247 flawless pieces
([msg 6891](https://groups.io/g/eternity2/message/6891)). That December he
disclosed the method part I could only hint at: he had found 467 "more than
50 times", by letting the backtracker place one mismatching edge at chosen
depths, the move known as
[edge slipping](/research/build/reduce/edge-slipping), gated by depth
([msg 7321](https://groups.io/g/eternity2/message/7321)).
The second scrutiny date, 31 December 2009, replayed the first as farce
minus the payout. Weeks of silence, then a news item ("Eternity II remains
unsolved") with no winner and, unlike 2008, no best-partial prize at all
([msg 7471](https://groups.io/g/eternity2/message/7471)). A member extracted
a written Q&A from Tomy: no organised activity in 2009, partial solutions
were not even scored, and "the year one prize was part of the initial
promotion only" ([msg 7488](https://groups.io/g/eternity2/message/7488)).
Verhaard confirmed he had not improved on his 467
([msg 7494](https://groups.io/g/eternity2/message/7494)). He also
republished his solver at shortestpath.se, its long-term home, after his
band's website folded
([msg 7439](https://groups.io/g/eternity2/message/7439),
[msg 7446](https://groups.io/g/eternity2/message/7446)).
Claims bloomed in the vacuum anyway: an E2Lab popup announcing a 471 on a
board that was never saved
([msg 7284](https://groups.io/g/eternity2/message/7284)), a hand-solve
implying at least 472, dismissed on sight
([msg 7383](https://groups.io/g/eternity2/message/7383),
[msg 7427](https://groups.io/g/eternity2/message/7427)). Neither produced a
board. Meanwhile Eternity I veteran doc_s_smith returned after three years
and turned the list into an algorithms workshop
([msg 7755](https://groups.io/g/eternity2/message/7755)), framing the
target the era could realistically chase: solving E2 itself being hopeless,
"the challenge is clear: beat 468 matching edges"
([msg 7803](https://groups.io/g/eternity2/message/7803)). When his
Monte-Carlo estimator disagreed with Brendan Owen's
[complexity model](/research/why/complex-theory) by five
orders of magnitude, Verhaard defended the model as "the finest work that
has ever been published about E2"
([msg 7810](https://groups.io/g/eternity2/message/7810)), a sentence the
next fifteen years would keep proving right.
## The contest dies (December 2010 → 2011)
Asked for a one-year extension of the final scrutiny date, Tomy answered in
one line: "No decsions [sic] will be made until the next scrutiny date"
([msg 8006](https://groups.io/g/eternity2/message/8006)). The deadline of
31 December 2010 passed in gallows humour, with members joking about
couriering a hypothetical last-minute solution to the post office
([msg 8197](https://groups.io/g/eternity2/message/8197)). Then the official
site went dark, "suspended pending final scrutiny confirmation"
([msg 8270](https://groups.io/g/eternity2/message/8270)); Verhaard posted
"I can honestly tell you that I didn't solve it"
([msg 8277](https://groups.io/g/eternity2/message/8277)).
The end arrived via France: a Tomy France press release (deadline passed,
no winner, no additional year, the $2M unclaimed) met with disbelief until
the source URL surfaced and the English page followed two days later
([msg 8339](https://groups.io/g/eternity2/message/8339),
[msg 8373](https://groups.io/g/eternity2/message/8373)). Monckton's emailed
reply to a member closed the door in person: the competition was over, and
"Extension of time is not allowed under the rules", rules which, members
immediately pointed out, explicitly allowed further scrutiny dates at the
promoter's discretion
([msg 8477](https://groups.io/g/eternity2/message/8477),
[msg 8478](https://groups.io/g/eternity2/message/8478)). Brendan Owen, the
man who had been posting since before the puzzle launched, said goodbye: "It
was great fun". He also asked Verhaard to congratulate his wife on the
highest score, first-party confirmation that 467 stood at the contest's
close ([msg 8429](https://groups.io/g/eternity2/message/8429)).
Two loose ends defined the aftermath. One was the claim vacuum: a member's
476 without the hint (the original message is missing from the archive;
it survives through its replies,
[msg 8247](https://groups.io/g/eternity2/message/8247)) escalated into
multiple hintless-480 claims and met the community's zero-knowledge wall
(Owen: publish the grid of piece ids); nothing was ever published
([msg 8676](https://groups.io/g/eternity2/message/8676)). The other was the
solution itself. The group owner laid out the custody arrangement: nobody,
not Tomy, not Monckton, knows the solution; it sits with independent loss
adjusters ([msg 8502](https://groups.io/g/eternity2/message/8502)). Owen
then gave the statement still quoted today: Alex Selby and Oliver Riordan
created a solution when Monckton paid them to generate a practically
impossible puzzle, and it lies "hidden amongst reams of printed text which
is locked in a safe" as insurance against legal challenge
([msg 8823](https://groups.io/g/eternity2/message/8823)). In April 2011 the
official site confirmed the prize had gone unclaimed
([msg 8846](https://groups.io/g/eternity2/message/8846)).
What replaced the money was a ladder. Community leaderboards appeared in the
group database ([msg 8222](https://groups.io/g/eternity2/message/8222),
[msg 8735](https://groups.io/g/eternity2/message/8735)); Owen, prompted by a
golden-ratio guess, proved that a backtracker's node counts peak at fraction
(1 − 1/e) of the board: 256 × 0.632 = 161.8, the closed form behind the
long-observed "magic of 161"
([msg 8125](https://groups.io/g/eternity2/message/8125)); and the first
exhaustive searches of Owen's
[9×9 benchmarks](/research/build/benchmarks) landed within an order of
magnitude of his theory's predictions
([msg 8793](https://groups.io/g/eternity2/message/8793),
[msg 8803](https://groups.io/g/eternity2/message/8803)). A new name did most
of that checking: Peter McGavin, who by 2011 was running the list's
complex-theory service desk and had computed the numbers that became
canon, about 14,702 expected solutions with the mandatory piece placed
([msg 8924](https://groups.io/g/eternity2/message/8924)).
## The quiet years (2012–2018)
Three full years, 2012 to 2014, produced fewer messages than a busy
quarter of 2008. Academia arrived in person: Tony Wauters posted his group's
peer-reviewed hyper-heuristic paper (best score 461/480 in an hour) and
stayed to answer questions
([msg 9017](https://groups.io/g/eternity2/message/9017)); a second paper,
on MILP and Max-Clique heuristics, followed in 2017
([msg 9683](https://groups.io/g/eternity2/message/9683)). The only prize
money of the era was a Czech retailer's 12,000 EUR re-run of the contest
([msg 9072](https://groups.io/g/eternity2/message/9072)), extended to July
2015 ([msg 9271](https://groups.io/g/eternity2/message/9271)); its expiry
then passed without a single mention on the list. When newcomer Dima asked
for the best known result in late 2014 and got the canonical answer, his
reply was the era in one line: "467 still, really? But that was 6 years
ago!" ([msg 9318](https://groups.io/g/eternity2/message/9318)).
Under the surface, two things matured. Theory became a document: McGavin
transcribed Owen's complexity model into LaTeX as complex_theory.pdf, in
the group's files ([msg 9188](https://groups.io/g/eternity2/message/9188)),
turning folklore into something newcomers could actually read. And speed
became a culture again: Arnaud Carré posted a 114 million recursions per
second single-core solver as a comparison standard
([msg 9233](https://groups.io/g/eternity2/message/9233)), Michael Field
designed an FPGA backtracker
([msg 9226](https://groups.io/g/eternity2/message/9226)), and a
nodes-per-watt ledger ran from Raspberry Pi clusters to an Xbox One X
([msg 9519](https://groups.io/g/eternity2/message/9519),
[msg 9816](https://groups.io/g/eternity2/message/9816)). The claim-filter
kept working too: a 2014 "eternity2 resolved ;)" died when verification
found one piece used five times
([msg 9286](https://groups.io/g/eternity2/message/9286),
[msg 9301](https://groups.io/g/eternity2/message/9301)).
Then, in September 2017, the community's shared benchmark fell. Peter
McGavin found the first new solution to Brendan Owen's hint-free 10×10,
the "monster" that had resisted everyone since 2008, by industrialising
exactly what the list had converged on: enumerate the ~20 million possible
first rows, rank them by complex theory's chance-of-solution per node, and
farm them to every core available, from home ODROIDs and $12 Orange Pis to
borrowed work servers, 400+ cores in all
([msg 9686](https://groups.io/g/eternity2/message/9686),
[msg 9688](https://groups.io/g/eternity2/message/9688)). The solution cost
about 2×10^17 nodes (roughly 180 core-years) and arrived on row-search
~92,907 against a predicted one-per-70,000, inside the theory's error bars:
in McGavin's words, "no new methods, just systematic persistence and the
law of large numbers". John Gilbert spoke for the list: "Amazing that it is
being truly just as hard as predicted"
([msg 9693](https://groups.io/g/eternity2/message/9693)). It was the
strongest validation Owen's theory ever received, and the proof that the
benchmark culture, not the record chase, had been the group's real
achievement of the decade.
The same thread raised a records question that never resolved: Dima
surfaced a 2009 tweet by TopCoder star Naohiro Takahashi claiming a 468,
never accompanied by a board, possibly scored under different contest rules,
and left permanently unverified
([msg 9694](https://groups.io/g/eternity2/message/9694),
[msg 9697](https://groups.io/g/eternity2/message/9697)). Two months later
McGavin placed all 256 real pieces on the 16×16 board with all 480 internal
joins matched, an "unframed E2" with gray edges buried in the interior, and
was scrupulously precise that it was not a solution to the puzzle
([msg 9736](https://groups.io/g/eternity2/message/9736),
[msg 9750](https://groups.io/g/eternity2/message/9750),
[msg 9757](https://groups.io/g/eternity2/message/9757)). The window closed
at low ebb: 2018 produced 42 messages, and to a newcomer's poll a veteran
answered, "anybody still crunching with backtracking ?? no way...."
([msg 9832](https://groups.io/g/eternity2/message/9832)).
## The migration (2019)
In October 2019 Yahoo announced it would erase all Groups content within
weeks: twelve years of piece counts, proofs, debunkings and solver source,
gone. Jef Bucas raised the alarm
([msg 9920](https://groups.io/g/eternity2/message/9920)), Robert Gerbicz
proposed groups.io, founder Ole Knudsen created the new group, and JSA paid
the transfer fee: "I can pay for the first 5 years"
([msg 2](https://groups.io/g/eternity2/message/2)). The automated transfer
completed on 8 November 2019 with roughly 2,000 members, messages, files and
photos intact ([msg 9934](https://groups.io/g/eternity2/message/9934),
[msg 9937](https://groups.io/g/eternity2/message/9937)). The
archive this chronicle is written from survived by days.
Two weeks later, Bucas launched the other half of the modern era's
infrastructure: [e2.bucas.name](https://e2.bucas.name), a board viewer whose
boards live entirely in the URL (nothing is transmitted to the server),
with a drop-down of best boards that quietly became the community's record
book ([msg 9955](https://groups.io/g/eternity2/message/9955)). Fittingly,
the year's records news was archival too: Verhaard confirmed his
astonishing 249-piece flawless partial was real, "no cheating… found it
after a week or so, on 1 computer"
([msg 9890](https://groups.io/g/eternity2/message/9890)).
## The record wave (2020–2021)
On 31 August 2020, a member relayed a Reddit post: a new best board with
only 12 breaks, 468/480, the first advance past Verhaard's 467 in twelve
years ([msg 10032](https://groups.io/g/eternity2/message/10032)). The author
was Joshua Blackwood, a complete outsider unknown to the list. Bucas
verified the board, added it to the viewer, and passed on the detail that
made veterans sit up: the author "says he can find a 468 every 4 days!"
([msg 10033](https://groups.io/g/eternity2/message/10033)). Three days later
Blackwood open-sourced the solver itself, carrying heuristic changes he
believed made it "twice as good" and pre-configured to hunt 469s
([msg 10037](https://groups.io/g/eternity2/message/10037)). His engineering
notes were as valuable as the code: SAT solvers, GPUs and pre-solved 2×2
caches had all been measured and discarded; the magic was the heuristic
schedule and a small set of allowed "break" depths
([msg 10056](https://groups.io/g/eternity2/message/10056),
[msg 10076](https://groups.io/g/eternity2/message/10076)).
The community did what it does with good code: it ran it. On 9 September
2020, Peter McGavin, the man who had solved the 10×10, posted: "Your
solver is fantastic!! I ran it for a few days on about a couple of hundred
cores and hit the jackpot. New record score of 469! Only 11 breaks!"
([msg 10045](https://groups.io/g/eternity2/message/10045)). Bucas rewrote
the solver in C, roughly doubling its speed, and a wave of further 469s
followed through November
([msg 10065](https://groups.io/g/eternity2/message/10065),
[msg 10067](https://groups.io/g/eternity2/message/10067)); the generator
behind the port was published as libblackwood
([msg 10078](https://groups.io/g/eternity2/message/10078)). Then, on
30 March 2021, a two-line message: the URL of a 470/480 board
([msg 10117](https://groups.io/g/eternity2/message/10117)). Blackwood later
confirmed the public repo was "the exact code used to find a 470"
([msg 10161](https://groups.io/g/eternity2/message/10161)), then retired: "I
have not written a line of code or run any algorithms since I found a 470"
([msg 10185](https://groups.io/g/eternity2/message/10185)).
One clarification matters for how these records are read. The contest's own
rules pinned only the starter piece, required at its specified location
and rotation; the four other clues were optional aids
([msg 11046](https://groups.io/g/eternity2/message/11046)). Every record
board from the 468 up sits in that same starter-only regime: piece 139 at
its mandatory square, none of the four optional clues at their official
positions, a claim verified at board level for the 470
([msg 10554](https://groups.io/g/eternity2/message/10554)). The 470 is the
same puzzle as the 469s long quoted as the ceiling, not an easier variant;
Blackwood had the clue tiles and "deliberately didn't use" them
([msg 10185](https://groups.io/g/eternity2/message/10185)). Boards that also
honour the four optional clues are a separate, stricter ladder (more on it
below). The [records page](/research/records) keeps the two regimes
explicitly apart.
Blackwood allowed himself one encore: reconfiguring his solver for zero
breaks, he pushed the consecutive-placement record from Verhaard's
long-standing 226 to 227, then 230, all on "a single PC in my basement"
([msg 10536](https://groups.io/g/eternity2/message/10536),
[msg 10544](https://groups.io/g/eternity2/message/10544),
[msg 10547](https://groups.io/g/eternity2/message/10547)).
## The modern era (2022–2026)
The post-wave years settled into a rhythm of invariants, calibration and
provenance. In May 2022 Al Hopfer, a regular since 2009, stated the
border condition now known on this site as the
[NS-1 balance](/research/why/border-balance): the outer rim
of the 14×14 interior "must have the same mixture (parity) of the internal
images on all the 56 border pieces"
([msg 10754](https://groups.io/g/eternity2/message/10754),
[msg 10757](https://groups.io/g/eternity2/message/10757)). In the same
thread, Carlos Fernandez was solving 14×14 quadrants in about four minutes
([msg 10802](https://groups.io/g/eternity2/message/10802)). Big interior
blocks had become routine; closing the frame had not. The 2008 prize score,
meanwhile, was formally demoted to a regression test: "last week I found
about 30 x 467. I use this target to calibrate the code I'm running",
wrote Bucas ([msg 11001](https://groups.io/g/eternity2/message/11001)).
The stricter ladder got its own record. A 2023 thread collecting bests that
respect **all five** clue placements climbed from 412 through 452 and 458
to Bruno Gauthier's 460 ([msg 11074](https://groups.io/g/eternity2/message/11074)),
which held for over three years until Benjamin Riotte's 464 in July 2026, a
461 rising to 464/480 with his own modified-Blackwood solver, with Igor Pejic
reaching the same 463–464 range independently
([groups.io](https://groups.io/g/eternity2/message/11919)). The
same year produced the era's canonical false positive: a complete,
480-matching 16×16 built by McGavin from pieces of one E2 set, one Clue-1
set and one Clue-2 set, a "480" that is not the puzzle, posted precisely
to make that point
([msg 11169](https://groups.io/g/eternity2/message/11169)).
Theory completed its long journey from posts to source code. In January
2024 McGavin restated the headline numbers: 14,702 expected solutions with
just the starter, 0.00000004 with all five hints, "very strongly
suggesting… a single, unique solution"
([msg 11193](https://groups.io/g/eternity2/message/11193)). The next
day he published complex_theory.c, his exact C implementation of Owen's
model ([msg 11197](https://groups.io/g/eternity2/message/11197)): the
reference this site's own
[complexity engine](/research/lab/experiments/joshua-blackwood/solver)
ports. In December, Bucas tied the record: "I restarted a few threads of
Joshua's code, and… another 470 appeared!"
([msg 11401](https://groups.io/g/eternity2/message/11401)), with Fernandez
posting a border-rearranged variation the same week
([msg 11403](https://groups.io/g/eternity2/message/11403)). Bucas's
crediting never wavered: "Joshua Blackwood is the current record holder. I humbly used his code"
([msg 11555](https://groups.io/g/eternity2/message/11555)). And his
wrapper_blackwood project put the point beyond opinion: a farm that varied
every parameter of Blackwood's algorithm and measured the results,
concluding that within the swept parameters the author's manual tuning was
already near-optimal; his notes
are published on this site with his explicit permission
([msg 11905](https://groups.io/g/eternity2/message/11905)).
And the founders came back, or didn't. In May 2025, Brendan Owen posted
for the first time in roughly fourteen years, answering a newcomer's
difficulty question ([msg 11500](https://groups.io/g/eternity2/message/11500)),
telling McGavin his years of work were "very impressive"
([msg 11521](https://groups.io/g/eternity2/message/11521)), and returning to
his own model with a refined calculation of edge-join probabilities
([msg 11546](https://groups.io/g/eternity2/message/11546)). The AI era
arrived first as farce: ChatGPT confidently informing a member that a
solution "was eventually found by a team of puzzle enthusiasts"
([msg 10993](https://groups.io/g/eternity2/message/10993)), then a
vibe-coded solver app whose owner conceded "it isn't 100% guarantee because
Base44 is AI", met with patient tutoring rather than ridicule
([msg 11755](https://groups.io/g/eternity2/message/11755),
[msg 11818](https://groups.io/g/eternity2/message/11818)). And in February
2026, the annual welcome note opened with a tribute to Kronjuvel: Ole
Knudsen, who created the group, ran it for the better part of two decades,
and has been missing from the list since 2023
([msg 11771](https://groups.io/g/eternity2/message/11771)).
The archive's current edge is 31 March 2026: a SAT thread still arguing
about 4 GB CNF encodings, a standing just-for-fun challenge to invalidate a
border partial, and one member's line for the anniversary: "slowly we are
progressing to 20 years of E2 next year ;-) Will we solve E2 in 2026...?"
([msg 11820](https://groups.io/g/eternity2/message/11820),
[msg 11823](https://groups.io/g/eternity2/message/11823)). The ledger after
nineteen years: the 480 has never been found; the 470 has stood since 2021;
and the ten missing edges are, as ever, [where the whole problem
lives](/research/records).
## Related
- [The hunt, a history (part I: 2000–2009)](https://eternity2.dev/research/community/hunt) — The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
---
# Contribute your research
> This wiki is the community's research home, and there is room in it for your work. Three ways to get it published, from a mailing-list post to a pull request, plus the small set of house rules that keep every page trustworthy.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/contribute/
- Updated: 2026-07-02
---
This wiki is not one person's blog with the door closed. It is meant to be the
community's research home: the place where nineteen years of Eternity II work
(records, methods, structural findings,
[dead ends](/research/build/dead-ends) included) gets written down, sourced,
and kept findable. Every page in [the lab](/research/lab) carries an author
byline and gathers on that researcher's [contributor
page](/research/people), today mostly
[Raphaël's](/research/people/raphael-anjou), because he is the one writing
things down, but the structure credits each contributor by name. Yours can sit
right beside his.
There is already a precedent. The page on
[Blackwood's algorithm](/research/lab/experiments/joshua-blackwood/solver) is built on Jef
Bucas's notes and parameter study, republished with his explicit permission,
his figures, and his name on every one of them. That is the model: the work
stays yours, the wiki gives it a permanent, citable home.
## Three ways in
**1. Post it and flag it: the lightest path.** Share your finding where the
community already talks: the [mailing list](https://groups.io/g/eternity2) or
the [Discord server](https://discord.gg/Ny5xs3q8w). If you would like it on the
wiki, just say so in the post. Someone (usually Raphaël) will write it up as a
page, run it past you, and publish it with full credit. That is exactly how
the Blackwood page came to exist from Jef's notes. You never have to touch the
repository.
**2. Open an issue or a pull request.** The wiki is
[an open repository](https://github.com/raphael-anjou/eternity2), and a
research page is nothing more than one MDX file under
`web/content/research/`. Adding the file *is* the registration: the sidebar,
search, and sitemap pick it up automatically. Each page starts with a small
frontmatter block:
```yaml
title: Your finding, in one line
description: >-
Two or three sentences that can stand alone in a search result.
kind: finding # or experiment, concept, reference...
updated: 2026-07-02
topics: [backtracking]
sources:
- label: Hopfer's statement of the NS-1 condition (2022)
url: https://groups.io/g/eternity2/message/10754
```
If the page is an **experiment** (a search run you measured), one more block is
required: the hardware it ran on. It is what lets a reader compare your result
to everyone else's, and the build refuses an experiment page without it.
```yaml
kind: experiment
hardware:
cores: 1 # logical cores used; sum across nodes for a cluster
cpu: "AMD Ryzen 9 7950X"
ramGb: 64
gpus: 0
accelerator: none # none | gpu | quantum | fpga | tpu
machine: "desktop, single box"
wallClock: "10 × 60 s" # budget per run × repeats
runs: 10
seedPolicy: "randomized corner permutation per run"
measured: true # true only for the standardized single-core bench
```
The page renders this as a spec card with one derived headline: **core-hours**
(cores × wall-clock hours), the true cost of the run. That single number is
what puts a one-core, one-minute search and a 400-core datacenter sweep in the
same table without one flattering the other. Set `measured: true` only for a
standardized single-core baseline (one core, a fixed time budget); a many-core
run is a `native run` and records its real kit and the best score it reached.
Not sure about the plumbing? Open an issue with your draft in plain markdown
and we will handle the rest. [Run it yourself](/research/build/run-it-yourself)
covers getting the site running locally if you want to preview your page.
**3. Share data and boards.** Not everything needs prose. A strong board, a
parameter sweep, a dataset from a long run: all of it is welcome. Every board
on this site is verifiable in the [viewer](/viewer), which speaks the
community's standard URL format natively, so anyone can check your result in
one click. Post the board or the data with a note on how it was produced, and
it can back a page or become one.
## The house rules
They exist for one reason: so that anyone landing on any page can trust what
they read.
- **Every claim links a source.** A mailing-list message, a repository, a
paper: something a reader can follow.
- **Every number is reproducible or labelled for what it is.** Deterministic
results come with a command that regenerates them exactly. Seeded,
stochastic, or heavy computations say so plainly ("won't reproduce exactly",
"~40 h on 8 cores") and ship the board they found so it can still be
verified.
- **Every experiment says what it ran on.** A measured search is meaningless
without its hardware, so an experiment page must carry the `hardware:` block
above (cores, kit, time budget). The build enforces it. A score with no
compute behind it is not a result anyone can compare.
- **French pages are written, not translated.** Each page exists in both
languages, and the French version is real French prose, not a machine pass.
If you only write one language, that is fine; a maintainer can arrange the
translation of the other.
- **Your name stays on your work.** Credit callouts, source links, figure
attributions: the Blackwood page shows what this looks like in practice.
Nothing gets absorbed anonymously.
## What the wiki gives back
In exchange for meeting those rules, your work gets infrastructure that a
forum post never has: a stable page in the docs shell with sidebar, search and
breadcrumbs; placement in the [topic hubs](/research) so it is found by
theme, not just by date; a raw-markdown export of every page (append `.md` to
its URL) so it stays machine-readable; and a permanent URL other researchers
can cite, years from now and not just this week.
The puzzle has resisted everyone so far. The least we can do is make sure
nobody's progress against it gets lost.
## Related
- [Run it yourself](https://eternity2.dev/research/build/run-it-yourself) — The whole site, the engine, and every result in this section run from one repository. Here is how to get it going, rebuild the WebAssembly engine, and reproduce the numbers.
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
---
# History: the big steps
> The Eternity II story at a glance, from the mailing list founded in 2000 to the 470 record that still stands. A scannable timeline of the turning points, each linking into the full two-part history and the message where it happened.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/history/
- Updated: 2026-07-11
- Source: The eternity2 mailing-list archive (groups.io): every milestone below links its original message — https://groups.io/g/eternity2
- Source: Wikipedia, "Eternity II puzzle" — https://en.wikipedia.org/wiki/Eternity_II_puzzle
---
Eternity II has an official story (a launch, a $2,000,000 prize, a deadline)
and a real one, which played out on a mailing list. This page is the real one,
compressed to its turning points: the moments that actually moved the puzzle,
each one a single line you can read at a glance.
For the whole account in prose, the history is told in two parts:
[part I](/research/community/hunt) covers 2000 to 2009, from the group's
founding to the first scrutiny prize; [part II](/research/community/hunt-part-2)
covers 2009 to today. Every milestone here links into them, and to the archive
message where it happened.
> **[Interactive: HistoryTimeline]** Rendered on the canonical page (link above); not shown in this markdown export.
## Where to go next
- The [record timeline](/research/records) plots the score itself over time,
the 467 of 2008, the long silence, the fast climb to 470, and the flat line
since 2021, with the full table and every board viewable edge by edge.
- The [full history](/research/community/hunt) tells the same story in prose,
with far more of the people, the arguments and the dead ends than fit on a
timeline.
- The [people behind it](/research/people) is the same story told by person:
who did what, and where to read it in their own words.
## Related
- [The hunt, a history (part I: 2000–2009)](https://eternity2.dev/research/community/hunt) — The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
- [The hunt, a history, part II: 2009–2026](https://eternity2.dev/research/community/hunt-part-2) — Seventeen years after the prize: the contest dies with its solution locked in a safe, 467 stands for a decade, the archive survives Yahoo's shutdown by days, and then an outsider from Reddit rewrites the record book. Every event sourced to its original message.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [Who's who of E2 research](https://eternity2.dev/research/people) — Two decades of Eternity II research were done by named people on a mailing list. This page is the gallery: who they are, what each of them contributed, and where to read it in their own words. A thank-you as much as an index.
---
# The lab
> The wiki's open notebook: structural findings and named search experiments, each credited to the researcher who ran it and reproducible from source. One corner of the community's wider research.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/
- Updated: 2026-07-10
---
This is the open notebook: original findings and experiments on Eternity II,
written up in full and credited page by page to the researcher who did the
work. It is not one person's blog. The
[experiments](/research/lab/experiments) are organised one section per
researcher, and each person's work also gathers on their own [contributor
page](/research/people). Today [Raphaël's](/research/people/raphael-anjou)
section is the full one, because he is the one writing things down, but the
structure is built for many hands and there is a section waiting for yours.
The community's methods, records and history live across the rest of the
research section.
The algorithms here are experiments, not breakthroughs. Some are original
ideas; others faithfully reimplement a known community technique to measure
exactly what it buys. Each write-up says which it is, what it attacked, where
it stopped, and what it left open, so anyone can pick up where it ended.
## Everything is reproducible
No result on this site is an unbacked claim. Deterministic computations ship
a runnable script and the exact output it produces. Searches that depend on
randomness or take hours ship the same script plus the board they found,
which you can load into the viewer and check edge by edge. The label on each
result tells you which kind it is.
## How the lab publishes
Every page here meets one shared editorial standard: what kind of contribution
it is, whether it publishes and at what tier, and how firmly its claim is backed.
[How the lab publishes](/research/lab/experiments/methodology) sets out that standard, so the
same rules apply no matter who writes the page.
## Add your own
The notebook is open. If you have a finding or a search worth writing down,
there is [a place for it here](/research/contribute), credited to you, sitting
beside the rest.
## Pages in this section
- [Experiments](https://eternity2.dev/research/lab/experiments) — The lab's named search experiments, one section per researcher. Each is a real run against Eternity II with its idea, its best board, and the questions it left open. Raphaël Anjou's notebook is here in full; the notebook is open to anyone else's.
---
# Experiments
> The lab's named search experiments, one section per researcher. Each is a real run against Eternity II with its idea, its best board, and the questions it left open. Raphaël Anjou's notebook is here in full; the notebook is open to anyone else's.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/
- Updated: 2026-07-13
---
Experiments are named search runs against Eternity II, each written up with
its idea, its best result, and the questions it left open. They are credited
page by page to the researcher who ran them, and gathered into a section per
researcher. This is one corner of the community's wider research; the methods,
records and history live across the rest of the [research
section](/research).
> **[Interactive: ExperimentAuthors]** Rendered on the canonical page (link above); not shown in this markdown export.
## What counts as an experiment here
Some are original ideas; some faithfully reimplement a known community
technique to measure exactly what it buys. Each write-up says which it is, what
it attacked, where it stopped, and what it left open, so anyone can pick up
where it ended. And every result is reproducible: deterministic runs ship a
script and its exact output; searches that depend on randomness ship the same
script plus the board they found, loadable in the [viewer](/viewer) and
checkable edge by edge.
Every experiment also states the **hardware it ran on**, in a fixed spec card:
the cores, the machine, the time budget, and one derived number, core-hours,
that says what the run actually cost. It is mandatory, because a score means
nothing without the compute behind it. Where a run is directly comparable it
wears a **standardized bench** badge (one core, a fixed minute budget, restarted
from randomized corners); a bigger run is a **native run** that records its real
kit and the best score it reached. That way a from-scratch minute on one core
and a datacenter sweep on four hundred cores sit in the same table, each judged
against what it actually spent.
## Pages in this section
- [How the lab publishes](https://eternity2.dev/research/lab/experiments/methodology) — The editorial standard for this open notebook: how a piece of Eternity II research goes from unpublished work to a published page. What kind of contribution it is, whether it publishes, at what tier, and where it lives. One shared standard, built to scale to many authors.
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Raphaël Anjou's experiments](https://eternity2.dev/research/lab/experiments/raphael-anjou) — A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- [Peter McGavin's engine](https://eternity2.dev/research/lab/experiments/peter-mcgavin) — Peter McGavin's own C backtracker, the fastest raw solver the community has measured. Fetched from the mailing list, built on an M1, and run on the real Eternity II. His code, his algorithm; run and written up here.
- [Joshua Blackwood's solver](https://eternity2.dev/research/lab/experiments/joshua-blackwood) — Joshua Blackwood's open-source C# solver, the one that found the standing 470 record. Built and run here as he published it, then with the five official clues pinned. His code, his algorithm; run and written up here.
- [Louis Verhaard's eii](https://eternity2.dev/research/lab/experiments/louis-verhaard) — Louis Verhaard's eii won the only prize the puzzle ever paid, but it shipped as a Windows binary with no source. What his method is, why the original will not run here, and where a faithful reimplementation of it lives.
---
# Joshua Blackwood's solver
> Joshua Blackwood's open-source C# solver, the one that found the standing 470 record. Built and run here as he published it, then with the five official clues pinned. His code, his algorithm; run and written up here.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/joshua-blackwood/
- Updated: 2026-07-15
---
This section is [Joshua Blackwood's](/research/people/joshua-blackwood) own
solver: the open-source C# program that found the 470 that still stands. The
algorithm and the code are his, public on GitHub. What is added here is the
running: built as he published it, measured single core, and then tested against
the real five-clue puzzle to see what its speed is actually buying.
## Pages in this section
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
---
# Blackwood's solver, decoded and run here
> Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver/
- Updated: 2026-07-17
- Topics: backtracking, speed
- Reproduce: `git clone github.com/jblackwood345/EternityII_Solver; dotnet build -c Release`
- Source: Joshua Blackwood's solver (GitHub, GPL-3.0, the original source) — https://github.com/jblackwood345/EternityII_Solver
- Source: Jef Bucas's notes on the Blackwood solver (wrapper_blackwood repo) — https://github.com/jfbucas/wrapper_blackwood/blob/main/doc/Notes/Notes.md
- Source: Jef's study announcement and permission to republish (groups.io message 11905) — https://groups.io/g/eternity2/message/11905
- Source: Blackwood's own tuning plan: the four levers and the 470 retune (groups.io message 10076) — https://groups.io/g/eternity2/message/10076
- Source: The repo re-published, "the exact code used to find a 470" (groups.io message 10161) — https://groups.io/g/eternity2/message/10161
- Source: Blackwood's 470 announcement (msg 10117) — https://groups.io/g/eternity2/message/10117
---
> **Whose work this is**
>
> The algorithm and the C# are **Joshua Blackwood's** ([EternityII_Solver](https://github.com/jblackwood345/EternityII_Solver), public on GitHub under the GPL-3.0). The decoding of the internals, along with the parameter study reported below, is Jef Bucas's, in his [wrapper_blackwood](https://github.com/jfbucas/wrapper_blackwood) project; the decoded sections rephrase and extend [his notes](https://github.com/jfbucas/wrapper_blackwood/blob/main/doc/Notes/Notes.md) with his explicit permission ([groups.io message 11905](https://groups.io/g/eternity2/message/11905)), and the figures are his, reproduced from the same notes. The final section is [Raphaël Anjou](/research/people/raphael-anjou) building and running Blackwood's code on one machine and writing up what it did; the three small edits to his source are listed there, each one his own README invites.
Joshua Blackwood holds the standing 470 record, found with code he made public.
His solver is, at heart, a classic depth-first backtracker running the same loop
as any other: place, check, back up. What makes it the engine behind the
community's best boards is a set of manually crafted heuristics layered on top: a
scoring rule that decides *which* pieces to try first, a depth-by-depth quota
schedule that prunes branches falling behind, a
[fill order](/research/build/backtracking/fill-order) tuned to be "not too tight
and not too loose", and a late-game allowance for accumulating mismatches. Every
one of those choices has numbers in it, and Jef Bucas's study asked the obvious
question: are Blackwood's numbers any good? They are, uncomfortably so. Because
the code is a real, buildable C# program, it can then be run exactly as its
author wrote it, which is what the last section here does.
## Score pieces by three colours
The solver scores every piece by the colours on its sides, and it privileges
exactly three of them, one border colour and two interior colours:
```csharp
heuristic_sides = new List() { 13, 16, 10 };
```
Pieces carrying those colours are sorted to the front of every candidate list, so
the search commits them early. Why these three? That is one of the two knobs
Jef's study turned (see below). His sampling suggests the best known scores come
from two distinct pattern sets:
> **[Figure]** interactive figure. Rendered on the canonical page (link above); not shown in this markdown export.
## The schedule: a quota curve over 256 depths
Prioritizing three colours only helps if the search is *forced* to actually place
them. So the solver carries a 256-entry array, one entry per depth, where each
value is the minimum number of those three colours that must already be on the
board to keep descending. Fall below the quota and the branch is cut on the spot:
backtrack, no discussion.
Blackwood filled that array by hand, as a piecewise-linear ramp:
```csharp
heuristic_array = new int[256];
for (int i = 0; i < 256; i++) {
if (i <= 16)
heuristic_array[i] = 0;
else if (i <= 26)
heuristic_array[i] = (int)(((float)i - 16) * (float)2.8);
else if (i <= 56)
heuristic_array[i] = (int)((((float)i - 26) * (float)1.43333) + 28);
else if (i <= 76)
heuristic_array[i] = (int)(((((float)i - 56) * (float)0.9)) + 71);
else if (i <= 102)
heuristic_array[i] = (int)(((((float)i - 76) * (float)0.6538)) + 89);
else if (i <= 160)
heuristic_array[i] = (int)(((((float)i - 102) / 4.4615)) + 106);
}
```
Free until depth 16, steep through the 20s, then flattening out to depth 160.
This is the "schedule" in schedule-and-break-index: a pre-committed timetable for
burning down the high-frequency colours, enforced as a hard prune.
> **[Figure]** interactive figure. Rendered on the canonical page (link above); not shown in this markdown export.
## The fill order: not too tight, not too loose
The search starts at the bottom-left of the board (the corner nearest the
mandatory centre piece) and runs a standard row scan all the way to depth 180.
After that, it intersperses the remaining border pieces (including the third
corner) every few steps among the interior placements.
That is a deliberate middle course. Spiralling in (border ring first) commits the
most constrained pieces too early; a pure row scan leaves them all to the end,
where they ambush you. Blackwood's order releases the border pressure
progressively. Why the order matters this much, and how to score one before
running it, is exactly what [complex theory](/research/why/complex-theory)
formalizes.
> **[Figure]** interactive figure. Rendered on the canonical page (link above); not shown in this markdown export.
## Breaks: buying the endgame with mismatches
Approaching the end of the search, the solver stops demanding perfection. A
budget of edge mismatches unlocks with depth, cumulatively: one break is allowed
from depth 201 onward (it can be spent at 201 or any later depth), a second from
206, and so on:
```csharp
break_indexes_allowed = new List() { 201, 206, 211, 216, 221, 225, 229, 233, 237, 239 };
```
Ten breaks in total, the last unlocking at depth 239. This is what makes the
[record boards](/research/records) reachable at all: a perfect 256 has never been
found, but a board that tolerates a handful of late mismatches is something a
backtracker can actually finish.
Two details sharpen the picture. First, the budget comes with a discipline: no
two breaks are ever allowed to touch: each mismatch must sit isolated among
matched edges. Blackwood spells out the consequence himself: any run that reaches
255 placements automatically completes to 256, and a 469 board is equivalently a
249-piece partial with seven holes
([groups.io message 10051](https://groups.io/g/eternity2/message/10051)). Second,
the ten-entry list above is itself a retune: planning the push from 469 to 470,
Blackwood cut the break budget from eleven depths to ten, a change he documented,
lever by lever, in his own tuning plan
([groups.io message 10076](https://groups.io/g/eternity2/message/10076)). The
shifted unlock points he sketched in that plan were never adopted (the published
470 code keeps the 469-era points and simply drops the eleventh), and that
tightened schedule is what found the 470.
## Restarts, randomness, and a 50-billion-node cap
A row-scan backtracker rarely climbs back up to its first rows, so a bad opening
can strand an entire run in an impossible region. Blackwood's answer is cheap and
effective: randomize the opening (the first corner and the bottom-row pieces are
shuffled on every attempt, so no two runs retrace the same prefix) and cap each
attempt at 50 billion nodes explored. Hit the cap, abandon,
[restart](/research/build/backtracking/restarts) with a fresh opening, and go
looking for a more fertile one.
## Make it fit in cache
The last ingredient is mechanical sympathy. Candidate pieces live in per-position
lookup tables, and each entry is a compact struct (piece number, rotation, the
two exposed sides, a break count and a heuristic count) sized so the working set
fits in CPU cache:
```csharp
public struct RotatedPiece
{
public ushort PieceNumber { get; set; }
public byte Rotations { get; set; }
public byte TopSide { get; set; }
public byte RightSide { get; set; }
public byte Break_Count { get; set; }
public byte Heuristic_Side_Count { get; set; }
}
```
This is the throughput half of the story: the same design that
[Peter McGavin](/research/lab/experiments/peter-mcgavin/backtracker) later pushed
to hundreds of millions of placements per second.
## From one 468 to a wave of records
The solver's history is as instructive as its internals. Blackwood arrived a
complete outsider: his 468, the first advance past
[Louis Verhaard's 467](/research/lab/experiments/louis-verhaard/eii) in twelve
years, reached the mailing list second-hand, relayed from a Reddit post
([groups.io message 10032](https://groups.io/g/eternity2/message/10032)). Three
days later he open-sourced the solver, with heuristics he reckoned made it "twice
as good" as the 468 run
([message 10037](https://groups.io/g/eternity2/message/10037)); Jef Bucas had it
running under Mono on Ubuntu almost immediately
([message 10038](https://groups.io/g/eternity2/message/10038)). Within a week,
Peter McGavin, running it "on about a couple of hundred cores," hit 469
([message 10045](https://groups.io/g/eternity2/message/10045)). Bucas then
rewrote the algorithm in C for roughly double the speed
([message 10065](https://groups.io/g/eternity2/message/10065)), a wave of fresh
469 boards followed within weeks, and the code generator behind the rewrite was
published as [libblackwood](https://github.com/jfbucas/libblackwood)
([message 10078](https://groups.io/g/eternity2/message/10078)). When Blackwood
posted his 470 in March 2021, it came from that same public code. His own words,
on re-publishing the repo after it had quietly gone private: "It is the exact code
used to find a 470"
([message 10161](https://groups.io/g/eternity2/message/10161)). One record,
open-sourced, became a community record machine.
## Jef's parameter study: wrapper_blackwood
All of the above is *how* the solver works. Jef Bucas built
[wrapper_blackwood](https://github.com/jfbucas/wrapper_blackwood) to ask whether
its numbers are *right*. The setup is a small
[distributed experiment](/research/build/faster/distributed-solving): a Python
server hands out jobs over HTTP, where each job is a variation of the solver's
parameters. Clients (one worker per core) fetch a job, generate the C# source
from templates with that variation baked in, compile it with Mono, run it, and
report the result back to the server for analysis.
Two parameters got the treatment:
- **The three prioritized colours.** Sampling many different sets of three
patterns, and recording how deep the algorithm got with each, his
[results page](https://github.com/jfbucas/wrapper_blackwood/blob/main/doc/batch00_edge_combos_stats.html)
maps every combination of one border colour and two interior colours to the
depth it reached, together with a heatmap of where in the board the search
spent its time. The best known scores concentrate on two distinct pattern sets.
- **The quota schedule.** Running many random variations of the 256-entry
heuristic array and plotting how far each one got (greener is better),
Blackwood's manually crafted curve lands right in the middle of the green
region.
> **[Figure]** interactive figure. Rendered on the canonical page (link above); not shown in this markdown export.
## What hand-tuning got right
That last result deserves emphasis. Blackwood filled his schedule by hand: five
linear segments, eyeballed coefficients. When a randomized sweep explored the
neighbourhood around it, the hand-tuned line sat squarely in the best-performing
zone. Jef's conclusion, and ours: the original parameters were near-optimal. The
decades-old advice holds. Before you redesign a record solver's heuristics, check
whether its author already found the local optimum by hand.
> **Blackwood's own negative results**
>
> Blackwood ran the same audit on himself. After the 469 he catalogued his dead ends on the mailing list ([groups.io message 10056](https://groups.io/g/eternity2/message/10056)): eliminating four colours early instead of three (no gain), saving colours for the endgame (worse), SAT solvers such as kissat, cryptominisat and Google OR-tools (poor), GPU acceleration, and caching all pre-solved 2×2 blocks (measured, then dropped). The only thing that ever paid was refining the heuristics themselves, worth another ~2x. It is the same conclusion as Jef's sweep, reached from the other direction: the schedule is the magic, not the raw technology.
## Replicate the study
One caveat on the parameter study, and it is Jef's own: the sample counts behind
these results are low, and he is not 100% sure of the conclusions. His notes say
it plainly: it would be good to attempt to reproduce these results, to validate
or invalidate them. He offers his samples, and the
[wrapper_blackwood](https://github.com/jfbucas/wrapper_blackwood) harness is
public: point a few machines at the server, re-run the sweeps, and post what you
find on the [mailing list](https://groups.io/g/eternity2). An independent
replication would confirm or refute the findings, and either outcome would be a
real contribution.
## Running it here
The study above audits Blackwood's parameters; this last section audits something
different: what his published program does when you build it yourself and hold it
to the real puzzle, single core, on my machine.
### Where the code came from
The source is his public repository,
[github.com/jblackwood345/EternityII_Solver](https://github.com/jblackwood345/EternityII_Solver),
under the GPL-3.0. It builds unmodified on .NET 8. It is not copied here: the
licence and good manners both say to link it, not vendor it, so it was cloned to
a scratch directory, run, and only the numbers kept.
### What was changed to run it
His program is written to run on one machine's full core count and to save only
near-complete boards. To measure it single core, and to see anything short of a
full solve, three edits were made, each of which his own README invites ("change
the number of cores", "change the save function"):
- **One search thread.** His `number_virtual_cores` constant defaults to 64. His
`Parallel.For(1, N)` runs `N-1` workers, so setting it to 2 gives exactly one
search thread.
- **A progress line.** Unmodified, in a bounded run it prints only "Solving...".
A single added line reports the deepest placement reached, so a run that does
not finish still yields a number.
- **A lower save threshold, and clue pinning** for the constrained run below.
### Left as published: fast, and aimed at the one clue that binds
Run as his code stands, single core for 60 seconds, it is fast and it gets far:
- **248 / 256 pieces placed**, a board that re-scores to **454 / 480** matched
edges.
- **~18.9 million nodes per second** (one node is one placement attempt).
His published program has no notion of clues at all: it treats every piece,
including the special ones, as ordinary and maximises raw matched edges. This is
less of a compromise than it first looks, because a legal Eternity II board has
exactly **one** binding constraint, the mandatory centre clue (piece 139 in its
centre cell, in its given orientation); the puzzle's other four "clues" were bonus
hints the maker released, not requirements. His celebrated **470** board is the
proof: in the [viewer](/viewer) it verifies as clues respected **1 / 5**, the one
respected being the centre, and one is all a legal board needs. Maximising edges
without clue logic is not a shortcut around the puzzle; it is a legitimate way to
attack the one that binds.
The last eight cells of that 454 board are a genuine dead end: no arrangement of
the eight leftover pieces completes them edge-perfectly. (For fun, handing that
board to [our ALNS](/research/build/local-search/local-search-alns) for 30
seconds filled all eight and pushed it to 462 by rearranging the region.)
### With the clues pinned: it stalls
Pinning the clue pieces into their cells forces the engine to respect them
instead of placing them wherever they fit. The same code, all five official
clues pinned, 120 seconds, single core:
- **Deepest placement: about 45 / 256.**
His scan-order heuristic has nothing to grip once the special pieces are fixed
mid-board: it thrashes near the bottom rows and never climbs. Pinning all five is
stricter than the puzzle strictly requires (only the centre clue is mandatory),
but it is the same constraint the other engines are held to here, and Blackwood
lands far below the
[Verhaard reimplementation's 438](/research/lab/experiments/louis-verhaard/verhaard-reimpl)
or [McGavin's 204](/research/lab/experiments/peter-mcgavin/backtracker) under the
same constraint.
> **A fair caveat**
>
> This clue pin is crude: it forbids every other piece at the five clue cells and reserves those pieces. A clue-aware redesign would search outward from the clues instead, and would do better. So 45 is a floor for "his published code with clues bolted on", not a verdict on the approach. His record 470 was found with this code plus a great deal of compute and luck, not in two minutes.
## Related
- [McGavin's C backtracker: the throughput story, built here](https://eternity2.dev/research/lab/experiments/peter-mcgavin/backtracker) — Peter McGavin's own C backtracker, the community's fastest: a 2007 optimization recipe compounded for two decades through generated code, lookup tables and counter tricks, then built on my M1 and pointed at the real 256-piece puzzle, where single core it drives past 200 of 256 pieces at ~109M placements/s.
- [Verhaard's eii: the solver that won the only prize](https://eternity2.dev/research/lab/experiments/louis-verhaard/eii) — The engine behind the 467, the only Eternity II score ever paid, read from Louis Verhaard's own mailing-list posts: forward pruning, comb-search fill orders, depth-gated edge slipping, a Markov-tuned slip schedule. And why his own source-less Win32 binary cannot be built or run on this machine at all.
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
---
# Louis Verhaard's eii
> Louis Verhaard's eii won the only prize the puzzle ever paid, but it shipped as a Windows binary with no source. What his method is, why the original will not run here, and where a faithful reimplementation of it lives.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/louis-verhaard/
- Updated: 2026-07-15
---
This section is about [Louis Verhaard's](/research/people/louis-verhaard) eii,
the solver that found the 467 and won the puzzle's only prize. Unlike McGavin
and Blackwood, he never released source, so his own program cannot be run here.
This page records his method and that gap; the runnable engine is a
[reimplementation](/research/lab/experiments/louis-verhaard/verhaard-reimpl),
which lives here beside the method it rebuilds, bylined to Raphaël Anjou because
the code is his even though the idea is Verhaard's.
## Pages in this section
- [Verhaard's eii: the solver that won the only prize](https://eternity2.dev/research/lab/experiments/louis-verhaard/eii) — The engine behind the 467, the only Eternity II score ever paid, read from Louis Verhaard's own mailing-list posts: forward pruning, comb-search fill orders, depth-gated edge slipping, a Markov-tuned slip schedule. And why his own source-less Win32 binary cannot be built or run on this machine at all.
- [Verhaard reimplementation](https://eternity2.dev/research/lab/experiments/louis-verhaard/verhaard-reimpl) — A from-scratch reimplementation of Louis Verhaard's eii method, since his own binary ships no source and will not run here. Set-composition swap-annealing under the 2×2-tiling metric; on the real five-clue puzzle it reaches 438 of 480, single core.
---
# Verhaard's eii: the solver that won the only prize
> The engine behind the 467, the only Eternity II score ever paid, read from Louis Verhaard's own mailing-list posts: forward pruning, comb-search fill orders, depth-gated edge slipping, a Markov-tuned slip schedule. And why his own source-less Win32 binary cannot be built or run on this machine at all.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/louis-verhaard/eii/
- Updated: 2026-07-17
- Topics: backtracking, local-search, speed
- Source: Louis Verhaard's eii solver details (the original, Win32 only) — https://www.shortestpath.se/eii/eii_details.html
- Source: The release: “I am stuck… my only hope is brute force” (groups.io message 5940) — https://groups.io/g/eternity2/message/5940
- Source: Decoder and internals documentation released; 467 found 40+ times (groups.io message 6275) — https://groups.io/g/eternity2/message/6275
- Source: The method disclosed: depth-gated edge slipping, clean score 247 (groups.io message 7321) — https://groups.io/g/eternity2/message/7321
- Source: JSA's independent benchmark: a 467 in 82 days with the public binary (groups.io message 6687) — https://groups.io/g/eternity2/message/6687
- Source: The solver's long-term home: shortestpath.se/eii (groups.io message 7439) — https://groups.io/g/eternity2/message/7439
- Source: Verhaard on his method (groups.io msg 6891) — https://groups.io/g/eternity2/message/6891
---
> **Whose work this is**
>
> The solver, called **eii**, and its documentation are **Louis Verhaard's**, hosted on his own site ([shortestpath.se/eii](http://www.shortestpath.se/eii/)). The winning entry was submitted under the name of his wife, Anna Karlsson; Verhaard's own words settle the credit: "My wife submitted my best solution last year and won 10 000 dollars" ([groups.io message 7451](https://groups.io/g/eternity2/message/7451)). This page, written up by [Raphaël Anjou](/research/people/raphael-anjou), reconstructs the machine from Verhaard's mailing-list posts of 2008–2010 and from JSA's public benchmark of the released binary, and records why the original binary cannot be run here at all. The runnable engine is a separate [reimplementation](/research/lab/experiments/louis-verhaard/verhaard-reimpl), filed here in this section beside the method, bylined to Raphaël because that code is his, not Verhaard's.
Louis Verhaard's eii is the solver behind the 467/480 that won the $10,000
runner-up prize at the first scrutiny date, the only money the Eternity II
contest ever paid out. The same score then stood as the record for twelve years,
until Joshua Blackwood's 468 in 2020. Like
[Blackwood's engine](/research/lab/experiments/joshua-blackwood/solver), it is at
heart a depth-first backtracker with hand-tuned heuristics layered on top. Unlike
Blackwood's, its internals were never open-sourced as code; what we have instead
is unusually good testimony: Verhaard documented the heuristics and search order
himself, discussed the design on the mailing list in first person, and shipped a
public binary that an independent user benchmarked all the way back to the prize
score. That missing source is why, of the three community engines studied in this
lab, his is the one that cannot be run at all; the last section here says exactly
why.
## The arc: stuck past 463, so release it
Verhaard's own origin story is disarming. His first high-score program only
placed perfectly matching pieces, and stalled around a 450: hundreds of
248-piece clean partials, no way forward. Only when he tried Bob Cousins'
(originally Dave Clark's) solver, which found a 458 within a minute, did he grasp
that the contest scored *matching edges*, so half-fitting placements count
([groups.io message 5767](https://groups.io/g/eternity2/message/5767)). He had, in
his own telling, never bothered to read the rules; Jef Bucas was still pointing
readers at that admission in Verhaard's documentation in 2021
([groups.io message 10582](https://groups.io/g/eternity2/message/10582)).
Through the summer of 2008 the community's publicly visible ceiling was 463
([groups.io message 5688](https://groups.io/g/eternity2/message/5688)), and in a
remarkable exchange Verhaard and Max compared notes as the only two known to be
past what they called the "don't-talk-about-limit"
([messages 5767–5787](https://groups.io/g/eternity2/message/5767)). Then, on 22
September 2008, with the first scrutiny date three months away, Verhaard
published the solver for anyone to run at fingerboys.se: "This because I am stuck
and my only hope to improve my best score is by using brute force"
([message 5940](https://groups.io/g/eternity2/message/5940)). The terms mirrored
eternity2.net's: you needed to own the real puzzle, and any prize would be split
50-50 between the best-scoring user and Verhaard. The community became his compute
farm.
It worked. On 6 January 2009 he released a decoder for the solver's `.eii` output
files, documented the internals (heuristics and search order), and disclosed the
number the farm had reached: 467, found more than 40 times by different users
([message 6275](https://groups.io/g/eternity2/message/6275)). Nine days later,
Tomy's announcement surfaced: no complete solution, and a $10,000 runner-up prize
to Anna Karlsson of Lund for 467 of 480
([message 6337](https://groups.io/g/eternity2/message/6337)). The list took about
an hour to decode "Lund + 467". The rest of the contest story belongs to
[the history page](/research/community/hunt): the silence from Tomy, the press
coverage, the aftermath. What follows here is the machine.
## Forward pruning against static thresholds
The base loop is a depth-first backtracker over a fixed
[fill order](/research/build/backtracking/fill-order), but one that prunes
forward: branches whose heuristic score falls below a threshold are cut before
they are explored. Verhaard described his thresholds as static and hand-tuned:
"mostly based on pure guesswork and only on limited amount on theory or
measurements," slow to tune but predictable over long runs
([groups.io message 5771](https://groups.io/g/eternity2/message/5771), quoted back
at him in [message 5772](https://groups.io/g/eternity2/message/5772)).
What are the thresholds *for*? The design he and Max converged on in that exchange
is the interesting part: you can only influence piece selection early in the
search, but what you want to maximize is the **tileability of the remaining
pieces** deep in the fill, around pieces 160–200, where the branching factor
collapses toward forced moves. So the early thresholds are tuned, semi-manually,
to answer: which heuristic score must early partials reach so that the survivors
carry a leftover piece set that still tiles well? Verhaard's verdict on Max's
description: "I think we work in a very similar way after all"
([message 5780](https://groups.io/g/eternity2/message/5780)).
## The dashboard: where the node distribution peaks
How do you know a heuristic is strong? The measure the two settled on (proposed
by Max, adopted by Verhaard) is where the peak of the node-depth distribution
sits. An unheuristic full search on E2 spends most of its time around depth 161,
Brendan Owen's baseline figure
([message 6112](https://groups.io/g/eternity2/message/6112)); both of their
heuristic searches had pushed the peak to just under 170
([message 5780](https://groups.io/g/eternity2/message/5780)). Max supplied the
interpretation: a peak at 170 is roughly equivalent to eliminating one interior
colour from the puzzle entirely; a "killer heuristic" that eliminates two seemed
out of reach ([message 5787](https://groups.io/g/eternity2/message/5787)).
## Comb search: a fill order for high scores
A full-solution search wants a row-by-row scan; a high-score search lives deeper
in the board, and wants a different frontier. When Brendan Owen posed exactly
that question, Verhaard revealed the shape of his answer: the best orders he had
found resemble a **comb search** (most rows searched horizontally, then the
remaining rows searched vertically), with the tooth length tied to the target:
"The lower the score you aim for, the longer the teeth of the comb become"
([groups.io message 6112](https://groups.io/g/eternity2/message/6112)). Max had
independently converged on nearly the same geometry (twelve rows of scanline,
then column scan) and reported his scores ran about one edge below "the results
that Louis solver accomplishes"
([message 6126](https://groups.io/g/eternity2/message/6126)).
## Edge slipping, gated by depth
The lever that actually bought the 467 is deliberate imperfection. Asked about it
directly a year later, Verhaard was precise: the 467 program searches "normally,"
but at certain depths it allows **edge slipping**: placing a piece with one
mismatching edge against an already-placed neighbour. It does not first build a
clean partial and then patch holes; the mismatches are budgeted *into the
descent*, unlocked at chosen depths. And the anatomy of the result is telling:
most of the 467 boards he examined had a clean score of only 247, with thirteen
slipped edges spent where the schedule allowed them
([groups.io message 7321](https://groups.io/g/eternity2/message/7321)). The 467
was no fluke, either: he found it more than 50 times
([same message](https://groups.io/g/eternity2/message/7321)).
The technique itself has [its own page](/research/build/reduce/edge-slipping): why
scheduled mismatches reach boards a clean search never could, and the counting
theory behind the cost of each extra matched edge. What belongs here is the
solver-side machinery: the per-depth mismatch budget is a **slip array**, one
entry per depth, and it is a tuned object, not a guess.
## The slip array's optimizer: a Markov chain
In January 2009, in the thread where Owen was extending
[complex theory](/research/why/complex-theory) to cover slips, Verhaard posted
the outline (with Java snippets) of the algorithm he used to optimize eii's
search order and slip array
([groups.io message 6423](https://groups.io/g/eternity2/message/6423)). The input
is a candidate search order plus a slip array. For every depth he estimates two
numbers from experimental runs (theory would do to start, he noted): the
probability that a random remaining piece fits perfectly, and the probability
that it fits with one slipped edge. From these he builds a Markov chain whose
state is *(depth, slipped edges so far)*, with transitions for a clean placement
and, where the slip array permits, for a slipped one. Running the chain end to
end yields the probability of reaching the bottom and the expected node count: a
cheap, closed-loop evaluator for any (order, slip-schedule) pair, memoized for
efficiency. He flagged its limit himself: the simple model ignores slip parity,
so it gets unreliable at very high target scores.
The family resemblance to what came a decade later is hard to miss: Blackwood's
break indexes are also a depth-gated mismatch budget, and his quota schedule is
also a pre-committed, per-depth curve, hand-tuned rather than chain-optimized.
The lineage runs through this solver.
## The clean-score sibling
The same heuristics powered a second program with a different late-game move:
instead of slipping an edge, it may *skip a square*, that is, leave a cell empty
and keep going. The name is his own
([message 7321](https://groups.io/g/eternity2/message/7321)). That is the
clean-score (no mismatches) hunter. With it, Verhaard filled 14 complete rows plus
two pieces, a 226-piece flawless partial and the record he knew of at the time
([groups.io message 6303](https://groups.io/g/eternity2/message/6303)). In a
one-day return in December 2009, after estimating ~2,000 248s per 249, he landed
three 249s in a row and stopped, putting a 250 at roughly 4,000 times harder
([message 7306](https://groups.io/g/eternity2/message/7306)). Asked about it a
decade later, he confirmed the 249 on his site is real: about a week of compute on
one machine ([message 9890](https://groups.io/g/eternity2/message/9890)).
## The public reproduction: 82 days to a 467
Because the binary was public, the 467 is the rare contest-era record with an
independent, quantified reproduction. JSA ran eii continuously on one PC and
logged the score distribution as it accumulated:
- **43 days:** 2,008,484 463s · 109,195 464s · 6,048 465s · 250 466s · nothing
higher ([groups.io message 6571](https://groups.io/g/eternity2/message/6571))
- **62 days:** 427 466s, arriving at roughly 5–6 per day, and still no 467
([message 6653](https://groups.io/g/eternity2/message/6653))
- **82 days:** two 467s, alongside 4,017,182 463s · 227,245 464s · 13,637 465s ·
625 466s ([message 6687](https://groups.io/g/eternity2/message/6687))
That last log is the cleanest public measurement of the exponential rarity ladder
near the top: four million 463s for every couple of 467s, and a 466→467 step that
took a single machine nearly three months. JSA's sign-off, "Congratulations to
Louis on a well-thought-out algorithm", doubles as the verification verdict.
## What it taught the community
Beyond the score, eii set a precedent: when you are stuck, release the solver and
let the community's machines hunt, with the prize split as the contract. The 467
was found by *users* of a published binary, 40+ times, before it won anything.
Twelve years later
[Joshua Blackwood repeated the pattern](/research/lab/experiments/joshua-blackwood/solver),
posting a 468 and open-sourcing the engine days later, and got the same reward: a
wave of community records on his own algorithm. The other lesson is methodological
and runs through this whole page: Verhaard tuned his solver against *models*
rather than raw vibes (the node-peak dashboard, the Markov-chain evaluator), at a
time when that discipline was rare.
## Where the code lives, and why it won't run here
The original home, fingerboys.se, was his band's website; when the band gave up
maintaining the site, the solver went offline with it for months. In January 2010
Verhaard republished it, unchanged, at
[shortestpath.se/eii](http://www.shortestpath.se/eii/)
([groups.io message 7439](https://groups.io/g/eternity2/message/7439)); that
address is still its home, and the first-party source behind the 467 row in
[the records table](/research/records). One operational footnote from the Q&A that
followed: the solver keeps no memory of past positions, so its thousands of
463–465 boards are not deduplicated
([message 7451](https://groups.io/g/eternity2/message/7451)).
What "the code lives there" hides is that no *source* lives anywhere. What
Verhaard shipped is `eii-1.0-win32.zip`: a Windows `eii.exe`, a readme, and some
`.bat` files. There are **zero source files**, and it is Win32 only. It does not
run on Apple silicon, and there is no Windows or emulation layer on this machine to
run it under. His artifact is, here, unrunnable, and there is nothing of his to
fetch and build. The comb-search fill orders and depth-gated slip schedule
described above became community canon precisely because they were recovered from
his posts and, where numeric, byte-exact from the strings inside `eii.exe`, since
that binary is the only surviving record of them.
So "running Verhaard" at all means rebuilding his method from that documentation.
That is what the
[Verhaard reimplementation](/research/lab/experiments/louis-verhaard/verhaard-reimpl)
does, and on the real five-clue puzzle it reaches 438 of 480, single core. That
engine is filed here beside this page, bylined to Raphaël, because it is a reading
of Verhaard's method rewritten as Raphaël's code. Naming that distinction is part of the
finding: of the three community engines studied here, his is the one that cannot
be run at all.
## Related
- [Verhaard reimplementation](https://eternity2.dev/research/lab/experiments/louis-verhaard/verhaard-reimpl) — A from-scratch reimplementation of Louis Verhaard's eii method, since his own binary ships no source and will not run here. Set-composition swap-annealing under the 2×2-tiling metric; on the real five-clue puzzle it reaches 438 of 480, single core.
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
- [Edge slipping](https://eternity2.dev/research/build/reduce/edge-slipping) — Louis Verhaard's prize-winning move: let the backtracker place a mismatched piece, but only at chosen depths near the end of the board. Each allowed slip costs one point of score and multiplies the number of target boards astronomically. That is why his 467 was found more than fifty times, and the direct ancestor of Blackwood's breaks.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
---
# Verhaard reimplementation
> A from-scratch reimplementation of Louis Verhaard's eii method, since his own binary ships no source and will not run here. Set-composition swap-annealing under the 2×2-tiling metric; on the real five-clue puzzle it reaches 438 of 480, single core.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/louis-verhaard/verhaard-reimpl/
- Updated: 2026-07-15
- Topics: speed, local-search
- Reproduce: `just experiments single-core-benchmark`
- Source: Runnable engine + committed results + scripts (this experiment's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark
- Source: Louis Verhaard's eii solver details (the original, Win32 only) — https://www.shortestpath.se/eii/eii_details.html
---
> **Whose work this is**
>
> The **method** is [Louis Verhaard's](/research/lab/experiments/louis-verhaard/eii). The **code** here is a from-scratch reimplementation by [Raphaël Anjou](/research/people/raphael-anjou). It is filed here, in Verhaard's section beside the method it rebuilds, as a reconstruction: a reading of his method, not his program. The byline stays Raphaël's because the code is his; the section is Verhaard's because the idea is.
Verhaard's own eii is a source-less Windows binary that
[will not run here](/research/lab/experiments/louis-verhaard/eii). The only way
to run his approach is to rebuild it, and this is that rebuild: the same engine
that appears as `verhaard` on the
[grid](/research/lab/experiments/single-core-benchmark).
## What it reimplements
Verhaard's documented method: **set-composition swap-annealing** under the
2×2-tiling metric. Pick a roughly 180-piece interior subset, swap-anneal its
composition until the number of achievable 2×2 sub-tilings is locally maximised,
front-load the worst performers, then search the rest under that scaffold. The
engine's numeric constants were recovered byte-exact from the strings inside
`eii.exe`, because that binary is the only surviving record of them.
Calling this "running Verhaard" is a stretch worth naming: it is a reading of his
method, not his code. That it is the only way to run his approach at all is part
of the finding.
## What it does on the real five-clue puzzle
Single core, 120 seconds, all five official clues pinned:
- **438 / 480** matched edges, a **fully placed** board (256 of 256 pieces).
- All five clues respected (a strict clue solution), 42 broken edges, at roughly
40 million nodes a second.
The run and its output board are committed in the engine's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark),
and `just experiments single-core-benchmark` reruns it, so the 438 is checkable
rather than asserted. This is the five-clue instance, a harder puzzle than the
corner-pinned variants the [leaderboard](/research/lab/experiments/single-core-benchmark)
scores, where the same engine reaches a best of 451.
It found 438 within about twelve seconds and then plateaued for the rest of the
budget: a genuine local optimum for that seed and time. It is the only one of the
three community reimplementations here that solves the real, five-clue puzzle well
rather than an easier version of it, but the three are not measuring the same
thing, so the comparison needs care.
[Blackwood](/research/lab/experiments/joshua-blackwood/solver) is measured on the
same footing, matched edges with the clues pinned, and stalls near 45.
[McGavin](/research/lab/experiments/peter-mcgavin/backtracker) is not: its figure
is a placement *depth* (about 204 of 256 pieces reached), not a matched-edge score,
so it cannot be read on the same axis as the 438 here, and the two numbers are not
directly comparable. What the three share is only the verdict that the constrained
puzzle is far harder than the unconstrained one; the numbers behind that verdict
are on different scales.
## Related
- [Verhaard's eii: the solver that won the only prize](https://eternity2.dev/research/lab/experiments/louis-verhaard/eii) — The engine behind the 467, the only Eternity II score ever paid, read from Louis Verhaard's own mailing-list posts: forward pruning, comb-search fill orders, depth-gated edge slipping, a Markov-tuned slip schedule. And why his own source-less Win32 binary cannot be built or run on this machine at all.
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Local search and ALNS](https://eternity2.dev/research/build/local-search/local-search-alns) — Destroy part of a board, rebuild it better, and let the algorithm learn which demolitions pay. Adaptive large-neighborhood search is the most reliable polisher this project has, and the cleanest demonstration of the wall where polishing ends.
---
# How the lab publishes
> The editorial standard for this open notebook: how a piece of Eternity II research goes from unpublished work to a published page. What kind of contribution it is, whether it publishes, at what tier, and where it lives. One shared standard, built to scale to many authors.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/methodology/
- Updated: 2026-07-15
---
This page is the editorial standard behind everything in the lab. It exists so
that a finding reaches a reader as an authored artifact, not a raw note, and so
the same rules apply no matter who writes it.
## Published, and unpublished
The ground rule everything rests on:
- **Unpublished work does not appear here.** Where a researcher keeps their work
in progress is their own business. It is simply not in this repository. There
is no half-finished draft sitting in the published notebook.
- **Published work lives in the lab.** Public, distilled, written for a reader.
It got here through review.
Publishing happens by pull request. A researcher opens a PR that adds or promotes
a page, and the review on that PR is where the decision to publish, and at what
tier, is made. Nothing enters the public record without a PR, and the PR is where
a second reader checks it against the standard below. This keeps review
collegial rather than a heavyweight gate, and it scales to many authors.
## Three axes describe every page
Keeping these separate is the whole method. It is easy to blur "polished,"
"reproducible," and "reviewed" into one vague notion of "done." They are not the
same thing.
1. **Contribution:** what kind of result is this?
2. **Tier:** how far through review has it come?
3. **Rigor and reproducibility:** how firmly is the claim backed?
Plus attribution: who did what.
## Axis 1: contribution
Not every result is a solver, and the notebook stopped pretending otherwise. A
page declares its contribution:
- **solver:** produces a competitive board by searching. Its score is a real
search result. Only these earn a row on the leaderboard.
- **analysis:** proves or computes a property of an existing board or of the
instance. It does not produce a new board.
- **reconstruction:** decodes or rebuilds the community's known work to extract
an insight. The number is theirs, decoded.
- **theory:** a mathematical property, law, or impossibility proof.
- **method:** a technique described for reuse, not a scored run.
- **measurement:** a benchmark or empirical observation about solvers or
instances.
- **negative:** a rigorously-run dead end. First-class, not a footnote.
- **tool** and **exposition:** a software artifact, or an explainer.
The load-bearing distinction is solver versus everything else. A meet-in-the-
middle band solver that proves an endgame is optimal is an analysis, not a
solver, and it does not belong on a score chart even though it emits a number.
Getting this axis right is what lets the leaderboard mean one thing.
## Axis 2: tier
Once a page is published, it carries a tier that says how firmly it has been
reviewed:
- **Technical report:** public, but the review confirmed only that it is sound
and fairly stated, not that it has been independently checked. The page
carries a "technical report" badge, the way a preprint is stamped "not yet
reviewed." A legitimate resting state: not everything needs to go further.
- **Reviewed finding:** a second reader, or the same author after a
cooling-off, confirmed it holds. Cited, treated as close to immutable;
corrections happen in place, annotated.
Two principles carry over from how research publishing works elsewhere. The tier
is stamped on the page, so a reader always knows what they are looking at. And
promotion turns on rigor, not on outcome: a sound method that found nothing gets
published as a negative result, while a striking result on an unsound method does
not.
## Axis 3: rigor and reproducibility
Every page states how firmly its central claim is established (proven, measured,
or conjectured) and how re-runnable it is. For anything quantitative the gate is
simple: a number publishes only when its exact configuration and seed are
archived and it can be re-run from the documentation. A benchmark score with no
re-runnable config does not get published. Cheaper contributions carry a lighter
bar: a negative result or an explainer needs sound reasoning and clearly stated
limits, not a reproduction script.
## Attribution, built for more authors
Each researcher's pages gather automatically on their contributor page, derived
from the byline, never hand-listed. Adding an author is adding an entry to the
registry and writing pages under their name. When a page has more than one hand,
a light contributor role (who ran it, who analysed it, who validated it, who
wrote it) records the split, so credit stays accurate as the lab grows.
## The decision, in short
When a piece of work is done:
1. **Name the contribution.** Solver, analysis, reconstruction, theory, method,
measurement, negative, tool, or exposition.
2. **Decide if it publishes.** Publish on rigor, not on excitement. A sound dead
end publishes. A number with no re-runnable config waits.
3. **Open a PR at a tier.** Technical report if it is written up but not yet
independently checked; reviewed finding once a second reader signs off.
4. **Place it.** Solvers and their kin under the experiments; structural results
under why it is hard; techniques under building a solver; cross-solver
measurements with the benchmarks.
5. **Chart it only if it is a solver.**
That is the whole standard. It is deliberately light, because the point is to
raise the floor on every page without slowing the notebook down.
## Related
- [The lab](https://eternity2.dev/research/lab) — The wiki's open notebook: structural findings and named search experiments, each credited to the researcher who ran it and reproducible from source. One corner of the community's wider research.
- [Experiments](https://eternity2.dev/research/lab/experiments) — The lab's named search experiments, one section per researcher. Each is a real run against Eternity II with its idea, its best board, and the questions it left open. Raphaël Anjou's notebook is here in full; the notebook is open to anyone else's.
- [Contribute your research](https://eternity2.dev/research/contribute) — This wiki is the community's research home, and there is room in it for your work. Three ways to get it published, from a mailing-list post to a pull request, plus the small set of house rules that keep every page trustworthy.
---
# Peter McGavin's engine
> Peter McGavin's own C backtracker, the fastest raw solver the community has measured. Fetched from the mailing list, built on an M1, and run on the real Eternity II. His code, his algorithm; run and written up here.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/peter-mcgavin/
- Updated: 2026-07-15
---
This section is [Peter McGavin's](/research/people/peter-mcgavin) own solver:
the C backtracker he posted to the mailing list, which by the community's
measurements is the fastest raw engine anyone has run. The algorithm and the
code are his. What is added here is the running: fetched from source, built on
one machine, and pointed at the real puzzle, with the numbers written up.
## Pages in this section
- [McGavin's C backtracker: the throughput story, built here](https://eternity2.dev/research/lab/experiments/peter-mcgavin/backtracker) — Peter McGavin's own C backtracker, the community's fastest: a 2007 optimization recipe compounded for two decades through generated code, lookup tables and counter tricks, then built on my M1 and pointed at the real 256-piece puzzle, where single core it drives past 200 of 256 pieces at ~109M placements/s.
---
# McGavin's C backtracker: the throughput story, built here
> Peter McGavin's own C backtracker, the community's fastest: a 2007 optimization recipe compounded for two decades through generated code, lookup tables and counter tricks, then built on my M1 and pointed at the real 256-piece puzzle, where single core it drives past 200 of 256 pieces at ~109M placements/s.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/peter-mcgavin/backtracker/
- Updated: 2026-07-17
- Topics: speed, backtracking
- Reproduce: `fetch genbody71.zip from groups.io msg 11749; gcc -Ofast -DG then without -DG`
- Source: Peter McGavin's genbody71.zip (groups.io msg 11749, the original source) — https://groups.io/g/eternity2/message/11749
- Source: Mike Field's optimization recipe: 'Brute force does not work' (groups.io message 3098, 2007) — https://groups.io/g/eternity2/message/3098
- Source: McGavin rediscovers Field's thread: 'I found the old thread' (groups.io message 11338) — https://groups.io/g/eternity2/message/11338
- Source: The 10×10 method and statistics: 180 core-years (groups.io message 9688) — https://groups.io/g/eternity2/message/9688
- Source: The 469 announcement: a few days on a couple of hundred cores (groups.io message 10045) — https://groups.io/g/eternity2/message/10045
- Source: Multi-core throughput table, Raspberry Pi to dual Xeon (groups.io message 11369) — https://groups.io/g/eternity2/message/11369
- Source: Single-core speeds across nine CPU/compiler combos (groups.io message 11643) — https://groups.io/g/eternity2/message/11643
- Source: The counter trick and the compiler playbook (groups.io message 11751) — https://groups.io/g/eternity2/message/11751
- Source: 295M placements/s measured on McGavin's code (groups.io message 11750) — https://groups.io/g/eternity2/message/11750
---
> **Whose work this is**
>
> The algorithm and the C are **Peter McGavin's** own auto-generated engine, posted to the mailing list. By his own account it is built on an optimization recipe Mike Field posted in 2007 ([message 3098](https://groups.io/g/eternity2/message/3098)), a debt he acknowledged twice ([message 11338](https://groups.io/g/eternity2/message/11338), [message 11780](https://groups.io/g/eternity2/message/11780)). One record is deliberately excluded from that lineage: his 469 came from running [Joshua Blackwood's solver](/research/lab/experiments/joshua-blackwood/solver), not his own ([message 10045](https://groups.io/g/eternity2/message/10045)). This page is [Raphaël Anjou](/research/people/raphael-anjou) building and running McGavin's code on one machine and writing up what it did; the only changes to his source are the two small ones described below.
From late 2010 to today, Peter McGavin has been the mailing list's theory desk
and its stopwatch. He is the person who typeset Brendan Owen's
[complex theory](/research/why/complex-theory) into a LaTeX paper
([message 9188](https://groups.io/g/eternity2/message/9188)), reimplemented it
as a C reference in 2024
([message 11197](https://groups.io/g/eternity2/message/11197)), solved the
community's hardest open [benchmark](/research/build/benchmarks), and set the
469 record that stood until Blackwood's 470. But under all of that runs a
quieter, twenty-year project: a plain row-scan backtracker in C, tuned until it
counts tile placements by the hundreds of millions per second. This page traces
that project through his own posted numbers (where the speed came from, what it
bought, and what he himself said it never could), then builds it here and points
it at the real puzzle.
## The recipe is from 2007, and he says so
In October 2007, answering the question "where have you heard of 70 million
pieces per second?", Mike Field posted a complete optimization playbook under
the pointed title "Brute force does not work"
([message 3098](https://groups.io/g/eternity2/message/3098)). Its ingredients: a
lookup table keyed by a cell's north and west colours, so finding candidate
pieces is one memory access; a **fixed** search order exploited ruthlessly
(placing a piece only updates its south and east neighbours); no loops at all,
but procedurally *generated* monolithic code, one straight-line block per cell;
minimal state (reconstruct the pretty output later, don't store it in the hot
path); piece sides packed into a single int; and reading the generated assembler
to hunt pipeline flushes and L1 misses. Field's code compiled down to about 33
instructions per cell and ran 60–80 million placements per second per core on a
2007 desktop, peaking near 100 million when the working set stayed in L1 cache.
That post is the genome of McGavin's engine. When he showed a snippet of his
"awful auto-generated source code" in 2024, it was recognisably the same
organism: a labelled block per cell (`cell_9_2_next:`), a `LookupNW` table
indexed by the north and west colours, a `tileFree` array, `register` hints, and
a `goto` back into the previous cell's block on exhaustion
([message 11337](https://groups.io/g/eternity2/message/11337)). He went looking
for the origin minutes later and posted the link: "I found the old thread"
([message 11338](https://groups.io/g/eternity2/message/11338)); in 2026 he
repeated the attribution: his optimised backtracker C code "is based on" Mike's
2007 post ([message 11780](https://groups.io/g/eternity2/message/11780)).
Field himself supplies the era's baseline. In 2011 he reported his own
backtracker at 75 million tiles placed per second per core on a 2 GHz AMD, about
26 clock cycles per tile, with over half the time stalled on memory access
([message 9003](https://groups.io/g/eternity2/message/9003)). That number,
roughly 40 to 80 million per core, is what a serious community engine did for the
next decade, McGavin's included.
## What he added, in his own words
The recipe was public; the compounding was McGavin's. The refinements he has
described on the list, roughly in the order they surface:
- **A code generator, not a program.** The per-cell C is regenerated for each
puzzle, hint set and
[placement path](/research/build/backtracking/fill-order). His 2026 workflow is
a compile/run/compile/run cycle: the first pass rebuilds `body.c`, the
straight-line cell code, and the second compiles the solver that embeds it.
Skip a step after changing an input file and the program "won't be doing
anything sensible"
([message 11751](https://groups.io/g/eternity2/message/11751),
[message 11782](https://groups.io/g/eternity2/message/11782)). In January 2026
he posted the whole thing to the list as `genbody71.zip`: 1,455 lines of
`genbody.c` plus a README, compiled with `-DG` to emit `body.c` and again
without it to build the solver that includes it. His own warning leads the
README: "This is experimental development code that evolved over several years
--- not nice, elegant code at all"
([message 11749](https://groups.io/g/eternity2/message/11749)).
- **A counter that costs almost nothing.** `ntpll` ("number of tile placements
per second long long", in his own gloss) is not incremented directly. A 16-bit
register is bumped on every placement, and `0x10000` is added to the 64-bit
counter each time it rolls over: a throwback to 32-bit machines where a 64-bit
increment wasted registers or hit slow RAM. He timed both approaches years ago;
the trick won. On early 64-bit machines it was even faster to install 32-bit
compatibility libraries and compile with `gcc -m32`
([message 11751](https://groups.io/g/eternity2/message/11751)).
- **Compiler archaeology.** Compiling with clang instead of gcc gave "a
significant speed boost"
([message 11330](https://groups.io/g/eternity2/message/11330)); on ARM, clang-15
beats clang-19 and gcc; add `-march=native` and `-mtune=native`; try icc and
icx; use profile-guided optimisation
([message 11751](https://groups.io/g/eternity2/message/11751)). None of this
changes the search. It changes how many searches a dollar of electricity buys.
- **The placement path as a measured choice.** Scan-row is best for E2, but
spiral-in wins on clue puzzles 1 and 3, border-first on 2 and 4, and spiral-out
on unframed variants; he checks by running the solver or by asking complex
theory, having found himself "poor at judging solving orders" by eye
([message 9703](https://groups.io/g/eternity2/message/9703),
[message 9713](https://groups.io/g/eternity2/message/9713)).
## The numbers, 2007 to 2026
Every figure below is from the archive, in its author's own units.
| When | Figure | Setting | Source |
| --- | --- | --- | --- |
| 2007-10 | 60–80M placements/s/core, ~100M peak | Mike Field's recipe, AMD X2 3800+ | [3098](https://groups.io/g/eternity2/message/3098) |
| 2011-02 | ~38M placements/s/core (2,300 × 10⁶/min) | McGavin counting 5x5 corners of Brendan's 10x10, 4 cores; 8672 is his own correction of the units (placements, not solutions) | [8672](https://groups.io/g/eternity2/message/8672) |
| 2011-06 | ~50M nodes/s, single core | AMD Phenom II, scan-row, Brendan's 8x8 | [8863](https://groups.io/g/eternity2/message/8863) |
| 2011-11 | 75M/s/core (~26 cycles/tile) | Field's baseline, 2 GHz AMD, memory-stalled | [9003](https://groups.io/g/eternity2/message/9003) |
| 2013-06 | 44.6M nodes/s sustained | one 683-billion-node 10x10 first-row test | [9167](https://groups.io/g/eternity2/message/9167) |
| 2014-04 | 67M placements/s, single core | Arnaud Carré's benchmark: all 4 solutions in under a minute | [9263](https://groups.io/g/eternity2/message/9263) |
| 2024-10 | 60–140M placements/s per backtracker | "depending on CPU type" | [11329](https://groups.io/g/eternity2/message/11329) |
| 2024-11 | 99M → 10,160M nodes/s per machine | Raspberry Pi 4 (4 processes) to dual Xeon Gold 6338 (128) | [11369](https://groups.io/g/eternity2/message/11369) |
| 2025-09 | 38–84M placements/s, single core | nine OS/compiler/CPU combos on Brendan's 8x8 | [11643](https://groups.io/g/eternity2/message/11643) |
| 2026-01 | ~225M placements/s, single core | Orange Pi 6 Plus on small puzzles; "halves on 16x16" | [11751](https://groups.io/g/eternity2/message/11751) |
| 2026-01 | 295M placements/s, single core | Joe running McGavin's code on a newer CPU, vs his own C# at 27–37M | [11750](https://groups.io/g/eternity2/message/11750) |
| 2026-02 | 44.0M vs 105.1M tiles/s | identical 2.12-trillion-node tree: 2010 Phenom II vs Ryzen 5 5600H | [11782](https://groups.io/g/eternity2/message/11782) |
Two readings of that table. First, the headline: the ~295M/s figure (the one our
[records page](/research/records) quotes) is real, but it is Joe's measurement,
made in January 2026 when McGavin shared his source and Joe ran it on hardware
newer than anything McGavin owns; Joe's own C# solver managed 27–37M/s on the
same machine, and the best speed he had seen mentioned on the list was 70–90M/s
([message 11750](https://groups.io/g/eternity2/message/11750)). McGavin's own
best figure is the Orange Pi's ~225M/s on small puzzles
([message 11751](https://groups.io/g/eternity2/message/11751)).
Second, the quieter and more instructive reading: single-core speed barely moved
for fifteen years. McGavin said so himself when he posted the 2025 table: speeds
on the latest CPUs "are only a little faster" than on his 2010 Phenom II
([message 11643](https://groups.io/g/eternity2/message/11643)). The engine was
already near the memory wall Field described in 2011. What actually grew was the
number of cores he could point at a problem.
## Hundreds of cores on a hobby budget
McGavin's chapter of the community's
[distributed-solving story](/research/build/faster/distributed-solving) is
charmingly domestic. For the 10×10 campaign he started with about 20 cores at
home, then added three octa-core Odroid XU4s and twenty-five quad-core Orange Pi
Lites at 12 dollars each (over 130 cores, each ARM core about a third the speed
of a PC core) and, when allowed, multi-core servers at work for a total of over
400 ([message 9688](https://groups.io/g/eternity2/message/9688)). The Orange Pis
ran off 12-port USB chargers (he permanently killed one charger by plugging in
twelve boards running 48 backtrackers, and cut back to eight per charger) with
WiFi networking: one cable per board, for power
([message 9690](https://groups.io/g/eternity2/message/9690)). Orchestration is
two shell scripts: one starts as many backtrackers as a node has cores, the other
launches it on ~20 nodes over ssh, each with 16 to 48 cores, though "well, they
are hyperthreads, strictly speaking"
([message 9753](https://groups.io/g/eternity2/message/9753)). At peak, "more than
400 backtrackers running at once"
([message 9751](https://groups.io/g/eternity2/message/9751)).
## What the speed bought
**Verified counts, first.** Fast full enumerations are what let the community
check theory against reality. In 2011, an overnight run counted
4,739,821,621,743 5×5 corner blocks of Brendan's 10×10, against a complex theory
estimate of 5.0077 × 10¹², "pretty close… if I do say so myself"
([message 8672](https://groups.io/g/eternity2/message/8672)). In 2014 his single
core swept Arnaud Carré's 256-piece benchmark tree (3,979,209,754 placements,
all 4 solutions) in under a minute
([message 9263](https://groups.io/g/eternity2/message/9263)). In 2026, two
different machines walked the same 2,120,424,701,160-node tree of an E2-like
puzzle and found the same lone solution: determinism as a feature, node counts as
a checksum ([message 11782](https://groups.io/g/eternity2/message/11782)).
**The 10×10, above all.** Brendan's set_1 10×10, a benchmark that had stood open
for a decade, fell in September 2017 to exactly this machinery: enumerate ~20
million candidate first rows, rank them by complex theory's
solutions-per-search-node, and let the farm test them one by one, about a
core-day each. The solution arrived on row-test ~92,907 of a predicted
one-in-70,000, after nearly 2 × 10¹⁷ nodes, about **180 core-years**, less than
0.5% of the whole tree
([message 9686](https://groups.io/g/eternity2/message/9686),
[message 9688](https://groups.io/g/eternity2/message/9688)), spread over about
four years ([message 9804](https://groups.io/g/eternity2/message/9804)). His
summary: "no new methods, just systematic persistence and the law of large
numbers." The [benchmarks page](/research/build/benchmarks) tells that story as
complex theory's strongest validation; here it stands as the throughput story's
high-water mark.
**And one record, on someone else's engine.** In September 2020, days after
Joshua Blackwood open-sourced his solver, McGavin ran it "for a few days on about
a couple of hundred cores and hit the jackpot. New record score of 469!"
([message 10045](https://groups.io/g/eternity2/message/10045)), with run
statistics to match (2,832 boards reaching 252 pieces, one 255, one 256:
[message 10049](https://groups.io/g/eternity2/message/10049)). Note what combined
there: [Blackwood's heuristics](/research/lab/experiments/joshua-blackwood/solver)
supplied the shape of the search; McGavin's farm supplied the placements. His own
C engine holds no E2 score record; his row-scan 226s of February 2020 equalled
[Verhaard's](/research/lab/experiments/louis-verhaard/eii) old consecutive-placement
mark, no more ([message 10523](https://groups.io/g/eternity2/message/10523)).
## What the speed could not buy
McGavin is also the archive's most consistent witness *against* raw speed. His
own numbers make the case. With scan-row order, complex theory puts E2's full
search tree at about 1.6 × 10⁴⁷ placements
([message 9710](https://groups.io/g/eternity2/message/9710)), some 9.3 × 10⁴² per
expected solution ([message 9713](https://groups.io/g/eternity2/message/9713)).
Even on his best 5-hint placement path (a far smaller tree, about 3.1 × 10⁴⁰
nodes) he computed 4.9 × 10³² years at 100 million nodes per second, and adding
billions of cores still leaves you "orders of magnitude longer than the age of
the Universe" ([message 11201](https://groups.io/g/eternity2/message/11201)). He
had drawn the conclusion long before, in 2013: "It seems clear to me that E2 will
not be solved by brute force. If it is to be solved at all, it will be by deep
analysis and/or clever insight, in my opinion"
([message 9117](https://groups.io/g/eternity2/message/9117)).
Which is why the biggest single speedup he ever reported was not a speedup at
all. For the unframed sub-solution hunts, he used complex theory as a chess-style
lookahead: estimate solutions-per-node at every leaf 13 or 14 plies ahead, cache
the repeated statistics, and steer into the best branch. Overall gain: a factor
of about 25 by depth 81 ([message 9751](https://groups.io/g/eternity2/message/9751)).
No compiler flag ever gave him 25×. Shaping the tree beat shrinking the
nanoseconds. That is this site's [core lesson](/research/why/prune-vs-speed),
reported here as one practitioner's twenty-year experiment: he built the fastest
engine in the community's history, measured everything, and concluded the gap to
480 was never on the clock.
## Running it here
All of the above is drawn from the archive. The rest of this page is that same
engine, built on my machine and pointed at the real puzzle, so the throughput
story carries a first-party measurement and not only a relayed one.
### Where the code came from
The source is `genbody.c`, attached as `genbody71.zip` to
[groups.io message 11749](https://groups.io/g/eternity2/message/11749). It is
1,455 lines of C, with a README, a piece file, and a hint file. It is not copied
into this repository: it stays on the list, where its author put it.
It is a two-pass program, and the README is candid about the style ("dreadful C
source code ... it really needs a lot of work to clean it up"). Compiled with
`-DG`, it generates a second C file, `body.c`, specialised to one puzzle.
Recompiled without `-DG`, it includes that generated file and runs the search.
That two-pass workflow is exactly the code-generator design described above, now
in front of me.
### What was changed to run it
Two things, both small:
- **Nothing, to build it.** It compiles clean on Apple silicon with `clang` (one
unused-variable warning). The POSIX status display it uses, `setitimer` and
`termios`, works on macOS with no shim.
- **The puzzle it reads.** The filenames were hardcoded to Joe's test puzzle. I
made them read from the command line so it could be fed the real Eternity II,
and wrote the official pieces and clues into his `.puz` / `.hnt` format from the
canonical puzzle file. His search path is built generically (clue cells first,
then a row scan), so no other change was needed.
### What it does on Joe's puzzle
Configured as shipped, on Joe's 18-hint 16×16 test puzzle, it is very fast and it
finishes:
- **~279 million tile placements per second**, single core.
- Solves the puzzle to completion in about 13 seconds (3.577 billion placements
to the first solution), identically on every run.
That is the number the community means by "McGavin is fast." It is real, and it
is on current hardware, not a seven-year-old machine: right in line with the
~295M/s Joe measured, and comfortably past McGavin's own ~225M/s Orange Pi.
### What it does on the real Eternity II
Fed the actual 256-piece puzzle, once with only the mandatory centre clue and
once with all five official clues: his solver only saves a board when it finds a
**complete** solution, and the real puzzle has never been solved, so it saves
nothing and runs without stopping. What it does report, live, is the deepest it
has placed:
| Puzzle | Deepest placement, 30 s | Rate | Solutions |
| --- | --- | --- | --- |
| Real E2, 1 clue | **205 / 256** | ~108 M placements/s | 0 |
| Real E2, 5 clues | **204 / 256** | ~109 M placements/s | 0 |
Two things are worth saying plainly. First, this is a **depth** (how far the
search reached before backtracking), not an edge score out of 480: his program
does not emit a partial board to re-score. Second, the rate on the real puzzle is
about 109 million placements a second, roughly 40% of its speed on Joe's puzzle,
because the real puzzle's constraints prune harder. The clue count barely moves
it: 205 with one clue, 204 with five. This is the same lesson his own posts
reach, now on my own hardware: the raw engine is superb at walking a tree and
says nothing, on its own, about where a high-scoring board hides.
### Single core, confirmed
The binary links only the system library, has no threading primitives in its
source, and holds one CPU at 100% (not 800%) with a single thread throughout. The
speed is one core's. That is the fair unit for comparing engines, and it is the
one the
[single-core benchmark](/research/lab/experiments/single-core-benchmark) uses to
put this engine next to
[Blackwood's](/research/lab/experiments/joshua-blackwood/solver) and the
[Verhaard reimplementation](/research/lab/experiments/louis-verhaard/verhaard-reimpl).
The three are not measuring the same axis (McGavin's number is a placement depth,
not a matched-edge score), so read the comparison with care.
One first-party postscript on the throughput itself. Rebuilt headless (its live
terminal display, it turns out, costs it ~2.7×) and pointed at both an easy and a
hard board, this C engine sets the bar a
[portable-Rust codegen backtracker](/research/lab/experiments/raphael-anjou/jit-backtracker)
was then measured against on the same M1: the Rust ties it on hard, deep boards like
the real puzzle (~105–110M each) and trails by ~2.3× on easy low-branching ones
(~287M vs ~122M). A useful calibration of how much of this engine's edge is portable
craft and how much is the shape of the board it runs on.
## Related
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Complex theory: counting the search before you run it](https://eternity2.dev/research/why/complex-theory) — Brendan Owen's complex theory estimates how wide the search tree is at every depth, and even how many solutions exist at all. Many in the community consider it the single most important thing to understand about Eternity II.
- [The community's benchmarks](https://eternity2.dev/research/build/benchmarks) — How a community forbidden from sharing the pieces built a shared test culture anyway: derived-count verification protocols, the Txibilis and beginner suites, node-count duels, full enumerations, and the one benchmark that is still standing open today.
---
# Raphaël Anjou's experiments
> A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/
- Updated: 2026-07-17
---
This is [Raphaël Anjou's](/research/people/raphael-anjou) notebook of Eternity II
search experiments. Some are original ideas; some faithfully reimplement a known
community technique to measure exactly what it buys. Every one records its idea,
the board it reached, and the questions it left open, and every board is real and
checkable in the [viewer](/viewer). The best reaches **463 of 480** matched edges;
the community's best on the same puzzle is 470. Methods and boards are here in
full, nothing withheld.
It is a notebook, not a single result, so it is worth knowing how it is laid out
before diving in.
## What's in this section
Two kinds of thing live here: the **apparatus** the experiments run on, and the
**experiments** themselves. The experiments come in three flavours: pipelines that
chase the whole-board score, studies that isolate one search decision at a time,
and exact solves that *prove* a small region rather than guess it. Start wherever
the question you care about lives.
- [The engines](/research/lab/experiments/raphael-anjou/engines) — The shared machinery the experiments run on: a constructive beam producer, a destroy-and-repair local search, and a family of CSP presets. Documented once here so each study can point back instead of re-explaining the machine. Start here if you want to know how a search works before reading what a study did with it.
- [Combination pipelines](/research/lab/experiments/raphael-anjou/pipelines) — The named runs that push the score. Each is a pipeline, not a single algorithm: build a board with one engine, then lift or finish it with another. The interest is in the division of labour between construction, repair and an exact endgame. This is where the record-approaching boards come from.
- [The DFS study](/research/lab/experiments/raphael-anjou/dfs-study) — Depth-first backtracking, dissected. A family of from-scratch backtrackers, each one change apart (fill order, heuristic, break policy), run on the same ten variants under a fixed budget, to price what each idea is worth.
- [The repair study](/research/lab/experiments/raphael-anjou/repair-study) — Its sibling, for destroy-and-repair local search: which region to destroy, how to rebuild it, when to keep a move, what board to start from. The loop the records actually reach the top with, taken apart one decision at a time.
- [Learning from strong boards](/research/lab/experiments/raphael-anjou/learning) — The third study, turned inward: instead of varying the search, vary what it is allowed to *know*. Five experiments mine the corpus of strong boards for structure and feed it back in, and all five hit the same wall.
- [Meet in the middle](/research/lab/experiments/raphael-anjou/meet-in-the-middle) — Exact endgame solves: enumerate a region from two ends and join on the seam to find the true best completion, with a proof nothing scores higher. These measure a small region exactly rather than chase the whole-board score.
- [Going fast](/research/lab/experiments/raphael-anjou/going-fast) — The other lever: not a smarter search, a faster one. A portable-Rust backtracker that generates and compiles per-puzzle Rust, taken rung by rung to McGavin-class throughput - a tie with the community's fastest hand-tuned C on hard, realistic boards, from a safe language. A speed result, on an axis of its own from the scores above.
The two constructive engines, the beam producer and the ALNS polish stage, do not
yet have their own write-ups; where a pipeline steers one, its own page says what
that engine does at the level the study needs. The CSP presets are the one engine
documented and measured in full.
## How to read the scores
The chart and table below carry two kinds of number that must never be conflated,
so they are drawn as separate groups.
- **Explore** is the best board each method reached in exploratory runs on 8
cores, wall-clock not logged. These are the headline numbers the write-ups
quote, and they top out at PALIMPSEST's 463.
- **Bench** is the standardized single-core benchmark: one core, sixty seconds,
every board re-scored by the same canonical scorer. Lower, and *directly
comparable* across methods in a way the explore numbers are not.
Not every experiment earns a row. The three studies vary a knob across a grid
rather than producing one board, so the [DFS](/research/lab/experiments/raphael-anjou/dfs-study)
and [repair](/research/lab/experiments/raphael-anjou/repair-study) studies carry
their own leaderboards instead of a chart row; only their single best bench result
appears here. The five [learning](/research/lab/experiments/raphael-anjou/learning)
experiments each reach one scored board, so they keep their rows as well as their
study home. The bench rows above are study results, not standalone pages, so they
appear only on the chart.
> **[Interactive: ExperimentScoreChart]** Rendered on the canonical page (link above); not shown in this markdown export.
The named experiments as a sortable table (click a column head to reorder by
score, name, method, or month). Author, month, rigor and reproducibility are read
from each experiment's own page, so this summary stays in step with them.
s.group !== "bench")} />
## Pages in this section
- [The reference engine behind this site](https://eternity2.dev/research/lab/experiments/raphael-anjou/engine) — The Rust-to-WebAssembly backtracker that runs every live demo and checks every number on this wiki. Not a record machine but a reference engine, ported four times and parity-tested byte-for-byte, built so the claims here can be re-run.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [The engines](https://eternity2.dev/research/lab/experiments/raphael-anjou/engines) — The shared engines under Raphaël Anjou's experiments. The named experiments are studies that run on these; this is the apparatus they share. The CSP presets and the Verhaard reimplementation are documented here; the constructive engines are not yet published.
- [The JIT backtracker: portable Rust that ties hand-tuned C on hard boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/jit-backtracker) — A safe, portable Rust depth-first backtracker, specialised at runtime by emitting and compiling per-puzzle Rust, taken from 43 to 123 million search-nodes per second on one core. Measured fairly against Peter McGavin's C on the same machine: a tie on hard, deep boards like the real Eternity II, and about 44% of its speed on easy ones. Every rung searches the identical tree; the whole gain is code, not algorithm.
- [Going fast: when a solver spends its budget on speed](https://eternity2.dev/research/lab/experiments/raphael-anjou/going-fast) — Some Eternity II engines pour their effort into walking the search tree as fast as possible; others spend it on judgement about where to walk. This is the case for the first kind - what raw throughput buys, the three different things people mean by "fast", and why the fastest engine ever built still cannot solve the puzzle.
- [Combination pipelines](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines) — The named search experiments that chase score. Each is a pipeline rather than a single algorithm: it builds a board with one engine, then lifts or finishes it with another. Each records its idea, its best board, and the questions it left open.
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [The hint study](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study) — Give a backtracker five correct pieces for free, in the puzzle's own clue geometry. It turns out not to help, and depending on the fill order it can hurt badly, because a pinned piece is a hard constraint a fixed fill order must satisfy on arrival. A family of fill paths, run on the same hinted boards, single core, measured against no hints at all.
- [The repair study](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study) — The sibling of the DFS study, for the other way people attack Eternity II: destroy part of a board, rebuild it, keep the change if it helps. One question, asked carefully. What does each decision in that loop buy: which region to destroy, how to rebuild it, when to keep a move, when to restart, and what board to start from?
- [Meet in the middle](https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle) — Exact endgame experiments that meet in the middle: enumerate a region from two ends and join on the seam, to find the true best completion with a proof rather than a heuristic's best guess. These measure a small region exactly instead of chasing the whole-board score.
---
# The DFS study
> One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/
- Updated: 2026-07-16
- Topics: backtracking, search-space, speed
- Reproduce: `just experiments dfs-study`
- Source: Runnable engine + committed results + scripts (this study's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/dfs-study
---
Every record-holding Eternity II solver (Blackwood, Verhaard, McGavin) is a
depth-first backtracker. What separates them from a first-week homework
backtracker is not their kind but a handful of decisions: which order they fill
cells in, which look-ahead they run, and whether they let an edge *break*. This
study takes those decisions apart. It builds a family of depth-first backtrackers
from scratch, each one exactly one change away from a sibling, and runs them all
on the same ten corner-pinned variants of the official puzzle, single core, sixty
seconds a run. The maximum score is 480 matched edges.
The point is not to win. The strongest variant here averages the low 430s, well
short of the community's 464 on these five clues, because sixty seconds on one
core is a small fraction of the compute the records took. The point is to isolate
*what each idea is worth* by changing one thing at a time and measuring the result
with the same canonical scorer for every board.
> **[Interactive: StartingPuzzleCarousel]** Rendered on the canonical page (link above); not shown in this markdown export.
## The family, and the leaderboard
Four families, laid out so that neighbours differ by a single decision.
**Baseline** is the rawest possible backtracker, alongside a hand-specialised twin
that prices the low-level engineering. **Path order** fixes everything but the
sequence in which cells are filled. **Heuristic** fixes the order and adds one
propagator at a time. **Break** is the elite axis: the depth-gated edge-break
mechanism the records rely on. The community record engines, McGavin's C and
Blackwood's C#, are themselves break backtrackers, so they sit in the break
family too, not in a category of their own. Whose code an engine is stays a
matter of labelling, not colour. Both appear on the leaderboard at their
corner-pinned score, badged where they collapse, and again on a fair unpinned
grid further down where they run as designed.
> **[Interactive: DfsStudyLeaderboard]** Rendered on the canonical page (link above); not shown in this markdown export.
## What the study found
- **Path order is the largest free lever, and the wrong order is catastrophic.**
Plain row-major averages 377; a strict border-first or spiral fill, with no
heuristic to rescue it, stalls near 67. Same engine, same budget, a swing of
more than 300 points from the fill order alone.
- **The most-constrained-cell heuristic (MRV) is what makes border-first viable.**
It lifts a stalled border-first from the sixties to a mean of 324, at a cost of
three orders of magnitude in node throughput. Node rate and score are different
axes, a distinction the study returns to throughout.
- **More propagation did not buy more score at this budget.** Forward-checking,
arc-consistency and per-colour reasoning land within a point of each other (322,
321, 321), a gap far inside the run-to-run spread, so heavier look-ahead neither
helped nor clearly hurt. It spends the sixty seconds proving small regions
rather than reaching deeper.
- **Breaks reach deeper than any strict search.** Strict backtrackers top out in
the low 200s (the fastest, NAIVE-CODEGEN, at 216); a depth-gated break budget
reaches past 245 and averages the low 430s, because it can push past a locally
unmatchable edge instead of backtracking out of it. The decisive factor is the
break *schedule*: unlocking breaks too early (Verhaard's ladder, mean 399)
scores well below the later Blackwood ladder (mean 431). Raising the per-cell
cap from one to two did not help at this budget, a null result reported as
measured.
Each of these has its own page: how the engine is built and what every raised
statistic means is on the [method page](/research/lab/experiments/raphael-anjou/dfs-study/method),
and the path, heuristic and break comparisons are worked through on the
[findings page](/research/lab/experiments/raphael-anjou/dfs-study/findings).
## How to read the numbers
Every board is re-scored by one canonical scorer, and no engine's self-reported
score is trusted. Throughput is reported in search-nodes per second and is
**never compared across families**, because a node that runs full arc-consistency
is not the same unit of work as a naive placement. Depth is the deepest placement
a variant reached, out of 256.
The break count deserves a precise definition, because it is easy to state
loosely. A board's *score* is its matched interior edges, and the gap
`480 − score` is the board's total unmatched-edge deficit. On a **completed**
board every unmatched edge is a genuine break, so there the score is exactly
`480 − #breaks`. Sixty seconds is rarely enough to fill the board, however, so
most break-variant boards here are partial, and their deficit is dominated by
edges that are simply still empty rather than broken. This study therefore reports
the **true break count**, meaning the interior mismatches the search actually
committed under its budget, tracked by the search itself rather than inferred from
the score. That number stays small even when the deficit is large. Every board
carries a bucas `.url` that opens in the [viewer](/viewer), so both the score and
the broken edges can be checked directly.
The whole apparatus (the engine workspace, the ten variants, the committed
per-run results and the grid scripts) lives under the study's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/dfs-study),
and `just experiments dfs-study` rebuilds the engine and reruns the whole grid.
## Pages in this section
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/method) — The engine behind the DFS study: one composable backtracker where a variant is a declared change over a parent, a shared IO layer every algorithm speaks, and the definitions of every statistic the study raises: node rate, depth, breaks.
- [What each idea buys](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/findings) — The three comparisons at the heart of the DFS study, worked through: fill order (row-major wins, a bad order is catastrophic), heuristics (MRV rescues border-first but costs throughput; more propagation did not help), and breaks (they smash the depth wall; the break schedule is the lever, not the per-cell cap).
## Related
- [Raphaël Anjou's experiments](https://eternity2.dev/research/lab/experiments/raphael-anjou) — A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Fill orders](https://eternity2.dev/research/build/backtracking/fill-order) — The order in which a backtracker visits the 256 cells is its one free choice: it costs nothing at runtime and moves the size of the search tree by orders of magnitude. Twenty years of community science, from the fixed-vs-dynamic wars and the strategy races to the magic 10×16 square and Verhaard's comb search, all answer the same question: which path through the board is cheapest?
---
# What each idea buys
> The three comparisons at the heart of the DFS study, worked through: fill order (row-major wins, a bad order is catastrophic), heuristics (MRV rescues border-first but costs throughput; more propagation did not help), and breaks (they smash the depth wall; the break schedule is the lever, not the per-cell cap).
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/findings/
- Updated: 2026-07-16
- Topics: backtracking, search-space, speed
- Source: Committed per-run results (results.jsonl) and per-family report (report.md) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/dfs-study/results
---
Three comparisons carry the [DFS study](/research/lab/experiments/raphael-anjou/dfs-study).
Each isolates one decision by holding everything else fixed. All scores are the
mean matched-edge count over the ten corner-pinned variants, single core, sixty
seconds. Throughput is search-nodes per second and is never compared across
families.
## What low-level specialisation buys: throughput, and only that
The two baselines run the *same* strict row-major algorithm: `NAIVE-CLEAN` as a
readable general engine, `NAIVE-CODEGEN` as a hand-specialised, row-major-only
16×16 hot loop. The specialisation delivers what it should on the axis it targets.
`NAIVE-CODEGEN` is meaningfully faster per node, by up to a third on some
instances. On *score* the two are a wash, a point or two apart at sixty seconds
and well inside the run-to-run spread, which is the fair reading rather than a
claim that specialisation *hurts*. At a fixed wall-clock a faster engine reaches a
different point of the same tree, and a backtracker's best partial does not move
monotonically with how fast it got there. The baselines are therefore a clean
*speed* comparison and are deliberately not presented as a *score* comparison. The
score effects worth studying are all on the axes below, where the search itself
changes.
## Fill order: row-major wins, and the wrong order is catastrophic
Fix the engine (strict, no heuristics) and change only the order cells are filled.
The six orders differ only in where the search sends its frontier, shown below.
> **[Figure]** The six fill orders, traced as the path the search walks — interactive: PathOrderDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
- **Row-major** is the strong baseline, averaging 377. Its damage zone stays
constant, since every new cell has the same two placed neighbours, so it reaches
deep before the strict wall.
- **A strict border-first or spiral fill collapses.** Filling the border ring
first, with no look-ahead, walks straight into the hardest corner and edge
constraints and stalls almost at once, near a mean of 67 at depth 66 or so. The
spiral pays the same closure tax.
- **Bottom-up row-major fares worse still on this clue geometry** (mean 226, but
as low as 18 on some variants), because the pinned clues sit in rows the
bottom-up fill reaches early and cannot satisfy.
Same engine, same sixty seconds, a swing of more than 300 points from the fill
order alone. This is the measured form of a piece of community wisdom: border-first
is only good *with* a heuristic to choose cells within the ring. On its own it is
one of the worst orders available.
## Heuristics: MRV rescues border-first, but throughput craters
Now fix the path near the frame and add one thing at a time. The largest lever is
**MRV**, filling the most-constrained empty cell next, chosen dynamically. It
turns the stalled border-first (mean 67) into a search that averages 324 (best
341) at depth 190 or so. It also recomputes the most-constrained cell over the
whole frontier at every step, so node throughput falls by three orders of
magnitude, from tens of millions of nodes per second to a few thousand. Each node
is worth far more, and far fewer of them are visited. Node rate and score are
different axes.
Adding heavier look-ahead on top did **not** buy more score at this budget.
- **Forward-checking** (reject a placement that empties a neighbour's domain)
averaged 322.
- **Arc-consistency** and the **per-colour supply check** averaged 321 each. With
a per-variant score spread of about 11 points over the ten instances, that
one-point gap is well within the noise: the three propagators are statistically
indistinguishable here, so the fair reading is that heavier look-ahead neither
helped nor clearly hurt, rather than that forward-checking won.
- A **rare-colour-first** value ordering was inert, no better than plain insertion
order, echoing the community's repeated negative result on within-bucket value
orders.
The lesson is not that propagation is useless. It is that at a small fixed budget,
on this instance, the cheapest useful prune (forward-checking) already captures
whatever benefit is available, and more expensive reasoning does not recover its
extra per-node cost within sixty seconds.
## Breaks: past the wall, and the schedule is the lever
Strict backtracking, whatever its order or heuristic, hits a wall well short of a
full board: row-major tops out around depth 208 of 256, and even the fastest
strict variant (NAIVE-CODEGEN) only reaches 216. The record engines get past it
by *breaking*: allowing a bounded number of interior edge mismatches,
released on a depth schedule, with a rule that no cell may carry too many broken
edges. The score of a full board is then `480 − #breaks`.
- **Breaks clear the wall.** A depth-gated break budget reaches past depth 245 and
averages the low 430s (break-1 means 431 of 480, best 435), a large gain over the
strict high-370s on the same instances and budget. This is the mechanism behind
the community's record backtrackers, not a different paradigm but a depth-gated
relaxation of the matching rule.
- **Allowing a second break per cell did not help here.** One-break and two-break
reach the same best board (435), and their means (431.3 against 428.5) sit within
the two-break variant's own spread, so neither leads on average. What does
separate them is consistency: the one-break variant is tightly clustered (never
below 427), while the two-break variant ranges down to 402. The double-break
geometry the community's 460-boards use appears to need more than sixty seconds
to pay off; at this budget the extra freedom mostly widens the branching without
reaching better boards. This is a null result, reported as measured rather than
the gain one might expect.
- **The schedule is the decisive lever.** Verhaard's slip ladder unlocks breaks
much earlier than Blackwood's (depth 193 against 201) and here scores markedly
worse (mean 399 against 431), because unlocking early spends the budget on
shallow breaks. When and how fast breaks open is a tuning decision rather than a
detail.
These break numbers sit alongside the sibling benchmark's from-scratch
Blackwood-style and Verhaard-style reimplementations, which reach the high 430s on
the same five clues. The agreement from an independent engine cross-validates the
break machinery here.
## Where the community engines sit: two grids
Blackwood and McGavin are the high end of this same family, and both build and run
on this machine, so this study ran them on both a pinned and an unpinned grid, with
every score canonically rescored from the engine's own board. The result is a
finding in its own right, and it has two halves.
**On the pinned grid, they collapse.** McGavin's C, built with its author's own ARM
flags (native tuning plus link-time optimisation), reaches depth 211 on the plain
centre-clue puzzle at about 85 million tiles per second, well ahead of our fastest
strict engine, yet pinning three corners collapses it to depth 21, a canonical
score of 13. Its generated scan path never visits the corners early, so a pinned
corner constrains its neighbourhood at once and dead-ends the fixed path almost
immediately. Blackwood's C# hardcodes its piece set and scan, so it cannot even
express an arbitrary corner pin; on the matching five-clue constraint its break
heuristic thrashes to depth 47, a canonical score of 75, because it is tuned for
the near-unconstrained one-clue instance where its 470 record was set.
> **[Figure]** The corner-pin collapse: cells placed out of 256 — interactive: PinCollapseDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
That collapse is the point. A record engine built around one clue configuration
does not transfer to another, and a fixed scan path cannot absorb an arbitrary pin.
Our from-scratch engines treat a pin as a pre-placed cell the scan simply skips,
which is why they, rather than the community binaries, are the stand-ins on the
pinned grid.
**On a fair unpinned grid, they run as designed.** Give every engine the official
pieces with only the one mandatory centre clue, and nothing dead-ends on a corner.
In sixty seconds on one core, McGavin reaches 392, our strongest break engine 344,
the Verhaard reimplementation 286, and Blackwood 214. Read those as what a single
core buys in a minute from a cold start, not as any engine's ceiling: Blackwood's
470 and McGavin's deep runs came from days on hundreds of cores, which this budget
cannot show. What the grid does show is that all four run properly once the pins
that break a fixed scan path are gone, which is exactly what the pinned grid denied
the two foreign engines. Both grids are on the study's leaderboard above.
## The through-line
One theme runs through all three comparisons: **node count is not score.** The
fastest variants per node (the codegen baseline and the strict row-major engine)
do not reach the best boards; the slowest per node (MRV with propagation) reaches
far better ones; and the breaks that win do so by changing *which* branch is
legal, not by visiting branches faster. Raw speed is a constant factor, while
where and whether the search is allowed to go is the exponential one.
## Related
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/method) — The engine behind the DFS study: one composable backtracker where a variant is a declared change over a parent, a shared IO layer every algorithm speaks, and the definitions of every statistic the study raises: node rate, depth, breaks.
- [Fill orders](https://eternity2.dev/research/build/backtracking/fill-order) — The order in which a backtracker visits the 256 cells is its one free choice: it costs nothing at runtime and moves the size of the search tree by orders of magnitude. Twenty years of community science, from the fixed-vs-dynamic wars and the strategy races to the magic 10×16 square and Verhaard's comb search, all answer the same question: which path through the board is cheapest?
---
# How the study is built
> The engine behind the DFS study: one composable backtracker where a variant is a declared change over a parent, a shared IO layer every algorithm speaks, and the definitions of every statistic the study raises: node rate, depth, breaks.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/method/
- Updated: 2026-07-16
- Topics: backtracking, speed
- Source: The engine workspace (dfs-engine, dfs-run) on the shared e2-core / e2-io library — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/dfs-study/engine
---
This page is the apparatus behind the [DFS study](/research/lab/experiments/raphael-anjou/dfs-study):
how the engine is built, why a new variant is cheap to add, and what every number
on the results means. None of it depends on the sibling
[single-core benchmark](/research/lab/experiments/single-core-benchmark)'s engine.
The whole point was to reimplement the family from scratch, keeping the software
engineering clean and the "what stacks on what" story explicit.
## A variant is a declared change over a parent
Every algorithm in the study is the *same* recursive depth-first backtracker,
parameterised by four independent choices:
- **path order**: the sequence cells are filled in (row-major, spiral,
border-first, Verhaard's comb, or dynamic most-constrained-cell);
- **value order**: the order a cell's candidate pieces are tried;
- **propagator**: the look-ahead run after each placement (none,
forward-checking, arc-consistency, per-colour reasoning);
- **break policy**: whether an edge may mismatch, and under what depth-gated
budget.
A variant is a small record naming those four choices, together with **the parent
it derives from and a one-line description of the single change it adds**. Adding
a variant means adding one record to the registry, with no new search code unless
the idea is a genuinely new strategy. The "what stacks on what" matrix on the
results page is generated from those descriptions, so it cannot drift from the
code that ran. That is what keeps an extensive study, with dozens of
one-change-apart variants, maintainable rather than a pile of copy-pasted solvers.
## One IO layer, and conversions between engines
Every algorithm in the study consumes one instance type and emits one output: the
best board, its canonical score, and a bucas URL. Around that sits a shared IO
layer with lossless converters between the formats the site's other engines
speak: the benchmark's site-schema JSON, the standalone community engines' CSV,
bucas URLs, and hint files. The study reads the *same* ten corner-pinned variants
the single-core benchmark uses, through this layer, so the two experiments are
directly comparable. A small `dfs-convert` utility exposes the conversions from
the shell, so any blog engine's output can be fed to any other.
## The scorer is the single source of truth
No engine's self-reported score is trusted. Every board, strict or broken, is
re-scored by one canonical scorer: matched, non-border, interior adjacencies,
counted right and down per cell. It is byte-for-byte the same formula the site's
scorer and the benchmark use, checked by a test that re-scores a known 469-board
and asserts 469. This is what lets scores from different variants, and from the
sibling benchmark, sit on one axis.
## The statistics the study raises
For every run the engine records, and the results carry all the way to the page:
- **score**: canonical matched edges (of 480). For a full board with breaks,
this equals `480 − #breaks`.
- **node throughput**: search-nodes per second, one node per attempted placement.
Reported per variant and **never compared across families**, because a node that
runs full arc-consistency is not the same unit of work as a naive placement.
Heavy propagation trades throughput for node quality, and the study measures
both axes rather than collapsing them into one. The slowest variants carry a
caveat worth stating plainly: the MRV engine picks the most-constrained cell by
scanning every empty cell's candidate list at each node, which is inherently
heavier than a fixed fill order. Two behaviour-preserving optimisations bring it
within a few times of the fast engines rather than the thousands it once was: the
most-constrained search stops counting a cell's candidates the moment they exceed
the best cell found so far (a minimum search never needs the exact count of a
cell that cannot win), and it skips re-checking the edges the candidate list is
already indexed on. What it still does not do is maintain each cell's candidate
count fully incrementally across placements, which a production CSP solver would;
that last step would need to track how a newly-used piece affects every cell's
count, and it is left out here to keep the engine legible. The ranking by score
does not depend on any of this, since throughput is a separate axis, but the MRV
node rate should be read as this clean engine's rather than as MRV's best
possible.
- **max depth reached**: the deepest placement the search made, out of 256, the
study's measure of how far a variant got. Strict backtracking hits a wall in
the low 200s (row-major 208, the fastest strict variant 216); breaks push well
past it, to 243 to 245.
- **depth at timeout**: where the search frontier stood when the clock struck, so
a variant that never completes still records where it was working.
- **number of breaks**: the interior edges the search actually *broke* on the best
board. This is zero for a strict variant, and for a break variant it is the count
its own budget accounting committed rather than the score deficit. On a completed
board it equals `480 − score`; on a timed-out partial the deficit also counts
still-empty edges, so the study reports the true break count instead. Every
board's bucas URL makes both checkable in the [viewer](/viewer).
- **backtracks**: retreats out of a cell after its candidates are exhausted.
## The two baselines, and what specialisation costs
The rawest variant, `NAIVE-CLEAN`, is a readable general engine: sentinel-free
per-cell candidate lists indexed by the two already-placed neighbours, a resolved
edge cache so no rotation is recomputed on the hot path, and no allocation inside
the search. Its twin, `NAIVE-CODEGEN`, is the *same algorithm* re-expressed as a
hand-specialised, 16×16-only, row-major hot loop, kept as a separate program so
the general engine stays clean. Running them head to head prices the low-level
engineering: on this puzzle it buys a modest, instance-dependent throughput gain
and, notably, no better score. The numbers are on the
[findings page](/research/lab/experiments/raphael-anjou/dfs-study/findings).
## Soundness of the propagators
Forward-checking, arc-consistency and the per-colour supply check are sound *by
construction*. Each only ever rejects a state in which some cell already has an
empty domain, a cell that no unused piece can fill, so none of them can remove a
branch that leads to a real completion. The arc-consistency revise is deliberately
conservative where it is imprecise: wherever it might be unsure it prunes *less*
rather than more, staying on the safe side. The propagators are sound only under
strict placement, because under a break budget a local look-ahead can prune a
branch the global budget could still rescue, so the break variants deliberately
run no propagator. The registry enforces this: a variant that pairs breaks with a
strict-only propagator fails to build. The engine tests check that constraint and
the score and break bookkeeping, though the soundness of the prune itself rests on
the argument above rather than on a test.
## Reproducibility
The engine workspace, the ten variants, the committed per-run results and the
grid scripts all live under the study's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/dfs-study).
`just experiments dfs-study` rebuilds the engine and reruns the whole grid; the
run is deterministic at a fixed seed, and the corner arrangement is the only
diversity axis.
## Related
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [What each idea buys](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study/findings) — The three comparisons at the heart of the DFS study, worked through: fill order (row-major wins, a bad order is catastrophic), heuristics (MRV rescues border-first but costs throughput; more propagation did not help), and breaks (they smash the depth wall; the break schedule is the lever, not the per-cell cap).
- [Fill orders](https://eternity2.dev/research/build/backtracking/fill-order) — The order in which a backtracker visits the 256 cells is its one free choice: it costs nothing at runtime and moves the size of the search tree by orders of magnitude. Twenty years of community science, from the fixed-vs-dynamic wars and the strategy races to the magic 10×16 square and Verhaard's comb search, all answer the same question: which path through the board is cheapest?
---
# The reference engine behind this site
> The Rust-to-WebAssembly backtracker that runs every live demo and checks every number on this wiki. Not a record machine but a reference engine, ported four times and parity-tested byte-for-byte, built so the claims here can be re-run.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/engine/
- Updated: 2026-07-17
- Topics: backtracking, speed
- Source: The eternity2 repository on GitHub — https://github.com/raphael-anjou/eternity2
---
> **Whose work this is**
>
> This page is about the engine behind the site you are reading, written by the person who wrote it, [Raphaël Anjou](/research/people/raphael-anjou). It is not a community record solver: those are studied elsewhere in this lab ([Blackwood](/research/lab/experiments/joshua-blackwood/solver), [McGavin](/research/lab/experiments/peter-mcgavin/backtracker), [Verhaard](/research/lab/experiments/louis-verhaard/eii)). This one sits among them as a peer, and a modest one: the record solvers hold the records; this one holds the receipts.
Where the [named experiments](/research/lab/experiments/raphael-anjou) each ask a
question and the [shared engines](/research/lab/experiments/raphael-anjou/engines)
are the apparatus those studies run on, this page is a third thing: the small
reference engine that powers the site itself, checks the numbers the other pages
quote, and animates every live demo. It scores no records. Its job is
verifiability and teaching.
## What it is
A small Rust crate compiled to WebAssembly, running live in your browser on every
interactive page of this wiki. It implements the classics, plainly: the official
16×16 piece set, a generator that builds solvable puzzles of any size (with an
optional real-E2-style mode that restricts border colours to the frame band),
nine cell-visit orders, a strict depth-first backtracker, a scorer, and a
break-tolerant variant of the search, a reimplementation of the break-index idea
from [Blackwood's solver](/research/lab/experiments/joshua-blackwood/solver),
built so the labs here can demonstrate it.
One design choice matters more than the algorithms: the solver is a step-able
machine, not a recursive function. Callers run it one bounded step at a time, one
placement or one backtrack, and read the board between steps. That is what lets a
web page animate a real search instead of a canned recording:
[the watch page](/playground/watch) is stepping this exact engine, not a video of
it.
## One engine, many ports
The site runs one engine: the Rust/WASM crate, the canonical reference. But the
repository keeps a whole collection of faithful reimplementations of it in other
languages: a pure TypeScript port (zero WASM), a C port, a C++ port, and smaller
studies in Python, Lua, COBOL, even Brainfuck. Each is validated byte-for-byte
against golden data the Rust crate emits: generated puzzles down to the RNG
output, all nine fill paths at several sizes, and full solver runs with exact
node, attempt, and backtrack counts.
That discipline exists for one reason: an interactive demo you cannot cross-check
is just an animation. Two independent implementations that agree to the last
backtrack are much harder to get wrong the same way twice, and eight are harder
still. The ports are a study exhibit, not build options (the site always runs
Rust), and they live together in the repo's `engine-ports/` collection. Some of
them (the Brainfuck backtracker especially) are there for joy.
## What it does not claim
This is not a record machine, and it would be misleading to dress it as one. The
strict solver carries none of the hand-tuned quota schedules or restart strategies
that make [Blackwood's solver](/research/lab/experiments/joshua-blackwood/solver)
the engine behind the community's best boards. The best score produced by
[the experiments here](/research/lab/experiments) is 463 of 480; the community's
best on the same puzzle is 470. The record solvers studied in this lab are simply
better at finding boards.
Throughput is a separate axis, and one this reference engine deliberately does not
chase - but a sibling experiment does. The
[JIT backtracker](/research/lab/experiments/raphael-anjou/jit-backtracker) asks how
fast a *portable* Rust search can go, and reaches
[McGavin-class throughput](/research/lab/experiments/peter-mcgavin/backtracker) on the
hard, deep boards that resemble the real puzzle - a tie with his hand-tuned C on the
same machine (his C stays ~2.3× faster on easy boards). That is a speed result, not a
solving one: it still plateaus where every strict backtracker does. Speed and score are
[different axes](/research/lab/experiments/raphael-anjou/going-fast), and this page's
engine optimises for neither - only for being checkable.
This engine's job is different: verifiability and teaching. When this wiki states
a node count, a feasibility number, or a score, the claim is checked by this
engine, and because it runs in your browser and its source is public, you can
check it too.
## How it powers the wiki
Every interactive element in this research section is this engine: the live DFS
demos, the fill-order races on [the paths page](/playground/paths), the
break-index lab on [the solver hub](/research/build/solvers), the board viewer's
scoring and verification, and the committed reference counts that research pages
quote. The engine's own test suite cross-checks against real community boards, so
a change that broke scoring or rotation conventions would fail loudly rather than
silently corrupting the site's numbers.
## Where to get it
Everything is in one repository,
[github.com/raphael-anjou/eternity2](https://github.com/raphael-anjou/eternity2),
and [run it yourself](/research/build/run-it-yourself) walks through building the
engine, running its tests, and reproducing the published results command by
command.
## What drives it
The engine does not grow on its own schedule; it grows when
[an experiment](/research/lab/experiments) needs something. The break-tolerant
solver exists because demonstrating break indices required one; the
frame-restricted generator exists because a lab needed puzzles that behave like
the real E2 border. That keeps the engine small, and it keeps every feature
attached to a question someone actually asked.
## Related
- [The engines](https://eternity2.dev/research/lab/experiments/raphael-anjou/engines) — The shared engines under Raphaël Anjou's experiments. The named experiments are studies that run on these; this is the apparatus they share. The CSP presets and the Verhaard reimplementation are documented here; the constructive engines are not yet published.
- [Run it yourself](https://eternity2.dev/research/build/run-it-yourself) — The whole site, the engine, and every result in this section run from one repository. Here is how to get it going, rebuild the WebAssembly engine, and reproduce the numbers.
- [Blackwood's solver, decoded and run here](https://eternity2.dev/research/lab/experiments/joshua-blackwood/solver) — Joshua Blackwood's record backtracker, decoded through Jef Bucas's notes (a colour quota schedule and a late-game mismatch allowance, tuned near-optimally), then built and run on my M1: left as published it flies to 248 of 256 pieces ignoring the clues; pin the five official clues and the same engine stalls near 45.
---
# The engines
> The shared engines under Raphaël Anjou's experiments. The named experiments are studies that run on these; this is the apparatus they share. The CSP presets and the Verhaard reimplementation are documented here; the constructive engines are not yet published.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/engines/
- Updated: 2026-07-15
---
The [experiments](/research/lab/experiments/raphael-anjou) are studies: each
one asks a question and reports the board it reached. But most of them are not
built from nothing. They run on a small set of shared engines, and the same
engine turns up under experiment after experiment. This page documents that
apparatus once, so the studies can point here instead of re-explaining the
machine each time.
- [The CSP presets, measured](/research/lab/experiments/single-core-benchmark/csp-presets) — A single constraint-propagation engine with interchangeable orderings and propagators: arc-consistency, colour-graph matching, several fill orders. Faithful reimplementations of known community techniques, run under a dozen presets on the ten variants so the exact cost of each knob is on the leaderboard rather than argued about.
- [The Verhaard reimplementation](/research/lab/experiments/louis-verhaard/verhaard-reimpl) — A from-scratch reimplementation of Louis Verhaard's eii method, since his own binary ships no source and will not run here. Set-composition swap-annealing under the 2×2-tiling metric; on the real five-clue puzzle it reaches 438 of 480, single core.
The two constructive engines these studies also lean on, the beam producer and
the ALNS polish stage, do not yet have their own write-ups here; those pages are
held back for now, and neither is on the benchmark leaderboard. Where an
experiment steers one of them, its own page says what that engine does at the
level the study needs, so nothing below depends on reading a page that is not
here yet. The two engines documented and measured in full are the CSP presets
and the Verhaard reimplementation.
## Why document the engines apart from the experiments
A study and its engine answer different questions. The engine answers *how the
search works*, the data structures, the fill order, the propagators, and it
is reused unchanged across many studies. The study answers *what happens when
you point that engine at a particular idea*: a learned prior, a fixed border, a
break mask. Keeping them apart means a reader who wants to understand PRIOR's
prior does not have to re-read how a beam works, and a reader who wants the
engine gets it in one place, current and complete.
Where these engines race head to head under a fixed budget, they appear on the
[single-core benchmark](/research/lab/experiments/single-core-benchmark). The
general theory behind each method lives in the
[build-a-solver](/research/build) section; these pages are the specific engines
as built and measured here.
---
# Going fast: when a solver spends its budget on speed
> Some Eternity II engines pour their effort into walking the search tree as fast as possible; others spend it on judgement about where to walk. This is the case for the first kind - what raw throughput buys, the three different things people mean by "fast", and why the fastest engine ever built still cannot solve the puzzle.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/going-fast/
- Updated: 2026-07-20
- Topics: speed, backtracking
- Source: McGavin: even at 100M nodes/s, the 5-hint tree is 4.9×10³² years (groups.io msg 11201) — https://groups.io/g/eternity2/message/11201
- Source: Razvan's epitaph for the 2025 speed thread: 'we will not make a dent' (groups.io msg 11657) — https://groups.io/g/eternity2/message/11657
---
Every Eternity II solver has a fixed budget of human effort and machine time, and it
spends that budget in one of two places. It can spend it on **speed** - walking the
search tree as fast as the hardware allows, trying billions of placements a second -
or on **judgement** - being cleverer about *which* placements to try, so it walks a
smaller, better tree. This page is the case for the first kind: what going fast
actually buys, what it doesn't, and how to talk about it without fooling yourself.
It is also the conceptual home of a specific result: a
[portable Rust backtracker that ties McGavin's hand-tuned C on hard boards](/research/lab/experiments/raphael-anjou/jit-backtracker)
- the fastest engine in the community's history - and runs at about 44% of its speed
on easy ones. That page is the engineering diary; this one is what the diary *means*.
## Three things people mean by "fast"
The single biggest source of confusion in twenty years of speed talk is that "fast"
names three unrelated quantities. Keeping them apart is most of the battle.
> **The three axes, and why they don't convert**
>
> **1 · Placements per second (a.k.a. pieces/s, nodes/s).** How fast the search *walks*. This is an engine-*and-board* number: McGavin's C does ~287 M on an easy board but ~105 M on a hard one; the [JIT engine here](/research/lab/experiments/raphael-anjou/jit-backtracker) does ~122 M on the easy board and ~110 M on the hard one - tying the C exactly where the puzzle is hard. Bigger is faster, but only ever *on the same board*. **2 · Matched edges out of 480.** How *good* a board is. This is where [records](/research/records) live - the ceiling is 470. It has nothing to do with axis 1: a slow engine can find a better board than a fast one, and routinely does. **3 · Aggregate placements per second.** A *fleet* number - many machines summed. The community's "~300 M/s" figure that sometimes gets attached to a single engine is actually the [Eternity 2 Syndicate swarm](/research/build/faster/distributed-solving): ~20 machines added together, not one core. A high number on axis 1 tells you nothing about axis 2, and axis 3 is not an engine speed at all. Every speed claim worth trusting says which axis it is on.
For the precise, sourced history of how the community settled its definition of a
"node" - pieces *placed*, chess-style, and why even that flatters scan-line fill
orders - see the measurement-discipline section of the
[solver-engineering ledger](/research/build/faster/solver-engineering). This page
takes that vocabulary as given and asks what the speed is for.
## What speed buys
Real things, and it is worth being concrete, because the case *against* speed only
lands once you respect the case *for* it.
- **Verified enumeration.** A fast, deterministic engine can walk a benchmark tree to
the last node and *count* it, turning theory into checked fact. The community's
[benchmarks](/research/build/benchmarks) - census protocols, the 10×10 that finally
fell after ~180 core-years - are throughput victories. You cannot verify what you
cannot finish.
- **Determinism as a checksum.** Two fast engines that walk the same tree must report
the same node count. That equality is how ports, rewrites and new hardware prove
they search the same tree before their speed means anything - it is the backbone of
the [JIT engine's optimization ladder](/research/lab/experiments/raphael-anjou/jit-backtracker),
where every rung halts at the identical node count.
- **More attempts per second under a heuristic.** Speed is a multiplier on judgement:
a good repair loop that runs twice as fast gets twice as many shots at a better
board in the same wall-clock. Speed does not *replace* judgement, but it amplifies
whatever judgement you already have.
## What speed cannot buy
The wall. And no engine-builder has been clearer about this than the person who built
the fastest engine. On his best five-hint placement path - already a far smaller tree
than the raw puzzle - McGavin worked out that even at 100 million nodes per second the
search would take about **4.9 × 10³² years**, and that throwing billions of cores at
it still leaves you "orders of magnitude longer than the age of the Universe"
([msg 11201](https://groups.io/g/eternity2/message/11201)). When the community's 2025
speed thread wound down, Razvan wrote its epitaph: however fast we can check, "we will
not make a dent" in the space ([msg 11657](https://groups.io/g/eternity2/message/11657)).
The arithmetic is unforgiving and it is the site's [core lesson](/research/why/prune-vs-speed):
a constant-factor speed-up, however hard-won, is *multiplied against* a number so
large that no constant factor matters. Doubling the walking speed of a search that
would take 10³² years gives you a search that takes 5 × 10³¹ years. Shrinking the
*tree* is the only lever with exponents on it.
## The controlled experiment
This is exactly why the [JIT backtracker result](/research/lab/experiments/raphael-anjou/jit-backtracker)
is framed as a speed experiment and nothing more. It answers a clean, bounded
question - *can portable, safe Rust reach hand-tuned-C throughput?* - and the answer is
board-dependent: on hard, deep boards like the real puzzle it **ties** the C, walking
the identical tree; on easy low-branching boards the C is about 2.3× faster. It settles
a smaller open question the community had left standing: whether the famous ~4× edge of
McGavin's engine over "typical" solvers was **codegen craft** or merely **newer
hardware**. Holding the hardware fixed and reaching his speed *on hard boards* from
portable code shows it was craft - the same craft, reproducible in a safe language, and
written down rung by rung. (It also shows where the craft pays most: on easy boards,
where per-node work is nearly free, his tighter codegen still wins.)
And then it stops, plainly, at the same place every fast engine stops: pointed at the
real puzzle it plateaus in the high-300s of 480, because a strict backtracker is a
superb tree-walker and a poor solver. The records belong to engines that spend their
budget the *other* way - on
[destroy-and-repair judgement](/research/lab/experiments/raphael-anjou/repair-study)
and [what a search is allowed to learn](/research/lab/experiments/raphael-anjou/learning) -
and they get there at a small fraction of the speed.
That is the whole trade. Speed is real, learnable, and worth mastering; the JIT engine
is proof you can master it in a safe language. But on Eternity II, the fast lane and
the winning lane are not the same lane. Which is why the interesting engineering
question is never only "how fast" - it is "fast *at what*, and is that the thing in the
way?"
## Related
- [The JIT backtracker: portable Rust that ties hand-tuned C on hard boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/jit-backtracker) — A safe, portable Rust depth-first backtracker, specialised at runtime by emitting and compiling per-puzzle Rust, taken from 43 to 123 million search-nodes per second on one core. Measured fairly against Peter McGavin's C on the same machine: a tie on hard, deep boards like the real Eternity II, and about 44% of its speed on easy ones. Every rung searches the identical tree; the whole gain is code, not algorithm.
- [McGavin's C backtracker: the throughput story, built here](https://eternity2.dev/research/lab/experiments/peter-mcgavin/backtracker) — Peter McGavin's own C backtracker, the community's fastest: a 2007 optimization recipe compounded for two decades through generated code, lookup tables and counter tricks, then built on my M1 and pointed at the real 256-piece puzzle, where single core it drives past 200 of 256 pieces at ~109M placements/s.
- [Solver engineering: the craft below the algorithm](https://eternity2.dev/research/build/faster/solver-engineering) — Every record solver runs the same depth-first backtracker. What separates them is the layer underneath: lookup tables, perfect hashes, cache-sized structs, generated code, compiler archaeology. That craft decides whether a node costs 26 cycles or 2,600. The community's twenty-year engineering ledger, technique by technique, and what it all bought.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
---
# The hint study
> Give a backtracker five correct pieces for free, in the puzzle's own clue geometry. It turns out not to help, and depending on the fill order it can hurt badly, because a pinned piece is a hard constraint a fixed fill order must satisfy on arrival. A family of fill paths, run on the same hinted boards, single core, measured against no hints at all.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/
- Updated: 2026-07-21
- Topics: backtracking, search-space, structure
- Reproduce: `just experiments hint-study`
- Source: Runnable generator + committed results + scripts (this study's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/hint-study
---
There is a folk answer to how you make Eternity II easier: hand the solver some
correct pieces. The natural next question is *how many*, and the community's own
[hint-geometry discussion](/research/why/hint-geometry) already sharpened it to
*where*. This study asks something more basic that both questions skip past: on our
own generated boards, do the hints help *at all*?
For a chronological backtracker, the answer is **no**. Seed one with the five
official clue pieces at their real board positions, then measure it against the very
same board with no hints, and it does worse on every fill order tested, without
exception. On a compact row-major sweep the five correct pieces cost only ten to
twenty matched edges; on the fill order a newcomer would reach for first, "get to
the clues and connect them up", they cost around three hundred and forty-five,
turning one of the best no-hint orders into the worst. The reason is simple once
seen: to a solver that fills cells in a fixed order, a pinned piece is not free
information but a **hard constraint it must satisfy the moment it arrives**, and
sometimes it cannot.
## See it: five hints, four orders
Each board below fills along a different order, on a loop. Nothing is being *solved*;
this is only the order the search would visit cells, made visible. The bright cell is
the one just placed, and the trail behind it is the recent **wavefront**, so you can
see how much boundary each order keeps open as it runs. That open boundary, the
frontier, is what a backtracker pays for: its branching grows with the frontier, and
a small frontier is also what leaves the search room to route around a hostile pin.
> **[Interactive: HintPathFill]** Rendered on the canonical page (link above); not shown in this markdown export.
A row-major sweep keeps one thin frontier and rolls it down the board, so when it
meets a pinned piece it can adjust the single row it is building. The hint-seeking
orders do the opposite. Spiralling in or out drags a whole ring as its frontier, and
tracing the hints scatters its wavefront across the board from the very first move,
committing everywhere before it can know whether the commitments are consistent. The
compact sweep survives the five hints; the hint-seeking orders are undone by them.
## What the study measures
Two paired comparisons, run on the same generated boards:
- **Do the hints help, and which order survives them?** Fix the hints (five, in the
shape of Eternity II's own five clues) and vary only the fill order, measured
against the identical board with no hints. This is the study's spine: it shows the
hints never help, and that how much they hurt is set by the fill order and its
open frontier.
- **Count.** Does adding *more* hints help? Some, but the naive way of measuring it
is *confounded*. A block of clustered hints banks a pile of correct edges for free
just by being pinned; that free "floor" flatters clustered layouts on raw score
while saying nothing about whether the board got easier to *finish*. The
[method page](/research/lab/experiments/raphael-anjou/hint-study/method) defines the
floor and the floor-immune metrics (solved-rate and earned score) that see past it.
Everything here is built and measured from scratch: our own parametric board
generator, our own family of backtrackers, our own canonical scorer, and a beam
solver as a non-backtracker contrast. No community board, puzzle, or engine is used;
only the *shape* of the five-clue arrangement is borrowed from the list discussion,
as a geometry to test.
## How to read the numbers
Every board is re-scored by one canonical matched-edge scorer that never counts a
border-facing (grey) seam, the same convention as the
[benchmark](/research/lab/experiments/single-core-benchmark) and the
[DFS study](/research/lab/experiments/raphael-anjou/dfs-study). The maximum is 480.
Each of the fifteen seeds is a genuinely distinct generated instance, not merely a
different solver seed, so the spread across seeds is real instance-to-instance
variance and is reported as such. No self-reported score is trusted.
The comparisons are worked through on the
[findings page](/research/lab/experiments/raphael-anjou/hint-study/findings); how
the boards are generated, why the colour recipe stays faithful across sizes, and
what every metric means are on the
[method page](/research/lab/experiments/raphael-anjou/hint-study/method).
## Pages in this section
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/method) — The apparatus behind the hint study: a parametric board generator faithful to Eternity II's colour recipe at every size, the family of fill-path backtrackers, the one canonical scorer, and the piece of arithmetic that keeps the count axis meaningful, the pinned-seam floor.
- [What the study found](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/findings) — The results, worked through: on these boards the five clue-shaped hints never help a backtracker, they range from a mild cost to a catastrophe, and the fill order decides how much damage they do; the scores are bimodal, not a smooth gradient; and the hint-count question is confounded by a free pinned-seam floor.
## Related
- [Raphaël Anjou's experiments](https://eternity2.dev/research/lab/experiments/raphael-anjou) — A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [Where you place the hints beats how many](https://eternity2.dev/research/why/hint-geometry) — On a 16×16 puzzle built like Eternity II, eighteen hints scattered across the board solve it in minutes, while the same puzzle needs eighty or more hints piled into contiguous rows to be as easy. Position, not count, is the lever, and it points straight at the endgame.
---
# What the study found
> The results, worked through: on these boards the five clue-shaped hints never help a backtracker, they range from a mild cost to a catastrophe, and the fill order decides how much damage they do; the scores are bimodal, not a smooth gradient; and the hint-count question is confounded by a free pinned-seam floor.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/findings/
- Updated: 2026-07-21
- Topics: backtracking, structure, search-space
- Source: Committed per-run results (results.jsonl) and the analysis script — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/hint-study
---
The study fixes five hints in the shape of Eternity II's own five clues, runs eight
fill orders on the identical boards, and asks a simple question: what are those five
hints worth? Every number below is over fifteen distinct generated instances,
single core, eight seconds a run, re-scored out of 480 by one canonical
scorer. The comparison is paired, every fill order sees the same boards, and the
per-instance spread is shown rather than averaged away, because it turns out to
matter more than any median.
## The hints don't help, and the fill order decides how much they hurt
Start with the distribution. Each dot is one board; the tick is the median.
> **[Interactive: HintStudyCharts]** Rendered on the canonical page (link above); not shown in this markdown export.
Two things are visible at once. First, the scores are **bimodal**: on most orders a
board either climbs into the 360s or stalls in the double digits, with little in
between. A
median drawn through that is a summary of a gap, not a centre, which is exactly why
the dots are shown. Second, the orders split hard: the three compact sweeps
(row-major, its bottom-up mirror, Verhaard's comb) sit high on most boards; the
fragmenting orders sit low throughout.
But the compact-beats-fragmenting split is not the real finding, because it invites
the wrong question, "which order best *uses* the hints?" The question that matters
is whether the hints help *at all*. The second chart answers it, and the answer is no.
It shows the **paired change** from adding the five hints: each order's score on a
board minus its score on the *same board with no hints*.
Every bar is at or below zero. The five clue-shaped hints do not help a single fill
order. On the compact sweeps they cost only ten to twenty points. On spiral-out they
cost about ninety. And on the two hint-seeking orders they are ruinous: the
`trace-hints` order, which draws a skeleton between the clues before filling, loses
about three hundred and twenty-five, and the `connect-hints-first` flood loses
roughly **three hundred and forty-five**, taking an order that scores among the
*best* of all eight with no hints to the *worst* with them. The more deliberately an
order chases the hints, the more they cost it. Handing the solver five correct
pieces, in the puzzle's own clue geometry, made every version of it worse.
## Why a correct hint hurts
A pinned piece is not free information to a chronological backtracker; it is a **hard
constraint the fixed fill order must satisfy on arrival**. When a compact sweep rolls
down to a pinned interior cell, the piece is already there, and the row it just built
has to match that piece's edges. Most of the time it can, at a small cost: the
sweep steers around the constraint and loses a few tens of points. But on some boards
the pinned piece contradicts what the frontier has committed to, and there is no
local repair: the search hits a wall it cannot pass and thrashes below it. That is
the stalled mode, and it is the hints that create it. On row-major, the boards that
collapse into the double digits are precisely the ones where a clue-shape pin lands
where the sweep cannot honour it; the same board with no pins climbs into the 360s
and 370s.
This reframes the fill-order result. The order still matters (a compact sweep
survives the hint constraints with a scar of ten to twenty points while
`connect-hints-first` is destroyed by them), but what the order is buying is not
"using the hints well." It
is **surviving them**. The frontier is why: an order that keeps a single tight
frontier has room to route around a bad pin; an order that has already fragmented
into five open blobs has committed everywhere at once and cannot.
That frontier relationship, across the eight orders, is strikingly clean, and it is
worth showing precisely because the frontier can be computed from an order's geometry
with no solver at all, then checked against the measured scores:
> **[Interactive: FrontierPhaseChart]** Rendered on the canonical page (link above); not shown in this markdown export.
The average open frontier an order holds predicts its median score closely across
these eight orders. It is a strong descriptive relationship, not a law proved on
eight points, and it stops settling the ranking among the fragmenting orders on the
right (spiral-in holds a larger frontier than spiral-out yet scores higher). But the
direction is exactly what the mechanism predicts: branching cost is multiplicative in
the frontier, so an order that keeps the frontier small keeps the room to absorb a
hostile pin.
For contrast, the beam solver, which is not a chronological backtracker and does not
pay the frontier cost the same way, reaches a median in the 450s on these same
hinted boards. The hints and boards are nowhere near unsolvable. It is specifically
the *chronological, fixed-order* backtracker that cannot turn five correct pieces
into progress.
## Count: a threshold, not a gradient, and a floor that hides it
The same question one level out: does adding *more* hints help? The answer is not a
smooth "more is better", and the raw number hides which part is real.
The lower chart above splits each layout's score into two parts. The **floor** is
the seams the pins complete for free, because both their endpoints are pinned to the
true solution; the **earned** part is what the search actually found. A solid
clustered block banks a tall floor (five 4×4 blocks pin a quarter of the whole
board's seams before the search takes a single step), while a spread lattice, whose
hints never touch, banks nothing. So a raw-score comparison hands clustered layouts
a hundred-point head start that says nothing about whether the search made headway.
Read the earned column across the spread lattices, whose floor is zero so earned
*is* the score, and a threshold appears. A sparse spread, four to sixteen hints,
earns almost nothing (in the twenties): the pins are just scattered constraints the
sweep keeps tripping over, exactly the five-clue effect. But keep adding them and
the picture flips. Twenty-five spread hints earn 127, and thirty-six, a six-per-line
lattice, **solve the board outright on most instances**. Below the threshold the
spread hints only get in the way; above it there are finally enough of them to
carve the board into pieces small enough for the sweep to finish. It is not a
gradient of help, but a wall the count has to clear.
The clustered blocks show the mirror image. Their raw score is mostly *floor*: five
4×4 blocks (eighty hints, a quarter of the board) do solve every instance, but they
have pinned so much of the board that they have half-solved it by hand. Strip the
floor away and the clustered layouts short of that extreme earn only double digits,
coasting on the free seams. So the real count story is not "more hints help" or
"more hints hurt", but: it takes a *lot* of correct hints, spread or clustered, to
move a chronological backtracker at all, and until you reach that amount the extra
pins are as likely to trip the search as to speed it.
The [method page](/research/lab/experiments/raphael-anjou/hint-study/method) works
the floor arithmetic through in full.
## Scattered versus contiguous: we get the opposite
The community's
[hint-geometry write-up](/research/why/hint-geometry) makes a sharper version of the
placement claim: eighteen hints *scattered* on a lattice solve a 16×16 E2-like
puzzle in minutes, while eighteen piled into *contiguous* top rows barely help, and
you need eighty or more contiguous hints to match the scattered eighteen. That
result was measured on one specific puzzle with one specific engine. We put its two
exact layouts, the scattered rows-{1,3,5} lattice and the eighteen-cell top block,
on our own generated boards and solvers to see whether the direction holds.
> **[Interactive: HintGeoComparison]** Rendered on the canonical page (link above); not shown in this markdown export.
It does not. On our boards the **contiguous** eighteen score *higher* than the
scattered eighteen for seven of the eight fill orders, and by a wide margin: a
row-major sweep reaches a median of 377 with the contiguous block and only 138 with
the scattered lattice. This looks like a flat contradiction, and it is worth being
precise about why it is not quite one.
The two studies measure different things. The community result is about *time to a
full solution*: scattered hints reach into the deep endgame where a backtracker
spends almost all of its time, so they prune the expensive part, while a contiguous
top block prunes only the cheap opening. Our number is *matched-edge score at a
short budget*, and at eight seconds none of these strict backtrackers reaches the
endgame at all. What a top-anchored contiguous block does buy, immediately, is a
large correct region for the row-major sweep to build against, so the score climbs
fast even though the hard part of the board is untouched. Scattered hints, by
contrast, fragment the early fill exactly the way the five-clue shape did. So the
two results are consistent once you separate "solves the whole board eventually"
from "scores well in the first eight seconds": scattered placement helps the former
and hurts the latter. The next section makes that split visible.
## Solving it: where scattered finally wins
At 16×16 nothing solves in eight seconds, so the score is always a snapshot of a
search still in its opening. To see the *endgame* effect the community reported, we
need a board that fully solves. An 8×8 built to the same colour recipe does, in well
under a second, so on it we can measure the quantity that actually matters: the
number of search nodes a row-major backtracker needs to reach a complete solution. Fewer nodes means
the hints did real pruning work. Below, a spread lattice and a matched contiguous
block, at rising hint counts.
> **[Interactive: SolveSpeedChart]** Rendered on the canonical page (link above); not shown in this markdown export.
This is the whole story in one chart, and it finally lines up with the community
claim. At four hints the scattered lattice is *worse* than useless, solving only two
of thirty boards, the same sparse-hints-trip-the-search effect from every other
axis, while the matched contiguous block solves almost all of them. But cross a
threshold around sixteen hints and the lines swap hard: a sixteen-hint scattered
lattice solves in about four thousand nodes, while the matched contiguous block
still grinds through nearly two million, a five-hundred-fold difference at the same
hint count. Push further and both layouts become easy (by thirty-six hints a quarter
of the board is pinned and either one solves in a few hundred nodes), so the
scattered advantage is a window, widest where the count is enough to reach the deep
search but not so large that the board is half-solved by hand. In that window,
spread hints reach into the part of the search a backtracker actually struggles with
and cut it off, exactly as Joe and Peter McGavin described; a contiguous block only
ever prunes the easy opening. The 16×16 score comparison looked like it contradicted
them only because at a short budget the search never lives long enough to reach the
region where scattered placement pays off.
## What it does and does not say
This is a result about **strict, chronological depth-first backtrackers** on
generated 16×16 boards built to Eternity II's colour recipe, at a short (eight
second) budget. Each of those scoping words earns its place. *Chronological*: the
result is specifically about solvers that fill cells in a fixed order and must
satisfy a pinned piece when they reach it, a beam solver, which does not, reaches
the 450s on the same boards. *Generated*: the boards share E2's colour recipe and
its five-clue *shape*, but they are not the official puzzle, and the study transfers
no number to it. *Short budget*: none of these strict backtrackers solves within
eight seconds, so this measures how far each reaches, not a race to 480; whether the
"hints hurt" effect survives at much longer budgets is untested here.
Within that scope the lesson is sturdy and, we think, counterintuitive: for a
chronological backtracker, five correct pieces placed in the puzzle's own clue
geometry are not a gift but a constraint, and can cost far more than they give. What
decides the damage is not the hints but the fill order that has to live with them
and the order pays for a hint the way it pays for everything else, in the size of the
frontier it keeps open. It is one more reading of
[why the puzzle resists](/research/why/hint-geometry): even correct information helps
only a solver built to receive it.
## Related
- [The hint study](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study) — Give a backtracker five correct pieces for free, in the puzzle's own clue geometry. It turns out not to help, and depending on the fill order it can hurt badly, because a pinned piece is a hard constraint a fixed fill order must satisfy on arrival. A family of fill paths, run on the same hinted boards, single core, measured against no hints at all.
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/method) — The apparatus behind the hint study: a parametric board generator faithful to Eternity II's colour recipe at every size, the family of fill-path backtrackers, the one canonical scorer, and the piece of arithmetic that keeps the count axis meaningful, the pinned-seam floor.
- [Where you place the hints beats how many](https://eternity2.dev/research/why/hint-geometry) — On a 16×16 puzzle built like Eternity II, eighteen hints scattered across the board solve it in minutes, while the same puzzle needs eighty or more hints piled into contiguous rows to be as easy. Position, not count, is the lever, and it points straight at the endgame.
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
---
# How the study is built
> The apparatus behind the hint study: a parametric board generator faithful to Eternity II's colour recipe at every size, the family of fill-path backtrackers, the one canonical scorer, and the piece of arithmetic that keeps the count axis meaningful, the pinned-seam floor.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/method/
- Updated: 2026-07-21
- Topics: backtracking, structure, search-space
- Source: The parametric generator and the fill-path engine (this study's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/hint-study
---
This page is the apparatus behind the [hint study](/research/lab/experiments/raphael-anjou/hint-study):
how the boards are generated, why they stay faithful to Eternity II across sizes,
the family of fill paths, and the piece that matters most for reading the
results correctly, the arithmetic of the *pinned-seam floor*, which is what
separates a real effect from a measurement artifact on the count axis.
## The boards: a faithful, size-parametric generator
Every instance is generated from scratch, seeded, and deterministic: the same
`(size, colours, seed)` produces the same board on any machine. A generated board
is a *solved* board, piece $i$ sits at cell $i$ at rotation zero, whose piece
IDs are then relabelled by a seeded permutation, so a hint for cell `pos` pins the
(relabelled) true piece, and a solver cannot simply walk the identity placement.
The colour recipe mirrors the official puzzle's structure exactly. On an $n\times n$
board there are
$$
E(n) \;=\; 2\,n\,(n-1)
$$
interior seams (the matched edges; $E(16) = 480$). These split into a **frame band**,
the seams joining two border pieces along the rim, and the deep interior. The
five *border colours* are confined to the frame band and never appear in the
interior; the *interior colours* appear both in the deep interior and on the
inward-facing edge of border pieces. This is the defining structural fact of a real
Eternity II board, and the generator reproduces it and is tested against it.
### Why the census is automatically balanced (a claim to state carefully)
It is tempting, and the earlier write-ups did this, to present Eternity II's
*even* colour census as a specially tuned gift: every colour appears an even number
of times, so a perfect solution's $\sum_c N_c / 2$ matches land with zero slack.
The even parity is real, but it is **not** a tuning achievement. It is forced.
A colour painted on $k$ interior seams appears on exactly $2k$ piece-edges, one on
each side of every seam. So for every colour $c$,
$$
N_c \;=\; 2\,k_c \quad\text{is even, for any seam colouring whatsoever.}
$$
Zero-slack parity is therefore automatic for *any* board built by colouring seams;
it says nothing special about Eternity II. The properties that are genuinely
load-bearing, and that the generator must get right, are three: border colours
confined to the frame band, the per-colour counts kept **balanced** (so no colour is
rare enough to over-constrain), and every piece **distinct up to rotation** (so a
pinned hint names a unique piece). The study's write-up is precise about this where
the earlier framing was not.
### Scaling the recipe without changing the difficulty
The generator and the solver are both size-parametric, the board can be $8\times8$
or $12\times12$ as readily as $16\times16$, which opens a natural follow-up: does the
placement effect *strengthen* as the board grows? Answering that cleanly needs a
colour recipe that does not change the puzzle's *difficulty* as the size changes.
Simply keeping the colour *counts* fixed while growing $n$ would make the puzzle
structurally *easier* at large $n$: with more seams and the same palette, each colour
repeats more often, so the average cell accepts more neighbours and the constraint
loosens. That would confound size with difficulty.
Instead the recipe holds the **per-colour multiplicity** roughly constant. Writing
$F(n)$ for the frame-band seam count and $E(n) - F(n)$ for the interior, the number
of border and interior colours is chosen as
$$
b(n) \;=\; \operatorname{round}\!\Big(\tfrac{F(n)}{12}\Big), \qquad
i(n) \;=\; \operatorname{round}\!\Big(\tfrac{E(n) - F(n)}{24}\Big),
$$
targeting the multiplicities Eternity II itself uses at $n=16$ (border $\approx 12$,
interior $\approx 24$). At $n=16$ this returns exactly the official recipe, five
border colours and seventeen interior.
Getting this right on small boards took one fix to the generator. Eternity II's
palette is *interior-dominant*, five border colours to seventeen interior, but the
default generator caps the border count at five and takes everything else as
interior, which on a small board inverts the ratio: at $8\times8$ the recipe wants
eight colours, and the cap would split them five border to one interior, a
near-uniform interior sea that behaves nothing like E2. The generator now accepts an
explicit border-colour count, and the recipe holds the interior at roughly three
times the border at every size ($8\times8 \to$ two border, six interior;
$16\times16 \to$ five, seventeen, unchanged). With that, small boards are faithful
*and* fully solvable, which is what the solve-speed comparison on the
[findings page](/research/lab/experiments/raphael-anjou/hint-study/findings) relies
on. The main path and count results on this page are all at $16\times16$; the
$8\times8$ board is used only where a full solve is needed.
## The hint geometries
Each layout is a pure function of the board size, so the same geometry can be drawn,
measured, and scaled consistently. The gallery below renders them all from the one
shared board primitive.
> **[Interactive: HintLayoutGallery]** Rendered on the canonical page (link above); not shown in this markdown export.
## The pinned-seam floor: keeping the count axis meaningful
Here is the subtlety that reshaped the study. Ask "do more hints help?" and the
obvious move is to compare final scores at different hint counts. But a hint does two
different things at once, and score conflates them:
1. it **removes a piece** from the search (the useful part, it prunes the tree);
2. it **may complete a seam for free**, if a neighbouring cell is also pinned.
The second effect is a pure bookkeeping gift. Define the **pinned-seam floor** of a
layout as the number of interior seams with *both* endpoints pinned:
$$
\text{floor} \;=\; \#\{\, \text{interior seams } (u,v) : u \text{ and } v \text{ both hinted} \,\}.
$$
Because the pins are true-solution pieces, every such seam is guaranteed correct
before the solver runs. A solid $k\times k$ clustered block contributes $2k(k-1)$ of
them; five $k{=}4$ blocks bank $5 \cdot 24 = 120$ correct seams, a **quarter of the
whole 480**, for free. A spread lattice, whose hints never touch, has a floor of
**zero**.
So a raw-score comparison systematically **flatters clustered layouts**: they start
a hundred-plus points ahead on bookkeeping alone, regardless of whether the board
became any easier to *finish*. This is the same family of error as counting rim
seams in a partial board, a floor that inflates the number without reflecting
progress. The toggle in the gallery above draws these banked seams so the free score
is visible.
The study therefore does not rank layouts by raw score on the count axis. It uses two
floor-immune metrics:
- **solved-rate**, the fraction of instances a path actually completes to 480;
- **reached depth**, how far past the pinned cells the search got, out of 256.
Both measure whether the search made *progress the pins did not hand it*. On the
path axis, where every compared layout shares the *same* hints and hence the same
floor, raw score is directly comparable and is used.
## Measuring against no hints, paired per instance
The question "what are the hints worth?" only has an answer relative to *not* having
them. So the path axis is run twice on every board: once with the five clue-shape
hints, once with none (`baseline_00`), and the reported effect is the **paired
difference**, the hinted score minus the no-hint score on the *same* generated
board. Pairing per instance removes the board-to-board difficulty variance, which
on these bimodal boards is large enough to swamp the effect if the two conditions
were compared across different seeds. A negative paired difference means the hints
made that fill order worse than it was with a blank interior, which is what the
findings report. All comparisons use the common set of seeds that ran to completion,
so every path and layout is aggregated over the identical instances.
## The fill paths, and why the frontier is the lever
The engine is the study's sibling [DFS backtracker](/research/lab/experiments/raphael-anjou/dfs-study),
run strict (no breaks, no propagation) so that the **fill order is the only thing
changing**. The orders tested are row-major, its bottom-up mirror, spiral-in,
spiral-out, border-first, Verhaard's comb, a clue-rows-first control, and the
study's own hint-seeking order, `connect-hints-first`.
Why does the order matter so much? A backtracker's cost is governed by the **open
frontier**: the set of already-filled cells still adjacent to an empty one. When the
next cell is placed against a frontier of size $f$, the number of partial boards the
search may have to consider grows multiplicatively in $f$, branching is exponential
in the frontier, not in the board. A single compact sweep keeps $f$ to about one row
($\approx n$); an order that opens blobs around $k$ scattered hints runs $k$ frontiers
at once, and
$$
\text{work} \;\sim\; \prod_{j} b^{\,f_j} \;=\; b^{\sum_j f_j},
$$
so fragmenting the fill into disconnected regions multiplies, not adds, the cost.
This is exactly why `connect-hints-first`, the order that *seeks* the hints, is the
worst performer: reaching the hints early is worth far less than keeping the frontier
small, and connecting scattered anchors does the opposite of keeping it small.
## How to read the numbers
Every board is re-scored by one canonical matched-edge scorer that never counts a
border-facing (grey) seam. The maximum is $E(n)$ ($480$ at $16\times16$). Each of the
fifteen seeds is a distinct generated instance, so across-seed spread is genuine
instance variance. Throughput, where reported, is search-nodes per second and is
never compared across different path orders, since a node under one order is not the
same unit of work as under another. The whole apparatus, generator, layouts,
per-run results, and the grid script, lives in the study's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/hint-study),
and `just experiments hint-study` reruns it.
## Related
- [The hint study](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study) — Give a backtracker five correct pieces for free, in the puzzle's own clue geometry. It turns out not to help, and depending on the fill order it can hurt badly, because a pinned piece is a hard constraint a fixed fill order must satisfy on arrival. A family of fill paths, run on the same hinted boards, single core, measured against no hints at all.
- [What the study found](https://eternity2.dev/research/lab/experiments/raphael-anjou/hint-study/findings) — The results, worked through: on these boards the five clue-shaped hints never help a backtracker, they range from a mild cost to a catastrophe, and the fill order decides how much damage they do; the scores are bimodal, not a smooth gradient; and the hint-count question is confounded by a free pinned-seam floor.
- [Where you place the hints beats how many](https://eternity2.dev/research/why/hint-geometry) — On a 16×16 puzzle built like Eternity II, eighteen hints scattered across the board solve it in minutes, while the same puzzle needs eighty or more hints piled into contiguous rows to be as easy. Position, not count, is the lever, and it points straight at the endgame.
---
# The JIT backtracker: portable Rust that ties hand-tuned C on hard boards
> A safe, portable Rust depth-first backtracker, specialised at runtime by emitting and compiling per-puzzle Rust, taken from 43 to 123 million search-nodes per second on one core. Measured fairly against Peter McGavin's C on the same machine: a tie on hard, deep boards like the real Eternity II, and about 44% of its speed on easy ones. Every rung searches the identical tree; the whole gain is code, not algorithm.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/jit-backtracker/
- Updated: 2026-07-20
- Topics: speed, backtracking
- Reproduce: `cargo run --release -p dfs-codegen --bin run_dfs_codegen_jit -- --puzzle P.json --chain2 --opt native`
- Source: Peter McGavin's genbody71.zip — the C reference this engine chases (groups.io msg 11749) — https://groups.io/g/eternity2/message/11749
- Source: rust-lang/rust#80630 — LLVM cannot lower loop+match to a computed goto (why the function-chain design exists) — https://github.com/rust-lang/rust/issues/80630
- Source: The eternity2 repository on GitHub — https://github.com/raphael-anjou/eternity2
---
> **What this experiment is, and is not**
>
> This is a **speed** experiment, not a solving one. It asks a single question: can a *safe, portable* Rust backtracker reach the throughput of the community's fastest engine - [Peter McGavin's hand-tuned C](/research/lab/experiments/peter-mcgavin/backtracker) - without leaving Rust or dropping to hand assembly? The answer turns out to depend on the board: on **hard, deep boards like the real Eternity II it ties his C**; on easy low-branching boards his C is about **2.3× faster**. It says nothing about *scores*: a fast walk of the tree and a high-scoring board are [different axes entirely](/research/lab/experiments/raphael-anjou/going-fast). The best this engine reaches on the real puzzle is in the high-300s of 480, exactly what a strict backtracker should reach - the record boards come from metaheuristics, not from walking faster.
Every fast Eternity II engine is, underneath, the same depth-first backtracker:
fill cells in a fixed order, try each piece that fits, recurse, back out on a dead
end. [McGavin's page](/research/lab/experiments/peter-mcgavin/backtracker) tells
the throughput story of the C engine that has held the community's single-core
speed crown for years. This page is the other side of that story: how close
*portable Rust* can get to it, and where the line really falls.
The result is board-dependent, and that turns out to be the interesting part.
Measured on my Apple M1, single core, both engines built headless with native
codegen and run back to back:
| Board | McGavin's C (headless) | This engine | Result |
| --- | --- | --- | --- |
| Easy (Joe's 18-hint 71.puz) | ~287 M/s | ~122 M/s | McGavin **~2.3×** |
| **Hard (deep 16×16, like real E2)** | **~102–105 M/s** | **~104–111 M/s** | **~tie** |
On the boards that actually resemble Eternity II - deep, densely constrained, where
the search spends its time backtracking - a safe, portable Rust engine **matches
hand-tuned C**. On easy, low-branching boards, where there is almost nothing to do
per node, McGavin's straight-line generated C is more than twice as fast. Both
numbers are real; the engineering below is what pulled Rust up to the tie on hard
boards.
Everything that follows is **code generation and data layout**, not a smarter
search: every rung walks the *identical* tree and, on a solvable test board, halts
at the *identical node count*. That invariant is the backbone of the whole
experiment, so it is worth stating first.
> **Measure the reference at full speed, or you will fool yourself**
>
> An earlier version of this work reported that we *beat* McGavin outright. That was wrong, and the reason is instructive. McGavin's engine ships with a live terminal display (`#define INTERACTIVE`): a status readout it redraws constantly. That display costs it about **2.7× of its throughput** - his headless speed is ~287M on the easy board, but with the display on it reads ~106M. The first comparison pitted our headless engine against his *throttled* one and produced a phantom win. Rebuilt headless (`-mcpu=native`, display off), the true picture is the table above: he wins the easy board comfortably, and we tie on hard ones. The lesson is general - always build the reference the way it runs at full speed before trusting any ratio.
## The rule of the game: same tree, every rung
A depth-first backtracker is deterministic. Given a board and a fixed fill order,
it visits exactly one sequence of nodes, on any machine, in any language. So there
is a hard test for whether an "optimization" is really just a speed-up and not a
silent change to the search: **the node count must not move.**
Throughout this work the oracle was a 60-hint solvable 16×16. Every version of the
engine solves it to 480 and reports **251,815 nodes** - the same number, to the
digit, from the slowest baseline to the fastest fused build. Any change that moved
that number was a bug in disguise and was reverted. That single discipline is what
lets the speed ladder below be read as an apples-to-apples comparison rather than a
collection of differently-behaving programs.
> **Correctness is a baseline invariant here, not a milestone**
>
> The engine checks all four edges of every placement against every already-fixed neighbour (placed piece or the frame rim), and it forbids a border/grey edge from facing into the interior. Those are not optimizations that were "added later" - they are the definition of a legal Eternity II placement, present in every version on this page. The ladder varies *only* how fast a fixed, correct search runs.
## The core idea: generate the program, don't interpret the puzzle
A generic backtracker pays, at every single node, for questions whose answers never
change during a run: *which cell am I on? where are its neighbours? which candidate
pool do I read?* A table-driven loop looks those up from arrays, per node, forever.
McGavin's C answers them **once, at build time**, by generating a program
specialised to one puzzle: `genbody -DG` writes a second C file with one
straight-line block per cell, neighbour addresses baked in as constants, and `goto`
chains threading the blocks together. That specialisation is the source of its
speed - same algorithm, run near the metal.
This engine does the same thing in Rust, at runtime:
1. `emit_program(&instance)` writes a **self-contained, puzzle-specialised Rust
program** - pieces, candidate buckets, pins and fill order all baked in as
`const` data, no external crates.
2. The wrapper compiles it with `rustc -O` (about 0.15 s for the plain build,
~1.7 s for the fused one).
3. The generated binary runs the search and prints its result.
It is McGavin's `genbody -DG → compile → run` flow, in Rust, invoked as a library.
Nothing exotic - just moving the puzzle's fixed facts out of the hot loop and into
the compiler's hands. Everything after this is squeezing constant factors out of
the generated code, and each squeeze is a small, self-contained change best shown
as a diff.
## The ladder, one change at a time
The diffs below are **simplified for reading**: the real engine emits its inner loop
as machine-generated Rust (constants like a cell's position and neighbours are baked
per cell, and the source is a few hundred lines per puzzle). Each diff shows the idea
of the change, not the literal generated text - run the engine with `--emit-src out.rs`
to see the actual code for a given board.
### 1 · Pack each candidate into one word
The first generated loop still chased a pointer: read a `u32` index, follow it into
an `oriented[]` array for the piece's `(id, rotation, edges)`, *then* test whether
the piece was already used. Three dependent loads to consider one candidate.
```diff
- let idx = pool[i]; // load an index …
- let (pid, rot, edges) = oriented[idx as usize]; // … chase it into a second array …
- if !used[pid] { /* consider */ } // … then test usage
+ let cand = pool[i]; // one contiguous load: pid<<48 | rot<<40 | edges<<8
+ let pid = (cand >> 48) as usize;
+ if free[pid] != 0 { let edges = (cand >> 8) as u32; /* consider */ }
```
Each candidate becomes a single packed `u64` stored directly in its `(up, left)`
bucket. The hot loop does **one** load, extracts the piece id, and tests usage
*before* unpacking the edges - McGavin's "check `tileFree` first" pattern. **+12 %**,
and it set up the packed representation the rest of the work depends on.
### 2 · Store board cells as one `u32`, not four bytes
The board stored each cell's four edges as `[u8; 4]`. Reading a neighbour's edge to
match against meant four separate byte loads. McGavin keeps each placed tile's edges
as **one `u32`** and pushes them to neighbours with shifts.
```diff
- let cell: [[u8; 4]; N]; // four byte loads to read one neighbour
- let up_edge = cell[up_pos][2]; // …and index arithmetic each time
+ let cell: [u32; N]; // one u32 per cell, URDL packed, empty = 0xFFFF_FFFF
+ let up_edge = (cell[up_pos] >> 8) as u8; // one load + one shift
```
This was the single biggest data-layout lever: **34 → 57 M nodes/s, +67 %**. Node
count unchanged. Representing the hottest data structure well mattered more than any
micro-optimization that came after it.
### 3 · Emit one function per cell - the pivot
Here is the move that made "match McGavin" plausible. The remaining tax was the
data-driven loop itself: `free_order[level]`, `cursor[level]`, `score_at[level]` -
indexed array loads *every node* for values McGavin has as compile-time constants.
The obvious way to bake them - one giant `loop { match level { …256 arms… } }` -
does not work: `rustc` takes over a minute to compile one enormous function, and
[LLVM cannot lower a `loop`/`match` to a computed goto](https://github.com/rust-lang/rust/issues/80630)
anyway, so it would not even reproduce McGavin's `goto` structure.
The way that *does* work is to emit **one small `#[inline(never)]` function per
cell**:
```diff
- // one generic loop, indexing arrays by depth on every node
- loop {
- let pos = free_order[level];
- let (up_pos, left_pos) = neigh[level];
- // …scan, place, advance level, or back out…
- }
+ // one function per cell; its position and neighbours are baked constants
+ fn cell_37(st: &mut St, left_arg: u8) -> bool {
+ const POS: usize = 138; const UP: usize = 122; // this cell's facts, as constants
+ for cand in POOL_UL[/* up*COLORS+left */] { // its exact candidate pool
+ // place …
+ if cell_38(st, right_edge) { return true; } // advance = call the next cell
+ // unplace …
+ }
+ false // exhausted = plain return (backtrack)
+ }
```
Advancing is a call to the next cell's function; backtracking is a bare `return`.
Roughly 256 *small* functions compile in about two seconds (a 200-function chain
compiles in ~1 s; the one giant function took >60 s). This is McGavin's per-cell
straight-line code, expressed as a chain of tiny functions Rust will actually
compile. **58 → 92 M nodes/s, +56 %** - the single largest lever in the ladder. Node
count unchanged.
### 4 · Stop re-deriving what you already loaded
Two smaller changes, same theme: never read from memory something you already have
in a register.
The **left neighbour** of a cell is, 94 % of the time (240 of 256 cells, all but
the row boundaries), exactly the piece the *calling* cell just placed. So the caller
passes its own right edge down as an argument, and the cell derives its left
constraint with no board read at all:
```diff
- let left_edge = (cell[LEFT] >> 24) as u8; // re-read the neighbour we just placed
+ fn cell_38(st: &mut St, left_arg: u8) -> bool { // caller handed us its right edge
+ let left = left_arg; // …no board read
```
And the **match gain** - how many new matched edges a placement adds - was re-reading
all four neighbours. But the up and left edges were *already loaded* to pick the
candidate pool, and in row order those neighbours are always placed (or a frame edge,
worth no gain). So up/left gain becomes two branchless `bool → u32` adds on values
already in hand; only genuinely-pinned down/right neighbours cost a read:
```diff
- let gain = matched(up) + matched(left) + matched(down) + matched(right); // 4 reads
+ let gain = u32::from(e_up == up) + u32::from(e_left == left) // 0 reads: cached
+ + need_down_read + need_right_read; // only if pinned
```
Together: **92 → 106 M nodes/s** - on the hard board this pulled level with McGavin's
headless C (~102–105 M there). Node count unchanged.
### 5 · A byte array for the used-set
The engine tracked placed pieces as a `u64` bitset: shift, mask, and, test. McGavin
uses a flat `unsigned char tileFree[256]`; his check is one byte load and a
compare-to-zero.
```diff
- if used[pid >> 6] & (1u64 << (pid & 63)) == 0 { /* free */ } // shift, mask, and, test
+ if free_pc[pid] != 0 { /* free */ } // one byte load + compare
```
**106 → 108 M.** Node count unchanged.
### 6 · Fuse cells to amortise the call - the win
The last thing standing between the function-chain and McGavin's `goto` was the
call itself: a `goto` back to the previous cell is a bare jump; a function `return`
restores callee-saved registers first. So **fuse several cells into one function** -
nest the second cell's scan *inside* the first's placement loop, the third inside
the second, and so on, so there is one `call` per *group* of placed cells instead of
one per cell:
```diff
- fn cell_37(st){ for c in pool { place; if cell_38(st, r) {return true} unplace } }
- fn cell_38(st){ for c in pool { place; if cell_39(st, r) {return true} unplace } }
+ fn cells_37_38_39(st){ // three cells, one function, one call in/out
+ for c37 in pool37 { place37;
+ for c38 in pool38 { place38;
+ for c39 in pool39 { place39;
+ if next_group(st, r) {return true}
+ unplace39 }
+ unplace38 }
+ unplace37 }
```
Fusion trades fewer calls for more register pressure, so there is an optimum, and it
is a shallow one. Sweeping the group size on `bench-hard`, three trials each, back to
back:
| group | 1 | 2 | 4 | 6 | 8 |
| --- | --- | --- | --- | --- | --- |
| nodes/s | ~102 M | ~107 M | **~110 M** | ~108 M | ~108 M |
Fusion clearly beats no fusion (group 1), but past group 2 the differences are within
run-to-run noise: the peak wanders between group 4 and 6 depending on the board and the
LLVM register allocator's mood, and it is never more than a couple of percent. The
`--chain2` default is group 6; group 4 edged ahead on this particular board. What
matters is the jump from group 1, not the exact winner. Node count is preserved through
the nesting either way.
## The ladder at a glance
Three anchors are re-verified today on the committed `bench-*.json` boards; the steps
between them are the development-time deltas from the build diary (each a self-contained
commit), which drift a few percent with machine state, so read the middle rows as the
*shape* of the climb, not lab-grade constants.
| Rung | Engine | Change | status |
| --- | --- | --- | --- |
| **naive-clean** | recursive | plain portable Rust DFS, the true baseline | **~44 M, verified** |
| **data-driven JIT** | codegen | generated program, table-driven inner loop | **~61 M, verified** |
| u32 board cells | codegen | one load + shift per neighbour | +~65 % (diary) |
| function-chain | codegen | one function per cell | +~55 % (diary) ★ |
| cached up/left + byte used-set | codegen | stop re-reading placed neighbours | +~15 % (diary) |
| **fusion (group 4–6)** | codegen | **one call per several cells** | **~108–110 M hard / ~122 M easy, verified** |
| *McGavin's C (headless)* | C | *same machine, for reference* | *~102 M hard / ~287 M easy* |
Two engines share this table: `naive-clean` is a separate recursive backtracker
(runnable with `run_dfs --algo naive-clean`), and everything from "data-driven JIT" down
is the codegen path this page is about (`run_dfs_codegen_jit`). The headline, stated
plainly, is a **~2.5× lift from the naive-clean baseline to the fused champion on the hard board**
(~44 M → ~110 M), and **~2.8× on the easy board** (~44 M → ~122 M) - landing, on the hard
board, level with McGavin's headless C. Three meta-lessons fall out of it, the transferable part:
1. **Representation beats micro-ops.** The two biggest single wins - u32 cells
(+67 %) and the function-chain (+56 %) - were both about the *shape* of the data
and the code, not about shaving instructions. No branch-fiddling came close.
2. **Reuse what you have already computed.** Cached-gain and left-as-argument were
pure "stop re-loading it" wins.
3. **Structure beats cycles.** Fusion attacked the *call structure*, not any single
cycle - and that is what brought Rust level with hand-tuned C on hard boards.
## Reproduce it
The engine and two committed benchmark boards live in the public repo under
`research/experiments/dfs-study/engine/crates/dfs-codegen` (`bench/bench-easy.json`,
`bench/bench-hard.json`, and a `bench/README.md` with the full recipe). The three
anchors of the ladder - the naive-clean baseline, the codegen floor, and the fused
champion - are each runnable directly, so anyone can re-run them on their own machine,
back to back, and see the same tree walked at different speeds. From
`research/experiments/dfs-study/engine`:
```bash
# naive-clean: the honest baseline (a separate plain recursive backtracker) ~44 M
cargo run --release -p dfs-run --bin run_dfs -- \
--puzzle crates/dfs-codegen/bench/bench-hard.json --algo naive-clean --seed 1 --budget-s 10
# data-driven JIT floor: the codegen path with no chain/fusion ~61 M
cargo run --release -p dfs-codegen --bin run_dfs_codegen_jit -- \
--puzzle crates/dfs-codegen/bench/bench-hard.json --budget-s 15 --opt native
# fused champion (group = 6): the engine this page is about ~108–110 M hard, ~122 M easy
cargo run --release -p dfs-codegen --bin run_dfs_codegen_jit -- \
--puzzle crates/dfs-codegen/bench/bench-hard.json --budget-s 15 --chain2 --opt native
# any fusion width, to walk the group sweep yourself
cargo run --release -p dfs-codegen --bin run_dfs_codegen_jit -- \
--puzzle crates/dfs-codegen/bench/bench-hard.json --budget-s 15 --group 4 --opt native
```
Point `--puzzle` at `bench-easy.json` and every configuration prints `score=480`
at the **same node count** (3,577,121,570) - the equality that proves the ladder is
a speed ladder and nothing more. On the non-terminating `bench-hard.json` they
print the same partial score at different `nps`. (Give it a real budget: a very
short one reports compile warm-up, not throughput. The easy board wants ≥30 s to
solve.) The three intermediate micro-steps in the ladder table below - packed
candidates, cached gain, the byte used-set - are not separate flags; they are the
commit sequence in `bench/README.md`, reproducible with `git checkout`.
To reproduce the McGavin comparison fairly, build **his** engine headless - comment
out `#define INTERACTIVE` near the top of `genbody.c` so the live display does not
throttle it - then read the true throughput:
```bash
# McGavin, headless + native: emit the puzzle-specialised C, then link and run
gcc -o genbody genbody.c -lm -Ofast -mcpu=native -DG # emits body.c for this puzzle
./genbody PUZZLE.puz HINTS.hnt
gcc -o solve genbody.c -lm -Ofast -mcpu=native # links body.c, runs the search
./solve PUZZLE.puz HINTS.hnt
# read the "Rate:" line (= placements / elapsed) on a hard board, or the final
# "tiles/second" summary on a solvable one — that is his real speed
```
With the display left on, the same binary reads roughly a third of that - which is
exactly the trap that produced the earlier false "we beat him."
> **First: are we even counting the same thing?**
>
> A speed comparison is meaningless unless both engines count the same events, so before trusting any ratio we read McGavin's source. His counter (`ntp`, the famous 16-bit-rollover trick) increments **once per piece committed to the board** - after the candidate has passed the colour-fit lookup and the used-check, at the moment it is placed (`genbody.c`, the `ntp++` right after `square[x][y].tile = t`). Our `st.nodes` does exactly the same: it increments after a candidate passes the edge-fit and used checks, as the piece is placed. Neither counts candidates that fail those checks; both count placements that will later backtrack. So McGavin's "tiles placed per second" and our "search-nodes per second" are **the same measurement** - committed placements, not attempts. (The one asymmetry: he counts the handful of forced hint placements and we bake those out - ≤18 on a count of billions, i.e. nothing.) The number to be wary of is a *third* engine's "M/s": that is where "placements attempted" vs "placements committed" can differ by an order of magnitude, which is the community's own long-standing caution about [what a "node" is](/research/build/faster/solver-engineering).
> **Measure carefully, or don't measure**
>
> Absolute nodes/s drifts with machine load and thermals, and - as the correction above shows - with how the *reference* engine is built. **Only same-state, back-to-back, headless, otherwise-idle ratios are trustworthy.** Every comparison on this page was taken with nothing else running, both engines built headless with native codegen and run seconds apart. The early "McGavin is 5× faster" folklore was also a phantom, in the other direction: it compared his run on an easy puzzle to ours on a hard one. Match the board, match the build, measure back to back - or the number means nothing.
## The easy-board gap is a real, open lever
Why does McGavin's C win the easy board by 2.3× yet only tie on hard ones? Because on
a low-branching board there is almost nothing to do per node - pick the one or two
candidates, place, move on - and his straight-line generated code, with each cell's
candidate list baked to a minimal-perfect-hash lookup, does that with the fewest
possible instructions. Our per-node work (index the pool, compute the gain, check the
rim) is cheap but not *nothing*, and when the tree is shallow-and-wide that overhead
shows. On a hard board the same nodes are dominated by backtracking and cache misses,
where the two engines converge. Closing the easy-board gap would mean adopting his
tighter candidate-list codegen - a concrete next step, not a wall. It is logged here
as open rather than papered over.
## What the speed is worth
Pointed at the real 256-piece Eternity II - six cores, fifteen minutes, the mandatory
centre clue pinned - this engine walks the tree at tens of millions of nodes per
second per core and plateaus, across seeds, in the high-300s of 480. That is not a
disappointment; it is the [whole point of the site's core lesson](/research/why/prune-vs-speed).
A strict backtracker is a superb tree-walker and a poor puzzle-solver: the search
space is so vast that no achievable speed makes a dent in it, which is precisely why
the records come from [heuristics and repair](/research/lab/experiments/raphael-anjou/repair-study),
not from raw throughput.
So the result here is deliberately narrow and, I think, worth stating plainly: a
*safe, portable* Rust engine can **match hand-tuned C on the hard, deep boards that
resemble the real puzzle**, walking the same tree on the same hardware - while the C
still wins by ~2.3× on easy ones. And even at parity it cannot solve Eternity II,
because speed was never the thing standing in the way. The longer argument about that
trade-off - why some algorithms spend their budget on speed and others on judgement -
is its own page: [going fast](/research/lab/experiments/raphael-anjou/going-fast).
## Related
- [McGavin's C backtracker: the throughput story, built here](https://eternity2.dev/research/lab/experiments/peter-mcgavin/backtracker) — Peter McGavin's own C backtracker, the community's fastest: a 2007 optimization recipe compounded for two decades through generated code, lookup tables and counter tricks, then built on my M1 and pointed at the real 256-piece puzzle, where single core it drives past 200 of 256 pieces at ~109M placements/s.
- [Going fast: when a solver spends its budget on speed](https://eternity2.dev/research/lab/experiments/raphael-anjou/going-fast) — Some Eternity II engines pour their effort into walking the search tree as fast as possible; others spend it on judgement about where to walk. This is the case for the first kind - what raw throughput buys, the three different things people mean by "fast", and why the fastest engine ever built still cannot solve the puzzle.
- [Solver engineering: the craft below the algorithm](https://eternity2.dev/research/build/faster/solver-engineering) — Every record solver runs the same depth-first backtracker. What separates them is the layer underneath: lookup tables, perfect hashes, cache-sized structs, generated code, compiler archaeology. That craft decides whether a node costs 26 cycles or 2,600. The community's twenty-year engineering ledger, technique by technique, and what it all bought.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
---
# Learning from strong boards
> A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/
- Updated: 2026-07-16
- Topics: learning, construction
---
Most of this notebook searches Eternity II from first principles: a beam builds a
board from an empty grid, a backtracker digs and retreats, a repair loop polishes
a full board by moves. Each reasons only from the rules of the puzzle and the
board in front of it. This study is the exception, and it is a study rather than a
loose set of runs because all five experiments in it ask one question from
different angles: **what can a search learn from the boards people have already
found, and how far does that carry it?**
The raw material is a corpus of strong boards, every arrangement the community and
this project have pushed above 400 of 480. The move, in every experiment here, is
to read that corpus for structure (where pieces sit, which pieces touch, which
local patterns recur, which agreements are traps) and feed the structure back into
a search as a bias. It is closer to imitation learning than to search design, and
it earns its own home because the same handful of ideas keep returning, each a
little sharper than the last.
The corpus itself is [published](/research/build/dataset): 7,658 distinct boards
scoring 400 to 469, released under CC0, with the diversity check that stops a
learned signal from merely re-encoding one board. The method-agnostic write-up of
each technique lives in the theory section on
[learning from strong boards](/research/build/learning); the pages here are the
lab side, the runs that put each technique on the real 16×16 board, with their
scores and the questions they left open.
## The five experiments, simplest signal to subtlest
The study reads as a sequence. Each experiment adds one turn of the screw over the
one before it, and the pages are ordered to be read in that order.
| # | Experiment | The learned signal | Reached (of 480) |
|---|---|---|---|
| 1 | [PRIOR](/research/lab/experiments/raphael-anjou/learning/prior) | where each piece tends to sit, a single position count | 460 |
| 2 | [KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring) | three signals (position, adjacency, 2×2 patch) voting together | 460 |
| 3 | [LODESTONE](/research/lab/experiments/raphael-anjou/learning/lodestone) | which pieces serve a scarce colour demand, a faint compass | 451 |
| 4 | [PALIMPSEST](/research/lab/experiments/raphael-anjou/learning/palimpsest) | the shared bad habits that cap boards, mining the trap not the structure | 463 |
| 5 | [REPLAY](/research/lab/experiments/raphael-anjou/learning/replay) | the one move a single record board used that our search could not | 460 |
**[PRIOR](/research/lab/experiments/raphael-anjou/learning/prior)** is the simplest
form the idea can take, a count: over the strong boards, how often does each piece
sit in each cell? Used as a construction tiebreak it builds a competitive board
from nothing and reaches 460. **[KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring)**
adds signals rather than trusting one, letting a position count, an adjacency
count, and a learned 2×2-patch quality vote on each placement, which carries a
build into a corner family no earlier search had cracked.
**[LODESTONE](/research/lab/experiments/raphael-anjou/learning/lodestone)** goes the
other way, to a single deliberately faint signal (a per-piece scarce-demand
weight) and, in doing so, finds the knife-edge that runs through the whole study:
the moment a learned signal stops being a tiebreak and starts being the objective,
the search collapses.
**[PALIMPSEST](/research/lab/experiments/raphael-anjou/learning/palimpsest)** is the
subtle turn. Every experiment before it trusts agreement between strong boards;
this one asks which agreements are *traps*, placements every board adopts but no
top board keeps, and steers a repair search to attack them. It produced this
project's best board, 463. **[REPLAY](/research/lab/experiments/raphael-anjou/learning/replay)**
learns from a single board rather than a statistic: rebuild a community record
exactly, and whatever you must add to your search to make it walk that path is the
ingredient it was missing. Here that was the double break, the move that lifted the
strict ladder from 458 to 460.
## The one wall all five reach
The through-line is what makes this a study and not a bag of tricks, and it is a
sobering one. Every signal here is safe only as a gentle bias: over-trust it and
the search falls apart (LODESTONE drops from 451 to 380 as its weight climbs;
PALIMPSEST's trap list helps as a steer and hurts as a ban). And even used
perfectly, none of these signals raises the ceiling. They reach the top of a
search's own range, fast and reliably, and stop, because the corpus they learn
from is made of boards that all hit the same
[rigidity wall](/research/why/rigidity-wall). Learning from strong boards is the
fastest way *to* the plateau and, on its own, no way *past* it. That failure mode
is the subject of the theory section's
[collapse capstone](/research/build/learning/when-learning-collapses), and it is
the note the whole study resolves on.
This is one of three studies in the [experiments notebook](/research/lab/experiments/raphael-anjou).
Its siblings, the [DFS study](/research/lab/experiments/raphael-anjou/dfs-study)
and the [repair study](/research/lab/experiments/raphael-anjou/repair-study), take a
single search paradigm apart one decision at a time; this one holds the paradigm
loose and varies what the search is allowed to *know*.
## Pages in this section
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [LODESTONE](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/lodestone) — A faint compass for a from-scratch search: nudge it to commit the rare pieces early, where they're needed. It doesn't raise the ceiling; it makes the search reliably reach the top of its own range.
- [PALIMPSEST](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/palimpsest) — Read every strong board to find the habits that quietly hold a board back, then break them. This experiment produced the project's best board: 463 of 480.
- [REPLAY](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/replay) — Rebuild the community's strict 460 boards exactly, and in doing so discover the move ordinary solvers can't make: paying two mismatches at a single cell.
## Related
- [Learn from strong boards](https://eternity2.dev/research/build/learning) — Most attacks on Eternity II search from first principles. A distinct family does the opposite: it mines the corpus of boards people have already found for structure, then feeds that structure back into the search. Position priors, learned move-ordering, anti-pattern mining, record decoding, and the failure mode where a learned signal collapses.
- [The dataset](https://eternity2.dev/research/build/dataset) — A public, CC0 dataset for Eternity II in two parts: fourteen benchmark instances to solve, and a corpus of 7,658 distinct strong boards to learn from. Every score is recomputed from the board itself, and the corpus is checked to be genuinely diverse rather than a thousand copies of one board.
- [Raphaël Anjou's experiments](https://eternity2.dev/research/lab/experiments/raphael-anjou) — A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [The repair study](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study) — The sibling of the DFS study, for the other way people attack Eternity II: destroy part of a board, rebuild it, keep the change if it helps. One question, asked carefully. What does each decision in that loop buy: which region to destroy, how to rebuild it, when to keep a move, when to restart, and what board to start from?
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
---
# KEYRING
> Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring/
- Updated: 2026-07-10
- Topics: construction, learning
- Reproduce: `just research-record-boards`
- Source: Beam search as bounded-width best-first construction (this project's concept page) — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/construct/beam-search.mdx
---
[PRIOR](/research/lab/experiments/raphael-anjou/learning/prior) trusted one
learned signal, a position count, to break its ties. KEYRING is the next step in
the [study](/research/lab/experiments/raphael-anjou/learning): what if one signal
is not enough? Building a board one piece at a time, the hard part is deciding
which piece to place next when several would fit, and a single rule of thumb, even
a learned one, tends to march the search into the same dead ends every time.
KEYRING carries three different learned hunches at once, a keyring of them, and
lets them vote, which keeps the search from over-trusting any one signal.
## How it works
From the library of strong boards, KEYRING learns three things. First, where
each piece likes to sit: how often a piece appears in each position across
good boards. Second, which pieces like to be neighbours: how often two pieces
end up touching. Third, which little 2×2 patches show up in good boards
versus bad ones.
It then fills the board with a beam search, keeping many partial boards alive
at once. When it has to choose the next piece, it scores each option by how
many edges it matches, nudged by the three learned signals together. A small
amount of randomness keeps the many parallel attempts from collapsing onto
the same path.
> **[Figure]** Interactive: the three-signal placement vote — interactive: KeyringDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The board
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## The result
KEYRING reached 460 of 480, and did it in a corner arrangement where no board
had reached that level before, so it's not just another route to a known
board but a genuinely new region. Across repeated runs it landed a strong
board far more often than the simpler single-signal version it grew out of.
It is not the project's top score (that's 463), but finding a high board in a
fresh family matters: the strong boards are known to be isolated from each
other, so each new family is its own foothold.
## Method
The three signals, made precise, then how they steer the beam.
The strongest of the three is the **2×2 patch prior**. Over the corpus, split
boards into *high* (score ≥ 460) and *low* (< 460), and for every 2×2 patch $q$
(four pieces with their rotations) score it by a Laplace-smoothed log-odds
ratio:
$$
\text{score}(q) = \log\big(\text{count}_\text{high}(q) + \alpha\big) - \log\big(\text{count}_\text{low}(q) + \alpha\big), \quad \alpha = 0.5
$$
A patch that shows up in strong boards and not weak ones scores positive; a
consensus-trap patch scores negative. (In the run that found the 460, the
corpus split 23 high boards against 1255 low.) The other two signals are
simpler counts: a **piece-in-position** frequency and a **piece-pair
adjacency** frequency, each tallied over the same strong-board set.
The build is a **beam search**: keep $W$ partial boards alive, and at each step
extend every beam by scoring each candidate placement as its matched-edge gain
plus a weighted sum of the three priors, then keep the top $W$. A little
injected randomness stops the $W$ beams from collapsing onto one path, which
is what let KEYRING reach a *new* corner family rather than re-deriving a known
board. Patches pack into a `u64` key (piece ≤ 8 bits, rotation 2 bits, ×4
cells = 40 bits) so the prior lookup is a hash hit.
As with [PRIOR](/research/lab/experiments/raphael-anjou/learning/prior), the beam
build reaches the high 450s on its own; the committed 460 board takes that
construction and a local-refinement tail on top, the same construct-then-refine
split the record pipelines use. The three signals are what carry the *construction*
into a fresh family; the last few edges are the refinement's.
One limit worth stating: the priors are learned from a corpus that is itself
sub-perfect, so they encode the community's ceiling as much as its wisdom:
the same double edge that [PALIMPSEST](/research/lab/experiments/raphael-anjou/learning/palimpsest)
turns into a feature by separating good consensus from traps.
## Reproduce
`just research-record-boards` verifies the committed 460 board's score
byte-for-byte from its stored Bucas string. The search is stochastic (beam
search with injected randomness), so a re-run will not reproduce the same board;
the board is the artifact of record. The engine is the shared beam
producer; the change
this page describes is the three learned ranking signals. Those signals are
mined from a corpus of strong boards, so the run is not reproduced from scratch
here.
## Open questions
Would grading the patch signal by degree, rather than treating patches as
simply good or bad, help further? Could the weights on the three signals
shift as the board fills, trusting structure more late in the build? And can
this new family be pushed past 460 with a longer refinement?
## Related
- [Learned move-ordering](https://eternity2.dev/research/build/learning/learned-value-ordering) — When several pieces fit, which do you place next? Instead of one rule of thumb, carry several signals learned from strong boards and let them vote. Voting keeps the search from over-trusting any one hunch and marching into the same dead end.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [Why basin-hopping looks impossible](https://eternity2.dev/research/why/sigma-cycles) — If you can't improve a great board by polishing it, maybe you can jump to a different great board. On every record pair tested, you can't, and the structural reason why is worth seeing.
- [Forbidden patterns](https://eternity2.dev/research/why/forbidden-patterns) — Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- [Beam search](https://eternity2.dev/research/build/construct/beam-search) — Keep the K most promising partial boards alive at once and grow them cell by cell. Beam search is the workhorse behind this project's from-scratch builders, and a clean illustration of why breadth alone stalls in the deep interior.
---
# LODESTONE
> A faint compass for a from-scratch search: nudge it to commit the rare pieces early, where they're needed. It doesn't raise the ceiling; it makes the search reliably reach the top of its own range.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/lodestone/
- Updated: 2026-07-10
- Topics: construction, search-space, learning
- Reproduce: `just research-record-boards`
- Source: Rare-colour geography (this project): where the scarce (N,W) demands live — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/why/rare-color-geography.mdx
---
[KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring) added signals;
LODESTONE, the third experiment in the
[study](/research/lab/experiments/raphael-anjou/learning), pares back to a single
deliberately faint one, and in doing so exposes the knife-edge every learned
signal here balances on. The strong boards quietly agree on something: as the
score rises, they increasingly satisfy a particular set of scarce demands, places
where only one or two pieces in the whole set can serve a cell's north-and-west
colours. LODESTONE asks whether telling a from-scratch search about those demands
helps it commit the right rare pieces before they're stolen.
## How it works
From the corpus of strong boards, LODESTONE builds a prior: for each piece,
how often it ends up serving one of those scarce north-west demands in a good
board, weighted up sharply for the rarest, so a piece that is the only possible
server gets the biggest boost. The beam search then ranks its candidate
placements by the usual matched-edge score plus a small multiple of this
prior, so among otherwise-equal moves it prefers to place the pieces the good
boards learned to spend early.
The crucial detail is the size of that multiple. The prior has to be a pure
tiebreaker, not part of the objective: it breaks ties between equally-matching
moves and nothing more.
> **[Figure]** Interactive: rare-colour attraction map — interactive: LodestoneRarityLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
At a tiny weight the prior gives a small, consistent gain: across five seeds
the median from-scratch score rises from 449 to 451, and, more usefully, the
spread tightens, from a 446–451 scatter to a reliable 450–451. It makes the
constructor land at the top of its range instead of sometimes stumbling.
Turn the weight up even slightly and it collapses (422, then 380) because
chasing the corpus's scarce demands then trades directly against matching the
edge in front of you. That failure is itself the finding: scarcity is a real
signal for *which* piece to prefer, but a weak one, and only safe as a
tiebreaker. LODESTONE is modest about its size: it improves construction
quality and consistency by a couple of edges, and does not touch the basin
ceiling that stops every method near the top.
## Method
The signal comes first, then the deliberately-tiny way it is used.
**The measurement.** A *scarce demand* is a cell whose north-and-west colours
can be served by only one or two pieces in the whole set. Over the corpus, the
count of scarce demands satisfied in at least half the boards rises
*monotonically* with score, roughly 1 → 2 → 3 → 10 as boards climb toward
458+. Strong boards don't just happen to place rare pieces well; they
increasingly satisfy the *same* scarce demands. That is a real, score-correlated
structure that no earlier constructor had wired in.
**The prior.** For each piece, weight how often it serves one of those scarce
north-west demands in a good board, boosted sharply for the rarest (a piece
that is the *only* possible server gets the biggest weight). The beam ranks
candidates by the usual matched-edge gain plus a small multiple of this weight.
**The knob is the whole story.** The multiple must be a pure *tiebreaker*:
it decides between equally-matching moves and nothing else. At a tiny weight:
median from-scratch score 449 → 451 across five seeds, and the spread tightens
from 446–451 to a reliable 450–451. Nudge the weight up and it collapses (422,
then 380), because chasing scarce demands then trades directly against matching
the edge in front of you. The collapse *is* the result: scarcity says *which*
piece to prefer, but weakly, safe only as a tiebreak.
## Reproduce
Seeded; the reproducible artifact here is the tiebreak *effect* (a tighter
spread and a +2 median across five seeds), not a single board. The engine is the
shared beam producer;
the change this page describes is a faint scarce-piece-early tiebreak. That
tiebreak weight is derived from the piece set and a corpus, so the run is not
reproduced from scratch here.
## Open questions
Could the prior be made position-aware without becoming part of the objective,
strong only in the regions where scarce demands actually concentrate? Does
combining it with KEYRING's three signals add a fourth useful vote, or just
more noise? And is the consistency gain worth more than the median gain, given
that a tighter search is easier to build a ladder on?
## Related
- [Corpus priors](https://eternity2.dev/research/build/learning/corpus-priors) — The simplest way to learn from strong boards: count where each piece tends to sit, or how often it serves a scarce demand, and use that count as a gentle tiebreak in construction. It has to stay a tiebreak; the moment it becomes part of the objective it collapses the search.
- [When learning collapses](https://eternity2.dev/research/build/learning/when-learning-collapses) — The sobering half of learning from strong boards. A learned signal reaches the top of a search's own range reliably, and then stops. Over-trust it and the search collapses; even used perfectly it does not raise the ceiling, because the ceiling is not a thing the corpus knows.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [The rare colors live on the frame](https://eternity2.dev/research/why/rare-color-geography) — Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
---
# PALIMPSEST
> Read every strong board to find the habits that quietly hold a board back, then break them. This experiment produced the project's best board: 463 of 480.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/palimpsest/
- Updated: 2026-07-10
- Topics: local-search, learning
- Reproduce: `just research-record-boards`
- Source: Ropke & Pisinger 2006, adaptive large-neighbourhood search (the destroy-and-repair frame this steers) — https://doi.org/10.1287/trsc.1050.0135
---
Every experiment in the [study](/research/lab/experiments/raphael-anjou/learning)
so far has trusted agreement between strong boards: where the good boards concur,
follow them. PALIMPSEST is the turn where that trust is examined, and it produced
this project's best board. When many independent searches all reach a high but not
perfect board, they tend to agree on a lot of placements. Some of that agreement
is genuine structure, and some of it is a shared bad habit: a local choice that
looks good and keeps every search stuck just short of the top. The idea behind
this experiment is to read the whole corpus of strong boards, separate the helpful
agreement from the trap, and attack the traps.
## How it works
Take every board anyone has found that scores reasonably well, and for each
pair of neighbouring positions count how often a given pair of pieces sits
there, weighted by how good the board was. Two patterns fall out. Some
adjacencies show up again and again in the very best boards: those are safe,
real structure. Others show up almost everywhere but never in the top boards:
those are the traps, the choices that feel right and cap the score.
The traps cluster in particular regions of the board rather than spreading
evenly. Knowing where they are turns the corpus into a map: which placements
to trust, and which to pull apart and rebuild. Grouping boards by how their
corners are arranged then points the search at the most promising family to
attack.
> **[Figure]** Interactive: the trap map, corner-family by corner-family — interactive: PalimpsestDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The board
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## The result
Used as a map to steer the search, this reached 463 of 480 matched edges, the
best board this project has produced. For comparison, the community's best on
this puzzle is 470, and a complete solution is 480.
One caveat worth stating: trying to use the trap list directly, by forcing the
search to avoid trapped placements, did not work on its own and tended to
make boards worse. The value was in reading the corpus to choose where to
focus, not in hard-coding its conclusions into the search.
## Method
For the reader who wants the exact procedure. It runs in two passes.
**1. Score-weighted, basin-separated consensus mining.** Over a corpus of
boards scoring 400–480, for every adjacent cell-pair $(i, j, \text{dir})$ and
every piece-pair that ever sits there, compute two frequencies rather than one:
$$
\begin{aligned}
p_\text{high}(\text{pair}) &= \frac{\#\{\text{boards with the pair, score} \ge 460\}}{\#\{\text{boards with score} \ge 460\}} \\[4pt]
p_\text{all}(\text{pair}) &= \frac{\#\{\text{boards with the pair}\}}{\#\{\text{all boards}\}}
\end{aligned}
$$
together with a *ceiling*: the maximum score of any board containing that pair.
The ceiling, not $p_\text{high}$, is what separates the two categories.
**Good consensus** is high $p_\text{all}$ with a ceiling that reaches the top
family, boards at or within a point of this project's best (463): patterns the
strongest boards keep, so trustworthy structure. **Consensus traps** are high
$p_\text{all}$ with a ceiling that stalls a couple of points short, no board
carrying the pair breaking past the low-460s: the agreed-upon wrong choice that
locks a whole family below the record. The cut sits between "reaches the top"
and "stalls just below"; on this small corpus (a 463 ceiling, boards clustered
from the high 450s to low 460s) it is set by hand rather than swept, and
$p_\text{high}$ (score $\ge$ 460) is reported alongside but is not the
discriminant. The single-frequency view (persistence alone) cannot tell these
apart; splitting by ceiling is the whole trick.
**2. Trap-destroy ALNS.** Take a 461 board (it lives *inside* the trap basin
by construction), locate the ~50 highest-ranked trap pairs it contains, and
find the wedge: the position whose removal breaks the most trap pairs while
preserving the good-consensus ones. Then perturb those trap cells, swap them
to non-trap pieces, introducing 4–8 deliberate mismatches, and feed the board
back to [adaptive large-neighbourhood search](/research/build/local-search/local-search-alns).
ALNS's worst-band destroy operator preferentially tears open exactly those
intentionally-broken regions and rebuilds them. Run: 30 minutes × 6 seeds.
The complexity is unremarkable: the mining pass is linear in the corpus size,
and the real cost is the ALNS search it steers. The contribution is *where* it
aims that search, not a new search.
## Reproduce
The verifier `just research-record-boards` recomputes the matched-edge score of
the committed 463 board from its raw Bucas edges and checks it equals the
claim, so the board is reproducible and checkable byte-for-byte in the viewer.
The search that *found* it is a stochastic run of the shared ALNS
engine (30 min × 6
seeds), steered by the trap map described above. It will not reproduce the same
board, which is why the artifact of record is the board, not a re-run. The
steering is read from a corpus of strong boards, so the run is not reproduced
from scratch here.
## Open questions
Why does rebuilding the trapped regions tend to land back on the same known
top board rather than a genuinely new one? Would a separate map per corner
family reveal structure that the combined map hides? And could a gentle
penalty for trapped placements help where a hard ban hurt?
## Related
- [Anti-pattern mining](https://eternity2.dev/research/build/learning/anti-pattern-mining) — The subtle idea in learning from strong boards: not every agreement between them is good. Some shared placements are real structure; some are a shared trap that caps every search just short of the top. Separate the two, and attack the trap.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [Local search and ALNS](https://eternity2.dev/research/build/local-search/local-search-alns) — Destroy part of a board, rebuild it better, and let the algorithm learn which demolitions pay. Adaptive large-neighborhood search is the most reliable polisher this project has, and the cleanest demonstration of the wall where polishing ends.
---
# PRIOR
> Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior/
- Updated: 2026-07-10
- Topics: construction, learning
- Reproduce: `just research-record-boards`
- Source: Beam search (this project's concept page): bounded-width best-first construction — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/construct/beam-search.mdx
---
PRIOR opens the [learning-from-strong-boards study](/research/lab/experiments/raphael-anjou/learning)
with the simplest thing the idea can be: a count. Most strong boards on this site
are found by taking an existing good board and improving it. PRIOR asks a harder
question: can you build a competitive board from an empty grid, with no board to
anchor to? The trick is to let the crowd of past good boards quietly guide the
construction without copying any single one of them, and the guidance is nothing
more than a tally of where pieces tend to sit.
## How it works
From the library of boards scoring well, PRIOR learns one simple thing: for
each position on the board, how often each piece shows up there. That gives a
gentle preference, a prior, for what tends to belong where.
Then it builds with a beam search, keeping many partial boards alive at once
and extending them cell by cell. When two options match the same number of
edges, the prior breaks the tie toward the piece that's more typical of
strong boards in that spot. A diversity rule keeps the many parallel attempts
from collapsing onto the same path. No single board is copied; the guidance
is statistical.
> **[Figure]** Interactive: the position-prior heatmap — interactive: PriorDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The board
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## The result
From scratch, PRIOR reaches the mid-450s, and with a sharper prior built from
only the very best boards it climbs higher. Followed by local refinement it
reaches 460, the board shown here. That a from-nothing build lands this close
to the records is the point: the structure of good boards is partly
learnable, and you don't need to start from one to get there.
It is not the project's top score (463), and the final lift here leans on a
refinement step, which the board's label notes. But as a clean-slate result
it's the strongest the project has, and it's the foundation the later
multi-signal builders grew from.
## Method
The prior is deliberately the simplest thing that could work: a count. Over
every corpus board scoring above a threshold (440 for the base prior), tally a
$256 \times 256$ matrix $M[p][c]$, how many strong boards place piece $p$ in
cell $c$. Normalized per cell, that is the preference used to break ties.
Three sharpenings, each a variant of the same tally:
- **Rotation-aware.** Split each piece by its four rotations: a
$1024 \times 256$ tensor ($256 \times 4$ rows). The prior now prefers not
just the right piece but the right orientation, 262,144 entries against the
base 65,536.
- **Per corner-family.** The strong boards fall into families by their four
corner pieces. Building a *separate* prior from just one family's boards
gives a signal the pooled matrix averages away, this is what let the
from-scratch build reach a fresh 460 rather than the crowded common basin.
- **Sharper threshold.** Rebuilding the prior from only the very best boards
(a higher cutoff) raises the ceiling of the construction, at the cost of a
thinner, noisier signal.
Construction is a [beam search](/research/build/construct/beam-search): keep $W$
partial boards, extend cell by cell, and when candidates tie on matched edges
let $M$ break the tie; a diversity rule keeps the $W$ beams apart. The 460 here
takes a from-scratch beam build to the mid-450s, then a short local-refinement
tail to 460.
**Novelty was checked, not assumed.** The resulting board is compared against
every 460-tier board already known by corner-permutation and Hamming distance
on (piece, position); a match only counts as a new basin when the corner
family differs or the Hamming distance is large. PRIOR's board passed that
test, it is a genuinely distinct basin, not a re-discovery.
## Reproduce
`just research-record-boards` verifies the committed 460 board's score exactly,
edge by edge, from its stored Bucas string, so the result is checkable even
though the search that found it is not deterministic (its final lift uses a
stochastic refinement tail). The board is the artifact of record. The search
that produced it is the shared beam
producer, with the
one change this page describes: the learned positional prior that breaks its
ties. That prior is a matrix mined from a large corpus of strong boards, which
is why the run is not reproduced here from scratch.
## Open questions
How sharp can the prior get before it overfits? Built from only a handful of
top boards the signal is strong but thin. Could a separate prior per
board-family capture structure the combined one averages away? And how much
of the final gap is the construction versus the refinement that follows it?
## Related
- [Corpus priors](https://eternity2.dev/research/build/learning/corpus-priors) — The simplest way to learn from strong boards: count where each piece tends to sit, or how often it serves a scarce demand, and use that count as a gentle tiebreak in construction. It has to stay a tiebreak; the moment it becomes part of the objective it collapses the search.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [GAUNTLET](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/gauntlet) — Run the same beam search across nine different scan orders, so it lands in different regions instead of always converging to the same one. The zigzag order found a brand-new 458 board.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- [Beam search](https://eternity2.dev/research/build/construct/beam-search) — Keep the K most promising partial boards alive at once and grow them cell by cell. Beam search is the workhorse behind this project's from-scratch builders, and a clean illustration of why breadth alone stalls in the deep interior.
---
# REPLAY
> Rebuild the community's strict 460 boards exactly, and in doing so discover the move ordinary solvers can't make: paying two mismatches at a single cell.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/replay/
- Updated: 2026-07-10
- Topics: backtracking, learning
- Source: Community strict-460 boards (records timeline): the witnesses REPLAY reconstructs — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/records.mdx
---
REPLAY closes the [study](/research/lab/experiments/raphael-anjou/learning) with a
different kind of learning. Every experiment before it mined a *statistic* from the
whole corpus; REPLAY learns from a *single* board, and it learns the one thing a
statistic cannot show, the exact move a record used that our search could not make.
The public fully-clued boards this project was measured against reach 460 (the
community record has since advanced to 464), but this project's own break-allowing
search stalled at 457 or 458 no matter what. REPLAY set out to reproduce those 460
boards exactly, piece for piece, to learn what they were doing that our search could
not. The answer turned out to be a single overlooked move.
## How it works
A break-allowing search normally lets a cell carry at most one mismatch as
it's placed. REPLAY relaxes that to allow two at certain cells, and reorders
how candidate placements are ranked so that the moves a known good board
actually used outrank cheaper-looking ones. With those two changes it can
walk the same path the witness board took.
Played back this way, the community 460 boards rebuild exactly, every piece
in place, and the score checks out. The reproduction is the proof that the
missing ingredient was real.
> **[Figure]** Interactive: the double-break cells — interactive: DoubleBreakDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
> **[Figure]** Interactive: replay the break schedule — interactive: DoubleBreakLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
Both community strict-460 boards replay exactly to 460. The discovery: each
contains four or five cells that pay two mismatches at once. A search that
allows only one mismatch per cell literally cannot reach those boards, which
is exactly why the project's earlier runs saturated at 457 to 458. Allowing
the double break lifts the strict ladder to 460.
It's a clean explanation of a long-standing plateau, and a caution: a
reasonable-looking rule (one break per cell) silently fenced off the very
boards we were chasing.
## Method
The replay is a depth-first search run in a deliberately constrained mode.
- **Prior-over-cost ordering.** Ordinary break-tolerant DFS ranks candidate
placements by immediate mismatch cost, cheapest first. REPLAY flips the
priority to *prior-over-cost*: candidates the witness board actually used are
ranked ahead of cheaper-looking ones, so the search is pulled down the known
good path instead of wandering off it. This is the `--prior-over-cost` flag
driving off the witness board's own schedule.
- **Exact tail.** The last 14 cells are solved exactly (`--exact-tail 14`)
rather than heuristically, so the endgame that ordinary runs fumble is closed
deterministically.
- **The relaxation that mattered.** The per-cell mismatch budget is raised from
one to two. That single change is what admits the witness boards at all.
Run on a fixed 460 frame, 8 threads, a 5-second restart cadence and a per-run
budget, both community strict-460 witnesses rebuild piece-for-piece and the
score checks out, the reproduction *is* the proof that the missing ingredient
was the double break.
**The finding.** Each witness contains four or five cells that pay two
mismatches at once. A search capped at one break per cell cannot represent
those boards, which is precisely why the project's earlier break-tolerant runs
saturated at 457–458. It is a search-completeness result dressed as a record
attempt: the wall was in the move set, not the compute.
## Reproduce
This one is seeded and closer to deterministic than the stochastic builders:
the DFS replay off a fixed frame and witness schedule rebuilds the 460 boards
reliably. The targets it reconstructs are the community's own strict-460
[records](/research/records), so the boards themselves are on the record
timeline; what this experiment adds is the replay that rebuilds them. The
replay needs two inputs beyond the puzzle, a border frame and the community
witness board it reconstructs; a runnable backing directory that ships both is
planned.
## Open questions
Does allowing two breaks per cell open a path to 461 and beyond, or just to
the known 460s? Are there boards needing a triple break? And can the
double-break cells be predicted from a partial board rather than discovered
by replay?
## Related
- [Decoding records](https://eternity2.dev/research/build/learning/decoding-records) — The most literal way to learn from a strong board: rebuild it exactly, piece for piece, until your search can reproduce it. What the reconstruction forces you to add is the ingredient your search was missing, and the reproduction is the proof.
- [Learning from strong boards](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning) — A study in five experiments of one idea: instead of searching Eternity II from first principles, mine the corpus of strong boards already found for structure and feed it back into a search. A position prior, a learned move-vote, a scarce-demand compass, an anti-pattern miner, and a record decode, ordered from the simplest signal to the subtlest, and the one wall all five reach.
- [LADDER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/ladder) — Throw hundreds of cheap short searches at the board, keep only the deepest starts, and promote the survivors through longer and longer rounds.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [No-good learning: remembering why you failed](https://eternity2.dev/research/build/reduce/nogood-learning) — A failed subtree is a theorem: this partial state can never extend. Store it and never re-enter. The community tried both flavours: chess-style transposition tables over the search frontier, and mined constraints about the puzzle itself. The full ledger of what memory buys at E2 scale, and the small boards where it genuinely pays.
---
# Meet in the middle
> Exact endgame experiments that meet in the middle: enumerate a region from two ends and join on the seam, to find the true best completion with a proof rather than a heuristic's best guess. These measure a small region exactly instead of chasing the whole-board score.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle/
- Updated: 2026-07-17
---
Heuristic search guesses; it never knows it has the best possible finish. The
experiments here take the opposite stance for a small region: enumerate it from
two ends, join the halves wherever their seam colours agree and their piece sets
don't overlap, and come away with the *exact* best completion plus a proof that
nothing scores higher. It is the classic
[meet-in-the-middle](/research/build/exact/meet-in-the-middle) time-for-space
trade, pointed at the puzzle's endgame.
These are not attempts on the score. They answer a different question from the
[combination pipelines](/research/lab/experiments/raphael-anjou/pipelines) and the
two search [studies](/research/lab/experiments/raphael-anjou): not *how high can a
heuristic climb*, but *what is the true best ending of this region, and where do
exact methods stop being affordable*. An exact answer about a small region is
worth more here than another near-miss on the whole board.
This is the first experiment of that kind; the section is named for the technique
rather than the one page, because more exact endgame ideas belong alongside it as
they land.
## Pages in this section
- [BANDSAW](https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw) — Solve a band of rows exactly by meeting in the middle, to find the true best ending and to measure how far ahead an endgame can be decided.
---
# BANDSAW
> Solve a band of rows exactly by meeting in the middle, to find the true best ending and to measure how far ahead an endgame can be decided.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw/
- Updated: 2026-07-10
- Topics: exact-methods
- Source: Meet-in-the-middle (this project's concept page): the two-halves-and-join technique — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/exact/meet-in-the-middle.mdx
---
Heuristic search guesses; it never knows it has the best possible finish.
BANDSAW is the opposite experiment: for a band of rows near the bottom, it
computes the exact best completion, with a proof that nothing scores higher.
The aim isn't speed, it's certainty, and the certainty doubles as a ruler
for how hard the endgame really is.
## How it works
Split the band into a top half and a bottom half. Enumerate every way to fill
the top half up to a small mismatch budget, keyed by two things: which pieces
it used, and the row of colors it leaves dangling at the seam. Enumerate the
bottom half the same way, but only from the pieces the top half didn't use.
Then join the two halves wherever their seam colors agree and their piece
sets don't overlap. That meet-in-the-middle join finds the exact best
completion without walking the whole tree.
Exact lower-bound tables, computed by working backwards column by column, let
it prune branches that already can't beat the budget, and it raises the
budget step by step until a round finds nothing new, which proves the best
score for that band.
> **[Figure]** Interactive: the meet-in-the-middle endgame tree — interactive: MeetInMiddleDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
On a 10×10 testbed BANDSAW settles the endgame exactly and pins the budget
where exactness stops being affordable: the search tree grows about
twenty-fold per extra mismatch, on both sides, so meeting in the middle stops
paying off at full board size. That negative is the useful part: it tells you
exactly where exact methods give out and heuristics must take over. The exact
pieces that survived, the suffix lower-bound tables and the pruned
branch-and-bound, became reusable instruments. A frame-free board scoring 437
came out of the same machinery.
## Method
The join is the idea; the pruning is what makes it affordable.
- **Meet in the middle.** Split the band into a top and bottom half. Enumerate
every top-half fill up to a mismatch budget, keyed by (piece-set used, seam
colour row). Enumerate the bottom half the same way, drawing only from the
pieces the top half left. Join two halves wherever their seam colours agree
*and* their piece-sets are disjoint. That join finds the exact best
completion without ever walking the full tree, the classic
[meet-in-the-middle](/research/build/exact/meet-in-the-middle) time-for-
space trade.
- **Suffix lower bounds.** Working backwards column by column builds exact
lower-bound tables, so a partial that already cannot beat the current budget
is pruned before it is extended.
- **Budget ratchet.** Raise the mismatch budget one step at a time and re-solve;
when a round finds nothing better, the previous best is *proven* optimal for
that band. That proof is why this page is tagged **proven**, not measured:
the result is a certificate, not a sample.
The measured ceiling: each half grows ~20× per extra unit of budget, so at full
16×16 board size the top-half table no longer fits, memory, not time, is the
wall. That negative is the deliverable: it pins exactly where exact methods give
out and heuristics must take over. The frame-free 437 board fell out of the same
machinery.
## Reproduce
Deterministic (`kind: exact`): the meet-in-the-middle solve and its optimality
proof reproduce byte-for-byte for a given band, and the 437 board is checkable
in the viewer. The MITM enumerator and suffix-bound tables are committed with
the research code.
## Open questions
Can the lower-bound tables scale to the full 16×16 endgame, or does the seam
state space grow too large? At what band size does memory rather than time
become the limit? And can the rare cases where the middle join does fire be
spotted in advance and finished exactly?
## Related
- [STAGED](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/staged) — Build the whole board from scratch with no pre-set frame, in stages, letting the border emerge last from whatever pieces are left.
- [Entropy and the area law](https://eternity2.dev/research/why/entropy-area-law) — Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
- [Meet in the middle](https://eternity2.dev/research/build/exact/meet-in-the-middle) — Enumerate two halves of a problem and join them on a shared interface, trading memory for an exponent cut in half. The classic Horowitz–Sahni trick, what it looks like on bands of the board, and what this project's BANDSAW experiment measured, including the one-sided method that beat it.
---
# Combination pipelines
> The named search experiments that chase score. Each is a pipeline rather than a single algorithm: it builds a board with one engine, then lifts or finishes it with another. Each records its idea, its best board, and the questions it left open.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/
- Updated: 2026-07-16
---
Every page here is one named experiment with an idea and a board, and nearly
every one is a *pipeline*: it builds a board with one engine, then lifts or
finishes it with another. That is what sets this group apart from the two studies
that take a single paradigm apart, the
[DFS study](/research/lab/experiments/raphael-anjou/dfs-study) and the
[repair study](/research/lab/experiments/raphael-anjou/repair-study). These chain
engines together, and the interest
is as much in the *combination*, the division of labour between construction,
repair and an exact endgame, as in any one stage.
They share the [engines](/research/lab/experiments/raphael-anjou/engines); what
differs is how each one composes and steers them: what it seeds, what it forbids,
what it tears up and rebuilds. Each page's method section walks its stages in
order. Several of the from-scratch builders lean on the beam producer, and two on
the ALNS repair loop, engines that do not yet have their own write-ups here (see
the [engines page](/research/lab/experiments/raphael-anjou/engines)); where a
pipeline does, its page says so plainly rather than pointing at a page that is not
up yet.
They are ordered by what they taught, not by score. A pipeline that ended lower
but explained why is worth more here than one that scraped a point and could not
say how.
## Pages in this section
- [GAUNTLET](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/gauntlet) — Run the same beam search across nine different scan orders, so it lands in different regions instead of always converging to the same one. The zigzag order found a brand-new 458 board.
- [CLOISTER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/cloister) — Fix a perfect border, then search the interior with the border's edges treated as hard constraints from the very first cell.
- [MIDDEN](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/midden) — Decide in advance not when a board may break, but where: confine every mismatch to a chosen shape of cells, and search for the best shape.
- [LADDER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/ladder) — Throw hundreds of cheap short searches at the board, keep only the deepest starts, and promote the survivors through longer and longer rounds.
- [MOSAIC](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/mosaic) — Tile the board into small blocks, solve each one to proven optimality, and glue them together, paying for the seams instead of forbidding them. From scratch, with no record to copy, it reaches 448.
- [STAGED](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/staged) — Build the whole board from scratch with no pre-set frame, in stages, letting the border emerge last from whatever pieces are left.
---
# CLOISTER
> Fix a perfect border, then search the interior with the border's edges treated as hard constraints from the very first cell.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/cloister/
- Updated: 2026-07-10
- Topics: local-search, backtracking
- Source: Rare-colour border geography (this project): why the frame is the natural cut — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/why/rare-color-geography.mdx
---
The border and the interior are usually solved together, which wastes effort:
the interior keeps proposing pieces that can't possibly meet the border
later. CLOISTER pins a perfect 60-piece frame first, then searches the 14×14
interior with the frame's inward-facing edges as real constraints from cell
one, so a doomed interior is rejected immediately instead of at the end.
## How it works
Start from a complete, perfectly-matched border. The interior search is a
break-allowing depth-first fill, but the cells along the inner edge of the
frame must match the frame, and that requirement is live from the first
placement, not checked only when the interior is finished. Exact endgames
score the last region, including how well it seals against the rim.
Because the border is fixed, the search collects something a post-hoc
interior can't: the handful of extra edges that come from the interior
actually fitting the rim it was built against, rather than being grafted onto
a rim later.
> **[Figure]** Interactive: how the frame anchors the interior — interactive: BorderAnchorDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
Try it live: fix a frame and watch the interior search run against it in your
browser.
> **[Figure]** Interactive: solve the standalone interior live — interactive: CloisterLiveLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
As a standalone interior solver CLOISTER reaches an interior score of 453
unhinted, in minutes, and confirms a real effect: an interior built against
its own rim attaches with a few more matched edges than the same-quality
interior attached after the fact. With all five official clues forced it
settles in the high 440s to low 450s.
It doesn't break the top records, and it saturates like everything else does
near the wall. But it cleanly isolates and measures the border-interior
coupling that whole-board search blurs together.
## Method
The search is a break-tolerant DFS over the interior, but the leverage is in
the framing and a two-phase campaign over *many* frames.
- **Frame as hard constraint.** A complete 60-piece border is fixed, and the
frame's inward-facing colours become hard constraints on the rim cells from
the very first interior placement, not a check deferred to the end. A doomed
interior is pruned immediately.
- **Exact tail.** The endgame region is solved exactly (a 14-cell exact tail),
including how it seals against the rim, so the last few cells are closed
deterministically rather than by heuristic.
- **Frame-breadth campaign.** Rather than one border, sweep a directory of
candidate frames in two phases: a cheap ~6-second hint-compatibility probe
drops frames that can't host the clues, then a 30-second tail recipe runs
over each surviving frame with 8 seeds. The best interior wins.
The measured effect is small but real, and it comes from one thing: the
border-interior coupling that whole-board search averages away, but that a
rim-first build gets to exploit while it still can. 453 unhinted; high-440s to
low-450s with all five clues forced.
## Reproduce
Seeded and close to deterministic given a frame: the DFS interior solve off a
fixed frame reproduces its score reliably, and the committed board is checkable
edge by edge in the viewer. The solve needs one input beyond the puzzle, a
pre-solved perfect border frame; a runnable backing directory that ships that
frame alongside the engine is planned.
## Open questions
How much of the rim-compatibility bonus can be collected without fixing the
border first? Does scanning many borders, rather than one, find an interior
that attaches better still? And where exactly does the strict-clue version's
ceiling come from?
## Related
- [STAGED](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/staged) — Build the whole board from scratch with no pre-set frame, in stages, letting the border emerge last from whatever pieces are left.
- [The rare colors live on the frame](https://eternity2.dev/research/why/rare-color-geography) — Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
- [MIDDEN](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/midden) — Decide in advance not when a board may break, but where: confine every mismatch to a chosen shape of cells, and search for the best shape.
- [Meet in the middle](https://eternity2.dev/research/build/exact/meet-in-the-middle) — Enumerate two halves of a problem and join them on a shared interface, trading memory for an exponent cut in half. The classic Horowitz–Sahni trick, what it looks like on bands of the board, and what this project's BANDSAW experiment measured, including the one-sided method that beat it.
---
# GAUNTLET
> Run the same beam search across nine different scan orders, so it lands in different regions instead of always converging to the same one. The zigzag order found a brand-new 458 board.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/gauntlet/
- Updated: 2026-07-10
- Topics: construction
- Reproduce: `just research-record-boards`
- Source: Beam search (this project's concept page): the construction this varies the scan order of — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/construct/beam-search.mdx
---
A beam search that always fills the board in the same order tends to discover
the same kind of board, no matter the random seed. The order you visit the
cells quietly decides which region you end up in, and across sixteen seeds,
a single scan order produced just one distinct corner-arrangement family.
GAUNTLET turns that into a tool: run the search across nine genuinely
different scan orders, and you sample genuinely different regions.
## How it works
GAUNTLET forks the constructive PRIOR beam and gives it a scan order it can
swap: row, row-reversed, column, column-reversed, zigzag, zigzag-reversed,
spiral-in, spiral-out, and diagonal, nine in all. Because the order of cells
changes which partial boards survive at each step, each order explores a
different trajectory rather than crowding into one. A sweep of 9 scans × 4
seeds produced **18 distinct corner-arrangement signatures**, against just
one across sixteen seeds of the single-scan beam it grew from.
Each full run produces a complete board; the strongest are then polished by a
30-minute ALNS lift. The finding is sharp: *scan order is a stronger
diversity axis than the random seed.* That is the lever, and it is what the
later KEYRING and PRIOR build on.
> **[Figure]** Interactive: the construction-order schedule — interactive: GauntletDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
Step through one scan to see how the visit order steers the beam.
> **[Figure]** Interactive: step through the build order — interactive: GauntletStepThrough. Rendered on the canonical page (link above); not shown in this markdown export.
Or race the nine orders against each other, live, in your browser.
> **[Figure]** Interactive: race the construction strategies — interactive: GauntletLiveRace. Rendered on the canonical page (link above); not shown in this markdown export.
## The board
The search is stochastic, so it won't reproduce this exact board, but the
board it found is fixed, bundled here, and checkable edge by edge.
> **[Interactive: RecordBoard]** Rendered on the canonical page (link above); not shown in this markdown export.
## The result
The zigzag order at seed 99, lifted, reached 458 of 480 in the corner
arrangement cp=(3,0,1,2), the first board at or above 458 ever found in that
arrangement in our database. It is at least 246 of 256 cells away from
anything we had at that level: a genuinely new basin, not another route to a
known one.
So GAUNTLET's contribution is reach rather than a record: it opened a new
region of strong boards. A second round (32 more lifts) topped out at 457
with no 461, which says the new family saturates like the others, but the
family itself was the prize.
## Method
The pipeline is three stages, and the parameters are the point.
**Stage 1: diverse construction.** Fork PRIOR's beam and give it a swappable
scan order over the 256 cells: `row`, `row_rev`, `col`, `col_rev`, `zigzag`,
`zigzag_rev`, `spiral_in`, `spiral_out`, `diagonal`, nine orders, crossed
with seeds `{1, 7, 42, 99}`. All share one prior (the `high459` matrix) at a
low softmax temperature (0.05), so the beam is prior-guided but not
deterministic. The order changes which partial boards survive each step, so
each (scan, seed) walks a different trajectory.
**Stage 2: lift.** Each strong build gets an ALNS lift (prior-escape +
`basic_lkh` destroy/repair operators), 5 minutes per build in the sweep, longer
for the finalists.
**Stage 3: cluster.** Group all outputs by corner-permutation and flag the
signatures absent from the existing database. That is how the 9×4 sweep
surfaced **18 distinct corner families** where sixteen seeds of a single scan
had surfaced one, the measured claim that *scan order is a stronger diversity
axis than the seed.*
The winning board, zigzag, seed 99, lifted, sits at cp = (3, 0, 1, 2),
Hamming distance ≥ 246/256 from anything previously at that level: a new basin
by the same corner-perm-plus-Hamming test [PRIOR](/research/lab/experiments/raphael-anjou/learning/prior)
uses.
## Reproduce
`just research-record-boards` verifies the committed 458 board's score exactly
from its stored Bucas string. The sweep is stochastic (nine scans × four seeds,
then a polish pass), so it will not reproduce the same board; the board is the
artifact of record. The engine is the shared beam
producer run across
nine fill orders; it uses the same corpus-mined prior as
[PRIOR](/research/lab/experiments/raphael-anjou/learning/prior), so the run is not
reproduced from scratch here.
## Open questions
Does this new family have the same rigid local structure as the others, or a
different shape? And can a longer refinement lift it past 458, the way a
fresh family sometimes has more room than a well-worn one?
## Related
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [Beam search](https://eternity2.dev/research/build/construct/beam-search) — Keep the K most promising partial boards alive at once and grow them cell by cell. Beam search is the workhorse behind this project's from-scratch builders, and a clean illustration of why breadth alone stalls in the deep interior.
---
# LADDER
> Throw hundreds of cheap short searches at the board, keep only the deepest starts, and promote the survivors through longer and longer rounds.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/ladder/
- Updated: 2026-07-10
- Topics: local-search
- Source: Successive halving / restart strategies (this project's restarts concept page) — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/backtracking/restarts.mdx
---
Most of a long search is wasted on starts that were doomed early. LADDER
spends almost nothing to find out which beginnings are worth pursuing. It runs
a flood of very short probes, keeps the few that got deepest, and only then
pays for longer runs, on those alone. It's tournament selection for search
starts.
## How it works
Round one is hundreds of five-second probes from different random seeds, each
trying to lay down a long run of perfectly-matched cells. Keep the deepest
prefixes, and throw out ones that are near-duplicates of each other so the
survivors stay diverse.
Promote those to a longer round with tighter quality gates, then promote the
best of those to a full-length run. Each rung spends more time on fewer,
better candidates, the way successive-halving tournaments allocate effort to
the contenders that keep winning.
> **[Figure]** Interactive: the rung-by-rung progression — interactive: LadderDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
> **[Figure]** Interactive: run the ladder search live — interactive: LadderLiveLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
LADDER produced a 451 board obeying all five official clues, with no guidance
from any known record, the first time the project escaped the band of 444 to
450 that unguided search kept landing in. The finishes are determined by their
prefix: once a strong enough opening is banked, the rest follows. A longer run
later touched 452, one edge more, though 451 is the board committed and
reproducible here.
It also mapped the limits: the supply of perfect openings runs out, and beyond
a point the rungs all converge to the same ceiling. So LADDER is a good way to
find the best start, not a way past the structural walls that stop every
method near the top.
## Method
The ladder is a recursive successive-halving tournament over search *starts*,
scored by prefix depth.
- **Round one: flood.** Hundreds of 5-second probes from different seeds, each
laying down as long a run of perfectly-matched cells as it can. Rank by the
depth of that perfect prefix; keep the deepest, and drop near-duplicates
(prefixes too similar to a survivor) so the kept set stays diverse.
- **Ratchet.** Each subsequent round pins the previous round's deepest banked
prefix (a depth-15 pin in the recorded run) and probes *beyond* it, pushing
the perfect-prefix frontier one notch further. The recursion stops when a
round gains fewer than four cells of depth, the supply of deeper openings has
run dry.
- **Finish.** From the three deepest distinct prefixes, run long exact-tail-14
rungs (300 s × 8 seeds each) to convert the banked opening into a full board.
The finding is that *the finish is determined by the prefix*: once a deep enough
perfect opening is banked, the endgame follows. That is why concentrating
compute on finding the best start, rather than spreading it across full runs,
escaped the 444–450 band unguided search kept landing in. A longer finishing
round later touched 452; the committed, reproducible board here is the 451.
## Reproduce
Seeded; the probe-and-promote structure reproduces the *behaviour* reliably
though the exact board depends on seeds. The committed 451 board is checkable in
the viewer. The climb needs no corpus, frame, or witness, it runs from the
puzzle alone, so a runnable backing directory for it is planned alongside the
other from-scratch experiments.
## Open questions
What's the real ceiling of prefix-first selection given more compute on the
early rungs? Could the diversity rule be smarter about which near-duplicates
to keep? And does combining the deepest prefixes from different families beat
promoting within one?
## Related
- [REPLAY](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/replay) — Rebuild the community's strict 460 boards exactly, and in doing so discover the move ordinary solvers can't make: paying two mismatches at a single cell.
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [Restarts and heavy tails](https://eternity2.dev/research/build/backtracking/restarts) — Run the same backtracker on the same puzzle twice and the runtimes differ by powers of ten. The community measured this in 2007; the CSP literature had already named it. The cure (cut off, reshuffle, restart) is why every record solver since has been a restart portfolio.
---
# MIDDEN
> Decide in advance not when a board may break, but where: confine every mismatch to a chosen shape of cells, and search for the best shape.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/midden/
- Updated: 2026-07-10
- Topics: local-search
- Source: Piece theft (this project): the spatial view of where boards fail that MIDDEN gates against — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/why/piece-theft.mdx
---
Most break-allowing solvers control when a mismatch is permitted: at certain
depths, past a certain fill. MIDDEN controls where instead. It fixes a mask, a
chosen set of cells, and rules that mismatches may only be paid inside it;
everywhere else must match perfectly. Then it searches over the shape of that
mask. Existing methods say when to take damage; this experiment asks where
damage should live.
## How it works
Pick a mask: a couple of rows, a couple of columns, a scattered lattice of
cells, or a set chosen by color. Run the search forcing perfect matches
outside the mask and allowing mismatches only inside it. Different mask shapes
lead the search into different parts of the space, so the mask becomes a
design knob rather than a fixed rule.
Comparing shapes shows which geometry of allowed damage lets a board grow a
long perfect run before it has to spend a mismatch.
> **[Figure]** Interactive: the damage-geometry break masks — interactive: MaskShapeDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
A dispersed lattice of allowed-damage cells extends the longest perfect run
markedly further than concentrating the damage in a row or two, pushing the
perfect wall from around 150 cells to the 170s. The mechanism is clear; what
stays open is the economics: turning a longer perfect run into a higher final
score once the endgame has to absorb the deferred damage.
So the mask is a real lever, the spatial complement to the usual timing
controls, with a measured effect on how far a board stays perfect, but not
yet a finished route to a record.
## Method
Every other break-tolerant search gates damage by *when*, a depth, a fill
fraction. MIDDEN gates by *where*. It is the first cell-set (WHERE) damage
control against everyone else's depth (WHEN) gates.
- **The mask.** A chosen set of cells (a 256-bit mask) is passed as
`--break-cells`. Inside the mask, mismatches are allowed; everywhere else must
match perfectly. The search is otherwise a standard break-tolerant DFS off a
fixed frame.
- **The sweep.** Mask *shape* becomes the design variable: a couple of rows, a
couple of columns, a dispersed lattice, or a colour-chosen set, swept over
shapes and densities, each 300 s × 8 seeds.
The measured result is a graded-density design rule: a **dispersed lattice** of
allowed-damage cells extends the longest perfect run from ~153 cells to
167–174 (**+21**), far more than concentrating the damage into a row or two.
The mechanism is clear; the open economics is the endgame, a longer perfect
run only helps if the tail can absorb the deferred damage, and the dispersed
mask that maximizes the wall does not by itself seal the finish. The natural
next step (composing a dispersed body-mask with an open tail) is exactly what
the MIDDEN-v2 sweep set out to test.
## Reproduce
Seeded off a fixed frame; the wall-extension effect reproduces across masks,
though the exact board depends on seeds, and the committed 452 board is
checkable in the viewer. Like [CLOISTER](/research/lab/experiments/raphael-anjou/pipelines/cloister)
it needs a pre-solved border frame (plus the break-cell mask, which is a literal
input); a runnable backing directory that ships the frame is planned.
## Open questions
Which mask shapes convert a longer perfect run into actual matched edges at
the end, rather than just deferring the damage? Can spatial masks be combined
with timing controls so a board is gated in both where and when it may break?
And is there a mask that mirrors where the best known boards actually carry
their mismatches?
## Related
- [LADDER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/ladder) — Throw hundreds of cheap short searches at the board, keep only the deepest starts, and promote the survivors through longer and longer rounds.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [CLOISTER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/cloister) — Fix a perfect border, then search the interior with the border's edges treated as hard constraints from the very first cell.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
---
# MOSAIC
> Tile the board into small blocks, solve each one to proven optimality, and glue them together, paying for the seams instead of forbidding them. From scratch, with no record to copy, it reaches 448.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/mosaic/
- Updated: 2026-07-10
- Topics: exact-methods
- Source: RC2 MaxSAT (the exact block solver): a core-guided weighted-MaxSAT algorithm — https://doi.org/10.3233/SAT190116
---
Eternity II has no local structure you can exploit globally, but a small
window of it, a 4×4 block, solves to perfect optimality in seconds. MOSAIC is
an experiment built on that single fact: if you can solve a block exactly, can
you compose sixteen exact blocks into a whole board?
## How it works
The 16×16 board is cut into sixteen 4×4 blocks, filled one at a time. Each
block is handed to an exact MaxSAT solver, which finds the best possible
placement of pieces in it. The trick is in how a block meets its
already-placed neighbours: those shared edges are not hard requirements but
*soft* targets the block is rewarded for matching. So a block can never become
impossible; it simply pays for any seam it can't match, and always completes.
The second idea fights piece theft directly. Before filling a block, MOSAIC
holds back the globally scarcest pieces, so the last blocks aren't starved of
the rare pieces their seams will demand. Tuning how much to reserve is the one
real knob; too little and the corner starves, too much and the early blocks
suffer.
> **[Figure]** Interactive: the block-assembly search — interactive: MosaicBlockLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
The window primitive does what it promises: a 4×4 block solves to its 24-edge
optimum in about thirty seconds, a 3×3 in eleven, confirming the puzzle
really is tractable in the small. Composed across the whole board, from
scratch and with no warm start, MOSAIC reaches 448 of 480. The reservation
sweet spot is around eight percent of the pool.
The shortfall is informative: almost all of it is in the last three blocks of
the bottom-right corner, where the pool finally runs thin: piece theft again,
now visible as a single bright spot on the board. Exact-in-the-small does
compose, but the composition order spends its freedom early and pays for it at
the end, the same shape every method here runs into.
## Method
The exactness is genuine, and so is the backtracking that stitches it
together.
- **Exact blocks.** Each 4×4 block is encoded as a *weighted MaxSAT* problem
and solved by RC2, a core-guided solver, to a provably-optimal fill. Shared
edges with already-placed neighbours are *soft* clauses (rewarded, not
required), so a block can never be infeasible, it pays for any seam it can't
match and always completes.
- **Backtracking over solutions.** MOSAIC is not a one-shot glue. Each block
level keeps an *enumerator* of MaxSAT solutions, best-first, via RC2 plus
blocking clauses that rule out already-seen fills. When a later block is
starved or a level exhausts, it backtracks and pulls the *next* solution of
the previous block (freeing a different set of pieces). It is a coarse
16-level backtracking search where each node is a whole optimal block.
- **Scarcity reservation.** Before filling, MOSAIC holds back the globally
scarcest pieces so the final blocks aren't starved of the rare pieces their
seams demand. That reservation fraction is the one real knob; the measured
sweet spot is ~8% of the pool.
The window primitive is real: a 4×4 block reaches its 24-edge optimum in ~30 s,
a 3×3 in ~11 s, the puzzle *is* tractable in the small. Composed from scratch
it reaches 448, with the residual shortfall concentrated in the last three
bottom-right blocks: [piece theft](/research/why/piece-theft) made visible as a
single bright spot.
## Reproduce
Deterministic: the MaxSAT block solves and the backtracking composition are
exact, so `kind: exact`, the 448 board reproduces and is checkable edge by
edge in the viewer. The block engine runs from the puzzle alone, with no corpus
or seed board, its only external dependency being a MaxSAT solver, so a runnable
backing directory for it is planned alongside the other exact experiments.
## Open questions
Would a non-row-major block order, spiralling in or solving the constrained
corner first, move the depletion off the hardest block? Could blocks overlap,
so seams are solved twice and reconciled? And does a faster (Rust) primitive
make a larger block size, with its stronger exact guarantee, affordable?
## Related
- [BANDSAW](https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw) — Solve a band of rows exactly by meeting in the middle, to find the true best ending and to measure how far ahead an endgame can be decided.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- [Which wall stops which method](https://eternity2.dev/research/why/walls-and-methods) — The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- [SAT and CSP encodings](https://eternity2.dev/research/build/exact/sat-csp-encodings) — Write the puzzle as clauses and hand it to an industrial solver: the obvious move, tried since 2008. Why complete solvers stall on the full board, and where their verdicts still earn their keep as impossibility proofs.
---
# STAGED
> Build the whole board from scratch with no pre-set frame, in stages, letting the border emerge last from whatever pieces are left.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/staged/
- Updated: 2026-07-10
- Topics: construction
- Reproduce: `just research-record-boards`
- Source: Beam search (this project's concept page): the staged construction primitive — https://github.com/raphael-anjou/eternity2/blob/main/web/content/research/build/construct/beam-search.mdx
---
Almost every solver starts by locking down the border, because the border is
the most constrained part and pinning it shrinks the search. STAGED is an
experiment in refusing that crutch. It builds the board in stages from the
top down with all 256 pieces free, and never commits to a border until the
very end, when the border simply falls out of what remains. The question it
answers: can you reach a strong board without ever anchoring to a frame?
## How it works
The build runs in four stages. The first two fill the top half of the board
with a fast plain search, banking the partial boards that survive. At the
handoff, a cheap admissible estimate throws out any partial that clearly
can't be finished well, so later stages only work on promising starts.
The third stage grows the next band of rows from those survivors. The fourth
is an exact finisher on the bottom rows that minimizes mismatches, and it
chooses the bottom border last, against whatever pieces are
still unused. So the frame is not designed up front; it emerges as a
consequence of everything above it.
> **[Figure]** Interactive: the stage-by-stage build — interactive: StageBuildDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
## The result
STAGED reaches 436 of 480 from scratch, with an emergent border and all five
official clues respected, built end to end with no frame to lean on. That's
well below the records, and that gap is the finding: it measures how
much the usual frame-first anchor is worth, and it showed the frame-free
machinery works at full scale.
Along the way it pinned down the anatomy of the very best boards: they are a
perfect block of most of the board plus a thin band of mismatches
concentrated in a few top rows. That shape is what later builders aim to
reproduce on purpose.
## Method
Four stages, each handing survivors to the next through an admissible filter.
1. **Top-half beams (stages 1–2).** Fill the top half with a fast plain search,
banking the partial boards that survive. All 256 pieces are free, no frame
is pinned.
2. **Admissible handoff.** At each stage boundary, a cheap *admissible* estimate
(an optimistic upper bound on the best possible finish) discards any partial
that provably can't be completed well. Because the estimate never
under-counts the achievable score, discarding is safe, it only removes
partials that cannot win.
3. **Band grow (stage 3).** Extend the surviving partials down the next band of
rows.
4. **Exact finisher (stage 4).** Solve the bottom rows exactly, minimizing
mismatches, and *choose the border last* from whatever pieces remain. The
frame is not designed up front; it falls out of everything above it.
The result is 436 from scratch with an emergent border, all five clues
respected, well below the records, and that gap is the measurement: it prices
what the usual frame-first anchor is worth. The by-product mattered more than
the score: STAGED pinned the anatomy of the best boards, a large perfect block
plus a thin band of mismatches in a few top rows, the target shape later
builders aim at deliberately.
## Reproduce
Stochastic (the top-half beams use randomized tie-breaking), so a re-run won't
reproduce the exact board; the committed 436 board is the artifact of record and
is checkable in the viewer. The engine is the shared beam
producer, run in
stages so the border emerges last rather than being fixed first. It needs no
corpus or seed board, so a runnable backing directory for the full staged
pipeline is planned alongside the other from-scratch builders.
## Open questions
Can a generator that deliberately spends its mismatches in the top rows, to
keep the rest perfect, reach the 450s frame-free? How much of the
436-to-record gap is the missing frame anchor versus the harder endgame? And
does a learned finish-quality estimate pick better survivors than the cheap
admissible one?
## Related
- [CLOISTER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/cloister) — Fix a perfect border, then search the interior with the border's edges treated as hard constraints from the very first cell.
- [BANDSAW](https://eternity2.dev/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw) — Solve a band of rows exactly by meeting in the middle, to find the true best ending and to measure how far ahead an endgame can be decided.
- [Beam search](https://eternity2.dev/research/build/construct/beam-search) — Keep the K most promising partial boards alive at once and grow them cell by cell. Beam search is the workhorse behind this project's from-scratch builders, and a clean illustration of why breadth alone stalls in the deep interior.
---
# The repair study
> The sibling of the DFS study, for the other way people attack Eternity II: destroy part of a board, rebuild it, keep the change if it helps. One question, asked carefully. What does each decision in that loop buy: which region to destroy, how to rebuild it, when to keep a move, when to restart, and what board to start from?
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/
- Updated: 2026-07-16
- Topics: local-search, search-space, speed
- Reproduce: `just experiments repair-study`
- Source: Runnable engine + committed results + scripts (this study's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/repair-study
---
There are two ways people actually attack Eternity II. One is to build a board
cell by cell and backtrack when it fails; the [DFS study](/research/lab/experiments/raphael-anjou/dfs-study)
takes that apart. The other is to hold a whole board and *repair* it: rip out a
region, rebuild it, keep the change if it helped, and repeat. Repair is the last
stage of the construct-then-refine pipelines behind this project's own
record-approaching boards, and it is the method the literature reaches for once a
board is too good for a single-piece move to improve. This study takes the repair
loop apart the same way its sibling took backtracking apart: one decision at a
time, on the same ten corner-pinned variants, single core, sixty seconds a run.
The maximum score is 480 matched edges.
The loop itself is short. One iteration, walked through below.
> **[Figure]** One destroy-and-repair iteration, step by step — interactive: RepairLoopDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
The point is not to win. Most variants finish in the 360s and 370s, and the one that
does best (446) does so by starting from a strong backtracked board rather than a
constructed one, which is itself the study's central lesson. Repairing a plain
greedy board on one core for a minute is a small thing next to the pipelines the
records use; what the study measures is *what each decision in the
loop is worth* by changing one at a time and measuring the result with the same
canonical scorer the DFS study and the rest of the site use.
## The family, and the leaderboard
Five families, laid out so that neighbours differ by a single decision, all
branching off one plain anchor loop (greedy start, destroy the mismatched cells,
greedy refill, keep the result unless it loses score, never restart). **Starting
board** changes where the loop begins. **Destroy operator** changes which cells
each iteration lifts. **Repair** changes how the hole is rebuilt. **Acceptance**
changes when a non-improving move is kept. **Restart** changes what happens once
the loop stalls.
> **[Interactive: RepairStudyLeaderboard]** Rendered on the canonical page (link above); not shown in this markdown export.
## What the study found
- **Starting from a strong backtracked board wins the whole study.** The variant
that spends its first twenty seconds running the DFS study's break-DFS, then
repairs the resulting low-440s board, finishes highest of all, at a mean of 446
(best 449). Repair adds only a handful of edges on top of that board, but the board it
starts from is a hundred points better than a greedy construction, and that
carries through. This is the construct-then-refine division of labour the records
use, reproduced end to end on one core in one minute.
- **Among the greedy-start variants, random destroy wins and targeting the breaks
backfires.** A geometry-blind destroy that lifts twelve *random* cells finishes
at a mean of 402 (best 409) and keeps improving deep into the run. Every operator
that targets the broken cells finishes below it, and the more precisely it
fixates, the worse it does: the mismatched-cells anchor lands at 366, a larger
connected-component destroy at 350. On a mediocre board there is improvement
available everywhere, so exploring beats attacking where it already hurts. That
inverts on a near-record board, where the few remaining mismatches are the only
thing left to fix.
- **The starting board sets the floor.** A greedy construction starts around 348;
a random one starts near 18, and although the loop lifts it a spectacular 308
points, it still finishes below where the greedy start *began*. Construction is
the lever, repair is the polish.
- **The acceptance rule is the other real lever.** Simulated annealing, which
steps downhill occasionally to leave a plateau, is the strongest acceptance rule
by a clear margin (mean 377, best 394), well
ahead of a strict hill-climb (361). How much a non-improving move is allowed is
worth a sixteen-point swing on the same loop.
- **The clever refinements buy nothing here.** A whole-band destroy is inert (zero
improvements on nine of ten instances). A bounded *exact* refill of a small hole
does not beat a plain greedy refill of the same hole, a clean negative result:
its locally-perfect rebuilds cost enough iterations that more, cheaper greedy
rebuilds reach just as far. Neither a random kick nor a revert-to-best on a stall
moves the score off the no-restart baseline.
Each of these has its own treatment: how the engine is built and what every
raised statistic means is on the [method page](/research/lab/experiments/raphael-anjou/repair-study/method),
and the destroy, repair, acceptance, restart and starting-board comparisons are
worked through on the [findings page](/research/lab/experiments/raphael-anjou/repair-study/findings).
## How to read the numbers
Every board is re-scored by the one canonical scorer, and no engine's
self-reported score is trusted. The loop maintains its score incrementally as it
places and lifts pieces, but the published number is always a fresh canonical
re-score of the output board. Throughput is reported in repair-iterations per
second and is **never compared across families**, because an iteration that runs
an exact refill is not the same unit of work as one that runs a greedy fill.
The axis to watch is the *stall*: the iteration at which the global best last
improved, against the total iterations run. When the first is a few thousand and
the second is hundreds of thousands, the run found its answer early and then
ground the same basin for the rest of the minute. That gap is the measured form
of a long-standing observation about this puzzle, that destroy-and-repair is a
superb basin-explorer and essentially never escapes the basin it lands in.
The whole apparatus (the engine, the ten variants, the committed per-run results
and the grid scripts) lives under the study's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/repair-study),
and `just experiments repair-study` rebuilds the engine and reruns the whole
grid. The board, scorer and IO layer come from a
[shared library](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/common)
the DFS study uses too, so a repaired board and a backtracked one are scored by
exactly the same code.
## Pages in this section
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/method) — The engine behind the repair study: one composable destroy-and-repair loop where a variant is a declared change over a parent, the shared IO and scorer it sits on with the DFS study, an incrementally-maintained mismatch map, and the definitions of every statistic the study raises.
- [What each decision buys](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/findings) — The five comparisons at the heart of the repair study, worked through: blind random destroy wins while every conflict-targeting operator loses; construction sets the floor; simulated annealing is the strongest acceptance rule; and the clever refinements (exact refill, restarts) buy nothing at this budget.
## Related
- [Raphaël Anjou's experiments](https://eternity2.dev/research/lab/experiments/raphael-anjou) — A notebook of Eternity II search experiments, organised into the shared engines they run on, the combination pipelines that chase the score, three studies that take one search paradigm apart a decision at a time, and exact endgame solves. Each has its idea, its best board, and the questions it left open. The best reaches 463 of 480.
- [The DFS study](https://eternity2.dev/research/lab/experiments/raphael-anjou/dfs-study) — One question, asked carefully: among depth-first backtrackers for Eternity II, what does each fill order, each heuristic, and the break mechanism actually buy? A family of from-scratch backtrackers, each one change apart, run on the same ten corner-pinned variants, single core, sixty seconds.
- [Local search and ALNS](https://eternity2.dev/research/build/local-search/local-search-alns) — Destroy part of a board, rebuild it better, and let the algorithm learn which demolitions pay. Adaptive large-neighborhood search is the most reliable polisher this project has, and the cleanest demonstration of the wall where polishing ends.
---
# What each decision buys
> The five comparisons at the heart of the repair study, worked through: blind random destroy wins while every conflict-targeting operator loses; construction sets the floor; simulated annealing is the strongest acceptance rule; and the clever refinements (exact refill, restarts) buy nothing at this budget.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/findings/
- Updated: 2026-07-16
- Topics: local-search, search-space, speed
- Source: Committed per-run results (results.jsonl) and per-family report (report.md) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/repair-study/results
---
Five comparisons carry the [repair study](/research/lab/experiments/raphael-anjou/repair-study).
Each isolates one decision by holding the other four fixed at a plain anchor loop:
greedy start, destroy the mismatched cells, greedy refill, keep the result unless
it loses score, never restart. All scores are the mean matched-edge count over the
ten corner-pinned variants, single core, sixty seconds. Iteration rate is never
compared across families.
## The starting board: the lever the loop cannot replace
Change only the board the loop begins from, and everything else about the run is
downstream of it. A greedy construction begins around 348 and the loop lifts it
to a mean of 366. A random board begins near 18 and the loop lifts it an enormous
308 points, to a mean of 326. That larger lift is not the loop doing better; it is
the loop doing the construction's job badly. The random start, after millions of
iterations, still finishes below where the greedy start *began*.
This is the study's first and firmest result: **construction is the lever, repair
is the polish.** A rarest-colour-first greedy construction (the Selby and Riordan
heuristic that a scarce colour should be spent where it is forced) starts a little
higher still and finishes at a mean of 373, the best of the *greedy* starts.
The clearest proof of the point is to start from a genuinely strong board rather
than a greedy one. The final starting-board variant hands the loop a board built
by the [DFS study's](/research/lab/experiments/raphael-anjou/dfs-study) break-DFS:
the run spends its first twenty seconds backtracking to a board in the low 440s,
then repairs it for the remaining forty. That start is a hundred points above a
greedy construction, and it carries straight through to the finish. It is the
highest-scoring variant in the whole study, at a mean of 446 (best 449), far above
the random-destroy loop's 402, and it settles the construct-then-refine question the
rest of the study circles: repair *does* add a few edges on top of a strong
backtracked board, but only a few, and the board it starts from decides almost
everything. This is exactly the division of labour the records use, and this study
reproduces it end to end on one core in one minute: a backtracker to build a good
board, then a repair loop to squeeze the last edges out of it.
## The destroy operator: attacking the breaks can backfire
Now fix the greedy start and change only which cells each iteration lifts. This is
the axis with the study's most surprising result.
> **[Figure]** The four destroy operators, on one board of broken seams — interactive: DestroyOperatorDiagram. Rendered on the canonical page (link above); not shown in this markdown export.
- **A geometry-blind random destroy wins the destroy comparison, at a mean of 402
(best 409).** Lifting twelve *random* cells instead of the broken ones, it keeps
sampling fresh regions of the board, accepts nearly three quarters of its moves,
and its best score keeps improving deep into the run: its last improvement lands
past three million iterations, where every conflict-driven variant has long since
frozen. It is the best of every variant that starts from a greedy board (only the
DFS-seeded start, a different lever entirely, finishes higher).
- **Every conflict-driven operator finishes below it, and the more it fixates on
the breaks, the worse it does.** The mismatched-cells anchor, which lifts up to
a dozen of the broken cells, averages 366 and stalls early: it rips out the same
clustered tangle, greedily rebuilds it to almost the same placement, and its best
score stops moving within a few thousand iterations. The component-plus-halo
operator, which lifts one whole connected tangle (a much larger hole), does worse
still, at 350, and its best last improves at iteration 211: a hole that size hands
the greedy refill a subproblem it cannot improve, so almost nothing is ever kept.
- **A whole-band destroy is inert.** Lifting the two rows with the most breaks
leaves the greedy refill a subproblem nearly as hard as the puzzle itself; on
nine of the ten instances it makes *zero* improvements over the whole minute, so
its reported score is simply its starting board.
The pattern is clean and worth stating plainly: on this mediocre greedy start,
the more precisely an operator targets the existing breaks, the worse it does, and
blind random destroy wins. This is the opposite of the natural intuition, and of
what works on a *near-record* board, where the whole board is close to optimal and
the community's tuning found conflict-driven and component operators most valuable
precisely because the few remaining mismatches are the only thing left to fix. This
study never reaches that regime. Sixty seconds on a mediocre board leaves broad
improvement available everywhere, and there an operator that keeps re-attacking the
same broken cluster, or that tears out a hole too large to rebuild well, both lose
to one that simply keeps trying fresh small regions. The finding is not
"conflict-driven destroy is bad" but "which operator wins depends on how good the
board already is, and on a mediocre board, exploring beats both fixating and
over-destroying."
## Repair: how the hole is rebuilt
Fix the destroy and change only the refill. Breaking the greedy refill's exact
score ties with a seeded coin, rather than deterministically, adds a little
exploration and helps slightly, to a mean of 365. The sharper comparison is greedy
against *exact* refill, and it has to be set up carefully to be a clean one-axis
test: the exact refill only pays off when the hole is small enough to search, so
the study pairs it with a small destroy (at most six mismatched cells) and compares
it against the *same* small destroy refilled greedily, so the only thing that
changes between the two is the refill. Set up this way, the exact refill runs on
every iteration rather than falling back to greedy, rebuilding each small hole to
its true optimum.
The result is a clean negative one. The exact refill (mean 360) does **not** beat
the greedy refill of the same small hole (mean 363); if anything it is a shade
behind. Two things explain it. The exact refill accepts nearly every iteration,
because a locally-optimal rebuild of a six-cell hole almost never lowers the score,
so the loop drifts sideways rather than climbing. And it buys that local optimality
at roughly a third of the iteration count, so over a fixed minute it explores less
of the board. On this puzzle, at this budget, a locally-perfect rebuild of a tiny
region is not worth what it costs: the greedy refill's cheaper, slightly worse
rebuilds, run more often, reach just as far. This is the same node-rate-is-not-score
question the DFS study's heuristic engines raise, and here the answer comes out on
the side of more, cheaper iterations, which is worth recording precisely because the
opposite is so often assumed.
## Acceptance: how much a non-improving move is allowed matters
Fix the loop and change only when a non-improving candidate is kept. Of the five
axes, this one moves the score the most after the destroy operator.
- **A strict hill-climb, which keeps only strict improvements, is the weakest, at
a mean of 361.** Refusing every sideways move locks it into the first basin it
finds; its accept rate is essentially zero.
- **Allowing equal-or-better moves (the anchor) does better, at 366.** The
sideways moves let it drift across the equal-score plateaus that dominate this
landscape.
- **A cooling simulated-annealing rule is the strongest acceptance by a clear
margin, at 377 (best 394), among the best of the greedy-start variants.**
Allowing occasional *worsening* moves early, then cooling, lets it leave a
plateau that pure sideways drift is trapped on, and its best score keeps
improving to around iteration twenty thousand rather than freezing in the first
few thousand. A late-acceptance rule, which compares against the score some tens
of iterations ago, lands between the two at 369.
So the acceptance rule is a real lever here, not a detail: strict to annealing is a
sixteen-point swing on the same loop (361 to 377). That is worth stating against a known
observation from the community's own ALNS tuning, that across a wide range of
annealing temperatures the final scores were essentially identical. The two are not
in conflict. That tuning was done on near-record boards, where the landscape is a
sea of equal-score plateaus and any temperature accepts almost everything; this
study's mediocre starting board still has real downhill structure to exploit, so
whether the rule will step downhill to escape a plateau genuinely matters. Which
acceptance rule wins, like which destroy operator wins, depends on how good the
board already is.
## Restart: neither perturbation moves the needle
Fix the loop and change only what happens once it stalls. Doing nothing lets the
run grind the same basin for the rest of the minute. A random *kick*, which
unplaces and randomly refills a couple of dozen cells when the best has not moved
for a while, and a *revert to the best board so far*, which re-attacks the
incumbent, are the two perturbations tested. Both land within a point of the
no-restart anchor (365 and 365, against the anchor's 366). This is a clean null
result: bolting a stall-detector and a perturbation onto the greedy-mismatch loop
does not help it, because the perturbation either throws away the structure that
made the board good (the kick) or returns to a board the same operator has already
stalled on (the revert). Neither turns the loop into a basin-escaper; on this
puzzle, escaping a basin needs a different move than perturb-and-repair, which is
exactly what the [rigidity wall](/research/why/rigidity-wall) and the
[σ-cycle structure](/research/why/sigma-cycles) explain.
## The through-line
Two themes run through the five comparisons. The first is that **the two decisions
that move the score are the destroy operator and the acceptance rule**, and both
move it in the same counterintuitive direction: the winning choices are the ones
that keep the loop *exploring* rather than *exploiting*. Random destroy beats every
operator that targets the breaks; annealing, which steps downhill to leave a
plateau, beats every rule that only moves sideways or uphill. Refining the refill
and bolting on a restart, the two decisions that try to be cleverer about a region
the loop is already stuck on, buy nothing.
The second is why exploring wins here and fixating wins on a near-record board:
**the repair loop is a basin explorer, not a basin escaper.** On a mediocre board
there is broad downhill structure everywhere, so the moves that cover more of it
win; on a near-record board there is none left but a single interlocking cycle, so
the moves that target it precisely win. Neither regime lets perturb-and-repair
escape a basin once it is in one. That is why the construct-then-refine pipelines
that reach this project's best boards divide the labour the way they do: a strong
constructive producer to choose a good basin, and repair as the last mile inside
it, never asked to climb out.
## Related
- [The repair study](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study) — The sibling of the DFS study, for the other way people attack Eternity II: destroy part of a board, rebuild it, keep the change if it helps. One question, asked carefully. What does each decision in that loop buy: which region to destroy, how to rebuild it, when to keep a move, when to restart, and what board to start from?
- [How the study is built](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/method) — The engine behind the repair study: one composable destroy-and-repair loop where a variant is a declared change over a parent, the shared IO and scorer it sits on with the DFS study, an incrementally-maintained mismatch map, and the definitions of every statistic the study raises.
- [Local search and ALNS](https://eternity2.dev/research/build/local-search/local-search-alns) — Destroy part of a board, rebuild it better, and let the algorithm learn which demolitions pay. Adaptive large-neighborhood search is the most reliable polisher this project has, and the cleanest demonstration of the wall where polishing ends.
---
# How the study is built
> The engine behind the repair study: one composable destroy-and-repair loop where a variant is a declared change over a parent, the shared IO and scorer it sits on with the DFS study, an incrementally-maintained mismatch map, and the definitions of every statistic the study raises.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/method/
- Updated: 2026-07-16
- Topics: local-search, speed
- Source: The engine workspace (repair-engine, repair-run) on the shared e2-core / e2-io library — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/repair-study/engine
---
This page is the apparatus behind the [repair study](/research/lab/experiments/raphael-anjou/repair-study):
how the engine is built, why a new variant is cheap to add, and what every number
on the results means. It is the deliberate twin of the
[DFS study's method page](/research/lab/experiments/raphael-anjou/dfs-study/method),
and it sits on the *same* shared library: the board, the piece set, the one
canonical scorer and the IO contract all come from a common `e2-core` / `e2-io`
crate that the backtracking study uses too. A repaired board and a backtracked
board are therefore scored by the identical code, which is what lets numbers from
the two studies sit on one axis.
## A variant is a declared change over a parent
Every algorithm in the study is the *same* destroy-and-repair loop, parameterised
by five independent choices:
- **starting board**: the board the loop begins from (random, greedy
construction, or a rarest-colour-first greedy construction);
- **destroy operator**: which cells each iteration lifts (random, the mismatched
cells, the worst row band, or one connected mismatch component plus a halo);
- **repair**: how the hole is refilled (greedy most-constrained-first, the same
with jittered tie-breaks, or a bounded exact refill for small holes);
- **acceptance**: whether a candidate replaces the working board (keep if not
worse, strict improvements only, simulated annealing, or late-acceptance);
- **restart**: what happens on a stall (nothing, a random kick, or a revert to
the best board so far).
A variant is a small record naming those five choices, together with **the parent
it derives from and a one-line description of the single change it adds**. Adding
a variant means adding one record to the registry, with no new loop code unless
the idea is a genuinely new strategy. The "what stacks on what" matrix on the
results page is generated from those descriptions, so it cannot drift from the
code that ran.
## The mismatch map, maintained incrementally
A backtracker builds a board up from nothing; a repair loop always holds a
*complete* board and edits it. The cells worth attacking are the ones touching a
broken edge, so the engine keeps a live count, per cell, of the broken interior
edges incident to it. Placing or lifting a piece updates only the edges around
that one cell, never the whole board, so a conflict-driven destroy operator can
ask "which cells touch a mismatch?" without rescanning. The running score is kept
the same way: each placement adjusts it by the handful of seams that changed. A
full rescan happens exactly once, when the starting board is built; from then on
the loop is incremental. This is what makes hundreds of thousands of iterations
in sixty seconds possible, and it is the same discipline the
[ALNS theory page](/research/build/local-search/local-search-alns) describes as
keeping destroy at cost proportional to the hole, not the board.
## The scorer is the single source of truth
No engine's self-reported score is trusted. The loop's incremental score is a
performance device; the *published* number is always a fresh canonical re-score
of the output board, through the same scorer the site and the DFS study use
(matched, non-border, interior adjacencies, counted right and down per cell). A
test asserts the incremental score and the canonical score agree after thousands
of random place-and-lift edits, so the fast path can be trusted to track the
truth rather than drift from it.
## The statistics the study raises
For every run the engine records, and the results carry all the way to the page:
- **final score**: the canonical matched edges (of 480) of the best board found.
- **lift**: final score minus the starting board's score. This isolates the
repair loop's own contribution from the construction it began with: a variant
that starts high can add little and still finish high, and the lift is what
separates the two.
- **the stall (last-best iteration vs total iterations)**: the iteration at which
the global best last improved, against how many iterations the budget bought.
This is the study's headline axis: when the first is far below the second, the
run found its answer early and then moved nothing.
- **accept rate**: the fraction of iterations the acceptance rule kept. A rate
near zero means the loop is proposing changes it almost always rejects, often
a sign it is re-attacking the same region.
- **mean destroy size**: cells lifted per iteration, so a large-hole operator is
not silently compared to a small-hole one.
- **iterations per second**: reported per variant and **never compared across
families**, because an iteration running an exact refill is not the same unit of
work as one running a greedy fill.
- **restarts**: perturbations fired, zero for a variant with no restart policy.
The page also draws a **convergence curve** for a few representative variants: the
best score so far, sampled every two hundred iterations, as a typical run
progresses. It is the clearest picture of the stall: the curve rises steeply,
then goes flat while the iteration count keeps climbing.
## Honesty about the greedy engine's speed
Like the DFS study's MRV engine, the repair loop here is written for clarity
first. The greedy refill rescans the remaining piece pool for every cell it
fills, rather than maintaining candidate lists incrementally, so its
iterations-per-second is this clean engine's rather than the best a tuned repair
kernel could reach. The ranking by score does not depend on it, since throughput
is a separate axis never mixed into the score comparison, but the iteration rates
should be read as this engine's, not as the ceiling for destroy-and-repair.
## What this study deliberately does not claim
This is a study of the *plain* repair loop, one decision at a time, at a small
fixed budget. It is not the record pipeline: the boards that reach the mid-450s
and beyond combine a strong constructive producer, longer runs, and repair as a
final polish, and several of the operator findings here would read differently on
a near-record board than they do on the mediocre greedy start the loop begins
from. Where a result depends on the starting board's quality, the findings page
says so. The community's own tuning found, for instance, that conflict-driven and
component operators are valuable precisely on near-optimal boards where the few
remaining mismatches are the only thing left to fix, the regime this study's
short runs never reach.
## Reproducibility
The engine, the ten variants, the committed per-run results and the grid scripts
all live under the study's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/repair-study).
`just experiments repair-study` rebuilds the engine and reruns the whole grid;
the run is deterministic at a fixed seed, and the corner arrangement is the only
diversity axis.
## Related
- [The repair study](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study) — The sibling of the DFS study, for the other way people attack Eternity II: destroy part of a board, rebuild it, keep the change if it helps. One question, asked carefully. What does each decision in that loop buy: which region to destroy, how to rebuild it, when to keep a move, when to restart, and what board to start from?
- [What each decision buys](https://eternity2.dev/research/lab/experiments/raphael-anjou/repair-study/findings) — The five comparisons at the heart of the repair study, worked through: blind random destroy wins while every conflict-targeting operator loses; construction sets the floor; simulated annealing is the strongest acceptance rule; and the clever refinements (exact refill, restarts) buy nothing at this budget.
- [Local search and ALNS](https://eternity2.dev/research/build/local-search/local-search-alns) — Destroy part of a board, rebuild it better, and let the algorithm learn which demolitions pay. Adaptive large-neighborhood search is the most reliable polisher this project has, and the cleanest demonstration of the wall where polishing ends.
---
# Single-core benchmark
> Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/single-core-benchmark/
- Updated: 2026-07-13
- Topics: speed, construction, backtracking
- Reproduce: `just experiments single-core-benchmark`
- Source: Runnable engine + committed results + scripts (this experiment's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark
---
Give every solver the same budget, one core and one minute, and which one wins?
The answer overturns the obvious intuition. Several solvers, ours and our
implementations of the community's two record backtrackers, each ran once on ten
corner-pinned variants of the official puzzle, single-threaded, 60 seconds a run.
Every board was re-scored by one canonical scorer; no engine's self-reported
score is trusted. The maximum possible score is 480.
> **Every engine here is our code**
>
> `blackwood_style` and `verhaard_style` are **our from-scratch implementations** of Joshua Blackwood's and Louis Verhaard's published algorithms, not the authors' own programs. Verhaard's `eii` only ever shipped as a Win32 binary, so a re-implementation (its constants recovered from `eii.exe`) is the only way to run it at all. Blackwood's real C# *is* [public](https://github.com/jblackwood345/EternityII_Solver), but it hardcodes its 256 pieces and its thread count, so it cannot read this grid's variants or be pinned to one core without editing it. Scores here measure our reading of each algorithm, not the authors' engineering, and should not be quoted as "Blackwood scores N".
The leaderboard below shows the methods that compete on score. A second family,
the CSP presets (one arc-consistency engine run under a dozen ordering knobs),
is a different category: its best preset reaches about 183, less than half a
contender's score, so it earns no leaderboard row. That preset sweep is a study
in its own right, on its [own page](/research/lab/experiments/single-core-benchmark/csp-presets).
## What a minute of one core buys
The two record backtrackers dominate this budget. Our Verhaard-style engine
reaches a mean of **440.8** (best 451) and our Blackwood-style engine **436.4**
(best 440), both exploring 20 to 40 million search-nodes per second. Nothing
else comes close: naive DFS lands at **365.6**, and the best CSP preset at about
**183**.
The gap between those two groups is the finding. All four families see the same
puzzle and the same 60 seconds, and they finish 250 points apart. What separates
them is not speed: the CSP engines are three orders of magnitude slower per
node than the backtrackers and still beat naive DFS on friendly variants,
because each node is pruned rather than merely visited. Node count and score are
not the same axis.
What this grid cannot tell you is where the ceiling is. The best board here (451)
is still 13 points short of the community's 5-clue record of 464, and the top
engine's mean sits about 24 short; the records were not set in a minute, they
came from farms and months. This measures
per-core efficiency at a fixed small budget, nothing more.
> **[Figure]** The leaderboard: mean score over ten corner variants, single core, 60 s — interactive: BenchmarkLeaderboard. Rendered on the canonical page (link above); not shown in this markdown export.
## Reading the table
The two backtrackers are the top of this board and they are stable there: the
Verhaard-style engine spans 437 to 451 across the ten variants, the
Blackwood-style engine 431 to 440. The naive baseline is the floor: a fast,
junk-filled 365 that shows what raw matched-edge count looks like without any
board quality.
Throughput units differ by family and are never cross-compared. The backtrackers
count search-nodes per second, the CSP engines the same but at 5 to 10 thousand,
because each of their nodes runs full arc-consistency. The CSP family trades
throughput for pruning, which is why it explores far fewer nodes yet still beats
naive DFS on good variants.
## What each contender is doing
- **Verhaard and Blackwood (backtrackers, ranks 1 to 2).** Depth-gated-break DFS
at 20 to 40 million nodes per second. They punch to 437 to 451 but hit a wall:
the endgame needs far more compute than 60 seconds allows. Blackwood found his
470 after about a month on a single PC, and called it "a stroke of luck"
([message 10194](https://groups.io/g/eternity2/message/10194)).
- **Naive DFS (the baseline).** Fills the whole board allowing every break. In
row-major order (`anjou-naive_rowmajor`) it reaches a high matched-edge count
(365) but a low-quality board full of breaks. Fast, and junk. It sits on the
board as a floor: the number a method has to clear to be worth
anything. (The visit-order sensitivity of naive DFS, and why the spiral variant
collapses to 78, is on the [CSP presets page](/research/lab/experiments/single-core-benchmark/csp-presets)
alongside the other same-engine ablations.)
## Where these numbers sit
This grid caps every engine at one core for 60 seconds, far below the compute
that produced the records below. It measures per-core efficiency and heuristic
quality, not peak reachable score. An engine that ranks high here reaches good
boards cheaply; the record numbers need many cores times hours.
Every variant here pins all 5 official clues (plus 3 corners, so 8 hints in
total), so the relevant ceiling is the **5-clue community record of 464**, not
the all-hints 470 (which uses more than the 5 clues). That 464 is what the
leaderboard's dashed line marks.
| reference | score | conditions |
|:--|--:|:--|
| Community 5-clue record | 464 | Benjamin Riotte, July 2026 (same 5 clues these variants pin) |
| This grid's best (Verhaard-style) | 451 | one core, 60 s (mean 440.8 over 10 variants) |
| Blackwood-style in this grid | 440 | one core, 60 s (mean 436.4 over 10 variants) |
| Community all-hints ceiling | 470 | Blackwood, ~1 month on a single PC, uses more than 5 clues |
## Method and reproducibility
Each of the ten variants is the official puzzle plus three pinned corner cells
(distinct corner-piece arrangements), so all ten share the 256-piece set and 5
clue hints but differ in three corner constraints. They are emitted as both
site-schema JSON (for the native engines) and CSV (for the standalone engines)
from one generator, so every algorithm sees identical instances. One run per
puzzle, fixed seed; the corner arrangement is the only diversity axis. Every run
emits a bucas `.url`, and the score is the canonical matched-edge count from the
same scorer, never the engine's self-report. There were zero failures across all
150 runs.
Every engine in the grid is our own open-source code and runs from this
repository, including the two written from the community's published algorithms;
no third-party solver is vendored here. The grid measures all
of them; the leaderboard above shows only the five that compete on score, and the
[CSP presets page](/research/lab/experiments/single-core-benchmark/csp-presets)
shows the rest. The crate workspace, the ten variants, the committed per-run
results and the grid scripts all live under the experiment's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark),
and `just experiments single-core-benchmark` builds the engines and reruns the
whole grid. The native family (the naive and CSP presets) is one binary selected
by preset; the two backtrackers are a binary each, driven by a small wrapper.
Both speak the same puzzle-in,
bucas-url-out contract, and every board is re-scored by the one canonical scorer.
## Companion findings
Profiling put 96.6 percent of the CSP family's time in one heavily pre-optimised
AC-3 loop. That engine is already at its performance ceiling; the benchmark
speeds are its real speeds, not an implementation gap. Separately, the
community's Blackwood constants do not
transfer across colour labelings: its published privileged colours scored 387 in
our labeling versus 435 for a re-fit, a 48-point gap that our Blackwood closes by
re-fitting only the labeling-relative colour IDs while matching every structural
element of the published spec.
## Open work
The grid runs our implementations. Two of the community's three record engines
can in principle be run directly, and that is the obvious next measurement:
- **Peter McGavin's C generator.** He posted the source to the list in January
2026 as `genbody71.zip`
([message 11749](https://groups.io/g/eternity2/message/11749)). It builds on
Apple silicon with `clang` and its generate pass emits a 10,363-line `body.c`
specialised to one puzzle. It is the fastest engine the community has
measured, at 295M placements/s on Joe's CPU
([message 11750](https://groups.io/g/eternity2/message/11750)), and it is
absent from this grid.
- **Joshua Blackwood's C#.** [Public and
GPL-3.0](https://github.com/jblackwood345/EternityII_Solver); it builds
unmodified on .NET 8. But it hardcodes all 256 pieces in `Util.cs`, fixes
`number_virtual_cores = 64`, and takes no arguments, so pinning it to one core
or feeding it this grid's variants means editing his source, at which point
the artifact is no longer purely his. Running it as published, on the plain
puzzle, is the faithful form of that measurement.
- **Louis Verhaard's `eii`** cannot be run at all: the download from his own
site ships `eii.exe` and no source, which is why our engine reconstructs it
from the binary.
Neither engine is vendored into this repository; both are fetched and run
locally when measured.
## Pages in this section
- [CSP presets, measured](https://eternity2.dev/research/lab/experiments/single-core-benchmark/csp-presets) — One constraint-propagation engine, run under a dozen ordering and propagator presets, on the same ten corner-pinned variants as the leaderboard. A study of what each knob buys, kept off the headline board because the best preset reaches less than half a contender's score.
## Related
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
- [Beam search](https://eternity2.dev/research/build/construct/beam-search) — Keep the K most promising partial boards alive at once and grow them cell by cell. Beam search is the workhorse behind this project's from-scratch builders, and a clean illustration of why breadth alone stalls in the deep interior.
---
# CSP presets, measured
> One constraint-propagation engine, run under a dozen ordering and propagator presets, on the same ten corner-pinned variants as the leaderboard. A study of what each knob buys, kept off the headline board because the best preset reaches less than half a contender's score.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/lab/experiments/single-core-benchmark/csp-presets/
- Updated: 2026-07-15
- Topics: backtracking, search-space
- Reproduce: `just experiments single-core-benchmark`
- Source: Runnable engine + committed results + scripts (this experiment's backing directory) — https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark
---
The [single-core benchmark](/research/lab/experiments/single-core-benchmark)
leaderboard shows the methods that compete on score. This page shows the rest:
a dozen presets of one constraint-satisfaction engine, plus the naive DFS in its
weaker visit order. They are not a dozen solvers. They are one arc-consistency
core wearing different heads, kept so the exact cost of each classic technique
can be read off a number instead of argued about.
> **[Figure]** CSP presets: mean score over ten corner variants, single core, 60 s — interactive: BenchmarkLeaderboard. Rendered on the canonical page (link above); not shown in this markdown export.
## What each knob buys
The presets vary two things: the order the engine assigns cells, and how hard it
propagates before committing. Holding the puzzle and the budget fixed, the score
gap between two presets is the value of that one knob.
- **Least-constraining-value ordering helps.** `anjou-gacolor_ac3_lcv` reaches a
mean of 114 against the plain `anjou-gacolor_ac3` at 76: choosing the value
that rules out the fewest neighbours is worth about 38 points here.
- **Some knobs are inert at this depth.** Three presets (`anjou-gacolor_ac3`,
`anjou-gacolor_ac3_ns1`, `anjou-verhaard_preferred`) produced byte-identical
boards on this puzzle. NS-1 and preferred ordering add nothing at 60 seconds:
the engine never searches deep enough for them to bite.
- **Border-first ordering is the strongest preset.** `anjou-border_first_lcv` and
`anjou-rare_color_first` top the sweep at about 183, by committing the
over-constrained frame before the free interior. It is still less than half a
contender's score.
## Why none of them competes
Every preset is bimodal. On a friendly corner arrangement the arc-consistency
search reaches 340 to 349; on a hostile one it collapses to about 55, pinned in a
bad basin it cannot escape inside 60 seconds. The mean sits low because the bad
corners drag it down, and no ordering knob fixes the basin problem. That is the
real story of this family: propagation makes each node well-pruned but expensive,
so the search is strong where the instance is forgiving and helpless where it is
not. The constructive engines on the leaderboard never enter that trap, which is
why a preset at 183 sits on a different page from the leaderboard's contenders,
which top out at 451.
## The naive DFS is here too, for one reason
Naive depth-first search with every break allowed is not a CSP preset, but its
visit order belongs to the same lesson. In row-major order
(`anjou-naive_rowmajor`, on the leaderboard as the baseline) it reaches 365; the
same DFS in spiral visit order (`anjou-naive_spiral`) collapses to a mean of 78.
The visit-order choice alone costs the naive search nearly 290 points, the same
kind of ordering sensitivity the CSP presets show, at a larger scale.
## Method and reproducibility
These presets ran in the same grid as the leaderboard: ten official-puzzle
variants, each with three pinned corner cells, single core, 60 seconds, fixed
seed, one run per variant. Every board is re-scored by the one canonical
matched-edge scorer, never the engine's self-report. The engine, the ten
variants, the committed per-run results and the scripts live under the
experiment's
[backing directory](https://github.com/raphael-anjou/eternity2/tree/main/research/experiments/single-core-benchmark),
and `just experiments single-core-benchmark` reruns the whole grid, contenders
and presets together.
## Related
- [Single-core benchmark](https://eternity2.dev/research/lab/experiments/single-core-benchmark) — Fifteen solvers, ours and our implementations of the community's two record backtrackers, each run once on ten corner-pinned variants of the official puzzle, single core, 60 seconds per run. The finding: node count is not score.
- [Arc consistency, from AC-3 up](https://eternity2.dev/research/build/reduce/arc-consistency) — Forward checking looks one move ahead; arc consistency makes every cell's candidate list defend itself against every neighbour's, to a fixed point. Mackworth's AC-3, the optimal refinements that followed, and what the whole family actually measured on this puzzle, including where it is unsound.
- [Fill orders](https://eternity2.dev/research/build/backtracking/fill-order) — The order in which a backtracker visits the 256 cells is its one free choice: it costs nothing at runtime and moves the size of the search tree by orders of magnitude. Twenty years of community science, from the fixed-vs-dynamic wars and the strategy races to the magic 10×16 square and Verhaard's comb search, all answer the same question: which path through the board is cheapest?
---
# Papers
> The academic literature on Eternity II and edge-matching puzzles, drawn from the project's research notes and the community reading list, and ranked by how useful each paper actually is if your goal is to write a solver.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/papers/
- Updated: 2026-07-15
- Topics: exact-methods, search-space
---
The short version: the complexity results tell you why it's hard, the SAT/CSP
papers explain why off-the-shelf solvers hit a wall, and the
constraint-propagation and large-neighborhood papers are the ones whose ideas
show up in the strongest solvers.
> **[Interactive: PapersView]** Rendered on the canonical page (link above); not shown in this markdown export.
---
# Who's who of E2 research
> Two decades of Eternity II research were done by named people on a mailing list. This page is the gallery: who they are, what each of them contributed, and where to read it in their own words. A thank-you as much as an index.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/people/
- Updated: 2026-07-02
- Source: Brendan Owen derives the 17+5 design as the hardest possible puzzle (groups.io message 1947) — https://groups.io/g/eternity2/message/1947
- Source: Louis Verhaard's first-person account of the 467: “it was my program that did the job” (groups.io message 6891) — https://groups.io/g/eternity2/message/6891
- Source: Peter McGavin solves Brendan Owen's 10×10 benchmark (groups.io message 9686) — https://groups.io/g/eternity2/message/9686
- Source: Joshua Blackwood's 470, the standing record (groups.io message 10117) — https://groups.io/g/eternity2/message/10117
- Source: Al Hopfer states the NS-1 border balance (groups.io message 10754) — https://groups.io/g/eternity2/message/10754
- Source: Brendan Owen returns to the list after fourteen years (groups.io message 11500) — https://groups.io/g/eternity2/message/11500
---
The [history pages](/research/community/hunt) tell the community's story in
order; this page tells it by person. Nearly everything this wiki knows
(every record, theory, tool and
[documented dead end](/research/build/dead-ends)) traces back to someone who
posted it on the [eternity2 mailing list](https://groups.io/g/eternity2),
usually for free, often for years. Consider this gallery an index and a
thank-you at the same time. Names appear as people signed their public
posts; every claim links to a message.
Many of the people below have their own contributor page: a name in **bold
with a link** opens it, gathering their profile, their sourced posts, and any
pages they have written here. [Raphaël Anjou](/research/people/raphael-anjou),
who maintains the wiki and runs the lab experiments, has one too, one
researcher among the others. The names without a link are documented right
here, on this page.
## The founders and launch analysts (2000–2007)
### [Brendan Owen](/research/people/brendan-owen)
The founder. Owen created the eternity_two group in October 2000, six and a
half years before the puzzle existed
([msg 384](https://groups.io/g/eternity2/message/384)), and set its
scientific tone: within two days of the January 2007 announcement he had
derived the expected-solutions formula that makes difficulty a design
parameter ([msg 38](https://groups.io/g/eternity2/message/38)), and on
launch weekend he digitized the real pieces and published the first
solution-count estimate
([msg 987](https://groups.io/g/eternity2/message/987)). His signature
results: the proof that a 17+5 colour split is the hardest possible 16×16
([msg 1947](https://groups.io/g/eternity2/message/1947)), the exact
search-tree model now called
[complex theory](/research/why/complex-theory)
([msg 5197](https://groups.io/g/eternity2/message/5197),
[msg 5209](https://groups.io/g/eternity2/message/5209)), the 9×9/10×10
[benchmark puzzles](/research/build/benchmarks) the community still races
on, and the closed-form
peak-depth proof 256 × (1 − 1/e)
([msg 8125](https://groups.io/g/eternity2/message/8125)). After a farewell
at the contest's close
([msg 8429](https://groups.io/g/eternity2/message/8429)) he returned in
2025, refining his own model
([msg 11500](https://groups.io/g/eternity2/message/11500),
[msg 11546](https://groups.io/g/eternity2/message/11546)).
### [Günter Stertenbrink](/research/people/gunter-stertenbrink)
An Eternity I veteran and the group's earliest provocateur of estimates: in
2001 he asked, years ahead of the fact, how one would design a puzzle with a
huge prize and only a ~1% chance of being solved in ten years
([msg 15](https://groups.io/g/eternity2/message/15)), and he greeted the
2005 press claims with "So we can conclude, the end of the universe is in
several years." ([msg 34](https://groups.io/g/eternity2/message/34)). For
two decades he checked the list's numbers, from exact-cover conversions in
2007 to the nodes-per-watt hardware ledger of the 2010s.
### [Dave Clark](/research/people/dave-clark)
Author of the distributed Eternity I solver ESolve, he rejoined in 2001
([msg 21](https://groups.io/g/eternity2/message/21)) and built eternity2.net,
the BOINC project that was the community's public face in 2007
([msg 756](https://groups.io/g/eternity2/message/756)). He shut it down with
full accounting: 1.6 TFlops, over 10^19 operations, no solution
([msg 3511](https://groups.io/g/eternity2/message/3511)). He open-sourced
his solver ([msg 3716](https://groups.io/g/eternity2/message/3716)) and left
the archive its best primary source on the puzzle's creation: his phone call
with Monckton describing how the solution was generated and vaulted
([msg 4177](https://groups.io/g/eternity2/message/4177)).
### Txibilis
Angel de Vicente, who signed Txibilis, built the standard suite of E2-like
benchmark boards ([msg 1886](https://groups.io/g/eternity2/message/1886))
and was the community's best hand-designer of
[fill orders](/research/build/backtracking/fill-order). His duel with
doc_s_smith's automated optimizer drove full-search node counts down by
orders of magnitude
([msg 2928](https://groups.io/g/eternity2/message/2928)). The benchmark
culture that later validated complex theory starts with him.
### doc_s_smith
Mid-duel, the list discovered who doc_s_smith was: Dietmar Wolz, finder of
most of the known Eternity I solutions
([msg 2972](https://groups.io/g/eternity2/message/2972)). His automated
strategy optimizer set benchmark records in 2007
([msg 2896](https://groups.io/g/eternity2/message/2896)); returning in 2010
he published a Java toolbox that turned the list into an algorithms workshop
([msg 7755](https://groups.io/g/eternity2/message/7755)) and gave the
post-contest era its working goal: "beat 468 matching edges"
([msg 7803](https://groups.io/g/eternity2/message/7803)).
### kubzpa
Author of the first serious solution-count paper, putting E2 near 15 million
solutions ([msg 3497](https://groups.io/g/eternity2/message/3497)), and of
the era's cleanest impossibility result: the parity argument that no board can
score exactly 479 by its interior seams
([msg 1640](https://groups.io/g/eternity2/message/1640)). It held for seventeen
months, until Verhaard pointed out the one loophole
([msg 6317](https://groups.io/g/eternity2/message/6317)): flip a border piece
whose two outward border edges share a colour, and the board reads as 479 while
those unscored rim edges go untouched. That is a technicality of the unscored
border, not a break in the interior-parity math, which still holds, as Verhaard
himself noted that 479 "cannot be achieved in another way"
([msg 6319](https://groups.io/g/eternity2/message/6319)).
### mjqxxxx
Michael Quist was the list's mathematical referee. He posted the first fully
rigorous counting framework for E2-like puzzles
([msg 1221](https://groups.io/g/eternity2/message/1221)), sharpened the
border-balance theory
([msg 2098](https://groups.io/g/eternity2/message/2098)), and his reviews
caught the flaws that made others' results solid. Kubzpa amended his
solution-count paper after his pass caught a flawed Monte Carlo run
([msg 3589](https://groups.io/g/eternity2/message/3589)).
## The prize years (2007–2010)
### [Louis Verhaard](/research/people/louis-verhaard)
The only person the puzzle ever paid. His eii solver, released publicly
"because I am stuck"
([msg 5940](https://groups.io/g/eternity2/message/5940)), found the 467
that won the $10,000 scrutiny prize, entered under the name of his wife,
Anna Karlsson, as he confirmed himself: "Anna is my wife… it was my program
that did the job"
([msg 6349](https://groups.io/g/eternity2/message/6349),
[msg 6891](https://groups.io/g/eternity2/message/6891),
[msg 7451](https://groups.io/g/eternity2/message/7451)). His methods became
canon: comb-search fill orders
([msg 6112](https://groups.io/g/eternity2/message/6112)) and depth-gated
edge slipping ([msg 7321](https://groups.io/g/eternity2/message/7321)). He
was also complex theory's staunchest defender: "the finest work that has
ever been published about E2"
([msg 7810](https://groups.io/g/eternity2/message/7810)), and his 467 stood
for twelve years.
### Yannick Kirschhoffer
Author of the Eternity II Editor, the cross-platform Java editor and solver
GUI released in February 2008
([msg 4544](https://groups.io/g/eternity2/message/4544)) that became the
community's standard board tool for years. He was still offering help with
its code when it resurfaced in 2012
([msg 9064](https://groups.io/g/eternity2/message/9064)).
### Fred
Posting as Eternity Blogger, Fred built E2Lab in a burst of near-daily
releases in autumn 2009 ([msg 7148](https://groups.io/g/eternity2/message/7148)),
an editor/solver whose deliberate removal of its own "magic button", "to
respect the game rules", says a lot about the list's ethics
([msg 7150](https://groups.io/g/eternity2/message/7150)). His blog hosted
the community's tables through the post-contest years.
### [Al Hopfer](/research/people/al-hopfer)
A regular since 2008 and the community's border theorist. His "balance
doctrine" for puzzle generation appears in 2009
([msg 6842](https://groups.io/g/eternity2/message/6842),
[msg 6844](https://groups.io/g/eternity2/message/6844)); in 2022 he stated
the exact condition this wiki calls the
[NS-1 border balance](/research/why/border-balance)
([msg 10754](https://groups.io/g/eternity2/message/10754),
[msg 10757](https://groups.io/g/eternity2/message/10757)), and backed it
with a fully documented 222-piece partial with completed border
([msg 10862](https://groups.io/g/eternity2/message/10862)).
## The long decade (2010–2019)
### [Peter McGavin](/research/people/peter-mcgavin)
The era's anchor, and arguably the puzzle's most consequential researcher
after Owen. His work has [its own page](/research/lab/experiments/peter-mcgavin/backtracker).
He computed the canonical ~14,702 expected solutions
([msg 8924](https://groups.io/g/eternity2/message/8924)), transcribed
complex theory into LaTeX
([msg 9188](https://groups.io/g/eternity2/message/9188)) and later into
exact C code ([msg 11197](https://groups.io/g/eternity2/message/11197)),
which this site ports. In 2017 he solved Owen's hint-free 10×10 in ~180
core-years, inside the theory's error bars, its strongest validation ever
([msg 9686](https://groups.io/g/eternity2/message/9686),
[msg 9688](https://groups.io/g/eternity2/message/9688)), and in 2020 he
held the record himself: "New record score of 469! Only 11 breaks!"
([msg 10045](https://groups.io/g/eternity2/message/10045)).
### Tony Wauters
Academia in person. He posted his group's peer-reviewed hyper-heuristic
paper (461/480 in an hour) and stayed to answer questions
([msg 9017](https://groups.io/g/eternity2/message/9017),
[msg 9023](https://groups.io/g/eternity2/message/9023)), and his group
followed with the 2017 MILP and Max-Clique work
([msg 9683](https://groups.io/g/eternity2/message/9683)).
### Michael Field
A fastest-solver veteran of the early years who became the group's hardware
realist: his FPGA backtracker design projected ~5G placements per second per
chip ([msg 9226](https://groups.io/g/eternity2/message/9226)), and his
capacity analyses of GPU and FPGA routes told the list what silicon could
and could not buy ([msg 9003](https://groups.io/g/eternity2/message/9003)).
### Arnaud Carré
Arrived in 2009 and reset the speed standard in 2014 with a 114.5 million
recursions per second single-core solver offered as a comparison baseline
([msg 9233](https://groups.io/g/eternity2/message/9233)), returning in 2018
for the benchmark races.
### Adam Miles
The GPU flank. Arriving in 2017, he moved from CPU bit-tricks to a DirectX
12 compute solver running on an Xbox One X
([msg 9811](https://groups.io/g/eternity2/message/9811)) and re-verified
Owen's 9×9 set 1 exhaustively on GPU: the same 2 solutions as the 2014 CPU
census, in 25.4 hours
([msg 9822](https://groups.io/g/eternity2/message/9822)).
### JSA
The community's verifier and, later, its rescuer. In 2009 he reproduced the
467 with Verhaard's public solver, ~82 days on one PC, logging exactly how
thin the air gets above 466
([msg 6687](https://groups.io/g/eternity2/message/6687)). When Yahoo
announced it would erase the archive in 2019, JSA paid the groups.io
transfer fee, offering "I can pay for the first 5 years"
([msg 2](https://groups.io/g/eternity2/message/2)), and he still writes the
group's welcome notes
([msg 11771](https://groups.io/g/eternity2/message/11771)).
### Ole Knudsen
Kronjuvel (Kron) was there from the earliest years, retro-claimed a
231-piece partial from October 2007
([msg 7563](https://groups.io/g/eternity2/message/7563)), and as group owner
created the new groups.io home during the 2019 migration. The 2026 welcome
note opens with a tribute to him; he has been missing from the list since
2023 ([msg 11771](https://groups.io/g/eternity2/message/11771)).
## The record wave and the modern era (2019–2026)
### [Joshua Blackwood](/research/people/joshua-blackwood)
The outsider who ended the twelve-year freeze. Unknown to the list, he
announced a 468 on Reddit in August 2020
([msg 10032](https://groups.io/g/eternity2/message/10032)), open-sourced his
solver days later ([msg 10037](https://groups.io/g/eternity2/message/10037))
along with rare negative results (SAT, GPU and 2×2 caches all measured and
discarded, [msg 10056](https://groups.io/g/eternity2/message/10056)), and
in March 2021 posted the 470 that still stands
([msg 10117](https://groups.io/g/eternity2/message/10117)), found with the
exact public code ([msg 10161](https://groups.io/g/eternity2/message/10161)).
His algorithm is [decoded on this wiki](/research/lab/experiments/joshua-blackwood/solver).
### [Jef Bucas](/research/people/jef-bucas)
The modern era's infrastructure. He raised the alarm that triggered the
archive migration ([msg 9920](https://groups.io/g/eternity2/message/9920)),
built the [e2.bucas.name](https://e2.bucas.name) board viewer that became
the community's record book
([msg 9955](https://groups.io/g/eternity2/message/9955)), rewrote
Blackwood's solver in C as libblackwood, roughly doubling its speed and
powering the November 2020 wave of 469s
([msg 10065](https://groups.io/g/eternity2/message/10065),
[msg 10078](https://groups.io/g/eternity2/message/10078),
[msg 10067](https://groups.io/g/eternity2/message/10067)), and tied the 470
in 2024 ([msg 11401](https://groups.io/g/eternity2/message/11401)), always
crediting Blackwood. His wrapper_blackwood parameter study is republished
[on this wiki](/research/lab/experiments/joshua-blackwood/solver) with his permission
([msg 11905](https://groups.io/g/eternity2/message/11905)).
### Carlos Fernandez
The board surgeon. He produced a 469 by a single-piece swap of McGavin's
record board ([msg 10074](https://groups.io/g/eternity2/message/10074)), a
border-rearranged 470 variant
([msg 11403](https://groups.io/g/eternity2/message/11403)), four-minute
14×14 quadrant solves
([msg 10802](https://groups.io/g/eternity2/message/10802)), and high rungs
of the five-clue ladder
([msg 11068](https://groups.io/g/eternity2/message/11068)).
### Bruno Gauthier
A speed veteran of the 2014 era whose Forth solver ran at 80–90 million
nodes per second ([msg 9265](https://groups.io/g/eternity2/message/9265)),
he held the strictest record on the books for over three years: 460/480 with
all five clue pieces at their official positions, from 2023
([msg 11074](https://groups.io/g/eternity2/message/11074)) until Benjamin
Riotte's 464 in July 2026.
### Benjamin Riotte
Holder of the strict five-clue record. In July 2026 he pushed the best board
respecting all five clue placements from Gauthier's long-standing 460 to
**464/480** (16 broken edges), with his own modified-Blackwood DFS
([groups.io](https://groups.io/g/eternity2/message/11919)). Igor Pejic
reached the same 463–464 range independently in the same thread. It was the
first movement on the strict-canonical line in over three years.
### [Marijn Heule](/research/people/marijn-heule)
The SAT world's contact point. In the long-running SAT thread, a
collaborator of Marijn Heule reported that Heule's group had re-implemented
and improved the encoding behind the 2008 SAT benchmark results
([msg 10969](https://groups.io/g/eternity2/message/10969)), the
state of the art of the exact-methods flank, which was still debating
4 GB CNF encodings on the archive's very last day
([msg 11822](https://groups.io/g/eternity2/message/11822)).
### [Raphaël Anjou](/research/people/raphael-anjou)
Maintains this wiki and runs the experiments catalogued in
[the lab](/research/lab); his write-ups gather on his
[contributor page](/research/people/raphael-anjou), one researcher among the
others.
### [William Millilaw](/research/people/william-millilaw)
Ran a dense two-week solver campaign in 2026, most of it recorded as
refutations of methods that do not crack the puzzle. Two of his findings sharpen
the ceiling: a replica freeze test, and a halo SAT-residual test showing the
record boards are strict local optima. We reproduced the second on the public
boards, and it lands on the
[rigidity wall](/research/why/rigidity-wall) alongside the integer-programming
proofs
([reproduction](https://github.com/raphael-anjou/eternity2/tree/main/research/topics/rigidity-sat-halo)).
### onesmallstep
Founded the community's Discord server in November 2021 and kept it alive
through the quiet years (reported on the community Discord, November 2021; no
public message link). A dedicated solver
in his own right (self-reported best around 466–467), he is also the reason
the wiki once mis-carried a "2025 470": a record *relay* misread as a
record *claim*, corrected here from the Discord archive itself.
### Reinout Annaert
The Discord era's methodical hunter: a self-reported best of 469, the
linear-run subculture (229- and 230-piece consecutive partials), and the man
who settled where the hint placements come from: "They directly come from
Tomy's Hint Puzzles" (reported on the community Discord, December 2024; no
public message link). On the mailing list he confirms storing solution figures
scoring above 467/480 ([msg 11549](https://groups.io/g/eternity2/message/11549)).
His suggestions also shaped this site's playground roadmap.
## And many more
No gallery this size is complete. Among the many who belong here: Alan
O'Donnell, who had the first working solver weeks after the announcement
([msg 64](https://groups.io/g/eternity2/message/64)); Max, Verhaard's
sparring partner in the run to 467, whose own best was 465
([msg 6348](https://groups.io/g/eternity2/message/6348)); istarinz, who
broke 558 million placements per second in 2008 and became the group's
exhaustive-search verification authority
([msg 6212](https://groups.io/g/eternity2/message/6212)); antminder, whose
assignment-repair hybrid averaged a 462 per day in 2008
([msg 5589](https://groups.io/g/eternity2/message/5589)); Pierre Schaus,
whose constraint-programming paper supplied that repair operator
([msg 5601](https://groups.io/g/eternity2/message/5601)); capiman26061973,
founder of the five-clue ladder
([msg 11037](https://groups.io/g/eternity2/message/11037)) and the
invalid-combination mining programme
([msg 7768](https://groups.io/g/eternity2/message/7768)); David Barr,
open-source GPU and browser solvers across a decade
([msg 9367](https://groups.io/g/eternity2/message/9367),
[msg 11121](https://groups.io/g/eternity2/message/11121)); Henk van der
Griendt, who found the prize announcement everyone else had missed
([msg 6337](https://groups.io/g/eternity2/message/6337)); and
juraj.pivovarov, the community's rapid-prototyping conscience
([msg 9411](https://groups.io/g/eternity2/message/9411)).
## Corrections welcome
This page will always be incomplete, and it may be wrong in places: a
misattributed result, a missing name, a preferred spelling. If any of it
concerns you or your work, please say so on the
[mailing list](https://groups.io/g/eternity2) or through the
[contribute page](/research/contribute): corrections land with the same
sourcing rules as everything else here, and credit is the whole point of
this page.
## Related
- [The hunt, a history (part I: 2000–2009)](https://eternity2.dev/research/community/hunt) — The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
- [The hunt, a history, part II: 2009–2026](https://eternity2.dev/research/community/hunt-part-2) — Seventeen years after the prize: the contest dies with its solution locked in a safe, 467 stands for a decade, the archive survives Yahoo's shutdown by days, and then an outsider from Reddit rewrites the record book. Every event sourced to its original message.
- [Records & solvers](https://eternity2.dev/research/records) — Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
---
# Records & solvers
> Eternity II has never been solved, but fifteen years of community effort have pushed the best board to 470/480. Who holds what, how they did it, and why some headline "480" boards are not actually the real puzzle.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/records/
- Updated: 2026-07-16
- Source: eternity2 mailing-list archive (groups.io): the record announcements; free account required to read — https://groups.io/g/eternity2
- Source: Wikipedia: Eternity II puzzle (the $2M prize, the 2010 deadline and Verhaard's 467) — https://en.wikipedia.org/wiki/Eternity_II_puzzle
- Source: Louis Verhaard's own account of his 467 solver — https://www.shortestpath.se/eii/eii_details.html
- Source: e2.bucas.name (Jef Bucas): the community board viewer; every linked board can be re-scored there — https://e2.bucas.name
---
This page tracks the score: who held the best board at each point, and how they
did it. For the story around the numbers, the turning points that moved the
puzzle, see the [history at a glance](/research/history); the strongest methods
live in the mailing list and Discord, not in journals.
> **[Interactive: RecordsView]** Rendered on the canonical page (link above); not shown in this markdown export.
## Related
- [History: the big steps](https://eternity2.dev/research/history) — The Eternity II story at a glance, from the mailing list founded in 2000 to the 470 record that still stands. A scannable timeline of the turning points, each linking into the full two-part history and the message where it happened.
- [Experiments](https://eternity2.dev/research/lab/experiments) — The lab's named search experiments, one section per researcher. Each is a real run against Eternity II with its idea, its best board, and the questions it left open. Raphaël Anjou's notebook is here in full; the notebook is open to anyone else's.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
---
# Reference numbers
> Exact counts of how many valid ways a small block can be filled at a given position of the official Eternity II board, under increasingly constrained rules: known-good numbers to check your solver's edge-matching and constraint code against.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/reference/
- Updated: 2026-07-11
- Topics: backtracking
- Reproduce: `just research-subgrid`
- Source: sylvogel's published subgrid counts (groups.io message 11879) — https://groups.io/g/eternity2/message/11879
- Source: This project's reproducible Rust generator + committed results (GitHub) — https://github.com/raphael-anjou/eternity2/tree/main/research/topics/subgrid-placement-counts
---
How many distinct, edge-matching ways can you fill a small block of the
official board? These counts are the ground truth to test a solver against: if
your edge-matching and constraint code disagrees with the numbers below on a
2×2 corner, the bug is in your code, not the table.
The upright figures here are computed from the official piece set by this
project's reproducible Rust generator (linked below), so anyone can rerun them.
The handful of italic figures are counts too large to enumerate exactly in
seconds (tens of billions to tens of trillions of fillings); those are
**sylvogel's published values**, reproduced here for completeness and credited
in the sources. Everything else is recomputed here from scratch.
> **[Interactive: ReferenceTableView]** Rendered on the canonical page (link above); not shown in this markdown export.
## Related
- [Known facts & numbers](https://eternity2.dev/research/build/known-facts) — The numbers every Eternity II researcher ends up re-deriving, collected in one place with their provenance: the puzzle definition, the clue placements, scoring conventions, the record table, search-space sizes and the structural counts.
- [The community's benchmarks](https://eternity2.dev/research/build/benchmarks) — How a community forbidden from sharing the pieces built a shared test culture anyway: derived-count verification protocols, the Txibilis and beginner suites, node-count duels, full enumerations, and the one benchmark that is still standing open today.
---
# Why it's hard
> Eternity II is not accidentally difficult. It was designed to resist cleverness, and the measurable structural walls (rigidity, entropy, forbidden patterns) explain why no search, however clever, has reached the end.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/
- Updated: 2026-07-01
---
Eternity II is not accidentally difficult. This is the science of why no
search, however clever, has reached the end: the puzzle's design, and the
structural walls that show up once you start measuring.
> **[Interactive: ScoringPrimer]** Rendered on the canonical page (link above); not shown in this markdown export.
## Engineered to resist cleverness
Eternity I fell in 2000 because Alex Selby and Oliver Riordan discovered the
puzzle had vastly more solutions than its designer believed, and aimed their
search at the most "solution-dense" regions. For Eternity II, the publisher
hired the winners: Selby and Riordan helped design and stress-test the new
puzzle so that no such statistical shortcut survives.
The visible fingerprints of that vetting: a single designed solution baked
into balanced color counts, no rotationally-symmetric pieces, no duplicate
pieces, and piece-count/color-count parameters sitting at the empirical
hardness peak (later confirmed by Ansótegui et al.). The puzzle isn't
accidentally hard. It was tuned to be.
One thing that looks like a design choice but isn't: the border uses its own
set of five motifs, separate from the interior's. That separation is automatic,
not engineered. Because the outer rim is solid grey, every border piece has its
grey edge fixed outward, so its coloured edges only ever meet other border
edges (sideways) or the interior (inward), and the two pools never touch. The
border motifs could be any five colours, even a relabelling of interior ones,
without changing the puzzle at all. They read as "rare" only because there are
fewer border edges to colour, not because the designers confined a scarce
resource to the frame to thwart solvers. (Thanks to Vasily V. on the
groups.io list for the correction.)
## The structural walls
Beyond the design story, the puzzle has measurable structure that explains
the gap between the best known board (470/480) and a full solution. Each of
these is published with the exact computation behind it.
## Pages in this section
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [Which wall stops which method](https://eternity2.dev/research/why/walls-and-methods) — The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- [Is this instance NP-complete, and how do I encode it?](https://eternity2.dev/research/why/how-hard-is-this-instance) — Edge matching is NP-complete as a family, but that says nothing about one fixed 16×16 board: a single instance is a constant, not a problem. What is true is the family's worst-case hardness and this instance's empirical hardness, and how to write the puzzle for a SAT, exact-cover, or ILP solver with small worked sketches.
- [Complex theory: counting the search before you run it](https://eternity2.dev/research/why/complex-theory) — Brendan Owen's complex theory estimates how wide the search tree is at every depth, and even how many solutions exist at all. Many in the community consider it the single most important thing to understand about Eternity II.
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [Designed to be unsolvable: the recipe](https://eternity2.dev/research/why/design-recipe) — Eternity II follows a recipe for the hardest possible edge-matching puzzle: compact shape, no symmetric or duplicate pieces, split palettes, flat frequencies, one expected solution. The community reverse-engineered every ingredient in the launch year.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [Why basin-hopping looks impossible](https://eternity2.dev/research/why/sigma-cycles) — If you can't improve a great board by polishing it, maybe you can jump to a different great board. On every record pair tested, you can't, and the structural reason why is worth seeing.
- [Where the mismatches live](https://eternity2.dev/research/why/mismatch-geometry) — A near-perfect board doesn't scatter its few errors evenly. It packs them into one band of five rows and leaves all the rest flawless. Which band is decided by the direction the search filled the board, and you can see the mirror on the real record boards.
- [Forbidden patterns](https://eternity2.dev/research/why/forbidden-patterns) — Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
- [Where you place the hints beats how many](https://eternity2.dev/research/why/hint-geometry) — On a 16×16 puzzle built like Eternity II, eighteen hints scattered across the board solve it in minutes, while the same puzzle needs eighty or more hints piled into contiguous rows to be as easy. Position, not count, is the lever, and it points straight at the endgame.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- [The rare colors live on the frame](https://eternity2.dev/research/why/rare-color-geography) — Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
- [The border balance](https://eternity2.dev/research/why/border-balance) — A solved board hides a simple bookkeeping law: every colour the border hands to the interior, the interior hands straight back. Break it and you know instantly the board is wrong; obeying it, though, guarantees nothing.
- [Entropy and the area law](https://eternity2.dev/research/why/entropy-area-law) — Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
---
# The border balance
> A solved board hides a simple bookkeeping law: every colour the border hands to the interior, the interior hands straight back. Break it and you know instantly the board is wrong; obeying it, though, guarantees nothing.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/border-balance/
- Updated: 2026-07-02
- Topics: search-space, structure
- Source: Border edge-type balance observed in the launch summer (angwin_uk, August 2007) — https://groups.io/g/eternity2/message/2073
- Source: mjqxxxx's stronger border condition: even counts, split equally into left and right edges (August 2007) — https://groups.io/g/eternity2/message/2098
- Source: Brendan Owen's paired edge-count observation on planted sets (June 2007) — https://groups.io/g/eternity2/message/422
- Source: Hopfer's 2022 statement of the multiset-equality condition (groups.io msg 10754) — https://groups.io/g/eternity2/message/10754
- Source: Hopfer's crisp restatement: same mixture of internal images on all 56 border pieces (msg 10757) — https://groups.io/g/eternity2/message/10757
---
Look only at the seam between the outer ring of border pieces and the first
ring of interior pieces. Each edge across that seam shows one colour, counted
once from the border side and once from the interior side. In any complete
solution the two tallies are identical: the multiset of colours the border
presents inward exactly equals the multiset the interior presents outward.
Call the imbalance the **deficit**, $\Delta$: half the total mismatch between
the two tallies. A finished, correct board has $\Delta = 0$. The four known
full solutions, across four different piece sets, all satisfy it exactly. So
$\Delta > 0$ is a certificate that a board can never be completed: a real,
cheap necessary condition.
## The law, in one line
Across the border↔interior seam, the colours the border shows inward and the
colours the interior shows outward are the same multiset. Writing $A[c]$ and
$B[c]$ for the two per-colour tallies:
$$
\Delta \;=\; \tfrac{1}{2} \sum_{c} \bigl|\, A[c] - B[c] \,\bigr| \;=\; 0 .
$$
## See it on a real board
A solved 8×8 starts in balance ($\Delta = 0$). Lift a border piece and watch
which colours fall out of balance, and the deficit climb. Then try the swap
button, and watch the catch.
> **[Figure]** Interactive: the NS-1 border-balance deficit — interactive: Ns1Lab. Rendered on the canonical page (link above); not shown in this markdown export.
## The catch, and why it matters
Swapping two border pieces leaves $\Delta$ at 0. The swap moves colours around
the seam without changing either tally, so the invariant is blind to it.
Worse, on a near-solution most of the remaining errors aren't on the border
seam at all. They sit interior-to-interior, where NS-1 never looks: on
469-class boards over 85% of the unmatched edges are invisible to it.
That is the whole texture of Eternity II in miniature. A check this clean
still only ever says "definitely broken", never "definitely fine". The puzzle
resists every cheap certificate of progress.
## Is it useful, then?
Yes, as a late-search pruner. Once a solver has closed the border ring,
enforcing $\Delta = 0$ rejects 10–28% of deep dead-ends for the price of one
pass over the 56 seam edges. It is necessary-but-loose: it throws away bad
states cheaply and leaves the hard part, the interior, untouched.
## Related
- [The rare colors live on the frame](https://eternity2.dev/research/why/rare-color-geography) — Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
- [Forbidden patterns](https://eternity2.dev/research/why/forbidden-patterns) — Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
---
# Complex theory: counting the search before you run it
> Brendan Owen's complex theory estimates how wide the search tree is at every depth, and even how many solutions exist at all. Many in the community consider it the single most important thing to understand about Eternity II.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/complex-theory/
- Updated: 2026-07-02
- Topics: structure, backtracking
- Source: Brendan Owen's peak-depth closed form, 256·(1−1/e) (groups.io msg 8125, 2010) — https://groups.io/g/eternity2/message/8125
- Source: Peter McGavin's 10×10 solve validating complex theory (groups.io msg 9686, 2017) — https://groups.io/g/eternity2/message/9686
- Source: Brendan Owen's solution-count estimates (groups.io message 5209) — https://groups.io/g/eternity2/message/5209
- Source: Peter McGavin's reference implementation and depth table (groups.io message 11197) — https://groups.io/g/eternity2/message/11197
- Source: Brendan Owen's "Backtracker estimates" tabulation for the clue puzzles and E2 (groups.io Databases) — https://groups.io/g/eternity2/databases
- Source: Brendan Owen's estimated-vs-actual validation, "NxM puzzles using Eternity II subset pieces" (groups.io Files, Brendan folder) — https://groups.io/g/eternity2/files/Brendan/NxM_actual_theory.pdf
- Source: Brendan Owen's heuristic-vs-node-count study, "Heuristics: 20×2 rectangle" (groups.io Files, Brendan folder) — https://groups.io/g/eternity2/files/Brendan/heuristics.pdf
---
> **Whose idea this is**
>
> Complex theory is due to Brendan Owen, one of the puzzle's vetters; Peter McGavin [implemented it in C](https://groups.io/g/eternity2/message/11197) with arbitrary-precision arithmetic and posted the numbers. We're writing it up here because a community member (Dan Karlsson) rightly pointed out it was missing, and because it underpins almost every good decision you can make about a solver, above all the choice of search order.
## The idea
Take a scan order and walk it one cell at a time. At each new cell, a random
unused piece matches its already-placed neighbours with some probability: a
product of per-edge colour-match chances. Multiply that by how many pieces are
left and you get the expected number of ways to extend the board one cell
further. Chain this down all 256 cells and you have a closed-form estimate of
how wide the search tree is at every depth and, at the last cell, how many
full solutions the puzzle has.
It is an average, not a true count: it assumes the 22 edge colours are drawn
independently, which they aren't (four edges are bolted to one rigid piece).
But calibrated against small puzzles where the real count is known, it lands
within about a factor of two. That is more than enough to see the shape.
## The headline numbers
| Clues placed | Expected solutions |
| ---------------------------------- | ------------------ |
| One clue (the centre piece only) | ≈ 14,702 |
| All five official clues | ≈ 1 |
With only the mandatory centre piece the puzzle has on the order of fifteen
thousand solutions; add the four other clues and the expected count drops to
about $4\times10^{-8}$: overwhelmingly, exactly one. This is the formal
reason the 5-clue puzzle has a single designed solution.
## The same numbers, checked against reality
Brendan tabulated the estimate not just for E2 but for the four smaller *clue
puzzles*, and this is where it earns trust. The clue puzzles are small enough
that their trees were searched exhaustively, so the estimate sits right next
to the true count. It lands within a factor of two, the calibration this page
keeps promising. The table also records the best-known fill order for each
puzzle, and they are not all the same: the order is a choice the search-space
shape rewards or punishes, not a property of the puzzle.
> **[Interactive: ClueEstimatesTable]** Rendered on the canonical page (link above); not shown in this markdown export.
The clue puzzles are four data points; Brendan checked the model far more
widely. His **"NxM puzzles using Eternity II subset pieces"** study plots the
estimated nodes-per-solution against the *actual* count for on the order of a
hundred smaller boards built from E2's own pieces, and on a log-log axis the
cloud hugs the diagonal across eleven orders of magnitude, from ten nodes to
$10^{11}$. That is the real basis for trusting the estimate on a board too
large to ever search: it has been right everywhere it *could* be checked. A
companion study even shows a dead-simple static score (the sum of per-cell
edge-match counts, squared) predicts a rectangle's total node count with an
$R^2$ of about 0.84, more evidence that the search cost is baked into the
board's structure before you place a piece.
## The funnel
Plot the expected width at each depth and three regimes appear. Their shape is
what the community calls the E2 funnel. Play the sweep below and watch the
counter: it climbs into the billions, then barely moves for a hundred cells
across the plateau, that flat crawl through an astronomically wide band is the
wall, before the last sixty pieces funnel it back down.
> **[Figure]** interactive: ComplexFunnelAnimated. Rendered on the canonical page (link above); not shown in this markdown export.
## Try it: the order decides the funnel
The same estimate, run live for different scan orders. The plateau peak (the
widest point the search must cross) is decided by the order alone, before a
single node is placed.
> **[Figure]** Interactive: the search-space funnel — interactive: ComplexFunnelLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Three regimes
- **Growth (depth 1–50).** Solutions multiply geometrically from one to
about $10^{27}$. Every placement is essentially free; nothing constrains you
yet.
- **Plateau (depth 50–200).** The tree is at its widest, about $10^{45}$
ways to extend, while the solution count barely moves. This is where
backtrackers spend roughly 99% of their time, matching Joe's empirical
finding that most time is spent below depth 150.
- **Collapse (depth 200–256).** The width falls from $10^{45}$ back to about
$10^{4}$. The last ~60 pieces are tightly constrained: each one placed
eliminates orders of magnitude of branches. The endgame is locally easy;
the hard part is reaching it.
## Why it changes how you search
If almost all the work is in the plateau, the goal isn't raw speed. It is
getting across the plateau to the funnel entrance (around depth 200), after
which the search chains down deterministically. And because complex theory
scores a scan order before you run it, you can compare orders by the height of
their plateau peak rather than by trial and error. That is the rigorous
version of a rule this site states everywhere: the fill order is a
first-class choice, and McGavin's bottom-left, left-to-right scan was picked
because complex theory said it was good.
## Bigger tiles
The same idea works if you place 2×2 or 3×3 tiles instead of single pieces: a
whole block of cells is committed at once, with its internal edges already
matched. The [search-path playground](/playground/paths) lets you do this for
real: pick a block shape (1×1, 2×1, 2×2, 3×3, …) and stamp blocks onto the
grid to build a block path. Race it and a dedicated macro-piece solver commits
one whole valid sub-assembly per block instead of one piece at a time, so the
search advances region by region. The plateau-peak estimate alongside still
scores the cell order your blocks imply, predicting the cost before you run a
single node.
## What it can't see
Complex theory is a first-moment estimate, so it is blind to one thing:
whether the many counted partial boards are genuinely distinct. The
[entropy and area-law results](/research/why/entropy-area-law) show that
distinctness collapses past ~80 cells, a second-order effect the
independent-edge model cannot capture. So use complex theory to choose orders
and read the tree's shape, never as a true count or a bound.
## Provenance and validation
The theory's paper trail runs through the mailing list. Brendan Owen posted
the completed model in April 2008
([msg 5197](https://groups.io/g/eternity2/message/5197),
[5209](https://groups.io/g/eternity2/message/5209)), and later proved a neat
closed form: for a scan order the node-count peak sits at depth
$256\,(1 - 1/e) \approx 161.8$
([msg 8125](https://groups.io/g/eternity2/message/8125)); the funnel above
peaks there empirically. Peter McGavin typeset the theory
([msg 9188](https://groups.io/g/eternity2/message/9188)), published the
14,702 expected-solutions figure as early as 2011
([msg 8924](https://groups.io/g/eternity2/message/8924)), and in 2017
delivered its strongest validation: solving Brendan's hint-free 10×10
benchmark by searching complex-theory-ranked first rows: roughly 180
core-years, landing inside the theory's predictions
([msg 9686](https://groups.io/g/eternity2/message/9686),
[9688](https://groups.io/g/eternity2/message/9688)). His 2024 reference C
implementation ([msg 11197](https://groups.io/g/eternity2/message/11197)) is
what this page's live estimator ports, line for line.
## Related
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [Entropy and the area law](https://eternity2.dev/research/why/entropy-area-law) — Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
---
# Designed to be unsolvable: the recipe
> Eternity II follows a recipe for the hardest possible edge-matching puzzle: compact shape, no symmetric or duplicate pieces, split palettes, flat frequencies, one expected solution. The community reverse-engineered every ingredient in the launch year.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/design-recipe/
- Updated: 2026-07-02
- Topics: structure
- Source: guenter stertenbrink poses the design problem: a puzzle with ~1% chance of falling in ten years (msg 15, February 2001) — https://groups.io/g/eternity2/message/15
- Source: Brendan Owen, Design the hardest puzzle: the full recipe and the 17.14 derivation (msg 1947, August 2007) — https://groups.io/g/eternity2/message/1947
- Source: Brendan Owen, Hardest Puzzle Estimates: the hardest-parameter table for every board size (msg 2164) — https://groups.io/g/eternity2/message/2164
- Source: Brendan Owen, Piece Frequencies: the flat distribution that killed the Eternity I method (msg 1667) — https://groups.io/g/eternity2/message/1667
- Source: Brendan Owen citing The Times: Selby and Riordan wrote the generator for Monckton (msg 3373) — https://groups.io/g/eternity2/message/3373
- Source: Dave Clark's account of a phone call with Christopher Monckton on how the puzzle was generated (msg 4177) — https://groups.io/g/eternity2/message/4177
- Source: The piece design-space census: 256 pieces from roughly 21,000 possible, symmetric forms avoided (msgs 8014–8034; census in msg 8025) — https://groups.io/g/eternity2/message/8025
- Source: Brendan Owen on the designers' solution locked in a safe (msg 8823) — https://groups.io/g/eternity2/message/8823
---
Eternity II is not a puzzle that happens to be hard. It is the output of a
recipe: a short list of design rules that, followed together, produce the
hardest edge-matching puzzle a given number of pieces can make, while still
guaranteeing that a solution exists. The remarkable part is that the recipe
was not leaked or published by the designers. The community reconstructed it,
ingredient by ingredient, within weeks of launch, mostly in one August 2007
post by Brendan Owen titled, fittingly, "Design the hardest puzzle"
([msg 1947](https://groups.io/g/eternity2/message/1947)).
This page walks through that recipe: what each rule is, what it costs anyone
trying to solve the puzzle, and where each claim comes from. The measurement
of the sharpest single ingredient (the colour count sitting exactly on the
hardness peak) has [its own page](/research/why/phase-transition); here it
takes its place among the others.
## The problem came before the puzzle
The design problem was stated on the mailing list six years before anyone
had to solve it. In February 2001, with Eternity I barely cold and Eternity II
still a rumour, guenter stertenbrink asked the group how one would design a puzzle
carrying a £5M prize so that it had only about a 1% probability of being
solved within ten years ([msg 15](https://groups.io/g/eternity2/message/15)).
Replies discussed scaling Eternity-I-style difficulty and even hiding
cryptographic problems in the edges.
That is exactly the tightrope a prize puzzle must walk. Make it too easy and
the prize is lost; that is what happened to Eternity I, which fell in 2000
because it had vastly more solutions than its designer believed. Make it
literally impossible and the contest is fraud. The target is a puzzle that
provably has a solution, positioned so that no realistic amount of computing
finds it within the contest window. Eternity II's designers had watched
Eternity I die, and the recipe below reads like a point-by-point response.
## The recipe, ingredient by ingredient
### Keep the shape compact
The first rule in Owen's derivation: use a compact board, the 16×16 square,
rather than anything elongated or irregular
([msg 1947](https://groups.io/g/eternity2/message/1947)). A compact shape
maximises the share of interior joins, where uncertainty is highest, and
leaves no thin arms or corridors that a solver could exhaust cheaply and use
as an anchor. Owen later verified the consequence experimentally: on
comparable 16×16 designs with unbalanced palettes there is always some region
that is cheaper to tile first (for a 2/19 split, starting in the middle is
over a hundred times cheaper than a row scan), but on E2's actual
parameters no such region exists. The design has, in his words, "no weak
areas to start tiling from"
([msg 5263](https://groups.io/g/eternity2/message/5263), design comparison
[msg 5243](https://groups.io/g/eternity2/message/5243)).
### No symmetric pieces, no duplicates
Every one of the 256 pieces is unique, and none is symmetric under rotation
([msg 1947](https://groups.io/g/eternity2/message/1947)). In 2010 the
community counted the design space to see how deliberate this is: with 5
frame and 17 interior patterns there are roughly 21,000 possible piece
designs, including forms like *aaaa*, *abab* and *aabb* that repeat under
rotation. The real set pointedly avoids all of them
([msgs 8014–8034](https://groups.io/g/eternity2/message/8014)).
The consequence is the absence of freebies. A duplicate pair would let any
solution be rewritten by swapping the two pieces, doubling the solution
count for free; a rotationally symmetric piece would collapse orientations
and shrink the decision space. Denying both keeps the expected solution
count pinned where the designers wanted it and hands the solver exactly
zero symmetry to exploit: every placement is a full, independent decision
among 4 orientations of distinct pieces.
There is a second, quieter reason to bar symmetric pieces, and Owen measured
it. A symmetric piece is not just structurally redundant, it is easier to
*place*, because it fits more contexts. He built a set of 289 pieces (all 17
of the 90-degree-symmetric forms, all 136 of the 180-degree-symmetric ones,
and 136 random asymmetric pieces), tiled a small rectangle every possible way,
and counted how often each piece appeared across all 759 million solutions.
The asymmetric pieces showed up **2.08 times** as often as the
180-degree-symmetric ones and **4.14 times** as often as the
90-degree-symmetric ones
([msg 2076](https://groups.io/g/eternity2/message/2076)). So symmetric pieces
are the *hard-to-tile* ones, and a puzzle that included them would hand a
solver exactly the uneven-tileability handhold that
[flat colour frequencies](#make-every-frequency-flat) are meant to remove.
Banning them keeps every piece about equally hard to place, with no easy ones
to save for last.
### Two palettes, strictly separated
The 22 colours split into 17 interior colours and 5 that appear only on the
joins between border pieces, never inside
([msg 1947](https://groups.io/g/eternity2/message/1947)). This turns the
frame into its own sub-puzzle whose difficulty can be tuned independently of
the interior, so that neither part is a soft entry point: the same
balancing act as the compact shape, applied to the palette. The five
frame-only colours are also the rare ones, quarantined at the rim where
their scarcity cannot create over-constrained interior cells that
propagation would feast on. That visible signature has
[its own page](/research/why/rare-color-geography).
### Make every frequency flat
When Owen digitised his set on launch day he found the edge-colour
distribution "as flat as could be": 24 edges of each of the 5 border-join
colours, 48–50 of each of the 17 interior colours
([msg 1054](https://groups.io/g/eternity2/message/1054)).
This single ingredient is the one that killed the Eternity I playbook.
Eternity I was cracked largely through piece-difficulty ordering: its piece
tileabilities varied enormously, so solvers could save the easiest pieces
for last and let the statistics carry them home. Two weeks after launch Owen
showed E2's flat distribution makes that approach worthless: when every
colour is equally common, every piece is about equally tileable, and no
ordering heuristic gets traction
([msg 1667](https://groups.io/g/eternity2/message/1667)). As doc_s_smith,
one of the people who actually solved Eternity I, put it in reply,
most-constrained-position selection became "our only other hope"
([msg 1722](https://groups.io/g/eternity2/message/1722)).
### Aim for exactly one solution
The last ingredient sets the colour counts themselves. Owen worked backwards
from the requirement "about one expected solution": setting the expected
number of interior tilings to 1 and solving for the interior colour count
gives
$$I = (196! \cdot 4^{196})^{1/392} \approx 17.14$$
Round to 17, add the separately-tuned 5 border colours, and you have
Eternity II's exact palette
([msg 1947](https://groups.io/g/eternity2/message/1947)). Owen backed the
derivation with simulations, defended 17+5 against the neighbouring 16+8
design from the same one-expected-solution family, since it better balances
the tileability of border and interior pieces and leaves no cheap way in
through the frame ([msg 2426](https://groups.io/g/eternity2/message/2426)),
and followed up with a table of hardest parameters for every board size, a
general recipe of which E2 is the 16×16 row
([msg 2164](https://groups.io/g/eternity2/message/2164)).
One expected solution is not an arbitrary aesthetic. It is the setting where
solutions are as scarce as they can be while still existing: the peak of
the phase transition, the point where search is provably at its worst. That
measurement, and the published analyses that later confirmed Owen's number,
live on the [hardness-peak page](/research/why/phase-transition). The flip
side proves the knob is real: a deliberately *loosened* 16×16 design
discussed on the list has around 10^42 expected solutions, though Owen
warned that even that one is no pushover
([msg 4968](https://groups.io/g/eternity2/message/4968)).
## Who actually designed it
Christopher Monckton invented the Eternity franchise and put up the prize,
but his original idea for the sequel was a 1001-piece three-dimensional
puzzle. Owen relayed this while noting that "the actual design is from Alex
and Oliver" ([msg 2697](https://groups.io/g/eternity2/message/2697)). Alex
Selby and Oliver Riordan are the two mathematicians who won Eternity I by
discovering it had far more solutions than intended; Monckton hired the
people who beat him. The group suspected this weeks before launch
([msg 716](https://groups.io/g/eternity2/message/716)), saw it confirmed in
an official Tomy leaflet (the inventor met the E1 winners on breakfast
television and asked them to work on E2's development;
[msg 901](https://groups.io/g/eternity2/message/901)), and finally sourced
it to a Times article: Selby and Riordan designed the puzzle-generating
program ([msg 3373](https://groups.io/g/eternity2/message/3373)).
That provenance explains the recipe's precision. The one team on earth with
first-hand experience of how a prize puzzle fails statistically was paid to
make sure the failure mode was gone. Every ingredient above (flatness, no
duplicates, balanced palettes, one expected solution) closes a door that
Selby and Riordan themselves had walked through in 2000.
## How the pieces were generated: a second-hand account
The only detailed account of the generation process in the archive is
second-hand and should be read as such. Dave Clark, founder of the
eternity2.net distributed project, recounted a long phone conversation he
had with Monckton on 26 July 2007: the puzzle was generated from entropy
typed at a keyboard by the contest judges (some 200 inputs), regenerated
until the judges were satisfied, then printed exactly once and vaulted.
Monckton described the random-number generator to him as using "gaussian
residues of powers of appropriately chosen primes", which Clark took to be
Selby and Riordan's own implementation
([msg 4177](https://groups.io/g/eternity2/message/4177)). The thread briefly
entertained the idea of attacking a cryptographically weak generator; a
reply pointed out that if the construction was Blum–Blum–Shub-like it would
be provably hard ([msg 4180](https://groups.io/g/eternity2/message/4180)).
Nothing came of the angle, but the account stands as the archive's best
primary-adjacent source on where the 256 pieces actually came from.
## The insurance policy
A puzzle designed never to be solved still has to prove it *can* be. The
Tomy launch material stated that nobody, not the inventor and not the
designers, knows the solution: the generator printed it "between pages of
random text whilst all parties were out of the room", and the output was
sealed in front of witnesses
([msg 901](https://groups.io/g/eternity2/message/901)). After the contest
ended unclaimed, Owen gave the statement still quoted today: "I am sure Alex
and Oliver created a solution when Chris paid them to generate a practically
impossible puzzle", and that it lies hidden amongst reams of printed text
locked in a safe, as insurance against any legal challenge to the contest's
good faith ([msg 8823](https://groups.io/g/eternity2/message/8823)).
That safe is the recipe's final ingredient. The designed solution is what
lets the puzzle sit at one-expected-solution instead of zero: existence
guaranteed by construction, discovery priced beyond any contestant's reach.
## What the recipe means if you attack the puzzle today
Every generic shortcut you might reach for was anticipated and
priced in nearly two decades ago. Piece-frequency ordering died with the
flat distribution. Symmetry and duplicate tricks have nothing to grab.
There is no soft region to tile first, no palette imbalance to lever, and
the colour count sits at the exact setting where search is worst. The
community's twenty years of records (467 in 2008, 470 in 2021, none since)
are the empirical readout of a design that worked precisely as intended.
That does not make the puzzle impossible: one solution certifiably exists,
printed and locked away. It means the remaining gap is not a tuning problem.
Anything that closes it will have to be an idea the designers could not
anticipate. That is, in the end, what this wiki is for.
## Related
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [The rare colors live on the frame](https://eternity2.dev/research/why/rare-color-geography) — Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
- [The hunt, a history (part I: 2000–2009)](https://eternity2.dev/research/community/hunt) — The community's story, from a mailing list founded seven years before the puzzle existed to the $10,000 scrutiny prize won under a borrowed name, with every event sourced to its original message. Part I of a growing chronicle.
---
# Entropy and the area law
> Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/entropy-area-law/
- Updated: 2026-07-01
- Topics: structure
- Reproduce: `just research-entropy-area-law`
- Source: Fekete's subadditivity lemma, the theorem behind the entropy limit h∞ existing at all — https://en.wikipedia.org/wiki/Subadditivity#Subadditive_sequences
- Source: Shannon entropy of a source, the h(n) quantities measured here — https://doi.org/10.1002/j.1538-7305.1948.tb01338.x
---
CartesianGrid,
Line,
LineChart,
ResponsiveContainer,
Tooltip,
XAxis,
YAxis,
} from "recharts";
Forget the use-each-piece-once rule for a moment and treat the 196 interior
pieces as reusable tiles. Count the valid all-matched patches and they grow
exponentially with size: there's no shortage of locally-valid ways to tile.
The matching grammar is rich, not restrictive.
We can measure that richness exactly. For a strip of width $n$, the growth
rate per cell is an entropy density $h(n)$, and the sequence decreases toward
the true two-dimensional value as the strip widens.
## The grammar's entropy, width by width
> **[Figure]** interactive: EntropyChart. Rendered on the canonical page (link above); not shown in this markdown export.
## Then the area law bites
Now put the use-each-piece-once rule back, and ask how often a random
color-valid $n \times n$ patch actually uses distinct pieces. Call it
$\rho(n)$. It collapses, and it collapses in the area, not the
perimeter:
$$
\rho(n) \;\approx\; \exp(-\alpha\, n^2), \qquad \alpha \approx 0.085.
$$
An area-law decay is brutal because area grows quadratically. The fraction of
realizable patches drops below one in a thousand at around eighty cells.
> **[Figure]** interactive: RhoChart. Rendered on the canonical page (link above); not shown in this markdown export.
## See the gap open
The same idea on the real pieces, exactly counted. Step the block size and
watch how many colour-valid blocks survive the use-each-piece-once rule.
> **[Figure]** Interactive: distinctness collapse and rho decay — interactive: EntropyScarcityLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why it matters
Eighty cells is not arbitrary. It's the size of the smallest moves that
separate the best known boards. The matching grammar stays rich up to about
that scale, then the distinctness rule collapses it. So the wall isn't in the
part that looks hard, matching colors; it's in the quiet rule that each piece
is used once, whose cost grows with area, on a board just large enough for it
to bite.
See [why basin-hopping is impossible](/research/why/sigma-cycles).
## The theorem, briefly
The per-width entropy has a well-defined limit. Joining an $n_1$-wide and an
$n_2$-wide strip side by side only adds a seam constraint, so the eigenvalues
satisfy
$$
\lambda_{n_1+n_2} \;\le\; \lambda_{n_1}\,\lambda_{n_2}.
$$
Taking logs makes $\log \lambda_n$ subadditive, and Fekete's lemma gives the
limit as an infimum, which is exactly why the curve above decreases:
$$
h_\infty \;=\; \lim_{n\to\infty}\frac{\log_{10}\lambda_n}{n} \;=\; \inf_n \frac{\log_{10}\lambda_n}{n}.
$$
The upper bound is the purely horizontal rate: ignoring vertical constraints
only adds patches, so
$$
0 \;<\; h_\infty \;\le\; \log_{10}\lambda_H \;=\; 1.6645,
$$
with $\lambda_H = 46.18$ the spectral radius of the horizontal
color-compatibility matrix. Positivity holds because the grammar supports
exponentially many chains, so the density is strictly between zero and
1.6645, measured near 0.67.
## Related
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [Forbidden patterns](https://eternity2.dev/research/why/forbidden-patterns) — Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- [Solution counting: measuring what you cannot find](https://eternity2.dev/research/build/analysis/solution-counting) — Nobody has ever seen a full Eternity II solution, yet the community knows, to within a factor of two, how many exist. This page is the story and the craft of that number: exact censuses on small boards, the expectation formula and its twenty-year convergence on 14,702, and the culled-search estimates that were trusted only when four independent runs agreed.
---
# Forbidden patterns
> Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/forbidden-patterns/
- Updated: 2026-07-01
- Topics: structure, search-space
- Reproduce: `just research-forbidden-patterns`
- Source: Forbidden patterns: article, source and results (GitHub) — https://github.com/raphael-anjou/eternity2/tree/main/research/topics/forbidden-patterns
---
Eternity II has 256 square tiles, each with a color on all four sides. 196 of
them are interior pieces, the ones with no grey border edge, so the only
ones that ever sit inside the board. Take a few of those, drop them into a
small shape, and rotate them however you like. Most of the time the colors
simply won't line up, no matter what you try.
The bigger the shape, the worse it gets. Two pieces side by side fail to fit
about 39% of the time. Add a third in an L and you're stuck 83% of the time.
Close up a 2×2 square and 99.72% of all placements are dead on arrival: only
about 1 in 358 works.
## One that fits, one that can't
> **[Figure]** interactive: BoardSvg. Rendered on the canonical page (link above); not shown in this markdown export.
> **[Figure]** interactive: BoardSvg. Rendered on the canonical page (link above); not shown in this markdown export.
## Try it: draw four pieces
> **[Figure]** Interactive: paint a patch and watch it forbid itself — interactive: ForbiddenPatchLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Where 99.72% comes from
A rough estimate explains why the number is so high. Two adjacent pieces
share one edge. Each interior piece carries colors from the 17-color interior
palette, so a random pair of half-edges matches with probability roughly
$$
\Pr[\text{one edge matches}] \;\approx\; \frac{1}{17} \;\approx\; 6\%.
$$
A 2×2 square has four internal edges to satisfy at once. Rotations give each
piece four chances, but the four edges are coupled, so a back-of-the-envelope
independence estimate puts the chance all four match at very roughly
$$
\Pr[\text{2}\times\text{2 feasible}] \;\sim\; 1 - \left(1 - \tfrac{1}{17}\right)^{\!c} \ \text{per rotation budget} \;\Rightarrow\; \lesssim 1\%,
$$
which is already under one percent. The exact exhaustive count lands at 0.28%
feasible, that is 99.72% forbidden
$(\,0.28\% = \tfrac{3{,}993{,}696}{1{,}431{,}033{,}240}\,)$. The estimate is
crude because the edges aren't independent and the colors aren't uniform, but
it gets the order of magnitude right and shows why closing a square is so
much harder than placing a single pair.
## The exact counts
Every distinct-piece placement of each shape, checked exhaustively, with no
sampling.
| Shape | Placements | Forbidden | Forbidden % |
| ---------------- | -------------: | ------------: | ----------: |
| Two side by side | 38,220 | 14,890 | 38.96% |
| Two stacked | 38,220 | 14,890 | 38.96% |
| L of three | 7,414,680 | 6,173,828 | 83.26% |
| 2×2 square | 1,431,033,240 | 1,427,039,544 | 99.72% |
Computed exactly from the official set; the run reproduces identically every
time (about twenty seconds). The
[article, source and results are on GitHub](https://github.com/raphael-anjou/eternity2/tree/main/research/topics/forbidden-patterns).
## Forbidden before you even pick a piece
The scarcity starts one level below whole pieces, at the colors. An interior
cell shows two of the 17 interior colors on any given corner, the two edges
that meet there, which is $17 \times 17 = 289$ possible ordered color pairs. Not
all of them exist. Back in 2008 the community noticed that **20 of those 289
pairs are voids**: no interior piece, in any rotation, presents that particular
pair of colors on adjacent edges
([msg 5027](https://groups.io/g/eternity2/message/5027)). So a partial board that
forces a cell to answer with one of those 20 corner-pairs is dead on the spot,
before a single piece is tried, and a fast solver can reject it with one table
lookup. It is the same lesson as the 2×2 count, pushed down to the smallest unit
that can be impossible: the constraints bite so early that whole categories of
local demand have no legal answer at all.
## Why it matters
A finished, correct board has zero forbidden patches: by definition,
everything matches. So counting the forbidden patches in a board tells you
roughly how far it is from a real solution, even when two boards have the
same number of matched edges. Weak boards are full of forbidden squares; the
best boards ever found have only a couple of dozen left.
It also shows, from another angle, why the puzzle shrugs off clever local
fixes. When 99.72% of small squares are impossible, the pieces that do fit
together are rare and specific. There's almost no room to shuffle things
around without breaking something. The good arrangements are scarce and
rigid.
## A second axis of progress
Counting forbidden squares turns the idea into a usable signal. Pick a real
record board: the number of forbidden 2×2 windows falls as the matched-edge
score rises. Two boards with the same edge-score can still differ here: the
one with fewer forbidden squares is structurally closer to a solution, which
is why some solvers track forbidden-patch count as a tiebreak.
> **[Figure]** Interactive: count feasible vs forbidden placements — interactive: ForbiddenCountLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Related
- [Entropy and the area law](https://eternity2.dev/research/why/entropy-area-law) — Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
- [Arc consistency, from AC-3 up](https://eternity2.dev/research/build/reduce/arc-consistency) — Forward checking looks one move ahead; arc consistency makes every cell's candidate list defend itself against every neighbour's, to a fixed point. Mackworth's AC-3, the optimal refinements that followed, and what the whole family actually measured on this puzzle, including where it is unsound.
---
# Where you place the hints beats how many
> On a 16×16 puzzle built like Eternity II, eighteen hints scattered across the board solve it in minutes, while the same puzzle needs eighty or more hints piled into contiguous rows to be as easy. Position, not count, is the lever, and it points straight at the endgame.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/hint-geometry/
- Updated: 2026-07-10
- Topics: structure, search-space, backtracking
- Source: Joe's hint-density study and Peter McGavin's 18-hint solve (groups.io thread 'A method to prune E2 search space by 17-30%+', msg 11725) — https://groups.io/g/eternity2/message/11725
- Source: Peter McGavin: 18 scattered hints solve the 16×16 E2-like puzzle in under 15 minutes (groups.io msg 11746) — https://groups.io/g/eternity2/message/11746
- Source: Peter McGavin: the optimised backtracker traces to Mike's 2007 post (groups.io msg 3098) — https://groups.io/g/eternity2/message/3098
---
Give a solver some correct pieces for free and the puzzle gets easier. The
obvious question is how many you need. The better question, it turns out, is
*where* they go. On a 16×16 puzzle built with Eternity II's exact colour
recipe, eighteen hints placed in the right spots solve it in minutes; the same
puzzle wants eighty or more hints, if you pile them into contiguous rows, to be
that easy. A fourfold difference in count, decided entirely by geometry.
## Two ways to spend the same eighteen hints
> **[Interactive: HintGeometryDiagram]** Rendered on the canonical page (link above); not shown in this markdown export.
The left board is the real layout Peter McGavin used: eighteen hints on a
regular lattice, every third column on a few odd rows. His plain scan-row
backtracker solved Joe's 16×16 E2-like puzzle (five border colours, seventeen
interior colours, the same distribution as the official puzzle) in under
fifteen minutes on a single core, walking a search tree of 41,160,067,167
placements. Dropping to fifteen hints still worked; the search just grew to
several hours. Pile eighteen hints into the first rows instead, the way a
top-down scan naturally accumulates them, and they buy almost nothing: the hard
part of the board is still completely open.
Joe had come at it from the other side, seeding whole contiguous rows from a
known solution, and needed far more before the puzzle fell:
18
scattered hints, solved in minutes
88+
contiguous-row hints for comparable ease
99%
of search time spent below depth 132
70%
of search time spent below depth 150
## Why position wins: the hints have to reach the endgame
The two right-hand numbers explain the left-hand ones. Joe instrumented his
backtracker over a billion iterations and found the work is not spread across
the board at all: 99% of it happens after depth 132 of 256, and 70% after
depth 150. Nearly all the pain is in the back half of the fill, and most of it
past the three-fifths mark.
A block of contiguous hints at the top is spent exactly where the search was
never going to struggle. It shortens an easy beginning and leaves the
expensive tail untouched. Scattered hints do the opposite: dotted through the
board, including down into the region the search reaches last, they pre-empt
the choices that would otherwise blow up deep in the tree. This is the same
fact the record boards wear on their surface. A near-perfect board packs all
its damage into the band of rows the search finished on, because
[whichever rows you fill last are where the puzzle makes you pay](/research/why/mismatch-geometry).
Hints only help to the extent they reach that band before the search does.
It also fits [why the interior gives no forced moves](/research/why/no-forced-moves):
with every interior cell still accepting scores of neighbours, a hint's value
is not local propagation but global constraint, cutting off whole subtrees the
search would otherwise have to walk. A hint far from the hard region cuts off
subtrees that were cheap anyway.
## What it does and doesn't say
This is a result about a 16×16 puzzle built to Eternity II's colour recipe, not
about the official puzzle, whose five fixed clues are a different, much smaller
gift in different places. What transfers is the shape of the lesson, and it is
the same one the [prune-versus-speed](/research/why/prune-vs-speed) argument
makes from the other direction: what matters is changing where the search
spends its effort, and the effort lives in the endgame. A handful of hints
aimed at that endgame is worth a great many aimed anywhere else.
> **Note**
>
> The counts, the 41-billion-node tree, and the depth statistics are Joe's and Peter McGavin's measurements, reported on the eternity2 groups.io list in January 2026; the scattered layout shown is decoded from Peter's posted board (msg 11746). The optimised backtracker Peter used traces back to Mike's 2007 post (msg 3098). These are community results on a specific E2-like puzzle, recorded here with attribution rather than re-derived.
## Related
- [Where the mismatches live](https://eternity2.dev/research/why/mismatch-geometry) — A near-perfect board doesn't scatter its few errors evenly. It packs them into one band of five rows and leaves all the rest flawless. Which band is decided by the direction the search filled the board, and you can see the mirror on the real record boards.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
---
# Is this instance NP-complete, and how do I encode it?
> Edge matching is NP-complete as a family, but that says nothing about one fixed 16×16 board: a single instance is a constant, not a problem. What is true is the family's worst-case hardness and this instance's empirical hardness, and how to write the puzzle for a SAT, exact-cover, or ILP solver with small worked sketches.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/how-hard-is-this-instance/
- Updated: 2026-07-08
- Topics: exact-methods
- Source: Demaine & Demaine 2007, "Jigsaw Puzzles, Edge Matching, and Polyomino Packing: Connections and Complexity" (Graphs and Combinatorics 23): edge matching is NP-complete — https://doi.org/10.1007/s00373-007-0713-4
- Source: Ansótegui, Béjar, Fernàndez & Mateu, How Hard is a Commercial Puzzle: the Eternity II Challenge (CCIA 2008) — https://repositori.udl.cat/bitstreams/0b6533fe-54e5-4070-85fe-80f7d35837d8/download
- Source: Knuth, The Art of Computer Programming, Volume 4B: Algorithm X and Algorithm C (exact covering with colours) — https://www-cs-faculty.stanford.edu/~knuth/taocp.html
- Source: Eternity II puzzle (Wikipedia) — https://en.wikipedia.org/wiki/Eternity_II_puzzle
---
This page answers a question that keeps surfacing on Math and CS Stack
Exchange and tops the dispute list on Wikipedia's Talk page for the puzzle:
edge matching is NP-complete in general, so does that tell us anything about
*this* fixed 16×16 board, and concretely, how would you hand Eternity II to a
solver? The short answers are: no, not directly, and here are three encodings.
## First, the category error
NP-completeness is a property of a *problem*, which in complexity theory means
an infinite family of instances indexed by a size parameter $n$. "Edge-matching
puzzles" is such a family: given $n$, a set of square tiles, and a colour
alphabet, decide whether the tiles tile an $n \times n$ frame with all abutting
edges agreeing. That decision problem is NP-complete
([Demaine & Demaine 2007](https://doi.org/10.1007/s00373-007-0713-4)), which
means both that a proposed solution is checkable in polynomial time and that
every problem in NP reduces to it.
A single fixed board is not a family. The Eternity II instance has a definite
answer, "yes, it is solvable" (the designers built it from a solution) or in the
5-clue competition form "yes, with exactly this arrangement". That answer is a
single bit. A single bit is a constant, and "is this constant NP-complete?" is
not a well-formed question: there is a constant-time algorithm that prints the
answer (`return true`), it just has the answer baked in. Asking whether one
instance is NP-complete is the same category error as asking whether the
number 17 is polynomial-time.
So the family is hard and the instance is a constant. What, then, is actually
true and useful about the difficulty of the board on your desk?
## What is actually true
Two separate statements, kept apart:
- **Worst-case hardness of the family.** Because the general problem is
NP-complete, no algorithm is known that beats exponential time in the worst
case as $n$ grows, and finding one would prove $\mathrm{P}=\mathrm{NP}$. This
bounds what any general edge-matching solver can promise. It says nothing
about how a *specific* input behaves.
- **Empirical hardness of this instance.** Hardness you can measure. The 16×16
board has on the order of $1.115\times10^{557}$ distinct piece-and-rotation
arrangements, and the puzzle appears to admit very few solutions (the 5-clue
form is designed for essentially one; see
[complex theory](/research/why/complex-theory) for the expected-count
estimate). A backtracking search therefore threads an astronomically wide
space toward a near-empty solution set, so it spends almost all of its time
exploring dead ends. That is why the puzzle is hard *in practice*, and it is a
claim about this board, not about the family.
The worst-case theorem and the empirical difficulty point the same way here,
but they are different kinds of statement and only one of them is a theorem. A
family can be NP-complete while a given instance is trivial (many are), and an
instance can be brutally hard to solve in practice even inside a family that is
polynomial. Ansótegui et al. made exactly the empirical case for Eternity II,
building solver benchmarks from it and measuring the difficulty directly rather
than appealing to the general theorem
([CCIA 2008](https://repositori.udl.cat/bitstreams/0b6533fe-54e5-4070-85fe-80f7d35837d8/download)).
> **The one-line version**
>
> "Is Eternity II NP-complete?" No: an instance has no complexity class. "Is the edge-matching problem NP-complete?" Yes. "Is this instance hard to solve?" Empirically yes, because the search space is ~$10^{557}$ wide and the solution set is nearly empty, so search drowns in dead ends.
## Encoding it: the setup
The rest is practical: how do you write the board so a solver can chew on it?
All three encodings below share the same skeleton. Number the cells
$c = 1 \dots 256$, the pieces $p = 1 \dots 256$, and the rotations
$r \in \{0, 1, 2, 3\}$. A *placement* is a triple $(c, p, r)$: piece $p$ dropped
into cell $c$ turned by $r$ quarter-turns. Every encoding has to say three
things:
1. each cell gets exactly one placement,
2. each piece is used exactly once,
3. wherever two cells touch, the colours on the shared edge agree.
The encodings differ only in how they express constraint 3, the colour match,
and that difference is the whole story. The sketches below use a 2×2 or 3×3
board so you can see all the clauses; the 16×16 is the same shape at scale.
## SAT / CNF
Introduce a Boolean variable $x_{c,p,r}$, true when piece $p$ sits at cell $c$
in rotation $r$. On the full board that is $256 \times 256 \times 4 \approx
262{,}000$ variables before any constraint. Then:
- **Exactly one placement per cell.** For each cell $c$, an at-least-one clause
over all its placements, $\bigvee_{p,r} x_{c,p,r}$, plus at-most-one clauses
forbidding any two, $\lnot x_{c,p,r} \lor \lnot x_{c,p',r'}$ for distinct
placements.
- **Exactly one cell per piece.** The mirror image: for each piece $p$, one
at-least-one clause over the cells it could occupy, plus at-most-one clauses
so it is placed only once.
- **Edge match.** For every interior adjacency and every placement whose exposed
edge shows colour $k$, forbid every placement in the neighbouring cell whose
facing edge is not $k$: a binary clause $\lnot x_{c,p,r} \lor \lnot
x_{c',p',r'}$ for each conflicting pair.
A 2×2 sketch makes the third family concrete. Cells $A$ (top-left) and $B$
(top-right) share a vertical edge; $A$'s east edge must equal $B$'s west edge.
For each placement $(A,p,r)$ showing east colour $k$, and each placement
$(B,p',r')$ whose west colour is not $k$, add $\lnot x_{A,p,r} \lor \lnot
x_{B,p',r'}$. Do the same for $A$/$C$ vertically and the other two interior
adjacencies. That is all constraint 3 is: a large pile of "these two placements
cannot both be true" binary clauses.
The catch is that the pile is enormous. The conflict clauses dominate, the
at-most-one encodings add their own blowup (naive pairwise is quadratic; ladder
or commander encodings trade it for auxiliary variables), and the result is a
formula with millions of clauses whose structure gives conflict-driven search
almost nothing to learn from. This has been tried since 2008 and complete SAT
solvers stall on the full board; Blackwood, among others, reported that SAT did
not help. The encoding is clean, the solver is not the bottleneck, the
*instance* is. See [SAT and CSP encodings](/research/build/exact/sat-csp-encodings)
for the benchmark history and where SAT verdicts still earn their keep as
small-region impossibility proofs.
## Exact cover and dancing links
The exact-cover view is tidier, and it hides a trap that trips up almost
everyone who reaches for Knuth's dancing links.
Set up 512 *items*: one per cell ("cell $c$ is filled") and one per piece
("piece $p$ is used"). Each *option* is a placement $(c,p,r)$, and it covers
exactly two items: cell $c$ and piece $p$. A set of options covering every item
exactly once is a board with every cell filled and every piece used once. That
is a clean exact-cover instance, and Knuth's **Algorithm X** with dancing links
solves exact cover beautifully.
Here is the trap. Constraints 1 and 2 are cover-once conditions, which is
exactly what exact cover expresses. But constraint 3, the colour match, is not a
cover-once condition at all: a shared edge is not "used once", it is "assigned a
colour that both neighbours agree on". Plain Algorithm X has no way to say this.
People encode it, run it, and get boards with mismatched edges, or they try to
bolt on extra items and find the cover-once semantics fight them. This exact
confusion has its own question on the CS Stack Exchange.
The fix is Knuth's own extension, **Algorithm C**, for XCC, exact covering with
colours (TAOCP Volume 4B). Alongside the *primary* items (covered exactly once)
you add *secondary* items that may be covered any number of times, *provided all
options covering a given secondary item assign it the same colour*. Give each
interior edge of the grid one secondary item. A placement that exposes colour
$k$ on a shared edge assigns colour $k$ to that edge's secondary item. Two
placements can then coexist across the edge only if they colour it identically,
which is precisely the edge-match constraint, now expressed natively.
A 3×3 sketch: 9 cell items and 9 piece items (primary), plus 12 interior-edge
items (secondary, one per horizontal or vertical join). The centre cell's
placements each touch four secondary edge items and must colour-agree with all
four neighbours; a corner cell's placements touch two. Algorithm C threads this
without ever emitting a mismatched board. The takeaway students miss:
**exact-cover-DLX for Eternity II needs Algorithm C, not Algorithm X.** The
[exact cover and dancing links](/research/build/exact/exact-cover-dlx) page
walks the full XCC construction and where DLX genuinely shines (small boards,
exhaustive solution counting) versus where it stalls on the 16×16.
## ILP and max-clique, briefly
Two more framings, useful mainly as pointers:
- **Integer linear programming.** Reuse the SAT variables as 0/1 integers
$x_{c,p,r}$. Constraints 1 and 2 become equalities $\sum_{p,r} x_{c,p,r} = 1$
per cell and $\sum_{c,r} x_{c,p,r} = 1$ per piece. The edge match becomes, for
each interior edge and each colour $k$, a linking condition tying the two
neighbours' colour-$k$ placements together (one clean form: a fresh binary
edge-colour variable $y_{e,k}$ with $\sum_k y_{e,k} = 1$ and each side's
placements implying the matching $y$). It is a feasibility ILP, no objective,
and the LP relaxation is weak, so branch-and-bound behaves much like the SAT
search. See [LP relaxations](/research/build/exact/lp-relaxations).
- **Max clique.** Build a graph whose vertices are legal placements and whose
edges join any two placements that are mutually compatible (different cells,
different pieces, and agreeing on any shared edge). A full board is a clique of
size 256. This is elegant on paper but the graph is huge and clique solvers
fare no better; it is worth knowing as a reduction, not as a practical attack.
## Where this leaves you
If you came asking whether the theory of NP-completeness makes this instance
provably hard, the answer is that it does not, and cannot: complexity classes
describe families, and this board is a fixed input with a fixed answer.
The general edge-matching problem is NP-complete, which caps what any solver can
promise as $n$ grows, but the difficulty you actually feel is empirical, a
$10^{557}$-wide space over a nearly empty solution set. Every encoding above
faithfully captures the puzzle; none of them makes it easy, because the hardness
lives in the instance, not in the choice of formalism. Once the board is
encoded, the leverage that remains is
[arc consistency](/research/build/reduce/arc-consistency) and smart search
order, which is where the rest of this section picks up.
## Related
- [Complex theory: counting the search before you run it](https://eternity2.dev/research/why/complex-theory) — Brendan Owen's complex theory estimates how wide the search tree is at every depth, and even how many solutions exist at all. Many in the community consider it the single most important thing to understand about Eternity II.
- [SAT and CSP encodings](https://eternity2.dev/research/build/exact/sat-csp-encodings) — Write the puzzle as clauses and hand it to an industrial solver: the obvious move, tried since 2008. Why complete solvers stall on the full board, and where their verdicts still earn their keep as impossibility proofs.
- [Exact cover and dancing links](https://eternity2.dev/research/build/exact/exact-cover-dlx) — Eternity II states cleanly as an exact-cover problem, and Knuth's Algorithm X with dancing links is the classic machine for those. Where it genuinely shines (small boards, exhaustive counting) and the two reasons it does not crack the 16×16: an unshrunk search tree, and no partial credit.
- [Arc consistency, from AC-3 up](https://eternity2.dev/research/build/reduce/arc-consistency) — Forward checking looks one move ahead; arc consistency makes every cell's candidate list defend itself against every neighbour's, to a fixed point. Mackworth's AC-3, the optimal refinements that followed, and what the whole family actually measured on this puzzle, including where it is unsound.
---
# Where the mismatches live
> A near-perfect board doesn't scatter its few errors evenly. It packs them into one band of five rows and leaves all the rest flawless. Which band is decided by the direction the search filled the board, and you can see the mirror on the real record boards.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/mismatch-geometry/
- Updated: 2026-07-02
- Topics: structure
- Source: Peter McGavin's 469-record announcement: the board whose eleven unmatched edges this page analyses (groups.io msg 10045, September 2020) — https://groups.io/g/eternity2/message/10045
---
Score a record board edge by edge and a striking thing shows up: the
mismatches are not spread around. McGavin's 469 has all eleven of its
unmatched edges in the top five rows (rows 0 to 4); rows 5 through 15 are
locally perfect.
The board is, in effect, a flawless 11-row slab with all the damage swept up
against one edge.
Now score this project's best from-scratch boards the same way. The damage is
again in a five-row band, but at the bottom. KEYRING's and GAUNTLET's
mismatches sit in rows 11 to 15, with rows 0 through 10 perfect. It is the
same picture flipped top to bottom.
## See it on the real boards
Pick a board. The shaded band is where its mismatches actually fall, computed
live with the engine's own scoring rule: community boards at the top, this
project's at the bottom.
> **[Figure]** Interactive: where mismatches are forced to live — interactive: MismatchGeometryLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why it flips: scan order
The flip is not a coincidence; it is the fingerprint of how each board was
built. A search that fills the board from the bottom up spends its perfect
placements early, low down, and is forced to absorb every accumulated
conflict in the last rows it reaches: the top. A search that fills top-down
does the exact opposite and piles the damage at the bottom. The mismatches
always end up crammed against the edge the solver finished at. Same puzzle,
same kind of board, opposite construction order.
## The same band, under a different objective
The pattern is not an artefact of how we score. Some solvers optimise a
different thing entirely: not matched edges on a full board, but the most pieces
you can place with zero conflicts, leaving holes instead of mismatches. Run that
objective and the holes land in the same place. Louis Verhaard's "Only seven
holes" board places 249 of 256 pieces conflict-free, and all seven holes sit in
rows 1 to 4, the top band again. Laurent Zamofing, reaching a similar ceiling in
2026 by recombining the community's record boards, noted the same thing: "the
unsolved residual always lands in that top band"
([msg 11901](https://groups.io/g/eternity2/message/11901)). Two objectives that
share nothing but the puzzle, and the leftover damage collects against the same
edge. It is the construction order, not the scoring rule, that decides where the
difficulty ends up. (More on this variant on the
[variants page](/research/build/variants).)
## What it tells us
This is the visible form of two deeper facts. First, the great boards really
are almost-complete: the gap to 480 is concentrated, not diffuse, which is
why integer programming finds them locally frozen everywhere except that one
band (the rigidity wall). Second, it says the endgame is the whole game:
whichever rows you fill last are where the puzzle makes you pay, so the order
you search in decides where the difficulty lands. That is the same lesson the
playground's path-racing makes you feel, here written into the structure of
every record board.
> **Note**
>
> The per-row counts and the highlighted band are computed in your browser from the real board edges (the same scoring rule the engine uses), not hand-placed. The underlying finding and the scan-order explanation are recorded in the project lab notebook.
## Related
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [STAGED](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/staged) — Build the whole board from scratch with no pre-set frame, in stages, letting the border emerge last from whatever pieces are left.
- [Which wall stops which method](https://eternity2.dev/research/why/walls-and-methods) — The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- [Why basin-hopping looks impossible](https://eternity2.dev/research/why/sigma-cycles) — If you can't improve a great board by polishing it, maybe you can jump to a different great board. On every record pair tested, you can't, and the structural reason why is worth seeing.
---
# No forced moves
> The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/no-forced-moves/
- Updated: 2026-07-01
- Topics: structure, search-space
- Reproduce: `just research-no-forced-moves`
- Source: No forced moves: article, source and results (GitHub) — https://github.com/raphael-anjou/eternity2/tree/main/research/topics/no-forced-moves
---
Bar,
BarChart,
CartesianGrid,
ResponsiveContainer,
Tooltip,
XAxis,
YAxis,
} from "recharts";
Every one of the 196 interior pieces has between 73 and 137 other pieces that
can legally sit to its right (counting right-hand partners, as the chart above
does). The typical piece has over a hundred options. Not a single piece is
ever pinned to one choice.
This is the flip side of forbidden patterns. There, almost every combination
of pieces is impossible. You'd think all those rules would corner pieces into
place. They don't: the constraints rule out combinations without ever
cornering an individual piece, so a solver never gets a free, forced move to
build on.
{data.forcedPieces}
pieces forced to a single option
{data.minPartners}–{data.maxPartners}
partners per piece (min to max)
{data.meanPartners}
average partners
## How many neighbours each piece allows
The 196 interior pieces, bucketed by how many right-hand neighbours each one
accepts. The whole distribution sits far from one.
> **[Interactive: PartnerHistogram]** Rendered on the canonical page (link above); not shown in this markdown export.
## See it on a real puzzle
The engine fills a few cells; then we count, live, how many pieces legally
fit the next one. It almost never drops to one.
> **[Figure]** Interactive: candidate counts, cell by cell — interactive: ForcedMovesLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why it matters
Put this beside forbidden patterns and the real shape of the difficulty
appears. Locally the puzzle looks loose: any piece fits next to plenty of
others, so there's nothing to propagate and no chain of forced moves to ride.
Globally almost every combination is illegal. The hardness lives in that gap:
lots of local freedom, almost no global consistency. A solver has to make a
long run of free-looking choices that only turn out wrong much later.
## Related
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [Piece theft, where solvers die](https://eternity2.dev/research/why/piece-theft) — A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [Arc consistency, from AC-3 up](https://eternity2.dev/research/build/reduce/arc-consistency) — Forward checking looks one move ahead; arc consistency makes every cell's candidate list defend itself against every neighbour's, to a fixed point. Mackworth's AC-3, the optimal refinements that followed, and what the whole family actually measured on this puzzle, including where it is unsound.
---
# Tuned to the hardness peak
> Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/phase-transition/
- Updated: 2026-07-02
- Topics: structure
- Reproduce: `just research-phase-transition`
- Source: Brendan Owen, Design the hardest puzzle: the 17+5 derivation (eternity2 mailing list, August 2007) — https://groups.io/g/eternity2/message/1947
- Source: Ansótegui, Béjar, Fernández & Mateu, On the hardness of solving edge matching puzzles as SAT or CSP problems (Constraints, Springer 2013) — https://link.springer.com/article/10.1007/s10601-012-9128-9
- Source: Ansótegui, Béjar, Fernández & Mateu, How Hard is a Commercial Puzzle: the Eternity II Challenge — https://repositori.udl.cat/server/api/core/bitstreams/0b6533fe-54e5-4070-85fe-80f7d35837d8/content
---
Hard search problems have a difficulty knob. Loosen it and there are many
solutions, so a search trips over one fast. Tighten it and there are none,
which is often easy to prove. In between sits a narrow band where solutions
are scarce but real, and that's where search blows up. People call it a phase
transition, like water freezing.
For edge-matching puzzles the knob is the number of colors. Too few and pieces
fit together countless ways; too many and they barely fit at all. The
[published analysis](https://link.springer.com/article/10.1007/s10601-012-9128-9)
puts the peak at around 17 interior colors, the setting where a puzzle this
size has about one solution. Eternity II uses 17.
The community had that number within weeks of launch: in August 2007,
Brendan Owen derived $I = (196! \cdot 4^{196})^{1/392} \approx 17.14$
interior colours from the "about one expected solution" criterion, exactly
the published parameters, years before the academic analyses
([Design the hardest puzzle, msg 1947](https://groups.io/g/eternity2/message/1947)).
## The difficulty wall
> **[Figure]** Interactive: the colour-count difficulty peak — interactive: PhaseTransitionLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Or measure it yourself
The chart above is precomputed. This one isn't: the engine solves fresh
puzzles live, one per colour count, and draws the peak from real runs in your
browser.
> **[Figure]** Interactive: solve across the difficulty peak live — interactive: PhaseTransitionLiveLab. Rendered on the canonical page (link above); not shown in this markdown export.
## The transition, measured on real pieces
The colour-count argument says *where* the peak is. In March 2008 Brendan Owen
went and watched a phase transition happen, directly, on the actual pieces. He
took a thin 2-by-L interior rectangle and tiled it with random sets of E2's
interior pieces, four hundred random sets at each length, and recorded how many
could be completed with no mismatch
([msg 4909](https://groups.io/g/eternity2/message/4909)).
> **[Interactive: RectangleTransitionChart]** Rendered on the canonical page (link above); not shown in this markdown export.
The shape is the phase transition in miniature. A 2-by-1 or 2-by-2 strip is so
short that random pieces often just fit: 58% solvable at length 1. Then it dies
completely. From length 4 through 15, not one random set out of four hundred
tiles the rectangle at any length: solutions are so scarce they effectively do
not exist. And then they come back. At length 16 one set in four hundred works,
by 19 it is 8.5%, by 21 it is 57%, and by 22 nearly nine in ten. The region
where solutions are vanishingly rare but not yet impossible is exactly the hard
band, and it is not a story or a model here, it is a count. It is also why a
14-cell-wide interior is so punishing: it sits in the steep part of that climb,
where a solution exists but almost no random arrangement is one.
## The split, straight from the pieces
Sorting the official set's colors by where they appear shows the design
plainly.
5 frame-only colors
{data.frameOnlyColors.map((c) => (
> **[Interactive: MotifSwatch]** Rendered on the canonical page (link above); not shown in this markdown export.
'{colorToLetter(c)}'
))}
These appear only on border and corner pieces, never in the interior.
They are the rare colors, kept to the edge.
17 interior colors
{data.interiorColors.map((c) => (
> **[Interactive: MotifSwatch]** Rendered on the canonical page (link above); not shown in this markdown export.
'{colorToLetter(c)}'
))}
The palette of the inside of the board, where almost all the matching
happens.
## The set in numbers
| | Count |
| ----------------- | ----: |
| Corner pieces | 4 |
| Edge pieces | 56 |
| Interior pieces | 196 |
| Interior colors | 17 |
| Frame-only colors | 5 |
## Why it matters
This is the clearest single sign that Eternity II was made hard on purpose.
Board size, piece count and color split all aim at the same target: a puzzle
with about one solution, placed at the worst possible spot for any search to
find it. The difficulty was chosen, the way a good exam is neither trivial nor
impossible.
How do we know the peak is real and not just a story? Two ways meet here. The
published analysis
[derives it](https://repositori.udl.cat/server/api/core/bitstreams/0b6533fe-54e5-4070-85fe-80f7d35837d8/content):
for framed edge-matching puzzles, the color count where you'd expect about one
solution falls near 17, and that is the hardest setting to search. And you can
watch a piece of it yourself in the demo above: build real puzzles, count the
work, and see it explode when colors are scarce. The impact is concrete. It
means the gap to a solution isn't a tuning problem you can grind away with a
faster machine; the puzzle was placed where search is worst on purpose, so
beating it needs a genuinely better idea, not just more effort.
See the difficulty measured live on the [Algorithms page](/algorithms).
## Related
- [Designed to be unsolvable: the recipe](https://eternity2.dev/research/why/design-recipe) — Eternity II follows a recipe for the hardest possible edge-matching puzzle: compact shape, no symmetric or duplicate pieces, split palettes, flat frequencies, one expected solution. The community reverse-engineered every ingredient in the launch year.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [Entropy and the area law](https://eternity2.dev/research/why/entropy-area-law) — Eternity II has two rules: edges must match, and each piece is used once. The first is generous. All the hardness lives in the second.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
---
# Piece theft, where solvers die
> A solver fills a few rows for free, then hits a wall in the middle of the board. Here's the mechanism: a scarce piece spent in the wrong place, rows ago.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/piece-theft/
- Updated: 2026-07-01
- Topics: structure, search-space
- Reproduce: `just research-piece-theft`
- Source: Régin 1994, “A Filtering Algorithm for Constraints of Difference in CSPs” (AAAI-94), the all-different reasoning that predicts starvation — https://cdn.aaai.org/AAAI/1994/AAAI94-055.pdf
---
Bar,
BarChart,
CartesianGrid,
ResponsiveContainer,
Tooltip,
XAxis,
YAxis,
} from "recharts";
Fill the board top-left to bottom-right and every new cell already knows two
of its colors: the north color from the piece above, the west color from the
piece to the left. The cell needs an unused piece that can show that exact
pair. Those demands are scarce, with only about three possible pieces on
average, and 47 of them have just one.
So a board that looks healthy, most pieces still in the box, can already be
doomed. Somewhere back up the board, the single piece that could ever serve
an upcoming cell was used for something else. When the solver finally reaches
that cell, there's nothing to place.
## How a cell dies
> **[Interactive: PieceTheftDiagram]** Rendered on the canonical page (link above); not shown in this markdown export.
## How many pieces can serve a demand
{data.uniqueServerDemands}
demands served by a single piece
{data.meanServers}
pieces per demand, on average
{data.occurringDemands}
distinct demands that occur
> **[Figure]** interactive: ServersChart. Rendered on the canonical page (link above); not shown in this markdown export.
## Starve a cell yourself
The engine fills a board to a cell with a single legal supplier; steal that
piece and watch the cell die with the box still full.
> **[Figure]** Interactive: where solvers die to piece theft — interactive: PieceTheftLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why it matters
A tempting fix is a global check: do the remaining pieces still cover the
remaining cells? It doesn't help. Globally the supply is fine; the failure is
one scarce piece misallocated, not a shortage. So a global lookahead sees
nothing wrong right up until the cell turns out to have no server, which is
why this wall resisted so many attempts to prune it early.
Set this beside [no forced moves](/research/why/no-forced-moves) and the trap
is complete. Every piece has dozens of places it could go, so the solver is
never told where a scarce piece must be saved, yet each scarce piece has
exactly one demand it must be saved for. Freedom to place, no guidance on
where to save.
## Related
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
- [PRIOR](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/prior) — Build a board from nothing, breaking ties by where pieces tend to sit in the strong boards we already have. It reaches a high score with no starting board to copy.
---
# Why a faster computer doesn't help
> The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/prune-vs-speed/
- Updated: 2026-07-01
- Topics: speed, search-space, backtracking
- Reproduce: `just research-prune-vs-speed`
- Source: Article, source and results on GitHub (research/topics/prune-vs-speed) — https://github.com/raphael-anjou/eternity2/tree/main/research/topics/prune-vs-speed
- Source: Peter McGavin: early doom-detection pruning is generally too expensive to be worthwhile (groups.io msg 11848) — https://groups.io/g/eternity2/message/11848
- Source: @95A31: a full battery of feasibility checks turned out to be useless on an 8×8 search (groups.io msg 11856) — https://groups.io/g/eternity2/message/11856
---
Picture the search as a tree. From the empty board you choose a piece for the
first cell; from there a piece for the second; and so on, 256 cells deep. The
number of leaves at the bottom is the branching factor raised to the depth,
an astronomically large number. To prove a region has no solution, a search
has to walk that tree.
Now there are two ways to do less work. You can go **faster**: a better
engine, more cores, hand-tuned inner loops. Or you can make the tree
**smaller**, pruning branches that can't lead to a solution, so the effective
branching factor drops. These sound similar. They are not even close.
## Constant versus exponential
A speedup is a constant divisor. Make the machine 1000× faster and you do
1000× less waiting, the same whether the tree is ten levels deep or ten
thousand. It buys you a fixed multiple, full stop.
A prune compounds. Shave even a few percent off the branching factor and you
save that fraction at every single level. Over 256 levels the savings multiply
into each other: cutting the branching factor from $b$ to $b'$ divides the
work by $(b/b')^{256}$. A 5% cut, applied all the way down, is worth
$(1/0.95)^{256} \approx 5\times10^{5}$, a five-hundred-thousand-fold
speedup's worth of work, from one cheap structural idea. That beats almost any
speedup a real machine can offer.
## Feel the gap
Trade a raw speedup against a small per-level prune and watch the prune win by
orders of magnitude.
> **[Figure]** Interactive: pruning power vs raw speed — interactive: PruneVsSpeedLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why this is exactly E2's curse
If pruning is the lever that matters, the hard puzzles are the ones you can't
prune. Eternity II was tuned to be precisely that. Four of its walls are,
underneath, all the same statement: there is nothing local to prune on.
- **[No forced moves](/research/why/no-forced-moves)**: every interior cell
still has 73 to 137 legal neighbours, so propagation almost never collapses
a cell to one choice. The branching factor stays stubbornly high.
- **[On the hardness peak](/research/why/phase-transition)**: the piece and
colour counts sit where there is about one expected solution, leaving no
solution-dense region to aim a statistical shortcut at, the trick that
cracked Eternity I.
- **[The area law](/research/why/entropy-area-law)**: the count of
genuinely-distinct partial boards collapses past ~80 cells, but no local
scoring signal can see that global collapse, so you can't prune toward it
cheaply.
- **[Rigidity](/research/why/rigidity-wall)**: even at a record board, the
move to a better one is huge and indivisible, with no gradient to follow and
nothing nearby to prune away.
## What it means for everything else here
This is the lens for the whole research section. A far faster engine makes
the same search cheaper, not smaller, and does not move the record. Every
experiment that did move the needle changed the shape of the search instead: a
different scan order, a learned prior over where pieces sit, a confined region
for the mismatches. And every dead end is, at heart, a prune that the puzzle's
global structure refuses to honour. Speed first feels productive; it is almost
never where the gap to 480 is hiding.
## The community landed here too, the hard way
The counterintuitive half of this is that even *legal* pruning often loses. A
check that detects a doomed partial board and backtracks early sounds like a
free win, but if the check costs more than the subtree it saves, a plain
backtracker that just barrels ahead is faster. Peter McGavin put the settled
view plainly on the groups.io list: methods that try to detect a doomed partial
placement and backtrack early "are generally considered too expensive to be
worthwhile". A newcomer running a decision-diagram solver, @95A31, then
confirmed it from scratch: after building a full battery of feasibility checks
he reported that "all the feasibility checks I implemented turned out to be
useless", with a complete 8×8 search still grinding through 953 billion nodes
over 17 hours. The lesson is not that pruning is bad, it is that a prune only
pays if it is *cheaper than the search it removes*, and on this puzzle almost
nothing local clears that bar.
> **Note**
>
> The tree numbers in the demo are illustrative: a branching factor and depth chosen to be E2-like and legible, not a measurement of a specific solver. The hardness curve and node counts, however, are real engine measurements on small puzzles, deterministic and reproducible with `just research-prune-vs-speed`. The principle itself, constant divisor versus exponential divisor, is exact.
## Related
- [Which wall stops which method](https://eternity2.dev/research/why/walls-and-methods) — The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- [Complex theory: counting the search before you run it](https://eternity2.dev/research/why/complex-theory) — Brendan Owen's complex theory estimates how wide the search tree is at every depth, and even how many solutions exist at all. Many in the community consider it the single most important thing to understand about Eternity II.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.
- [Tuned to the hardness peak](https://eternity2.dev/research/why/phase-transition) — Eternity II uses 22 colors. They split 17 interior to 5 frame-only, and that 17 is exactly where this kind of puzzle is hardest to solve.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [Arc consistency, from AC-3 up](https://eternity2.dev/research/build/reduce/arc-consistency) — Forward checking looks one move ahead; arc consistency makes every cell's candidate list defend itself against every neighbour's, to a fixed point. Mackworth's AC-3, the optimal refinements that followed, and what the whole family actually measured on this puzzle, including where it is unsound.
---
# The rare colors live on the frame
> Five of Eternity II's 22 colors appear only along the border ring, each on exactly 24 edges, never once in the interior. A structural split that shapes how every solver treats the frame.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/rare-color-geography/
- Updated: 2026-07-01
- Topics: structure
- Reproduce: `just research-rare-color-geography`
- Source: Brendan Owen derives the 17+5 colour design as the hardest possible puzzle (groups.io message 1947) — https://groups.io/g/eternity2/message/1947
---
Bar,
BarChart,
CartesianGrid,
Legend,
ResponsiveContainer,
Tooltip,
XAxis,
YAxis,
} from "recharts";
The 22 colors don't play equal roles. Sort them by where their edges sit and
they split into two clean classes: five rare colors fenced off to the border,
and seventeen common colors that do almost all their work inside the board.
## Where the rare colors are allowed
Pick one of the five rare colors and watch the board: all 24 of its edges light
up around the frame ring, and the 14×14 interior stays completely blank. Switch
to a common color and the interior fills instead. That empty middle, for every
rare color, is the whole finding.
> **[Figure]** interactive: RareColorRing. Rendered on the canonical page (link above); not shown in this markdown export.
## The five rare colors
> **[Interactive: RareSwatches]** Rendered on the canonical page (link above); not shown in this markdown export.
{data.edgesPerRareColor}
edges each
{data.rareEdgesTotal}
rare edges, all on the frame
0
rare edges in the interior
## Frame vs interior, color by color
> **[Figure]** interactive: FrameInteriorChart. Rendered on the canonical page (link above); not shown in this markdown export.
> **[Figure]** Interactive: the rare-colour border geography — interactive: RareColorLab. Rendered on the canonical page (link above); not shown in this markdown export.
## Why it matters
The split itself is structural, not a defensive trick. Because the outer rim
is solid grey, every border piece keeps its grey edge facing out, so its
coloured edges only ever meet other border edges or the interior: the border
and interior colours live in separate pools no matter what. The five border
colours read as "rare" simply because there are far fewer border edges to
colour, and they could be relabelled to any five values (even reusing interior
ones) without changing the puzzle at all. (See the [design note on
why](/research/why#engineered-to-resist-cleverness), with thanks to Vasily V.
on the groups.io list for the correction.)
What is real, and what matters for solving, is the consequence: the 14×14
interior is left to seventeen common colours with no rare, highly-constraining
signal anywhere in it, which is a large part of why interior search has so
little to grip. The border, where the colour pool is small and matches are
tight, is correspondingly the part every strong board gets essentially perfect.
Eternity II's colour counts were balanced to remove the statistical handholds
that sank Eternity I; this geography is where you feel the result.
See [why the interior gives no forced moves](/research/why/no-forced-moves).
## Related
- [Designed to be unsolvable: the recipe](https://eternity2.dev/research/why/design-recipe) — Eternity II follows a recipe for the hardest possible edge-matching puzzle: compact shape, no symmetric or duplicate pieces, split palettes, flat frequencies, one expected solution. The community reverse-engineered every ingredient in the launch year.
- [Forbidden patterns](https://eternity2.dev/research/why/forbidden-patterns) — Almost every small patch of pieces you could build is impossible. For a 2×2 square, 99.72% of the ways to place four pieces can never be made to match.
- [The border balance](https://eternity2.dev/research/why/border-balance) — A solved board hides a simple bookkeeping law: every colour the border hands to the interior, the interior hands straight back. Break it and you know instantly the board is wrong; obeying it, though, guarantees nothing.
- [CLOISTER](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/cloister) — Fix a perfect border, then search the interior with the border's edges treated as hard constraints from the very first cell.
---
# The rigidity wall
> Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/rigidity-wall/
- Updated: 2026-07-10
- Topics: structure, local-search, exact-methods
- Source: Local rigidity theorem (project paper, 2026-05-16): ≥13 MIP proofs of halo-optimality across 3–4 basins, out to halo-4 — https://github.com/raphael-anjou/eternity2/tree/main/research
- Source: Land & Doig 1960, branch and bound, the exact method behind every region solve — https://doi.org/10.2307/1910129
- Source: benj39100: GPU annealing on the strict-5-clue board hits the same frozen core; best score grows with distance from the incumbent (groups.io msg 11902) — https://groups.io/g/eternity2/message/11902
- Source: William Millilaw's independent halo SAT-residual and replica freeze tests, reproduced here on the public record boards (all UNSAT out to a two-cell halo) — https://github.com/raphael-anjou/eternity2/tree/main/research/topics/rigidity-sat-halo
---
Here is the intuition almost everyone starts with: get close, then tidy up
the last few mismatches. Swap a couple of pieces, rotate a few, and surely
the score creeps up to 480.
It doesn't. On every top board we have tested, the last mismatches are
locked. Move anything nearby and the score stays the same or drops. The good
boards aren't almost-solutions waiting for a polish; they're isolated points
with nowhere better to go next door.
## Watch the region grow
> **[Interactive: RigidityHalo]** Rendered on the canonical page (link above); not shown in this markdown export.
## Try to beat it yourself
A real perfect small board. Swap any two pieces, or let the engine try every
swap. Nothing beats it; that's rigidity, live.
> **[Figure]** Interactive: swap pieces and watch the score refuse to rise — interactive: RigidityLab. Rendered on the canonical page (link above); not shown in this markdown export.
## How this is known
The result does not rest on trying a lot of swaps and giving up. It comes from
an exact question put to an integer-programming solver: take a region of the
board, free every piece inside it, and find the single best way to fill it back
in. The solver searches that whole local space exactly, not by sampling.
The answer keeps coming back the same: the best arrangement is the one
already there. We proved it for region after region, across three different
record boards, and out to a radius of four cells, which is a block of dozens
of cells at once. Every time, no improvement exists.
## Method: the MIP proof
The exact question, stated precisely. Around each defect cell take the
*halo-$r$* region: every cell within $r$ steps. Free all pieces inside it,
keep the rest of the board fixed as a boundary, and solve an integer program
for the best legal refill. Binary variables $x_{c,p,\theta} \in \{0,1\}$ place
piece $p$ at rotation $\theta$ in cell $c$; one-piece-per-cell and
one-cell-per-piece constraints make it an assignment problem; the objective
maximizes matched edges, counting the fixed boundary. A 64-cell region is about
18,700 binary variables. Branch and bound then either finds a strictly better
refill or proves, with a dual bound, that none exists.
It never finds one. The proofs, region by region:
| Region | Board | Cells | Solver time | Result |
|---|---|---:|---:|---|
| halo-2, per-component (×8) | 3 basins (469, 459, 458) | 2–32 | seconds each | all Δ = +0, **proven** |
| halo-1 joint | McGavin 469 | 37 | 895 s | Δ = +0, **proven** |
| halo-1 joint | Local 459 | 59 | 1800 s | Δ = +0, **proven** |
| halo-3, per-component | McGavin 469 | 42 + 47 | 929 s | Δ = +0, **proven** |
| halo-4, component 0 | McGavin 469 | 57 | 1200 s | Δ = +0, **proven** |
Across four basins and ≥13 regions the answer is invariant: the incumbent
board is locally MIP-optimal, out to halo-4 for McGavin's 469. The one region
left with a gap, the top-four rows, 64 cells, still gives a *sound* result:
a dual bound of 123 against the incumbent's 116, so even the unclosed case
cannot exceed the wall by much (it implies a board-wide bound of ≤ 476).
The badge says *proven* rather than *conjectured* for the regions actually
closed; the general statement over *all* boards remains a conjecture, supported
by every region tested to date.
## Why it matters
This reframes the whole gap to 480. The distance from the best known board to
a solution is not a pile of small fixes waiting to be found. If it were, this
kind of local search would have found them. The barrier is that the good
boards sit at the bottom of their own little valleys, and the valley walls
are exact, not approximate.
It also tells you what won't work. Polishing, hill-climbing, and most
local-repair heuristics are trying to walk uphill from a frozen point. There
is no uphill. Reaching a solution needs a move that rearranges a large region
all at once, or a different starting point entirely, not a better polish.
## Someone hit the same wall from the annealing side
The MIP proof approaches the wall analytically. In June 2026 another researcher
ran headlong into it empirically. Working the strict five-clue board with a GPU
simulated-annealing solver (4096 parallel replicas), benj39100 reported two
things that read like a restatement of this page. First, the best score a run
could reach *increased with distance from the current best board*: to climb
from 429 toward 432 the winning perturbations had to migrate steadily further
out, because near the incumbent there was no improving move to find. Second,
across all their best boards the same roughly forty-two broken edges formed a
frozen "hard core" that local search never touched, so they had to add an
explicit term that *rewarded* prying that core open. A locally frozen incumbent
with no nearby gradient, and a small locked set of defects nothing local will
move: that is the rigidity wall, seen from a completely different method.
## And again, from a SAT solver
A third researcher reached the same conclusion with a third tool. William
Millilaw, working the ceiling independently, ran two tests. The first was a
freeze test: perturb the root of a top board, re-optimize, and see which cells
come back. On his best boards 93 to 100 percent of the cells returned to exactly
where they were, a frozen core the search could not move. The second was a
decision question for a SAT solver. Free the cells around the mismatches, demand
that every freed edge match, and ask whether any arrangement of those pieces
satisfies it. His solver answered UNSAT.
We [reproduced that SAT test](https://github.com/raphael-anjou/eternity2/tree/main/research/topics/rigidity-sat-halo)
on the public record boards, with our own encoder and a positive control to make
sure a matched region comes back satisfiable. Every record board we tried, from
Verhaard's 467 up to Blackwood's 470, is UNSAT out to a two-cell halo: no local
rearrangement of a record board's own pieces closes a single mismatch. Integer
programming optimizes and bounds; annealing runs into the wall by hand; SAT
decides and returns a refutation. Three methods, one answer.
His earlier snake-and-local-search work put a number on why. The largest patch
any of his repair operators could rewrite in one move was about 48 cells, while
two of the known good boards, both near the ceiling, differ across roughly 225
cells. A move that can only touch 48 cells cannot cross a 225-cell gap, so no
sequence of them reaches a different basin. He looked for a chain of small
improving moves that would tunnel out and found none: zero improvements across
twenty-eight million four- and five-move combinations at the plateau. That is
the same wall stated as a budget. Local repair rewrites too little at once to
leave the valley, which is why escaping needs a move that rearranges a large
region all at once, exactly as the proof section above concludes.
> **The next piece of the picture**
>
> If you can't improve a board locally, maybe you can hop from one good board to another. That fails too, for a related reason: [see why basin-hopping is impossible](/research/why/sigma-cycles).
*The proofs use integer programming on regions of each board; they run for
minutes per region on a solver, so they aren't reproduced live here. The
boards themselves are bundled and checkable in the viewer.*
## Related
- [Why basin-hopping looks impossible](https://eternity2.dev/research/why/sigma-cycles) — If you can't improve a great board by polishing it, maybe you can jump to a different great board. On every record pair tested, you can't, and the structural reason why is worth seeing.
- [Which wall stops which method](https://eternity2.dev/research/why/walls-and-methods) — The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [PALIMPSEST](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/palimpsest) — Read every strong board to find the habits that quietly hold a board back, then break them. This experiment produced the project's best board: 463 of 480.
- [MIDDEN](https://eternity2.dev/research/lab/experiments/raphael-anjou/pipelines/midden) — Decide in advance not when a board may break, but where: confine every mismatch to a chosen shape of cells, and search for the best shape.
- [LP and ILP relaxations: half a piece everywhere](https://eternity2.dev/research/build/exact/lp-relaxations) — Write Eternity II as an integer program, drop the integrality, and a linear solver reaches zero error in seconds, with 30% of one piece and 20% of another sharing a corner. Eighteen years of community campaigns measured where the fractional comfort ends: a plateau at 420–440 edges the moment the pieces must be whole, an ILP wall at 8×8, and an academic best of 461-in-an-hour. What the optimizer's road teaches, and where LP still earns its keep.
---
# Why basin-hopping looks impossible
> If you can't improve a great board by polishing it, maybe you can jump to a different great board. On every record pair tested, you can't, and the structural reason why is worth seeing.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/sigma-cycles/
- Updated: 2026-07-02
- Topics: structure, local-search
- Source: Peter McGavin's 469-record announcement: the board the cycles are computed against (groups.io msg 10045, September 2020) — https://groups.io/g/eternity2/message/10045
---
Two top boards look totally different, but they're related: you can turn one
into the other by picking up a set of pieces and shifting each one into the
spot the next was using, all the way around a loop. Mathematicians call that
loop a cycle. Do the whole loop and you arrive at the other board.
Here's the catch. Between the best boards, that loop is enormous and there's
only one of it. Going from a 459 board to McGavin's 469 means one
single interlocking cycle of 154 cells, spanning the entire interior of the
board.
## One loop, all or nothing
> **[Interactive: SigmaCycleDiagram]** Rendered on the canonical page (link above); not shown in this markdown export.
## Try it on a real small puzzle
This is no diagram: the live engine finds two real solutions, computes the actual
cycles between them, and scores every step you take.
> **[Figure]** Interactive: apply part of the swap cycle — interactive: SigmaCycleLab. Rendered on the canonical page (link above); not shown in this markdown export.
## How we know
We computed the exact cycles between three different record boards and then
tried every partial move: apply just some of the loop and see what the score
does. The result is blunt. Every partial application makes the score worse,
with drops ranging from a few points to over a hundred. There is no subset of
the loop that helps, not one.
That's what makes it a wall. To move from one good board toward a better one
you'd have to commit to shifting up to 154 cells at once, with no improving
step along the way to guide you there. Every search that works in steps is
blind to a move like that.
## The cycles, measured
- 459 to 458: 14 cycles, largest 85 cells.
- 458 to McGavin 469: one giant cycle of 80 cells, spanning rows 1 to 14.
- 459 to McGavin 469: one giant cycle of 154 cells.
- **In every case, every smaller piece of the cycle scores worse: the
property held in all three basin pairs tested, though three pairs is a
small sample and it is not proven to hold in general.**
## Why it matters
Together with [the rigidity wall](/research/why/rigidity-wall), this closes
off the two obvious escape routes at once. You can't climb out of a good
board locally, and (on every record pair measured here) you can't hop to a
neighbour either, because the nearest better board is one indivisible move of
80 to 154 cells away. This is a plausible part of why the 470 record has stood
since 2021: the moves that would beat it appear too large for any
step-by-step search to find.
*Cycles are computed exactly from pairs of record boards by composing their
piece permutations; the subset tests enumerate sub-cycles and rescore. The
animation above is a schematic of the mechanism, not the 154-cell cycle
itself.*
## Related
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [Where the mismatches live](https://eternity2.dev/research/why/mismatch-geometry) — A near-perfect board doesn't scatter its few errors evenly. It packs them into one band of five rows and leaves all the rest flawless. Which band is decided by the direction the search filled the board, and you can see the mirror on the real record boards.
- [KEYRING](https://eternity2.dev/research/lab/experiments/raphael-anjou/learning/keyring) — Build a board from scratch, ranking each next piece by three signals learned from past strong boards. Reached 460 in a board family no earlier search had cracked.
---
# Which wall stops which method
> The research section has two halves: the structural walls that make Eternity II hard, and the algorithms built to climb them. This page is the bridge: each method lined up against the wall it actually attacks, and the score where that wall stopped it.
- Canonical page (with interactive figures/demos): https://eternity2.dev/research/why/walls-and-methods/
- Updated: 2026-07-02
- Topics: structure
- Source: Ansótegui, Béjar, Fernández & Mateu, How Hard is a Commercial Puzzle: the Eternity II Challenge — https://repositori.udl.cat/server/api/core/bitstreams/0b6533fe-54e5-4070-85fe-80f7d35837d8/content
---
Every cell of the map below is grounded in this project's published work. The
four walls themselves are corroborated by the published literature; the
experiment results are this project's own work, each written up on its own page; they
are not claimed to be externally verified.
## The four walls, in one line each
Every one of them is a version of the same statement: there is nothing local
to prune on.
- **[No forced moves](/research/why/no-forced-moves)**. Every interior cell
keeps 73–137 legal pieces, so the branching factor never collapses.
- **[The hardness peak](/research/why/phase-transition)**. With ≈17 interior
colours the puzzle sits at the phase transition: about one expected solution,
the worst place to search.
- **[The area law](/research/why/entropy-area-law)**. Genuinely-distinct
partial boards collapse past ~80 cells, but no local score can see that
global fact.
- **[Rigidity](/research/why/rigidity-wall)**. Records are locally frozen;
the step to a better board is one giant, indivisible swap with no gradient
to follow.
## The map
Columns are the walls; rows are the methods, best score first. A filled mark
(●) means the method is fundamentally working against that wall. The community
ceiling on this puzzle is 470; the full solution is 480.
| Method | Best | [Forced moves](/research/why/no-forced-moves) | [Hardness peak](/research/why/phase-transition) | [Area law](/research/why/entropy-area-law) | [Rigidity](/research/why/rigidity-wall) | New basin |
| --- | --- | :-: | :-: | :-: | :-: | --- |
| [PALIMPSEST](/research/lab/experiments/raphael-anjou/learning/palimpsest) · from the corpus | 463/480 | | | | ● | no |
| [PRIOR](/research/lab/experiments/raphael-anjou/learning/prior) · from scratch | 460/480 | | ● | | ● | no |
| [KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring) · from scratch | 460/480 | | | | ● | new family |
| [REPLAY](/research/lab/experiments/raphael-anjou/learning/replay) · decode & replay | 460/480 | | | | ● | no |
| [GAUNTLET](/research/lab/experiments/raphael-anjou/pipelines/gauntlet) · from scratch | 458/480 | | | | ● | new family |
| [CLOISTER](/research/lab/experiments/raphael-anjou/pipelines/cloister) · anchor & constrain | 453/480 | | | ● | ● | no |
| [MIDDEN](/research/lab/experiments/raphael-anjou/pipelines/midden) · anchor & constrain | 452/480 | | | ● | ● | no |
| [LADDER](/research/lab/experiments/raphael-anjou/pipelines/ladder) · concentrate effort | 451/480 | ● | ● | | | new family |
| [LODESTONE](/research/lab/experiments/raphael-anjou/learning/lodestone) · from scratch | 451/480 | ● | | | | no |
| [MOSAIC](/research/lab/experiments/raphael-anjou/pipelines/mosaic) · solve a piece exactly | 448/480 | ● | ● | | | no |
| [BANDSAW](/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw) · solve a piece exactly | 437/480 | ● | ● | | | no |
| [STAGED](/research/lab/experiments/raphael-anjou/pipelines/staged) · from scratch | 436/480 | ● | | ● | | no |
### How to read it
Almost every method ends up against rigidity, the wall that says the great
boards are isolated islands. The build-from-scratch and corpus methods (PRIOR,
KEYRING, PALIMPSEST, GAUNTLET) try to reach a new island by steering
construction with learned signal; they top out at 458–463 and a couple of them
do reach genuinely new families, but none crosses to the ceiling. The
concentrate and exact methods (LADDER, BANDSAW) instead attack the search
itself (the high branching factor and the unsearchable peak) and pay for it
at the endgame. The anchor methods (CLOISTER, MIDDEN) localise the damage but
hit the area-law wall in the interior. No single wall is the whole story, and
no method gets through all four.
## Where each one stopped, and why
The ceiling is never arbitrary. For each method, its own write-up records the
exact reason the score stopped climbing, quoted here in one line.
- **[PALIMPSEST](/research/lab/experiments/raphael-anjou/learning/palimpsest)** (463/480). Reached
463, the project best, by reading the corpus to find which shared choices
are traps and steering a 15-basin sweep around them. Forcing the search to
avoid the traps directly made boards worse: the value was in where to look,
not a hard rule.
- **[PRIOR](/research/lab/experiments/raphael-anjou/learning/prior)** (460/480). Plateaus at 460:
the learned position prior gets a from-scratch build into the 460 class
fast, but the corpus signal alone is not enough to leave it.
- **[KEYRING](/research/lab/experiments/raphael-anjou/learning/keyring)** (460/480). Three learned
signals voting (position, adjacency, 2×2 patch) reached 460 in a corner
arrangement no board had cracked before, a new family, but the patch
signal is marginal and the polish still caps at 460.
- **[REPLAY](/research/lab/experiments/raphael-anjou/learning/replay)** (460/480). Replays the
community's strict-460 boards exactly and reveals the move ordinary search
misses: 4–5 cells that take two mismatches at once, unreachable for a search
that allows at most one.
- **[GAUNTLET](/research/lab/experiments/raphael-anjou/pipelines/gauntlet)** (458/480). Running the
beam in nine scan directions opened a brand-new 458 family (scan order is a
stronger diversity axis than the random seed), but a second round topped
out at 457 with no 461: the new family saturates like the others.
- **[CLOISTER](/research/lab/experiments/raphael-anjou/pipelines/cloister)** (453/480). As a
standalone interior solver it confirms a real rim-compatibility bonus, but
that bonus cannot be retrofitted after the fact (the same rigidity as the
full board), so it settles in the low 450s.
- **[MIDDEN](/research/lab/experiments/raphael-anjou/pipelines/midden)** (452/480). Choosing where
(not when) the board may break extends the perfect run from 153 to 167–174
cells, but the dispersed geometry still fails at the endgame: nothing
absorbs the last damage.
- **[LADDER](/research/lab/experiments/raphael-anjou/pipelines/ladder)** (451/480). Floods cheap
probes and promotes the deepest, reaching a 451 strict board with no record
to copy, the first escape from the universal 444–450 band, but the supply
of perfect openings runs out and the rungs all converge to one ceiling.
- **[LODESTONE](/research/lab/experiments/raphael-anjou/learning/lodestone)** (451/480). A
scarce-demand prior used only as a tiebreaker lifts the from-scratch median
by two (449→451) and tightens the variance, but any larger weight collapses
it: scarcity is a real but weak signal, and it never touches the basin
ceiling.
- **[MOSAIC](/research/lab/experiments/raphael-anjou/pipelines/mosaic)** (448/480). Composes exact
4×4 block solutions with soft seams, reaching 448 from scratch, but the
shortfall lands almost entirely in the last three corner blocks, where the
piece pool runs thin: the same piece-theft, now a single bright spot.
- **[BANDSAW](/research/lab/experiments/raphael-anjou/meet-in-the-middle/bandsaw)** (437/480). Solves an
endgame band to proven optimality, and in doing so measures the exactness
wall: the search tree grows about twentyfold per extra allowed mismatch, on
both sides, so meeting in the middle stops paying at full size.
- **[STAGED](/research/lab/experiments/raphael-anjou/pipelines/staged)** (436/480). Builds the whole
board with no pre-set frame and an emergent border, reaching 436, well
below the records. That gap is the finding: it measures exactly how
much the usual frame-first anchor is worth.
## The shape of the gap
Read down the table and the lesson of the whole project is visible at a
glance: the methods that move the score change the shape of the search, whether
a scan order, a learned prior, or a confined region, never its raw speed. And
every one of them stops at a wall that is global, not local. The ten edges
from 470 to 480 are not a polishing problem; they are on the far side of all
four walls at once.
A separate 2026 campaign by William Millilaw reached the same conclusion from
the diversity side. Sweeping the full method list, he found that almost
everything (adaptive local search from scratch, snake placement, the top of a
trained generator's distribution) keeps rediscovering the same handful of
basins, and only parallel tempering reliably produced genuinely new ones,
boards far apart from the known set. Even that plateaus quickly. His reading is
the one this table keeps making: the bottleneck is not the score a method
reaches but the number of distinct basins it can find, and no method in the
standard arsenal finds enough of them.
> **Note**
>
> Each saturation line is distilled from the project's lab notebook (one entry per experiment). The four walls are corroborated by the published literature: the 17-colour phase transition by Ansótegui, Béjar, Fernández & Mateu, "How Hard is a Commercial Puzzle: the Eternity II Challenge"; the missing gradient and deep basins by the Eternity II local-search literature.
## Related
- [Why a faster computer doesn't help](https://eternity2.dev/research/why/prune-vs-speed) — The single most important idea in hard combinatorial search: shrinking the space you search beats searching it faster, by an exponential margin. Eternity II is engineered so you can barely shrink it at all.
- [The rigidity wall](https://eternity2.dev/research/why/rigidity-wall) — Every record board we have is frozen in place. You cannot nudge your way from a great board to a perfect one, and we can prove it.
- [Experiments](https://eternity2.dev/research/lab/experiments) — The lab's named search experiments, one section per researcher. Each is a real run against Eternity II with its idea, its best board, and the questions it left open. Raphaël Anjou's notebook is here in full; the notebook is open to anyone else's.
- [No forced moves](https://eternity2.dev/research/why/no-forced-moves) — The usual way to crack a logic puzzle is to find a spot where only one piece fits, place it, and repeat. That lever doesn't exist here: every interior piece has between 73 and 137 possible neighbours, and not one is ever pinned to a single option.