Settlers / Research

73 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Every search switch under the learned leaf

How the retest was measured

Each study took one family of switches from the hand-written-leaf program and measured it against the depth-2 search with the learned tables as its leaf, v2:{"depth":2,"leaf":{"tables":"artifacts/ntuple/hex-portfolio-main.bin"}}. Every contrast is a swapped pair on the same deterministic boards: the candidate in slot 0 and the control in slot 1, then the other way round. The effect is half the per-board difference of the two contrasts, which cancels the turn-order seating term between two adjacent searches. Two lineups filled the other seats. The builder table, an ETA and a fast builder, repeats the lineup of the hand-written-leaf cohorts. The population table, a third tables search and an ETA, answers trades, blocks, and robber moves with an opponent that trades and blocks back.

The stages were fixed in advance: development on boards 0 to 63 (the boards of the earlier cohorts), confirmation on fresh boards 800 to 863, the population table on the same fresh boards, and an extension to boards 864 to 927 for any switch whose builder-table result was positive but unresolved. A switch counted as worth the browser only with a pooled effect of at least +0.03 wins per game, an interval excluding zero, and a decision time within a quarter of the control's.

Switches that only change a hand-written leaf term do nothing when the tables alone are the leaf, so those were measured on top of a blended leaf. An early screen showed that the full blend loses about 0.2 wins per game to the tables alone, so every leaf-term arm moved to the quarter blend before its cohorts ran, and the full-blend registrations that had not started were withdrawn.

The evidence

Every number below was recomputed from the archived runs, not copied from the study reports, with analysis/tables_retest.py. Each cell pools every stage a switch ran at that table (after-the-fact pooling over the boards named), with a normal 95% interval over boards. Negative values mean the switch loses games. Time is the candidate's mean decision time over the control's.

StudySwitchLeafBuilder tablePopulation tableTimeReading
BargainingRefuse asks (stubborn seat)tables+0.006 (−0.008 to +0.020), 192 boards+0.006 (−0.023 to +0.035), 128 boards0.98×No clear effect
BargainingTwo-for-one asks offtables−0.108 (−0.152 to −0.065), 128 boards−0.037 (−0.093 to +0.018), 64 boards0.93×Hurts at the builder table
BargainingAll bargaining offtables−0.165 (−0.213 to −0.117), 128 boards−0.098 (−0.152 to −0.044), 64 boards0.63×Hurts at both tables
BargainingCounters offtables−0.094 (−0.135 to −0.052), 128 boards−0.119 (−0.176 to −0.062), 64 boards0.71×Hurts at both tables
BargainingThreat pricing offtables−0.006 (−0.051 to +0.039), 128 boards+0.014 (−0.041 to +0.069), 64 boards0.96×No clear effect
Search settingsBeliefs from public counts onlytables−0.143 (−0.188 to −0.097), 128 boards−0.053 (−0.107 to +0.002), 64 boards1.18×Hurts at the builder table
Search settingsTwo deeper-level scenarios (inert at depth 2)tables+0.000 (+0.000 to +0.000), 128 boardsnot run1.00×No clear effect
Search settingsLeaf-guided opponentstables+0.031 (−0.006 to +0.068), 192 boards+0.033 (−0.009 to +0.075), 128 boards1.85×No clear effect
Search settingsNever propose tradestables−0.172 (−0.211 to −0.133), 128 boards−0.143 (−0.197 to −0.089), 64 boards0.91×Hurts at both tables
Search settingsReciprocity scalingtables−0.073 (−0.115 to −0.031), 128 boards−0.082 (−0.141 to −0.023), 64 boards1.05×Hurts at both tables
Search settingsOne sampled world (default 4)tables−0.109 (−0.150 to −0.069), 128 boardsnot run0.34×Hurts at the builder table
Search settingsEight sampled worldstables+0.045 (+0.009 to +0.081), 192 boards−0.007 (−0.045 to +0.031), 128 boards1.75×Helps at the builder table only
Search settingsSixteen dice scenariostables−0.026 (−0.072 to +0.019), 128 boardsnot run2.73×No clear effect
Search settingsFour dice scenarios (default 8)tables−0.392 (−0.444 to −0.339), 128 boardsnot run0.68×Hurts at the builder table
Leaf blendExpected hidden pointsblend 0.25+0.003 (−0.017 to +0.023), 192 boards−0.009 (−0.037 to +0.019), 128 boards1.00×No clear effect
Leaf blendEndgame race leafblend 0.25−0.037 (−0.066 to −0.008), 128 boards+0.000 (−0.036 to +0.036), 64 boards1.01×Hurts at the builder table
Leaf blendBlend: hand-written terms at 0.25tables+0.066 (+0.024 to +0.107), 192 boards+0.064 (+0.021 to +0.108), 128 boards1.42×Helps at both tables
Leaf blendBlend: hand-written terms at 0.5tables+0.016 (−0.024 to +0.057), 192 boards−0.019 (−0.064 to +0.027), 128 boards1.34×No clear effect
Leaf blendTables scale 2000tables+0.029 (−0.008 to +0.066), 192 boards+0.034 (−0.011 to +0.079), 128 boards0.96×No clear effect
Leaf blendTables scale 500 (default 1000)tables−0.097 (−0.139 to −0.054), 128 boardsnot run1.07×Hurts at the builder table
OpeningSearch-ranked opening instead of the plannertables+0.061 (+0.009 to +0.113), 192 boards+0.001 (−0.059 to +0.061), 128 boards1.31×Helps at the builder table only
OpeningPlanner coverage 0.4 and balance 0.5tables+0.020 (−0.014 to +0.053), 192 boards+0.004 (−0.034 to +0.042), 128 boards1.01×No clear effect
OpeningPlanner without scarcity markuptables−0.005 (−0.063 to +0.053), 128 boards−0.012 (−0.073 to +0.049), 64 boards0.97×No clear effect
Card timingBuy cards only when stalledtables−0.033 (−0.075 to +0.009), 128 boards−0.053 (−0.110 to +0.004), 64 boards0.70×No clear effect
Card timingHold knights for the unblocktables−0.027 (−0.060 to +0.005), 128 boards−0.051 (−0.098 to −0.004), 64 boards0.98×Hurts at the population table
Card timingKnight schedule in the leafblend 0.25−0.001 (−0.021 to +0.019), 128 boards+0.004 (−0.026 to +0.034), 64 boards1.00×No clear effect
Hand managementProject production when discardingtables+0.012 (−0.010 to +0.034), 192 boards−0.002 (−0.023 to +0.019), 128 boards1.00×No clear effect
Hand managementSpend down over the discard limittables−0.006 (−0.042 to +0.030), 128 boards−0.047 (−0.098 to +0.004), 64 boards1.19×No clear effect
Hand managementNo seven-risk termblend 0.25−0.007 (−0.041 to +0.027), 128 boards+0.012 (−0.040 to +0.063), 64 boards0.99×No clear effect
RoadsBlock the leader's expansionblend 0.25+0.010 (−0.017 to +0.037), 192 boards−0.033 (−0.060 to −0.006), 128 boards1.31×Hurts at the population table
RoadsBlock the leader's expansiontables−0.003 (−0.011 to +0.005), 128 boards+0.002 (−0.007 to +0.011), 64 boards1.14×No clear effect
RoadsGate the road race on productionblend 0.25−0.060 (−0.095 to −0.024), 128 boards+0.021 (−0.026 to +0.069), 64 boards0.94×Hurts at the builder table
RoadsGate the road race on productiontables−0.019 (−0.051 to +0.014), 128 boards−0.023 (−0.060 to +0.013), 64 boards0.98×No clear effect
PredictionOpponents play their predicted buildtables−0.005 (−0.047 to +0.037), 128 boards−0.004 (−0.060 to +0.052), 64 boards1.01×No clear effect
PredictionThreat from the predicted sitetables+0.004 (−0.032 to +0.040), 128 boards+0.047 (−0.004 to +0.098), 64 boards0.98×No clear effect
PredictionDeny predicted sitesblend 0.25−0.018 (−0.045 to +0.010), 128 boards+0.023 (−0.009 to +0.056), 64 boards2.00×No clear effect
RobberLate leader biastables+0.011 (−0.005 to +0.027), 192 boards+0.015 (−0.007 to +0.036), 128 boards1.01×No clear effect
RobberLate leader bias plus rob what you needtables+0.035 (−0.015 to +0.085), 64 boards−0.039 (−0.090 to +0.012), 64 boards1.01×No clear effect
RobberRobber blocks the leadertables+0.007 (−0.035 to +0.049), 128 boards+0.031 (−0.030 to +0.092), 64 boards0.98×No clear effect
RobberRob the card this seat needstables+0.016 (−0.008 to +0.041), 192 boards+0.009 (−0.018 to +0.036), 128 boards0.99×No clear effect
RobberRobber by threat, spare partnerstables−0.005 (−0.042 to +0.033), 128 boards+0.043 (−0.005 to +0.091), 64 boards0.98×No clear effect

The study reports have the stage-by-stage cohorts, run IDs, and diagnostics: bargaining, search settings, leaf blend, opening, card timing, hand management, roads, prediction, and robber.

What changed between the two leaves

Three things held under both leaves. Counters are worth about a tenth of a win per game under either (−0.096 on the hand-written leaf's swapped pair, −0.094 and −0.119 here). The seating correction is essential: several arms that looked like gains in one seating were the turn-order term, and the search-ranked opening repeated that lesson in a quieter form, reading +0.135 on the reused development boards and +0.024 on fresh ones. And the tactical switches built from player advice (block the leader, rob what you need, hold knights, spend down before a seven, cut off expansion) do not help a search that already looks two turns ahead. The one reversal is the robber's victim rule: stealing the card this seat needs was refuted under the hand-written leaf (−0.045) and is inconclusive under the tables (+0.016 at the builder table, +0.009 at the table of searches).

Two things changed. The search's information inputs matter more under the stronger leaf. Under the hand-written leaf the number of sampled worlds showed no effect at depth 2 and public-count beliefs cost about 0.07 wins per game; under the tables one sampled world costs 0.109 and public-count beliefs 0.143. The dice scenarios were the largest single lever under both leaves (the earlier arm cut them to two, this one to four). And the learned leaf profits from a dose of the hand-written terms, as long as the dose is small: a quarter helps at both tables, a half is noise, and the full weight loses.

What this means for the browser

The browser player runs the hand-written leaf in WASM at depth 3 under one second, with the opening planner and bargaining on by default. Nothing measured here argues for changing that configuration:

  • Keep all five bargaining switches on. Counters and two-for-one asks carry the strength, threat pricing supplies the stated reasons for refusals, and accepting asks costs nothing.
  • Keep eight dice scenarios and four sampled worlds. The cheaper settings are the obvious way to buy depth in a one-second budget, and they are the most expensive losses measured.
  • Leave every switch of the robber, timing, hand, road, prediction, and endgame families off.

For a browser player that loads the learned tables, the evidence favours the quarter blend, "leaf":{"tables":…,"hand":0.25}, with the tables scale left at 1,000 (half the scale stops the seat trading and loses 0.1 wins per game). The blend costs 1.4 times the decision time, so under a fixed one-second budget it will reach depth 3 less often; that trade has not been measured, and the blend has not been confirmed through the protocol arena.

The study of search settings at the browser's own budget was stopped before its registered stages finished. Its unregistered screens agreed that four dice scenarios are ruinous and suggested that two sampled worlds and a per-decision transposition cache might pay at depth 3; those screens are not findings. Its work is kept on the research/browser-budget branches of research and server.

What could still be wrong

Every result is engine-arena evidence at depth 2 against one builder lineup and one population lineup; the browser plays at depth 3 against people. The pooled numbers are after-the-fact pooling across stages whose own rules were fixed in advance, and the development boards 0 to 63 were already used by the earlier program, which inflated at least one switch. Effects near +0.03 are below what these cohorts resolve; a switch reported as no clear effect may still be worth a few hundredths of a win. And the blended leaf is a new configuration: its gain holds at both tables here but has not been seen through the protocol arena or against people.

Agent notes

Every number in this report comes from python3 analysis/tables_retest.py --markdown TABLE.md, which rebuilds each swapped pair from the registered experiments and the archived runs of the program's studies and pools each switch after the fact per leaf and lineup. The studies were run by GLM-5.3 and Kimi K3 through opencode and by Claude Fable 5.1 and Claude Opus 5 through Claude Code; each study report names its own attribution. This synthesis was written by Claude Opus 5 (claude-opus-5) through Claude Code.