Settlers / Research

73 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Reading the next build under the learned leaf

The question

The hand-leaf study of these mechanisms (reports/build-prediction.mdx) measured three switches of the build predictor in v2/predict.rs, each off by default, and found nothing that survived fresh seeds. The search now evaluates with learned n-tuple tables, which change both what a good position looks like and what the rival seats do, so every switch was retested under the tables. Two of the three act under the tables alone:

SwitchWhat it does
predict.threatThe bargaining space threat keys on the other seat's predicted target site instead of a frontier scan, so only a seat whose predicted line ends on a contested site counts as threatening
predict.opponentsThe lookahead's opponent model executes the predicted build, trading toward it or saving, instead of following the builder's greedy step-by-step policy
weights.denial 0.5A leaf term discounting open sites opponents are predicted to take, in proportion to the site's value and the opponent's time to afford the line

The denial weight is a leaf term and is inert under the tables alone, so that arm runs on a blended leaf. Its first registrations used the full blend (leaf hand 1.0 in both seats); the patterns/leaf-blend screen then measured that blend 0.207 wins per game below the tables alone, so all six registrations were withdrawn before any of them ran and the arm was registered again on the quarter blend (hand 0.25), which the screen measured +0.074 above the tables. Everything below reports the quarter-blend arm, and the blend itself is measured by that study, not here.

How the cohorts ran

Each arm is a swapped pair of registered deterministic cohorts: half A seats the candidate in slot 0 and the control in slot 1, half B the reverse, both on the same 64 seeds and four rotations (256 games per half). The seating-corrected effect is the per-seed half-difference of the two slot-0-minus-slot-1 contrasts and the seating term their half-sum (analysis/prediction_tables.py, the study's diagnostics script). Two lineups: builders (eta in slot 2, fast in slot 3, the earlier study's lineup) and population (a second tables-baseline search in slot 2, eta in slot 3, so three searches answer trades and blocks). Stages: development on seeds 0 to 63, confirmation on seeds 800 to 863, which nothing had used, and population on seeds 800 to 863. The predicted threat ran its population stage first, because it only changes trades and the builders lineup trades less. All 18 halves completed 256 of 256 games.

Predicted space threat

Stage, seedsHalf A experiment, runHalf B experiment, runEffect (95% interval)Seating term
Population, 800 to 8632a3e7ddb-a627-4e0a-a5ea-0f1934e6d166, 98b12730-0997-4f13-932a-671536080240078dc121-0ee7-44f2-80c6-25b150fe5c14, 5d4485eb-5a7d-4525-8d5e-83b059633c39+0.047 (−0.004 to +0.098)−0.027
Development, 0 to 637a37ead7-5990-4ced-8ea6-cc139c14d3cb, 2fb51aa4-ec78-441a-ac88-4f3a1a1b077de8ed1b4d-9b55-41c0-b81a-72ca3446a6f0, 7c07bce7-df9f-4bfb-b95c-73f1cc1a0f54+0.014 (−0.037 to +0.064)+0.150
Confirmation, 800 to 863721e6d22-cda9-4bbd-920f-a4c97caa1293, 97c2c3c7-f23c-484f-8316-d8ef40f35728d9e60b56-b687-4452-8a44-8e7a342a1f97, e6694d7e-7ca2-4a3c-b63a-7f98ad345a59−0.006 (−0.057 to +0.045)+0.100

The single-seating contrasts inside each pair were +0.164, +0.137, +0.094, +0.105: mostly the seating term, which the pair design removes. The pooled builders-lineup estimate over 128 seeds, an after-the-fact pooling, is +0.004 (−0.032 to +0.040). The extension rule asks for a positive mean and an interval that comes within 0.03 of zero; the mean is positive but the nearest endpoint sits 0.032 away, so no extension runs, and at this effect size 64 more seeds cannot separate +0.004 from zero.

The switch does change the seat's trading: per game the candidate made about 1.5 more offers (16.3 against 14.8 on the population lineup), drew more counters (4.37 against 3.81), and declined fewer offers (11.7 against 13.1), at a decision time within 5 percent of its control in every half.

Predicted opponent model

Stage, seedsHalf A experiment, runHalf B experiment, runEffect (95% interval)Seating term
Development, 0 to 630f311f07-69b0-43a2-b00d-e65260e78c68, cd0eef08-680b-4639-9dba-8eb9e98ff79af7646e35-5a6e-4923-92eb-8ed00c0dbc55, 4dc942bd-5343-444d-aef8-9e89b60cc004−0.010 (−0.066 to +0.046)+0.096
Confirmation, 800 to 863131bd98f-9842-4da4-9ce6-4f818f018982, 596e8dd0-957e-49e8-bd30-47df5d0a6312906aeeee-fb36-4c4c-8e4c-bfb515c8e3b9, 26be7683-0b30-4cdb-a27a-1900a11adcc4+0.000 (−0.063 to +0.063)+0.055
Population, 800 to 863832402a0-649a-4afe-a6b5-0120eafe42f7, 3e1db769-42a7-48e3-842e-311d45776d327bed7450-6301-44bb-81cb-bafb69a54399, 15282f9d-520d-4290-8931-bf0739e88983−0.004 (−0.060 to +0.052)−0.012

Pooled over the 128 builders-lineup seeds: −0.005 (−0.047 to +0.037), so no extension runs. The candidate's own behaviour matches its control within noise on every counter (offers 14.7 against 14.9, trades 3.51 against 3.53, development bought 3.99 against 4.03, robber shares equal), which is what a switch that only changes the model of other seats should look like, and its decision time is within 5 percent of the control's.

Denial leaf term on the quarter blend

Stage, seedsHalf A experiment, runHalf B experiment, runEffect (95% interval)Seating term
Development, 0 to 6339f22891-3b02-4f8b-9d2a-54a68ecf8d65, 3bbbfea8-c08e-4910-9926-566ba647d534ad623e9c-d09c-4328-8f90-868cc0cc719c, b0c63544-f1ba-4a4c-924d-03ef381c4dbb+0.002 (−0.040 to +0.044)+0.123
Confirmation, 800 to 8633b404013-a489-4f0f-8c72-d5139030896a, 106ab3c6-89e2-47f5-a17b-50f4e91ecbc8f454b374-037c-4ea9-9dcc-34b9222d3fe6, 591c905a-0d87-4216-8b35-df28b3256109−0.037 (−0.071 to −0.003)+0.057
Population, 800 to 8631fda980d-5933-4182-8629-bb38219163ef, 5463cc1a-fd0f-467b-a33a-63992e8ebe0829355633-9630-484b-893c-5fa59b8bd158, ec5e3a3e-e989-40a6-bc77-701a2bbfc519+0.023 (−0.009 to +0.056)−0.004

The confirmation interval excludes zero below: on fresh seeds the term costs wins, repeating the hand-leaf study's fall from +0.061 on the development boards to −0.025 on fresh boards. Pooled over the 128 builders-lineup seeds: −0.018 (−0.045 to +0.010), so no extension runs. The candidate did lose slightly fewer frontier sites per game than its control on the builders lineup (1.78 against 1.84, the sites_lost counter), so the mechanism moves the board in the direction it claims without paying for itself.

The cost is decisive on its own. The denial seat's mean decision time was about double its control's in every half (1257 against 600 ms, 922 against 454, 199 against 105, and the mirrors), because the predictor runs at every leaf the term evaluates.

Both leaves side by side

SwitchHand leaf, seeds 0 to 63Hand leaf, fresh seedsTables, seeds 0 to 63Tables, seeds 800 to 863Tables, population lineup
Predicted threat+0.039 (−0.006 to +0.084)+0.031 (−0.024 to +0.087)+0.014−0.006+0.047
Predicted opponents+0.043 (−0.007 to +0.093)−0.039 (−0.100 to +0.022)−0.010+0.000−0.004
Denial 0.5+0.061 (+0.009 to +0.112)−0.025 (−0.071 to +0.020)+0.002−0.037+0.023

The hand-leaf numbers are the earlier study's mirrored-pair estimates; the denial row there was the pure hand leaf and here is the quarter blend, the strongest blend the patterns/leaf-blend screen measured. The pattern is the same under both leaves: development seeds flatter every arm, fresh seeds erase it, and only the population lineup gives the threat switch anything positive-leaning to hold.

Does the predictor read the tables?

From one observer seat's public knowledge on the population lineup (unregistered diagnostic, seeds 800 to 815, 64 games, --predict-diag 0):

OpponentPredictions matching the next completed build
eta builder256 of 434 (59.0%)
tables-baseline search348 of 1739 (20.0%)

Under the hand leaf the predictor scored 52.7 percent against eta and 28.5 percent against the hand-leaf search. The learned leaf is a less predictable opponent, not a more predictable one: it buys more development cards and settles its trades by table-specific patterns a greedy line does not follow. Builder-directed terms have even less to read at a table of searches than before, and the search is the seat that wins.

What could still be wrong

The population stage seats two of the four seats at the tables baseline, so half the table is the control and gains are diluted; a lineup of distinct searches could read differently. All effects are measured at depth 2 with a one-switch arm, and the switches could interact: the threat switch needs a prediction worth keying on, and the denial term needs the same prediction, so testing them together is a different hypothesis. The confirmation refutation of the denial term rests on one 128-game pair of halves; its interval barely excludes zero (−0.003), and the population reading (+0.023, −0.009 to +0.056) points the other way, so the honest statement is that the term is somewhere between −0.07 and +0.06 rather than decisively negative. Wall-clock decision times were measured on a machine shared with nine other agents; the within-run comparisons cancel the load, but absolute numbers would differ elsewhere. Nothing here tests a protocol cohort, a human, or the browser build's depth-3 one-second budget directly.

What this means for the browser

The browser player evaluates with the hand-written leaf at depth 3 under a one-second budget, so the hand-leaf measurements carry the direct evidence and the tables measurements say whether the mechanism generalizes beyond that leaf.

weights.denial is not worth enabling. It lost fresh seeds under the hand leaf (−0.025), lost fresh seeds under the quarter blend (−0.037), and roughly doubles decision time, which a one-second budget cannot absorb. The other two cost nothing, but neither earns its switch: under the hand leaf the corrected fresh-seed effects were −0.039 for the predicted opponents and +0.031 for the predicted threat, both crossing zero, and under the tables every interval crosses zero with the worth-the-browser rule unmet (it asks for at least +0.03 pooled with the interval excluding zero). The predicted threat is the only arm worth a follow-up, and only on a population of trading opponents: its two population-leaning readings, +0.031 under the hand leaf and +0.047 under the tables, both cross zero, so a useful cohort would need several hundred seeds per half to resolve an effect that small.

Agent notes

Eighteen registered deterministic cohorts under adaptation/prediction-tables, each 64 seeds, four rotations, 256 games, depth 2, six threads per tournament, played on the engine arena of the server worktree research/prediction-tables with the frozen tables artifacts/ntuple/hex-portfolio-main.bin. Every half completed 256 of 256 games; none was censored or invalid. analysis/prediction_tables.py computes each pair's seating-corrected effect, seating term, and pooled candidate and control diagnostics from the retained games.jsonl; reproduce any cohort with just engine-run EXPERIMENT. Run 83989250-85e2-4fe5-a9fd-c2a1c5aedef7 (experiment 2a3e7ddb-a627-4e0a-a5ea-0f1934e6d166) was interrupted after 6 games when the agent's session tooling killed the detached process; it is filed as interrupted and half A was rerun from scratch as 98b12730-0997-4f13-932a-671536080240. Six registrations for the denial arm on the full blend (hand 1.0) were withdrawn before running when the patterns/leaf-blend screen measured that blend far below the tables alone: 2324c88f-9407-455f-a76c-daafd52bf7e1, adb816b7-c2f2-4a47-a545-60b035870e45, d160d8e0-2d2a-45cb-ba12-f5fa82956d3a, 6f0f27f8-286c-424a-9786-7b43875e62cb, c62b41d6-5c90-447a-8196-6acb9b6200f1, and ff104571-dcc2-48f0-9dd5-0bae29a7cff5; the quarter-blend replacements are the three denial pairs above. The prediction-accuracy table comes from an unregistered diagnostic kept under runs/screens-prediction-tables/, 64 games with --predict-diag 0. Pooled 128-seed estimates are after-the-fact poolings named as such; no extension cohort ran, because no arm met its condition. Attribution: GLM-5.3 (baseten/zai-org/GLM-5.3) through opencode.