Settlers / Research

73 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Blending the learned leaf, and the endgame switches under it

The arms

The learned tables won their cohort as the only leaf (the learned leaf report); the search can also add the hand-written terms on top of them. LeafConfig exposes two knobs for that. leaf.hand scales the hand-written terms that are added to the tables' value (0 by default), and leaf.scale sets how many hand-written points one win is worth (1000 by default, so a win is ten points). The scale does not change the search's preference between moves; it changes how the tables' win probability compares with thresholds that are fixed in points: the trade margin and threat premium in bargaining, the deepen_gap stop rule, and the hand terms when they are blended.

ArmSeat changeWhat it tests
hand 0.25 / 0.5 / 1.0leaf.hand set on the tables seatDo the hand-written terms still carry information the tables missed, and at what weight do they stop helping
scale 500 / 2000leaf.scale halved or doubled on the tables seatDo the fixed thresholds that compare against the leaf value sit at the right win probability
raceendgame.race on top of the hand 0.25 blendThe race leaf from the endgame report, inert under the tables alone, measured on the blend that may become the default
hidden pointsendgame.hidden_points on top of the hand 0.25 blendExpected hidden victory points, same condition

Both endgame switches are hand-leaf terms, so they are inert under the tables alone; they were measured with leaf.hand 0.25 in both the candidate and the control seat. They were first registered on the hand 1.0 blend; when an unregistered screen showed the full blend much weaker than the tables alone, the unstarted hand 1.0 registrations were withdrawn and the arms were re-registered on the 0.25 blend, which is the configuration a switch must justify itself on. One full-blend pair (hidden points) had already run and is reported below as what it is.

The design

Every contrast is a swapped pair on the same seeds: half A seats the candidate in slot 0, half B seats it in slot 1, and the effect is the per-seed half-difference of the two slot-0-minus-slot-1 contrasts, which cancels the arena's seating term (the seating report); half their sum is the term. The builder lineup L1 puts eta and fast in slots 2 and 3, the population lineup L2 puts the tables baseline in slot 2 and eta in slot 3, so trades, blocks, and robber choices come from opponents that trade and block. Stages: development on seeds 0 to 63, confirmation on 800 to 863 (never used before), population on 800 to 863 in L2, and, when the pooled L1 effect stayed positive with its interval within 0.03 of zero, an extension on 864 to 927 in both lineups. Every cohort is 64 seeds, 4 rotations, 256 deterministic games, and every one completed all 256.

The stage 1 pair of the hand 0.25 arm reproduces the orchestrator's unregistered screen exactly (+0.074, +0.011 to +0.138): the same seats on the same deterministic boards are the same games, so the registered pair confirms the screen's arithmetic rather than independently replicating it. The fresh seeds of the later stages are the test.

The blend arms

Seating-corrected effect per stage, wins per game; pooled rows are named after-the-fact pooling.

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled L1Pooled L2
hand 0.25+0.074 (+0.011 to +0.138)+0.035 (−0.034 to +0.105)+0.049 (−0.008 to +0.105)L1 +0.088 (+0.006 to +0.170), L2 +0.080 (+0.015 to +0.146)+0.066 (+0.024 to +0.107)+0.064 (+0.021 to +0.108)
hand 0.5+0.045 (−0.028 to +0.118)+0.051 (−0.018 to +0.120)−0.027 (−0.094 to +0.040)L1 −0.047 (−0.113 to +0.019), L2 −0.010 (−0.072 to +0.053)+0.016 (−0.024 to +0.057)−0.019 (−0.064 to +0.027)
hand 1.0half A only, −0.117 (−0.233 to −0.002); pair withdrawn, screen −0.207 (−0.284 to −0.130)withdrawnwithdrawn

The quarter-strength blend is the only arm whose pooled intervals exclude zero, in both lineups. Its stage 2 number (+0.035) crosses zero, and the extension seeds came back at +0.088, so the pooled L1 estimate rests on 192 seeds and the pooled L2 on 128. The dose-response is monotone in the negative direction: the full blend loses to the tables even from the seat that wins the seating term, half strength is inside the noise, and a quarter is a real improvement.

The scale arms

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled
scale 500−0.059 (−0.121 to +0.004)−0.135 (−0.191 to −0.078), refutedwithdrawn after the refutation−0.097 (−0.139 to −0.054) over 128 seeds
scale 2000+0.051 (−0.009 to +0.111)−0.010 (−0.069 to +0.050)+0.035 (−0.024 to +0.095)L1 +0.047 (−0.026 to +0.120), L2 +0.033 (−0.034 to +0.101)+0.029 (−0.008 to +0.066) over 192 L1 seeds; +0.034 (−0.011 to +0.079) over 128 L2 seeds

Halving the scale makes every fixed threshold twice as large in win probability, and the arm shows it in play: the halved seat completed 2.3 player trades per game against the control's 3.7 and lost ground everywhere. Doubling the scale makes the thresholds relatively cheaper, the doubled seat traded more (5.1 against 4.1) and bought more development cards (4.07 against 3.88), and the effect on wins never separated from zero in six pairs. The default scale of 1000 is on the right side of both changes; the doubling is free in decision time but buys nothing.

The endgame arms

Both were measured on the hand 0.25 blend in both seats, with the blend alone as the control.

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled L1Under the hand-written leaf
race−0.039 (−0.076 to −0.002)−0.035 (−0.080 to +0.010)+0.000 (−0.036 to +0.036)not triggered, pooled L1 negative−0.037 (−0.066 to −0.008)+0.018 (−0.018 to +0.053)
hidden points+0.025 (−0.010 to +0.060)−0.010 (−0.043 to +0.024)−0.025 (−0.067 to +0.016)L1 −0.006 (−0.042 to +0.031), L2 +0.008 (−0.029 to +0.044)+0.003 (−0.017 to +0.023)+0.006 (−0.030 to +0.041)

The race leaf, which replaced the opponent terms with the own-minus-leader margin and the leader's expected points, is a small loss under the blend: its stage 1 interval excludes zero on the wrong side and the pooled L1 estimate does too, so the registered rule reads refutation. Under the hand-written leaf the same switch read +0.018 with an interval crossing zero (the endgame report), so the mechanism does nothing there and slightly hurts the blend. Hidden points is the same story at smaller size: nothing under the hand-written leaf, and nothing under the blend either: its extension was triggered by the letter of the rule (pooled L1 +0.008 after two stages) and the extension seeds brought the pooled 192-seed L1 estimate down to +0.003, with the population lineup slightly negative. The hidden-points pair on the full hand 1.0 blend, which ran before the withdrawal, read −0.029 (−0.075 to +0.016), also nothing.

What the blend changes at the table

Per seat per game over both halves of stage 1 (512 games each side):

ArmCitiesDevelopment cardsSettlementsRoadsPlayer tradesDiscardsMean pointsMean decision
tables (control of the blend arms)1.234.031.815.373.885.207.47667 ms
hand 0.251.503.952.055.454.295.927.95929 ms
hand 0.51.623.742.085.734.707.628.10880 ms
hand 1.01.863.342.126.195.689.028.11895 ms

The hand-written terms push toward cities and settlements, as expected: at a quarter strength the blend builds a quarter more cities and settles more without giving up development cards; at half strength it starts trading development cards for cities; at full strength it abandons the cards the tables were winning with (3.34 against 4.27 for its control) and carries two discards more per game, and the points it gains do not become wins. The quarter-strength blend keeps the tables' style and corrects its build choices.

Decision time is the blend's cost. Compared inside each cohort, where both seats shared the same machine load, the hand 0.25 seat spent 929 ms per decision against its control's 667 ms, about 40 percent more, and the hand 0.5 seat about the same. The two evaluators are both evaluated at every leaf. The scale arms cost nothing (the doubled-scale seat even ran marginally faster, 573 ms against 599 ms, inside noise). The endgame switches cost at most a few percent, and only because the race leaf's tempo estimate runs at the leaf.

Every cohort

All cohorts are 256 deterministic games; every one completed. "Contrast" is the registered slot-0-minus-slot-1 paired contrast of that half; the seating term is half the sum of a pair's two contrasts.

ArmStageHalfCandidate wins / control winsContrast (95% interval)ExperimentRun
hand 0.251 (0-63, L1)A138 / 87+0.199 (+0.107 to +0.291)337fa8bf-6dac-4798-ac5d-ef898b9511fc0a0e710b-9012-41fe-a947-393274a76c94
hand 0.251 (0-63, L1)B109 / 122+0.051 (-0.051 to +0.153) (control minus candidate)90b1cbe0-525c-440b-988e-f2ae7664e2806df6bf94-1853-46dc-b169-c8a2e5d7e7ba
hand 0.252 (800-863, L1)A120 / 102+0.070 (-0.032 to +0.173)643e51c4-d1e7-4545-8b2b-3f4cb64fe22a02c6ea23-1d12-45c2-9b4c-f3b4d3c0ef61
hand 0.252 (800-863, L1)B116 / 116+0.000 (-0.096 to +0.096) (control minus candidate)805db13b-33b8-4b74-aa92-69398e726fa96c737da9-c10a-46e5-9c53-b91c882e599a
hand 0.253 (800-863, L2)A92 / 74+0.070 (-0.012 to +0.153)f97d1156-4294-4ad2-8331-234e206d72a0d016da61-ed26-4072-b3ff-98c8afb95705
hand 0.253 (800-863, L2)B83 / 76-0.027 (-0.115 to +0.061) (control minus candidate)4924410e-b8e9-4752-9e0a-198bb826daaa8fcd82d8-09a4-4f85-968b-7f20120a83dc
hand 0.25ext L1 (864-927)A141 / 98+0.168 (+0.047 to +0.289)c0c8534f-9256-4576-a468-1959395cedfe532b7ed6-b77c-4ab8-ac9f-37537a9a90f7
hand 0.25ext L1 (864-927)B120 / 118-0.008 (-0.119 to +0.103) (control minus candidate)fb7327b6-8b65-44f0-bbb0-6ed513c1ab20d828c9b6-e8e6-45bc-8ae8-08a9f7b80b26
hand 0.25ext L2 (864-927)A107 / 75+0.125 (+0.031 to +0.219)9449ea85-5571-4a39-8ec3-f78a74c5c3c57697562f-eff1-4032-a32a-535ce8c323b5
hand 0.25ext L2 (864-927)B84 / 75-0.035 (-0.124 to +0.054) (control minus candidate)7287289b-adfa-4a85-8644-9570d8c6c10b9bae1e81-3ef3-41ad-aa6b-55876d2c2167
hand 0.51 (0-63, L1)A133 / 95+0.148 (+0.046 to +0.251)5f6a5f32-2629-4f89-8fc6-e3e8088e1e364a6bfee0-c5ca-4689-8c8b-0cac8df16a0d
hand 0.51 (0-63, L1)B104 / 119+0.059 (-0.045 to +0.163) (control minus candidate)9023a299-3644-4748-b349-4d6ff03b8ac539a0cdba-b3ab-4ed1-813d-537b3a60b103
hand 0.52 (800-863, L1)A128 / 103+0.098 (-0.010 to +0.205)f453131d-e920-4739-8f1c-fe0a2d57bdd781267946-c11f-4a2b-9c9a-bd31c7b102c0
hand 0.52 (800-863, L1)B116 / 115-0.004 (-0.115 to +0.107) (control minus candidate)9b7d17c5-f145-4467-a452-98fc6d1c913e281c1452-fd54-4b8f-85fb-f1192b6fa15f
hand 0.53 (800-863, L2)A87 / 83+0.016 (-0.079 to +0.110)ff2f050d-e848-4301-83ce-202d1abd15229bc75aea-a683-4d77-a99b-286a16083352
hand 0.53 (800-863, L2)B77 / 95+0.070 (-0.029 to +0.169) (control minus candidate)82f2db3f-6d33-49eb-b8d9-6ba914b23217807ae9e0-580f-4630-b4f6-d1c566eea0b7
hand 0.5ext L1 (864-927)A120 / 114+0.023 (-0.084 to +0.131)245f9d8a-5ebd-4cc3-9cb6-2f3a7e99b6ce86c15927-05fb-4b44-bdc2-9a631f070f0d
hand 0.5ext L1 (864-927)B105 / 135+0.117 (+0.002 to +0.233) (control minus candidate)a39dce2b-1918-4f22-b1c1-9584946b9e2f822ec329-678c-4dfd-9f2c-d347f5216438
hand 0.5ext L2 (864-927)A86 / 80+0.023 (-0.068 to +0.115)b4d35e5e-fbc9-4937-b236-b53f001d00f6b4e9e09c-4c38-4480-9cf6-23516107d05f
hand 0.5ext L2 (864-927)B79 / 90+0.043 (-0.051 to +0.137) (control minus candidate)905ad279-73cf-4145-a42f-b0ef6561192e32aa4be1-e50c-44e2-b7d2-8731f30257b4
hand 1.01 (0-63, L1), half A onlyA93 / 123-0.117 (-0.233 to -0.002)f2f21530-bf8d-400d-bfec-8614e6e0a6070d12835e-5262-4daf-9f2f-6a9b7288f2d6
scale 5001 (0-63, L1)A117 / 108+0.035 (-0.046 to +0.116)af67389d-785d-4bf0-894f-436836d8c2dba1830777-dcba-4e54-b394-833fc04adc0e
scale 5001 (0-63, L1)B91 / 130+0.152 (+0.052 to +0.253) (control minus candidate)fe58e64c-b9bb-493b-a4a3-411c5aff3c9f5d969f99-0067-4244-bb11-b88a3805bf63
scale 5002 (800-863, L1)A107 / 117-0.039 (-0.143 to +0.065)1a694630-8483-4512-849b-cfc7ca54bae1e0efc71e-f9d0-4d77-a57e-988bb4607654
scale 5002 (800-863, L1)B87 / 146+0.230 (+0.147 to +0.314) (control minus candidate)07a0544f-008e-4dae-a3c9-e6c00985e6f0b8662419-bf13-4333-b02f-bee9037ed7cb
scale 20001 (0-63, L1)A141 / 94+0.184 (+0.089 to +0.278)d5ed6196-0d94-4513-b8c4-6e8c83e39967d45f668c-558f-435e-9957-38e33e7ac864
scale 20001 (0-63, L1)B101 / 122+0.082 (-0.016 to +0.180) (control minus candidate)f10eb77e-ebb8-4711-90d8-a28b5e9ead0b75a28eed-d4f4-480a-95d7-8cec86ebdfec
scale 20002 (800-863, L1)A118 / 108+0.039 (-0.069 to +0.147)ad20b42d-9937-4515-acb1-7669b0f20aa8f9eb5200-ee12-4d9e-bcc3-fde9a73d24d4
scale 20002 (800-863, L1)B111 / 126+0.059 (-0.056 to +0.173) (control minus candidate)91718fce-02e4-4248-a9e1-3e43bd76c1181e9155d6-7a37-4f4a-92f2-d911b4cdaed0
scale 20003 (800-863, L2)A87 / 85+0.008 (-0.089 to +0.104)a1c94c34-3de8-4b37-a06a-792d93d66a5c7b35abeb-7384-4b24-b5d4-b7ad8da5bc13
scale 20003 (800-863, L2)B91 / 75-0.062 (-0.148 to +0.023) (control minus candidate)6dd18a29-bf0b-46e2-8ca4-bc815e07480dbbc95080-0a94-4b8c-aad9-a298f2531349
scale 2000ext L1 (864-927)A135 / 100+0.137 (+0.029 to +0.244)9d6eb9ca-7db0-4154-b64e-e15af957cac300a7596e-cbad-4576-8be7-0ea9849655d3
scale 2000ext L1 (864-927)B110 / 121+0.043 (-0.057 to +0.143) (control minus candidate)a2e97981-b6a7-4d3b-8a3b-ac416a6bfbac691cd581-79b3-46a8-89ba-4682c878cc3d
scale 2000ext L2 (864-927)A100 / 76+0.094 (+0.001 to +0.187)5e1b1d2c-8839-4bfb-8d65-0c3295df027abf588ca4-2392-4311-9783-9ecf88062a8b
scale 2000ext L2 (864-927)B83 / 90+0.027 (-0.072 to +0.126) (control minus candidate)4e2462ab-e524-439e-8242-c385d511bc0fb04b06d7-eb29-4be3-b759-bee7e517fb9d
race1 (0-63, L1)A129 / 100+0.113 (+0.024 to +0.203)2539e502-82fd-4435-8037-36576576149ea2305e39-c6ae-47cc-a48b-11a98a6a889f
race1 (0-63, L1)B92 / 141+0.191 (+0.099 to +0.284) (control minus candidate)afbb74be-1797-488e-bf6d-f507a2f7d21162a9c08a-525d-4647-8039-74762ece8915
race2 (800-863, L1)A123 / 104+0.074 (-0.037 to +0.186)2d7728ec-73da-4373-9bab-173c7d0fd23389995f9e-f3f9-476a-8d60-cea1c872882c
race2 (800-863, L1)B95 / 132+0.145 (+0.035 to +0.254) (control minus candidate)2ebd63d7-5030-44a7-aa22-e2c5b7eecf75a6160ac7-4b29-45fb-bede-7c7581ee49c3
race3 (800-863, L2)A84 / 83+0.004 (-0.086 to +0.094)583bb9b9-93eb-46f0-b881-3c04a8a09a43a84adf0d-5d3e-4508-bae6-4895bb988a5c
race3 (800-863, L2)B81 / 82+0.004 (-0.075 to +0.083) (control minus candidate)062e4f27-86af-4f43-a785-403963a72faacfa68fe3-5a0a-4b8c-a272-b18eabcf8d7d
hidden points1 (0-63, L1)A137 / 97+0.156 (+0.064 to +0.249)d7a3e2a6-12d7-49a3-bdea-73cb384d849ff71f0519-603f-4735-8b7c-216a027b4bfa
hidden points1 (0-63, L1)B102 / 129+0.105 (+0.012 to +0.199) (control minus candidate)0cee9b32-0a0e-45e9-9417-00816fed3055ffb5282f-aca3-4b7d-b6e2-0043cbb3c8a9
hidden points2 (800-863, L1)A125 / 104+0.082 (-0.025 to +0.189)c1fa5380-eb76-4b2e-b0d5-a5bbcc1191d621e1875d-8d52-4b46-a236-bc6b0f1c9269
hidden points2 (800-863, L1)B102 / 128+0.102 (-0.009 to +0.212) (control minus candidate)fc9a2726-8e79-4922-9454-7f0adf034d56ebca2e51-db2d-4d95-bdc0-0f90abb85e03
hidden points3 (800-863, L2)A83 / 84-0.004 (-0.092 to +0.084)707228f9-bf2c-480c-97bf-85e7915acbd6b7620d3e-e3c8-4f76-bd09-9cf142d9664f
hidden points3 (800-863, L2)B74 / 86+0.047 (-0.044 to +0.138) (control minus candidate)dcb44409-a705-4ab1-81b4-1d2a42b7711114d5ecc2-f876-4b89-bfc7-86afabbe2c5b
hidden pointsext L1 (864-927)A130 / 107+0.090 (-0.011 to +0.191)cb65e133-a18f-4d33-b4ea-afb2e2112cfe48063a4d-6412-4cbd-9241-bf8be44de2cd
hidden pointsext L1 (864-927)B103 / 129+0.102 (-0.003 to +0.206) (control minus candidate)4de83cf0-edca-45c8-8ee4-ae2025495157c4037bf3-c3f0-4d66-9c1d-b6eacb8eee14
hidden pointsext L2 (864-927)A91 / 82+0.035 (-0.046 to +0.117)e612fc8c-038c-4115-a7fc-ae0b625fa2e3eee597b6-f2e5-4d72-898d-014f7b7c5567
hidden pointsext L2 (864-927)B85 / 90+0.020 (-0.073 to +0.112) (control minus candidate)59099138-66bf-4389-b184-b5eb99e3bcec8d3434f3-4ff9-4f8e-b446-157887f3aa2b
hidden points1 on hand 1.0 (0-63, L1)A95 / 104-0.035 (-0.148 to +0.078)e7b09434-e4a4-409c-9bbc-6eeb56c3f65ab827db8d-00f4-43ba-95d3-6b726da9ce42
hidden points1 on hand 1.0 (0-63, L1)B101 / 107+0.023 (-0.101 to +0.148) (control minus candidate)f1d399f2-c7dd-4235-99d8-e2f997d6b00666738182-65d1-4db7-9163-81008c59d149

Withdrawn before running, with reasons in the daily log: the hand 1.0 half B (b6f5f82e-9a8a-4fd5-b9f5-bc7357fcaa9c) and the race halves on the full blend (47a866d0-b46b-49d4-bf92-fd0b520bc19f, 348e041a-9c73-43d8-93dc-80ce0ae5ec10), because the full blend measured much weaker than the tables alone and a switch on top of it says nothing about a configuration anyone would use; and the scale 500 stage 3 halves (7bd62467-9637-490f-a3cf-d9fe1a4a1bca, 436be4ef-b624-4d11-b3cc-6de6851facbc), because stage 2 refuted the arm.

What could still be wrong

Everything here is engine-arena evidence at depth 2 on deterministic boards: two fixed builders for L1, one extra search and one builder for L2. No protocol cohort has confirmed the quarter-strength blend, and nothing tests it at depth 3, against negotiating opponents, or with the tables the search would train next. The blend costs about 40 percent more decision time under machine load; the wall-clock numbers are distorted by a machine shared with other agents, and only the within-cohort ratios are meaningful. The 0.25 weight sits between two tested neighbours and was not itself tuned, so it is a coarse reading of where the optimum is; a cheaper subset of the hand terms might buy the same gain for less time. The extension rule was applied by its letter to three arms, and the pooled rows are after-the-fact poolings across stages that were registered separately. The stage 1 pair of hand 0.25 is the same games as the orchestrator's screen by construction.

What this means for the browser

The browser build runs the hand-written leaf in WASM at depth 3 under a one second budget and cannot load the tables, so leaf.hand and leaf.scale are not browser switches; they configure the server-side ntuple-leaf policy, and the quarter-strength blend is a candidate for that policy, worth a protocol cohort against ntuple-leaf before any default changes. The endgame race and hidden-points switches do run under the browser's hand-written leaf, and each now has two swapped-pair readings: under the hand-written leaf +0.018 and +0.006 (seeds 64 to 127), under the tables blend −0.037 pooled over 128 L1 seeds (negative) and +0.003 pooled over 192, all with intervals crossing zero except the race reading under the blend, which is negative. Neither switch reaches the +0.03 wins per game bar under either leaf, so neither is worth enabling in the browser; both stay off. The blend result does say the hand-written terms carry complementary information next to a learned evaluator, which is a reason to keep them sharp in the browser's own leaf, not a change to make today.

Agent notes

Study: patterns/leaf-blend under Learned pattern values. All arms are seat specs on the unchanged expectimax-v2 search: the baseline is v2:{"depth":2,"leaf":{"tables":"/home/keshav/settlers/research/artifacts/ntuple/hex-portfolio-main.bin"}}, the blend arms add "hand":0.25 (or 0.5, 1.0) or replace the scale with "scale":500 (or 2000) inside the same leaf object, and the endgame arms add "endgame":{"race":true} or "endgame":{"hidden_points":true} on top of the 0.25 blend in both search seats. The knobs are documented in the server's docs/expectimax.md under "Learned n-tuple leaf" and "Endgame"; no server code changed in this study.

Reproduction: python3 -m harness.engine run EXPERIMENT --threads 6 in the research checkout with SETTLERS_SERVER_DIR pointing at a server worktree at dev 38ae537. Effects and seating terms: analysis/seating_pair.py HALF_A_HALF_A HALF_B, per-arm diagnostics and pooled estimates: analysis/leaf_blend.py --pair LABEL A B. Lineup L1 seats eta and fast in slots 2 and 3; L2 seats the tables baseline in slot 2 and eta in slot 3. Every run directory retains its manifest, games.jsonl, summary, and inventory under runs/, and a compact record under records/runs/.

All studies · Learned pattern values · Log 2026-09-11