Population and self-play
Does an approach remain strong when the table stops resembling its training partners?
01Train diverse players
02Freeze the field
03Test unfamiliar tables
The decision at the table
A player can beat one style and lose to another. If those relationships form a cycle, one average win rate hides the structure. Four-player outcomes also depend on combinations of opponents, not just isolated pairs.
The mechanism to test
Train against diverse frozen opponents, keep a pool of earlier policies, and evaluate on withheld mixtures. Use a cross-play matrix to detect specialization, then whole-table tests to expose coalition effects. A robust player balances average performance with vulnerability to its worst matchups.
Proposed experiment
Compare training against a fixed baseline, self-play against only the latest policy, and a frozen population mixture under the same training budget. Hypothesis: the population-trained policy transfers better to withheld tables. Freeze the checkpoint and evaluation mix before the final cohort.
Show a matchup heatmap with game counts and uncertainty linked from every measured cell. Report mean and worst-family outcomes, completion, compute budget, and training/evaluation separation. In four-player studies, a cell identifies a full lineup, not an unspecified opponent.
First cross-play
Four depth-2 searches that differ only in inclination presets (builder, raider, gambler, trader) split 256 deterministic games almost evenly: 66, 59, 69, and 62 wins. No preset dominates; the sliders change style without changing strength against equals. See the ablation report.
What could disprove it
A finite test pool cannot establish unexploitable play or a Nash equilibrium. Repeatedly tuning against a published matrix turns the test set into training data. Leave a fresh final cohort and record negative transfer.
Agent notes
Use the shared experiment design to freeze candidate versions, full lineups, budgets, sample size, primary contrast, and stopping rules before collecting evidence. This is a draft study brief, not a preregistration. No run IDs exist for this proposal.
The deployable policy reads only its own observation and recipient-visible events. The diagram is a conceptual schematic. Build new measured exhibits from retained artifacts using the visual publishing guide.
CICERO is an example of combining language and strategic planning in Diplomacy. It is an architectural reference, not evidence that these methods transfer to Catan.
All approaches · Player’s guide · Experiment program
Tracked investigations
| Updated | Investigation | Status | Finding and next step |
|---|---|---|---|
| 2026-09-09 | Cross-play of four search styles | Active | Four inclination presets split 256 deterministic games almost evenly (66, 59, 69, 62 wins); no preset dominates. Next: Cross-play at depth 3 and against the frozen reference, then a heatmap of pairwise contrasts. Log 2026-09-09 |