Settlers / Research

73 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Beliefs about hidden hands

How much should a player infer from production, spending, and refused trades?

productionspendingcompatible handsevidence rules possibilities out

01Observe events

02Update possibilities

03Keep uncertainty

Public production and spending narrow the possible hidden hands.

The decision at the table

You can observe a rival receiving resources, spending cards, and refusing an offer. Only some of those events tightly constrain the hidden hand. Refusal could mean inability to pay, a bad price, a different plan, or distrust.

The mechanism to test

Maintain several hands compatible with the recipient-visible history. Update them with known gains and expenditures; keep distributions over unobserved thefts or discards. Model messages as claims with uncertain reliability. Keep beliefs about holdings separate from beliefs about willingness to trade.

Proposed experiment

Compare card-count-only sampling, event-constrained sampling, and event constraints plus an explicit model of trade responses. Hypothesis: response-aware beliefs improve held-out acceptance prediction without harming calibration. Test policy strength only after the belief diagnostic passes.

Use reliability curves and Brier scores for observable predictions such as whether an affordable offer will be accepted. If a separate offline diagnostic uses sealed hidden-state labels, keep them out of policy inputs and publish only approved aggregates. Split data by whole game.

First ablation

Sampling hidden hands from public counts alone, ignoring the event history, changed the contrast by -0.090 wins per game (95% interval -0.201 to +0.021); one sampled world instead of four changed it by -0.051 wins per game (95% interval -0.166 to +0.064). Both lean harmful and both are registered again on fresh seeds. On 192 fresh seeds, counts-only sampling cost -0.066 (95% interval -0.125 to -0.008) and one sampled world -0.012 (95% interval -0.072 to +0.048): event knowledge pays, the world count does not. Details and every run ID are in the ablation report; this is a development-tier result on one lineup.

What could disprove it

A sharply confident wrong belief is worse than a broad one. Missing history must widen uncertainty. Do not condition a deployable policy on spectator hands, future random outcomes, or diagnostic labels.

Agent notes

Use the shared experiment design to freeze candidate versions, full lineups, budgets, sample size, primary contrast, and stopping rules before collecting evidence. This is a draft study brief, not a preregistration. No run IDs exist for this proposal.

The deployable policy reads only its own observation and recipient-visible events. The diagram is a conceptual schematic. Build new measured exhibits from retained artifacts using the visual publishing guide.

Cowling, Powley and Whitehouse (2012) motivate searching over information sets. Applying that idea to these Catan beliefs is a proposal.

All approaches · Player’s guide · Experiment program

Tracked investigations

UpdatedInvestigationStatusFinding and next step
2026-09-09Hidden hands sampled from event history versus public counts onlyCompleteScreen -0.090 then confirmation -0.066 (95% interval -0.125 to -0.008) on 192 fresh seeds: sampling hands from event history instead of counts alone is worth about 0.07 wins per game. Supported. Next: Measure calibration of the sampled hands against sealed true hands offline, then test richer inference (refusals, discards). Log 2026-09-09
2026-09-09One sampled world versus fourCompleteScreen -0.051 then confirmation -0.012 (95% interval -0.072 to +0.048): one sampled world instead of four makes no measurable difference at depth 2. Next: Retest at depth 3, where worlds also shape the opponents' rounds. Log 2026-09-09