Settlers / Research

73 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Liquidity study audit trail

The primary result uses only the fresh 80-game cohort. Earlier outcomes were never pooled into its denominator. The strategic estimator stayed unchanged throughout protocol repairs. Completed experiment files and run inventories were retained.

RunPlannedAttemptedCompleteIncompleteUnplayedInterpretation
d63fca0341013Startup timeout; legacy runner omitted SetReady.
74e9d98a41013Confirmed sockets connect; the unchanged legacy startup still failed.
0bd1bc9341013Added native startup diagnostics; connected sockets still did not become ready.
7459585e41013Harness diagnostics identified readiness rejection being misreported as a connection timeout.
1201e0ae42112Readiness repaired. The second game hit the server communication cap; stopped.
fda25b9d44400Four-game smoke after the shared 16-offer budget; no strength inference.
83f0efb58032177First evaluation stopped at archive backpressure; 77 games deliberately left unplayed.
5a18c7f044400Excluded smoke: a build syntax error let an unguarded command sequence use the previous executable. Actual binary hash retained; not validation of the retry fix.
4486b5cc44400Guarded build and four-game smoke of the final retry-capable runner.

Changes made before the fresh evaluation

  • The shared native runner marks its own connected seat ready.
  • All policies get the same 16 offer attempts per game, including attempts lost to a race. Trade acceptance remains available.
  • Storage backpressure and transport ambiguity retry the identical command envelope with bounded backoff. Rule errors still stop the run.
  • The harness reports startup readiness accurately, hashes research-owned executables and Cargo.lock, excludes generated target trees, and strips research-spectator theft details and directed chat from new public tapes.
  • just smoke-liquidity depends on a successful format check, tests, Clippy and release build, preventing a failed rebuild from silently launching an old binary.

These adaptations define the tested baseline. No estimator weights were tuned against the failed evaluation or smoke outcomes.

What is retained

Each linked run record names its experiment, manifest digest and inventory digest. Verified local bundles live in artifacts/liquidity-study/; their immutable receipts are in records/artifacts/. The earlier diagnostic tapes predate the public-event filter and can include research-spectator theft details. They remain local diagnostic evidence and are not used for the public charts or replay. Credentials, participant observations and private process logs are not included in the bundles. There is no remote backup or public download URL.

Small connection diagnostics reused already failed lobbies without starting games. They produced no additional match outcomes. A runtime syntax failure is recorded above instead of being hidden as an unreported tuning attempt.

Return to the experiment.