♔ Do the bots really play in period?
Each era bot vs. the historical record it was trained on — plus measured
playing strength and the era classifier's accuracy on held-out games. Receipts, not vibes.
Method. Each bot played 150 self-play games at its serving temperature. The same
analyzer computed identical, move-sequence-based metrics on the bot games and on random
samples of the historical era corpora (Lumbra's Gigabase, OTB games) — no ECO tags, no
cherry-picking. Resignation and draw agreement — the social layer policy networks
lack — are both adjudicated from the model's own win-probability head, with
era-specific manners tuned against the historical record: the Soviet school resigns
promptly and agrees draws readily; the Classical era is moderate in both; Romantics
play on toward the mate and almost never agree to a draw.
historical corpus
era bot (self-play)
Loading…
Honest residuals. One systematic gap remains.
First moves lean 1.e4 beyond the historical rate: the bots are fine-tuned from
Maia-2, which was trained on modern online chess — one epoch of era data moves its
opening prior a long way, but not all the way. The two larger gaps documented at
launch — draw rates 8–10 points below history, games ~25 plies too long —
are now modeled rather than missing: the bots resign and agree draws with
era-specific manners driven by their own win-probability head, tuned against the
record. Draw rates land within ~3 points and average game length within ~5 plies
of history in all five eras. The draw-culture curve is itself a finding: it rises
for 150 years (12% → 25% → 28.75% → 31.75%) and then plateaus in the
engine age — Sofia rules and fighting-chess culture flatten it exactly as games
stretch to their historical maximum length — and the bots track every step. Tuning those social constants revealed they interact:
prompt resignation eats would-be draws, so the pairs were tuned jointly. The
remaining first-move lean doesn't affect the in-period style signal above.