Analysis
Comparisons, failure patterns, evaluation results, and the games that best explain what each model learned.
Post-training research
A research project for agents that play complete games of Catan. This is the public home for game replays, analysis, and notes on the training pipeline.
The replay browser is live now. Analysis and a clearer training-pipeline walkthrough will live here as they are ready to share.
Filter complete games by player type, judge which are worth watching, and inspect every decision after the game.
Open replay browser →Comparisons, failure patterns, evaluation results, and the games that best explain what each model learned.
How trajectories become training data, how models are post-trained, and how full-game behavior is evaluated.
During play, agents receive only information a real player is allowed to know. Opponent resources and unplayed development cards remain hidden.
The replay viewer is intentionally omniscient because every game shown here is already over: its purpose is post-hoc analysis, not live assistance.