Skip to content
AI.info

Research

Evaluating In Silico Creativity: An Expert Review of AI Chess Compositions

Overview Research area: Generative AI, computational creativity, and human evaluation of AI-generated artifacts, applied to the domain of chess puzzles (tactical problems and endgame compositions). Te

Evaluating In Silico Creativity: An Expert Review of AI Chess Compositions
arXiv
2510.23772
Published
2025-10-27
Authors
Vivek Veeriah, Federico Barbero, Marcus Chiam, Xidong Feng, Michael Dennis, Ryan Pachauri, Thomas Tumiel, Johan Obando-Ceron, Jiaxin Shi, Shaobo Hou, Satinder Singh, Nenad Tomašev, Tom Zahavy

AI summary

Overview

Research area: Generative AI, computational creativity, and human evaluation of AI-generated artifacts, applied to the domain of chess puzzles (tactical problems and endgame compositions).

Technical level: Intermediate. The AI methods are described in plain terms in this paper, but the puzzle analysis relies on chess notation (FEN strings, move sequences, tactical terminology) that assumes some chess literacy. The referenced "technical paper" holds the full methodological detail.

Scope: A human-expert evaluation study in which three titled chess experts review a curated booklet of AI-generated chess puzzles, selecting favourites and explaining what makes them creative, beautiful, or counter-intuitive.

What This Paper Is About

Generative AI can produce enormous volumes of output, but volume is not the same as creativity. This paper asks whether an AI system trained to generate chess puzzles can produce positions that expert humans judge as creative, novel, aesthetic, and counter-intuitive — rather than merely plausible or previously seen.

To answer that, the authors compiled a booklet of selected AI-generated puzzles, sent it to three internationally recognised chess experts and authors on chess aesthetics, and reported their selections and commentary alongside the puzzles themselves.

Key Contributions

  1. An expert human evaluation of AI-generated chess puzzles. Three world-renowned experts — International Master for chess compositions Amatzia Avni, Grandmaster Jonathan Levitt, and Grandmaster Matthew Sadler — reviewed a booklet of AI-generated puzzles, chose their favourites, and explained what made them appealing in terms of creativity, challenge, and aesthetic design.

  2. A generative pipeline for chess puzzles combining supervised learning and reinforcement learning. Neural networks (an Auto-Regressive Transformer, Discrete Diffusion, and MaskGit) were trained on a dataset of 4 million Lichess puzzles to model the distribution of puzzle positions, then further trained with RL using a custom reward.

  3. A reward-and-filter selection process built on two chess-specific criteria. The reward function combined a uniqueness check (ensuring only one winning move, similar to Lichess) with a counter-intuitiveness check (the position should be solvable by a strong chess engine but not a weak one).

  4. A themed puzzle booklet and a framework for studying creativity in chess. The appendix organises selected puzzles by aesthetic theme — including Sacrifice, Underpromotion, Attacking Withdrawal, Knight on the Rim is Dim, Sacrifice Pieces to Stalemate, Novotny, Interference, Unprotected Position, XRay Attack, Paralysis, Bristol, King on Tour, and Switchback — several with references to established composition literature.

Main Findings

  • Experts reacted positively overall. Feedback noted an innovative fusion of aesthetic themes and the "over-the-board" vision of the positions. Levitt described the positions as "a pioneering step in this human-AI partnership" that demonstrate potential to reach prize-winning level, though not yet at that level.

  • Criticism was specific and substantive. Reviewers said some positions were trivial, the collection overall lacked the profundity and complexity of traditional endgame studies, and certain puzzles were unrealistic.

  • Expert recommendations for improvement. Increase the complexity and depth of positions, incorporate problems with more complex sidelines and robust counter-play, and produce more surprising theme combinations.

  • One puzzle earned unanimous praise. The position in Figure 1 (FEN 1r1r2k1/Q2p1R1p/2p2R2/1p3pB1/1P4q1/8/5K2/8 w) was agreed by all three experts to be beautiful. The winning move, 1.Rg6+!, sacrifices both rooks at once to open the a1–h8 diagonal and prepare a slow repositioning of the misplaced queen on a7. All three experts described the move as "unorthodox" and by "no means natural or obvious sacrifice."

  • Amatzia Avni's selections (Puzzles 2, 3, 4). Puzzle 2 involves the retreat 1...Bc3 requiring long calculation, with the black king advancing into danger. Puzzle 3 features 1.Re7!, a quiet follow-up 4.Qh5!!, and an alternative line failing at 7...Qg7. Puzzle 4 features the under-promotion 2.d8=N!, where the tempting 2.d8=Q+ is met by the paradoxical 2...Bf8!!.

  • Jonathan Levitt's selections (Puzzles 5, 6). Puzzle 5 was described as a position that "could easily be from a real game," very close to endgame study material, with the counter-intuitive start 1.cxd4 Nxg2. Puzzle 6 was called "an elegant endgame position" close to endgame study standard, centred on 2...Bf5! and the slow 3...Kd8!!.

  • Matthew Sadler's selections (Puzzles 7, 8, 9). Puzzle 7 combines under-promotion (2.exd8=N) with a smothered mate motif that Sadler said he had never seen combined with an under-promotion. Puzzle 8 is built on 1...Rxh2 2.Kxh2 Qh6+ 3.Kg2 Qh4, where "obviously winning" lines such as 4.Rf3 actually lead to a very unexpected stalemate. Puzzle 9 adds a diversion to a typical smothered mate motif.

  • Creativity in chess is highly subjective. The experts rarely agreed on which puzzles were most compelling; the authors state that even very strong chess experts frequently differ in their assessments of a puzzle's creative merit, influenced by skill level, prior exposure to similar patterns, and individual aesthetic preferences.

  • The selection pipeline combined ranking and theme detection. Positions were first ranked by the reward function, then processed by aesthetic theme detectors. The detectors were imprecise alone but their effectiveness was greatly enhanced by the initial reward-based ranking. The authors manually reviewed the top 50 samples for each theme, a process validated with FIDE players in the 2200–2300 rating range.

Methodology in Plain English

Training the generators. The team trained three types of generative neural networks — an Auto-Regressive Transformer, Discrete Diffusion, and MaskGit — on a dataset of 4 million chess puzzles from Lichess, so the models would learn the distribution of those puzzle positions. Each position was written as a Forsyth-Edwards Notation (FEN) string, and the network learned to predict the next character in the string given the characters before it. At generation time, the network samples a puzzle character by character, starting from the first character of the FEN.

Improving generators with reinforcement learning. The generative network was then trained further with RL under a custom reward. The reward had two parts: a uniqueness check, similar to the one Lichess uses, to confirm there is only one winning move; and a counter-intuitiveness check, to confirm the position can be solved by a strong chess engine but not a weak one. The best samples under this reward were used to iteratively retrain the network toward higher-reward puzzles.

Selecting and packaging puzzles. Roughly 4 million chess positions were generated from the models and filtered with a hybrid approach: first ranked by the reward function, then passed through aesthetic theme detectors. The top 50 samples per theme were reviewed manually, with validation against FIDE players rated 2200–2300.

Expert review. Selected puzzles were compiled into a booklet and sent to the three experts, who were asked to pick their favourites and explain their appeal in terms of creativity, level of challenge, or aesthetic design. The paper presents those puzzles with the experts' analysis. In the appendix, the booklet is organised by themes drawn from the chess composition tradition, several of which are attributed to prior literature (Avni, 1991; Levitt and Friedgood, 1995; Persson, 2024).

Why This Matters

Impact on research. The paper offers a template for evaluating machine creativity through sustained, structured expert judgment rather than automated proxies alone. It documents that expert assessments of creative merit diverge substantially, which is directly relevant to how the field designs and reports creativity benchmarks. It also shows a route beyond imitating known patterns: generative models combined with a domain-specific reward can surface unusual combinations of established motifs — such as the under-promotion-plus-smothered-mate idea that Sadler said he had not seen before.

Real-world applications.

  • Chess composition and training tools: AI-generated puzzles that combine themes in unfamiliar ways could feed puzzle books, training sets, and composition-support software for players and composers.
  • Human-AI co-creation: the authors explicitly frame the work as a step toward "creative puzzle co-creation with human experts," with experts guiding selection and the model generating candidates.
  • Creativity evaluation methodology: the booklet-plus-expert-review format is portable to other domains where judging novelty and beauty requires trained human taste.
  • Generalisation beyond chess: the authors state they plan to extend the approach first to other board games and then to broader problem-solving domains.

Industry relevance. For teams building generative systems, the work illustrates a practical pipeline — distribution learning, reward-based filtering, detector-assisted selection, then expert curation — for domains where "looks plausible" is not the same as "is good." It also illustrates the value of partnering with domain professionals during evaluation rather than relying only on automatic metrics.

Future Directions

  1. Increase complexity and depth. The experts' primary recommendation was to raise the complexity of generated positions, adding more complex sidelines and robust counter-play, rather than producing positions that are trivial or where a complex solution yields a minimal advantage.

  2. More surprising theme combinations. Reviewers wanted to see additional novel fusions of aesthetic themes; the under-promotion combined with a smothered mate in Puzzle 7 is presented as evidence that the system can already find such fusions.

  3. Human-AI co-creation of puzzles. The authors propose extending the methodology to support creative puzzle co-creation with human experts, using expert judgement as part of the generation loop rather than only as a final evaluator.

  4. Generalisation beyond chess. The stated plan is to generalise these results first to other board games and then to broader problem-solving domains.

Target Audience

  • AI and machine learning researchers working on generative models, reinforcement learning from custom rewards, and computational creativity evaluation.
  • Chess composition and puzzle communities, including titled players and study composers interested in how AI-generated positions compare to human-composed studies.
  • Human-computer interaction and human-AI collaboration researchers studying expert-in-the-loop evaluation of generative systems.
  • Practitioners building creative generative tools in other domains who need a worked example of expert review as a validation stage.

Note: the version of the paper provided is truncated partway through Appendix A.13, so the full contents of the booklet beyond that point are not available in this content. Full methodological detail is stated to reside in a separate technical paper referenced by this one.

Authors’ abstract

The rapid advancement of Generative AI has raised significant questions regarding its ability to produce creative and novel outputs. Our recent work investigates this question within the domain of chess puzzles and presents an AI system designed to generate puzzles characterized by aesthetic appeal, novelty, counter-intuitive and unique solutions. We briefly discuss our method below and refer the reader to the technical paper for more details. To assess our system's creativity, we presented a curated booklet of AI-generated puzzles to three world-renowned experts: International Master for chess compositions Amatzia Avni, Grandmaster Jonathan Levitt, and Grandmaster Matthew Sadler. All three are noted authors on chess aesthetics and the evolving role of computers in the game. They were asked to select their favorites and explain what made them appealing, considering qualities such as their creativity, level of challenge, or aesthetic design.

Read the original paper