Research
AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
Overview Research area: Multi-agent LLM systems, multi-turn communication, and cooperative problem solving under information asymmetry. The paper is a workshop contribution ("Multi-Turn Interactions i

- arXiv
- 2512.03466
- Published
- 2025-12-03
- Authors
- Xavier Cadet, Edward Koh, Peter Chin
AI summary
Overview
Research area: Multi-agent LLM systems, multi-turn communication, and cooperative problem solving under information asymmetry. The paper is a workshop contribution ("Multi-Turn Interactions in Large Language Models"), arXiv:2512.03466v1 [cs.MA], dated 03 Dec 2025, by Xavier Cadet, Edward Koh, and Peter Chin of Dartmouth College, Hanover, NH 03755.
Technical level: Intermediate. The environment itself is simple to describe, but the paper assumes familiarity with LLM agent prompting, evaluation over seeds, and feedback-loop design.
One-sentence scope: The paper introduces AsymPuzl, a minimal two-agent puzzle environment where Alice and Bob each hold complementary, incomplete views and must exchange messages to solve the puzzle together, and uses it to measure how feedback granularity changes LLM cooperation.
What This Paper Is About
Existing multi-agent LLM research emphasizes open-ended role-play and dialogue rather than controlled evaluation, and rarely allows systematic manipulation of task difficulty. The authors ask how LLM agents adapt their communication when neither agent has enough information to solve a task alone, and whether the feedback the agents receive helps or hurts that coordination. To answer this, they build a deliberately minimal puzzle so that any failures can be attributed to communication and coordination rather than to reasoning about a complicated scenario.
Key Contributions
- The AsymPuzl environment: a two-agent asymmetric puzzle-solving testbed in which Alice sees correct positions and shapes but unknown colors, while Bob sees correct shapes and colors but unknown positions, and both must communicate to converge on the ground truth.
- A systematic feedback-granularity study: six feedback modes (No feedback, Own, Own detailed, Joint, Both, Both detailed) applied to 5-piece puzzles across 30 seeds at temperature 0.0.
- An empirical analysis of communication strategies: the paper shows that strong models (GPT-5, Claude-4.0) converge by sharing complete information in two turns, while weaker models ignore partner messages or over-correct their hypotheses.
- A difficulty-scaling evaluation: the same setup is run on puzzles of size 3, 5, 10, and 20 with adjusted turn budgets, plus a single-agent full-information control condition to verify the task is solvable when one agent holds everything.
Main Findings
-
Strong models solve reliably at size 5. GPT-5 and Claude-4.0 achieved 100.0% completion on 5-piece puzzles across all six feedback modes (30 seeds, temperature 0.0). The paper states they converge by sharing complete information in two turns.
-
Feedback generally helps, but detailed cross-agent feedback can hurt. Providing individual feedback about an agent's own working hypothesis raised completion. For GPT-4o, "Own detailed" improved completion from 43.3% to 63.3%. However, telling an agent which positions of the other agent's hypothesis are wrong (Both detailed) dropped Claude-3.5 from 83.3% under "Both" to 36.7%. The authors attribute this to information overload without context, since Alice does not see Bob's working hypothesis.
-
Different models favor different feedback modes. OSS-120B ranged from 53.3% with no feedback to 100.0% with Own detailed. GPT-4o peaked at 80.0% under "Both." Claude-3.5 peaked at 86.7% under Own detailed.
-
Two models fail entirely. GPT-3.5-turbo and Llama 3.2-11B scored 0.0% in every one of the six feedback conditions on 5-piece puzzles over 30 seeds, with Wilson 95% confidence intervals of 0.0–11.4 for those zeros.
-
Performance degrades with puzzle size under "Both" feedback. With sizes 3, 5, 10, and 20 (max turns 6, 20, and 40 for sizes 3, 10, and 20 respectively), GPT-4o fell from 60.0% (size 3) to 80.0% (size 5) to 60.0% (size 10) to 16.7% (size 20). Claude-3.5 followed a similar pattern: 50.0%, 83.3%, 56.7%, 16.7%. GPT-5, Claude-4.0, and OSS-120B held up better at the largest size, scoring 100.0, 100.0, and 93.3 respectively.
-
Smaller puzzles are not always easier for weak models. The paper notes the lower performance by some models on the 3-piece puzzle is explained by miscommunication leading to wasted turns under a tighter turn constraint.
-
Agents over-correct or under-correct. In an optimal solution Alice should modify each position once and Bob at most once, since a position may already be correct. GPT-5 and Claude-4.0 were nearly optimal, typically needing only two turns. GPT-3.5-turbo tended not to modify positions at all despite the puzzle being unsolved, while Llama 3.2-11B tended to modify positions more than 4 times on average.
-
Feedback also reduces the number of turns needed. Comparing No Feedback to Joint feedback, the paper reports that joint feedback increased the success rate and that GPT-5, Claude-4.0, and GPT-4o tended to solve puzzles in fewer turns.
-
Communication volume differs by role and vendor. Under feedback about each other's completion status, Bob used more tokens than Alice for most models, which the authors attribute to Bob needing to convey both shape and color while Alice can enumerate shapes. Claude models were the most verbose (Claude-4.0 averaged 225.6 output tokens as Agent A and 311.9 as Agent B). For GPT-5, close to 40% of generated output was dedicated to the message to the other agent (38.2% for Agent A, 37.4% for Agent B). Llama 3.2-11B dedicated the smallest share (9.1% for Agent A, 13.4% for Agent B).
-
The task is solvable by a single agent with full information. In a control experiment over 10 seeds with one attempt and no feedback, most models reached 100.0% completion. Claude-3.5 scored 100.0/86.7/93.3 and GPT-3.5-turbo 100.0/100.0/96.7 for sizes 5, 10, and 20, while Llama 3.2-11B scored 80.0/96.7/96.7.
-
Communication failures take distinct forms. The appendix documents successful collaboration in two turns (Claude-4.0), no cooperation where both agents repeat their own messages and ignore each other (Llama 3.2-11B), and miscommunication where one agent assumes what the other can see and the other asserts information it cannot guarantee (GPT-4o).
Methodology in Plain English
The authors generate a puzzle with N positions, each carrying a shape and a color. They split it into two complementary clues: Alice gets the correct positions and shapes with colors hidden, and Bob gets the correct shapes and colors with positions scrambled. Both start with a working hypothesis equal to their own clues.
Interaction is turn-based. On each turn, an agent receives the puzzle instruction, its original cues, its current working hypothesis, the message history, feedback from the previous turn, and the required output format. It then produces a message for the partner plus a list of actions to update its own hypothesis. The environment applies those actions, compares both hypotheses to the ground truth, and generates feedback. The agents must end their output with a valid JSON object containing a "message" field and an "actions" list; the partner sees only the message, never the actions.
Difficulty is controlled two ways: by puzzle size and by feedback mode. The turn limit is set to twice the number of elements, so a 5-piece game is cut off after 10 turns, and sizes 3, 10, and 20 get 6, 20, and 40 turns. Since full information sharing solves any size in two turns, this budget leaves margin for error and correction.
Models evaluated come from three vendors: OpenAI (GPT-3.5-turbo, GPT-4o, GPT-5, OSS-120B), Meta (Llama 3.2-11B), and Anthropic (Claude-3.5, Claude-4.0). All were run with a maximum of 4,096 output tokens, temperature 0.0, 30 seeds (the seed controls puzzle generation and initial clues), and a history length of 1, meaning agents see their own previous message and the latest message from the other agent. The implementation uses LangChain to query agents and vLLM to host open-source models. Metrics are success percentage within the turn limit and average number of actions per position, the latter serving as a proxy for whether changes are error corrections or random guessing. The search space for N positions with two attributes per position is stated as (N!)^2; each agent alone faces N! possible permutations, while with full information the puzzle is solvable in O(N) linear time.
Why This Matters
The paper argues that even in a simple cooperative task, LLM communication strategies diverge and depend on the granularity of the feedback signal. Because the environment is deliberately minimal, it reduces confounding factors and isolates communication rather than puzzle-solving ability, which the single-agent control confirms is not the bottleneck for most of the tested models. This gives the community a controlled testbed for studying coordination mechanisms, in contrast to role-play and open-ended dialogue settings that dominate current multi-agent LLM work.
Real-world applications the paper's framing points toward (the paper does not deploy or test any of these):
- Distributed decision-making, where separate parties hold complementary and incomplete information and must reconcile it through messages.
- Human-AI cooperation, where the human and the model each see part of the picture and coordination depends on what gets communicated.
- Multi-agent LLM pipelines in which different agents hold different context and must pass information explicitly rather than assuming shared state.
- Design of feedback and reporting interfaces for agent systems, since the results show that more detailed feedback is not automatically better.
Industry relevance: the failure modes documented here are operationally significant. Models that ignore partner messages (GPT-3.5-turbo, Llama 3.2-11B at 0.0% across every condition) will not function as cooperative components, and the finding that detailed cross-agent feedback cut Claude-3.5 from 83.3% to 36.7% shows that feedback pipelines need deliberate design rather than maximal information disclosure. The result that strong models solve the task in two turns by full information sharing also suggests a concrete protocol target: agents should default to broadcasting everything they know rather than trickling information out.
Future Directions
- Noisy and ambiguous views. The limitations section proposes giving, for example, Alice more shapes than required while Bob has the correct number of color-to-shape pairs, forcing them to determine which information is relevant; this ambiguity could also be applied in the opposite direction.
- Communication constraints. Restricting communication bandwidth, limiting the number of operations per turn, or otherwise constraining what agents may exchange.
- Persistent agent state. The current design re-injects all relevant state every turn (instructions, format, cues, working hypotheses, feedback) to isolate communication strategy. Future work would move toward persistent agent state and an internal scratch pad carried across turns.
- Scaling beyond two agents. Extending the testbed to three or more agents, and adding noisy views, as the conclusion suggests.
Target Audience
Researchers and practitioners working on multi-agent LLM systems, agent coordination, and multi-turn interaction design. It is most useful for those who need a controlled, reproducible benchmark for communication behavior rather than another open-ended role-play setting, and for engineers building multi-agent pipelines who want evidence about how feedback granularity and information-sharing protocols affect cooperative success. Readers looking for strong absolute performance numbers across all model tiers should note that two of the seven evaluated models scored 0.0% in every reported condition.
Authors’ abstract
Large Language Model (LLM) agents are increasingly studied in multi-turn, multi-agent scenarios, yet most existing setups emphasize open-ended role-play rather than controlled evaluation. We introduce AsymPuzl, a minimal but expressive two-agent puzzle environment designed to isolate communication under information asymmetry. Each agent observes complementary but incomplete views of a symbolic puzzle and must exchange messages to solve it cooperatively. Using a diverse set of current-generation and open-source LLMs, we show that (i) strong models such as GPT-5 and Claude-4.0 reliably converge across puzzle sizes on the solution by sharing complete information in two turns, (ii) weaker models often ignore partner messages or over-correct their hypotheses, and (iii) feedback design is non-trivial: simple self-feedback improves success rates, while detailed joint feedback can hurt performance. These findings show that even in simple cooperative tasks, LLM communication strategies diverge and depend on the granularity of feedback signals. AsymPuzl thus provides a testbed for probing the limits of multi-turn cooperation and opens avenues for studying coordination mechanisms.