Skip to content
AI.info

Research

Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning

Overview Research area: Evaluation of Large Language Model code generation, multi-agent systems, auction theory, and combinatorial optimization (logistics). Technical level: Intermediate. The paper is

Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
arXiv
2511.20613
Published
2025-11-25
Authors
Panayiotis Danassis, Naman Goel

AI summary

Overview

Research area: Evaluation of Large Language Model code generation, multi-agent systems, auction theory, and combinatorial optimization (logistics).

Technical level: Intermediate. The paper is readable without deep expertise, but its benchmark draws on concepts from auctions, constraint optimization, and multi-agent systems (marginal cost, opportunity cost, NP-hardness, admissible heuristics).

Scope: The paper introduces a reasoning-driven code-generation benchmark — the Auction, Pickup, and Delivery Problem (APDP) — and reports a tournament in which 40 LLM-coded agents compete against 17 human-coded agents across 12 double all-play-all tournaments (almost 40k matches).

What This Paper Is About

Existing coding benchmarks mostly measure whether generated code passes unit tests and is syntactically correct, which understates the difficulty of real-world problems that require planning, optimization, and strategic interaction. This paper builds a benchmark from a real logistics problem in which agents bid for parcel-delivery tasks in a reverse first-price sealed-bid auction and then route a fleet of vehicles to maximize profit. The goal is to test whether state-of-the-art LLMs, including "vibe coding" workflows, can produce code that competes with solutions written by graduate CS students before the advent of LLMs.

Key Contributions

  1. A new reasoning-driven benchmark. The Auction, Pickup, and Delivery Problem combines competitive multi-agent systems, auctions, and constraint optimization, and requires designing and coding advanced planning and optimization algorithms. The APDP benchmark will be open-sourced, with code available at https://panayotisd.github.io/apdp_bench/.
  2. A human-versus-LLM tournament. The authors evaluate a range of state-of-the-art LLMs (GPT-5 Thinking, Gemini 2.5 Pro, Claude Opus 4.1, DeepThink R1) against 17 human-coded agents developed before the advent of LLMs, 12 of which were written by students.
  3. Large-scale empirical results. 12 double all-play-all tournaments and almost 40k matches (38,304 in total) show the top 5 spots are consistently won by student-coded agents, the majority of LLM agents (33 out of 40) are beaten by very simple baseline agents, and giving an LLM the best human solution to improve makes the solution significantly worse.
  4. Diagnostics of LLM failure modes. The paper documents semantic bugs (timeout violations, tasks won but not delivered, capacity violations), along with cases where LLMs could not resolve the same bug across repeated prompting cycles and had to be restarted from scratch.

Main Findings

  • Human-coded agents dominate the top. In the aggregate table across the 12 tournaments, the top 5 spots are consistently won by student-coded agents, led by Student 1 with 108.167 average wins per tournament and a winrate of 0.9658.
  • The best LLM agents still rank below several students. The highest-ranked LLM-coded agent, LLM(O, IR, 1) with 95.417 average wins and a winrate of 0.8519, places sixth, behind five students but ahead of Students 6 through 12.
  • Most LLM agents lose to simple baselines. 33 out of 40 LLM-coded agents are beaten by very simple baseline agents such as the expected cost fixed bid (ExpCostFixedBid, winrate 0.689), the Honest agent (0.4516), ModelOpponent (0.7195), and RiskSeeking (0.7359).
  • The weakest LLM agents fall far below simple heuristics. The lowest-ranked LLM agent, LLM(D, CR, 2), records 5.667 average wins per tournament and a winrate of 0.0506, versus Naive at 36.75 average wins and 0.3281.
  • Improving the best human solution makes it worse. When the best-performing LLM was given the winning student solution as input and asked to improve it, the resulting agent performed significantly worse, dropping to 10th place.
  • Few syntactic errors, many semantic errors. The authors observed a minimal number of syntactic errors but a significant number of semantic bugs, including not respecting time-out limits despite being given template code that implements them, failing to pick up and/or deliver won tasks, and violating capacity constraints.
  • Bug fixing is laborious. For agents that consistently timed out, multiple (5–15) cycles of prompting the LLM with the error failed to resolve the same bug, and the only solution found was to re-start from scratch. Significant manual effort was required to achieve the 40 bug-free agents evaluated.
  • Simple variants were often solved, but with suboptimal design. In the Reactive variant, all LLMs solved the test case correctly on the first try with minimal syntax errors. In the Deliberative variant, 3 out of the 4 LLMs (all but GPT-5) initially failed to implement an admissible heuristic for A* despite explicit instruction, and one Claude agent switched from an admissible to an inadmissible heuristic while trying to fix a bug. DeepThink R1 was the only LLM to implement a Minimum Spanning Tree based heuristic and was one order of magnitude faster than both GPT's and Claude's agents.
  • In the Centralized variant, LLMs produced syntax-bug-free code but often made suboptimal design decisions, such as an agent using only one vehicle from its fleet.

Methodology in Plain English

The authors took a graduate course assignment (the Intelligent Agents course at EPFL, which uses the Logist platform, all code in Java) and turned it into a benchmark. In the APDP, several transportation companies compete in a market. Tasks (parcels, defined by source, destination, and weight) are auctioned one at a time via a reverse first-price sealed-bid auction, where the lowest bid wins and each agent has a fixed time to bid. After the auction, each agent has a fixed time to plan pickup and delivery routes for two vehicles characterized by different starting locations, capacities, and cost per kilometer, subject to capacity, delivery, pairing, and precedence constraints. Profit equals total revenue (sum of winning bids) minus total transportation cost (distance driven times cost per kilometer).

Because the problem is open-ended and NP-hard with no closed-form optimum, the authors score agents head-to-head rather than against a ground truth. Agents are evaluated in 12 double all-play-all tournaments — 4 network topologies (Switzerland, France, Great Britain, the Netherlands), 3 tournaments per topology. In each tournament, every agent meets every other agent in 1v1 matches, and each pairing is played twice with the companies swapped, yielding 3192 matches per tournament (57 × 56) and 38,304 matches overall. Each agent therefore plays 112 matches per tournament, and 50 tasks are auctioned per match.

On the human side, the authors used 17 pre-LLM agents: 12 student agents from the 2020 class (the top 8 from a single-elimination tournament plus 4 additional agents with the highest number of wins against the baselines) and 5 simple baseline agents built by members of the Artificial Intelligence Laboratory at EPFL (Naive, ExpCostFixedBid, Honest, ModelOpponent, RiskSeeking). On the LLM side, 4 models (GPT-5 Thinking, Gemini 2.5 Pro, Claude Opus 4.1, DeepThink R1) were each prompted twice with each of 5 prompting strategies — two author-written prompts (A1 and A2), Iterative Refinement (IR), LLM as a Critic (CR), and a GPT-5-generated optimized prompt (GEN) — producing 40 LLM-coded agents. The author prompts were written to contain the same information students received in the course, so the comparison reflects what a modern student using an LLM for the project, or a vibe-coding user, would have.

Why This Matters

Impact on research. The paper argues that prevailing benchmarks emphasizing unit-test pass rates and syntactic correctness understate the difficulty of real-world problems requiring planning, optimization, and strategic interaction. It provides a data-contamination-resistant comparison point (human solutions written before LLMs existed) and a benchmark for reasoning-driven code synthesis where pass/fail labeling is not feasible.

Real-world applications:

  • Logistics and delivery operations for companies whose core business is pickup and delivery routing (the paper names Amazon, DLH, and FedEx as companies sitting on this class of problem).
  • Ride-pooling and dial-a-ride services, which the paper cites as a domain where pickup and delivery problems arise.
  • Meal delivery routing.
  • Supply-chain management for manufacturing companies such as Huawei and Tesla.

Industry relevance. The findings temper claims that LLMs are ready to automate software engineering in open-ended domains; the paper notes that some prior work considered fully automating software engineering a step toward "AGI." The authors also highlight that developers in one cited study (Becker et al., 2025) believed they were 20% faster with AI tools but were actually 19% slower. The paper points to a July 2025 AtCoder World Tour Finals 2025 Heuristic event where an OpenAI system finished second against human coders, but notes that system was reportedly custom rather than a public model, and that there was no direct competition between opponents within the problem environment — a gap that APDP closes.

Future Directions

  • Building benchmarks, datasets, and open-source baselines that stress reasoning-driven code synthesis in real-world scenarios, as the paper explicitly aims to facilitate.
  • Closing the contamination and scope gaps the paper identifies in existing benchmarks, whose limitations include data contamination, limited scope that does not reflect real-world and open-ended tasks, and lack of adaptability and creative testing.
  • Testing whether prompting LLMs to include debug prints, as a human developer would, improves iterative refinement — the paper flags this as an interesting avenue because the feedback signal in its IR strategy was weak.
  • Investigating more broadly whether LLMs can be brought to graduate-level performance on tasks requiring strategic planning, opponent modeling, and constraint optimization, given the observed failure modes in admissible heuristic design and suboptimal design decisions.

Target Audience

Researchers and practitioners in LLM evaluation and code generation; multi-agent systems and auction researchers; logistics and operations researchers; and instructors designing graduate-level AI courses and assignments. It is also useful to engineers and technical decision-makers assessing how far LLM and vibe-coding workflows can be trusted on open-ended optimization and strategic-planning problems.

Authors’ abstract

The rapid proliferation of Large Language Models (LLMs) has revolutionized AI-assisted code generation. This rapid development of LLMs has outpaced our ability to properly benchmark them. Prevailing benchmarks emphasize unit-test pass rates and syntactic correctness. Such metrics understate the difficulty of many real-world problems that require planning, optimization, and strategic interaction. We introduce a multi-agent reasoning-driven benchmark based on a real-world logistics optimization problem (Auction, Pickup, and Delivery Problem) that couples competitive auctions with capacity-constrained routing. The benchmark requires building agents that can (i) bid strategically under uncertainty and (ii) optimize planners that deliver tasks while maximizing profit. We evaluate 40 LLM-coded agents (by a wide range of state-of-the-art LLMs under multiple prompting methodologies, including vibe coding) against 17 human-coded agents developed before the advent of LLMs. Our results over 12 double all-play-all tournaments and $\sim 40$k matches demonstrate (i) a clear superiority of human(graduate students)-coded agents: the top 5 spots are consistently won by human-coded agents, (ii) the majority of LLM-coded agents (33 out of 40) are beaten by very simple baselines, and (iii) given the best human solution as an input and prompted to improve upon, the best performing LLM makes the solution significantly worse instead of improving it. Our results highlight a gap in LLMs' ability to produce code that works competitively in the real-world, and motivate new evaluations that emphasize reasoning-driven code synthesis in real-world scenarios.

Read the original paper