Skip to content
AI.info

Research

End-to-end PDDL Planning with Hardcoded and Dynamic Agents

Overview Research area: Artificial Intelligence — automated planning, neuro-symbolic reasoning, and LLM-based multi-agent systems. Technical level: Advanced. The paper assumes familiarity with PDDL, c

End-to-end PDDL Planning with Hardcoded and Dynamic Agents
arXiv
2512.09629
Published
2025-12-10
Authors
Emanuele La Malfa, Ping Zhu, Samuele Marro, Sara Bernardini, Michael Wooldridge

AI summary

Overview

Research area: Artificial Intelligence — automated planning, neuro-symbolic reasoning, and LLM-based multi-agent systems.

Technical level: Advanced. The paper assumes familiarity with PDDL, classical planners, and formal optimisation notation, though the high-level architecture and empirical findings are accessible.

Scope: The paper presents and evaluates a fully automated, LLM-driven agentic pipeline that turns natural-language planning specifications into PDDL domains and problems, solves them with external symbolic planners, and translates the resulting plans back into natural language.

What This Paper Is About

Automated planning produces provably correct plans, but it requires expert-authored PDDL models and cannot cope with ambiguous or underspecified natural-language requirements. Large Language Models handle natural language well but are unreliable at long-horizon planning, often hallucinating steps or producing infeasible sequences.

This paper builds a framework that connects the two: an LLM orchestrator converts a human specification into PDDL, a pool of specialised agents iteratively repairs and refines that PDDL using solver and validator logs, an external planner solves it, and a final module translates the plan back into natural language — all without human intervention.

Key Contributions

  1. A fully automated end-to-end multi-agent planning pipeline. The system moves from free-form natural-language specification to validated plan with no human intervention, with an LLM orchestrator dynamically creating multi-agent workflows on the fly rather than relying on a fixed, pre-defined workflow.

  2. Two distinct agent categories. Hardcoded agents are informed by logs and error traces and have pre-defined goals (fixing PDDL syntax, checking temporal constraints, adapting to a target solver, generating the natural-language plan, terminating early). Dynamic agents have no predefined goal and instead adapt to the specific domain, revising the latent planning abstraction. The paper argues hardcoded agents mitigate longstanding obstacles such as ambiguity resolution, while dynamic agents increase workflow flexibility.

  3. Broad empirical evaluation. Experiments span more than ten domains and tasks, including the Google Natural Plan benchmark, Planbench, Tower of Hanoi, Blocksworld, and Sokoban, across GPT-{4o, 5-mini, 5.4} and Gemini-{2.5, 3}-flash. The framework was successfully tested with Fast Downward, LPG, POPF, VAL, and uVAL.

  4. Identification of constraint satisfaction as the weak point of PDDL+LLM methods. The paper frames CSP as the setting where neuro-symbolic techniques still struggle, and issues a call for better evaluation methods to mitigate the brittleness of LLMs used as judges when both ground truth and generated plan are expressed in natural language.

Main Findings

  • Large gains on long-horizon planning for GPT-5.4. On Planbench, the approach yields a perfect success rate (100%) on three out of four tasks and more than 90% on depots. Vanilla GPT-5.4 is far weaker on the same tasks, with reported vanilla values of 21.3 ± 1.8, 49.3 ± 2.4, 33.3 ± 2.4, and 7.3 ± 2.4 on the Planbench tasks shown in Table 1.

  • Frontier Gemini models are already near-saturated on Planbench. With Gemini-3-flash, vanilla and agentic approaches achieve comparable near-perfect rates, averaging 97% with a negligible gap of 0.1% between the two methods (standard deviation across task success rates of 1% for both methods).

  • The method helps smaller models most. Average Planbench success for GPT-5-mini is 91.6 ± 1.6 with agents versus 76.6 ± 23.2 for vanilla. For Gemini-2.5-flash it is 89.7 ± 6.8 with agents versus 74.7 ± 14.9 for vanilla.

  • Google Natural Plan is treated as constraint satisfaction, and behaves differently. The paper reports it is already saturated by base frontier models, matching vanilla GPT-5.4 and Gemini-3-flash. For GPT-5-mini the agents raise success rates to 93.3%, 53.3%, and 8.0% (average 51.5 ± 34.8) on calendar scheduling, meeting planning, and trip planning, versus 88%, 24%, and 2% (average 38.0 ± 36.4) for vanilla.

  • Beats Exploration Walk where that method scores zero. On floortile and childsnack, Exploration Walk achieves 0% with GPT-4o. The agentic framework with GPT-4o reaches 43.7% on floortile and matches 0% on childsnack. With GPT-5-mini, vanilla scores 26.0% and 0%, while the agents score 68.7% and 34.0%. With Gemini-2.5-flash, vanilla scores 10.5% and 5.2%, while the agents score 53.8% and 53.3%. Exploration Walk cannot be applied to GPT-5-mini or Gemini-2.5-flash (or newer) because it requires access to logits.

  • Tower of Hanoi and Blocksworld improvements. On the Tower of Hanoi, the framework achieves 91% success rate (+6% on average) versus 73% (+14%) for vanilla GPT-5-mini. On Blocksworld with problems requiring exactly 2−4, 6−8, and 10−12 actions ("Easy", "Medium", "Hard"), the framework matches GPT-5-mini on Easy and Medium but outperforms it on Hard (+7%).

  • Sokoban results are mixed and model-dependent. Focusing on hard scenarios with corridor length between 90 and 100, the published GPT-5-mini baseline of roughly 13% is raised to 34% by the framework. GPT-5 goes from 34% vanilla to 40% with the method. Vanilla Gemini-{2.5, 3}-flash score 78% and 76%, but drop to 20% and 68% respectively with the technique.

  • The dynamic agent aids efficiency rather than raw success. Removing it leaves success rates comparable to the full system but increases agent invocations per problem — on Gemini-2.5-flash with Google Natural Plan, the average rises from 3.2 to 4, nearly exhausting the budget; on Planbench it rises from 2.2 to 2.8. With all agents available, hardcoded and dynamic agents are invoked at roughly 50% each.

  • Benchmark quality problems were found. In Google Natural Plan, some calendar scheduling instances admit multiple valid solutions despite a single labeled answer, and trip-planning ground truth ambiguously counts travel days as time spent in both locations, while GPT-5-mini and GPT-5 count travel as time spent only in the origin. Planbench is described as providing well-defined tasks and solutions.

  • LLM-as-a-judge is unstable. Using GPT-5.4 as planner and GPT-5-mini as judge over three evaluations, the maximum variation in success rate for vanilla is 12% (meeting planning) with an average of 6.8% across the seven tasks; for the agentic method the maximum is 10% (depots) with an average of 3.7%.

Methodology in Plain English

An orchestrator agent reads the human specification and produces a structured JSON representation of the environment, agents, goals, and constraints. JSON is used as the intermediate representation because it flexibly captures heterogeneous data and is widely represented in LLM pre-training corpora.

The orchestrator then authors an initial PDDL domain and problem, which are handed to an external solver (for example Fast Downward or POPF2) to produce an initial candidate plan. A validation tool (VAL or uVAL) produces logs and error traces.

Those logs drive an iterative refinement loop, in which the orchestrator picks the most suitable agent from a pool. Hardcoded agents each have a fixed purpose: adapting PDDL to a target solver's syntax, repairing syntax errors flagged by the validator, checking temporal consistency, checking that natural-language constraints match the PDDL, detecting inconsistencies between constraints, goal and plan, converting the final plan into natural language, and terminating the computation early when the task is already solved. A further set of less-used agents handles asynchronicity, multi-agency enforcement, and hallucination detection. The dynamic agent, by contrast, has no fixed prompt goal: it receives a "task profile" generated on the fly by the orchestrator and rethinks the planning abstraction for that specific instance.

The paper formalises this as an optimisation problem: the framework maps a specification to a PDDL domain and problem, the solver maps those to a candidate plan, and an oracle scores the candidate as right or wrong. With a finite pool of hardcoded agents, agent selection reduces to a standard multi-armed bandit problem where each agent is an arm with unknown expected reward. Adding dynamic agents changes the problem: for each agent an inner optimisation searches its configuration space (such as prompts or generated programs), so the difficulty shifts from statistical estimation to structured search and no longer reduces to a bandit formulation.

Evaluation uses a budget of four for plan generation and refinement, GPT-5-mini as the judge comparing natural-language plans to natural-language ground truth, and results averaged over three independent runs of 50 problems each.

Why This Matters

The work argues that PDDL and LLMs are complementary: PDDL helps LLMs on planning tasks, while it reduces their effectiveness on constraint satisfaction tasks. It also supplies empirical support for position papers advocating more rigorous evaluation of LLM planning, and treats the brittleness of LLM-as-a-judge as an open problem for the field rather than a solved detail.

Real-world applications:

  • Logistics and supply-chain planning, where correctness and cost-optimality matter and expert-authored domain models are expensive to maintain.
  • Scheduling assistants — calendar scheduling, meeting planning, and trip planning, the three tasks in the Google Natural Plan benchmark.
  • Robotics and embodied agents, where natural-language goals must become valid action sequences before execution.
  • Digital assistants and multi-agent coordination, particularly settings requiring asynchronous or multi-actor plans, which the framework supports through its asynchronicity and multi-agency agents.

Industry relevance: the framework is planner-agnostic and was tested with Fast Downward, LPG, POPF, VAL, and uVAL, so it can be layered onto existing symbolic planning infrastructure rather than replacing it. It requires no human intervention at any stage, and it depends only on API-accessible LLMs — a deliberate contrast with techniques that need model logits, which are unavailable in recent commercial APIs.

Future Directions

  • Scaling orchestration beyond the current agent pool and budget of four refinement steps.
  • Incorporating multimodal inputs, extending the framework past text-only specifications.
  • Applying the framework to real-world and embodied settings, beyond the benchmark domains evaluated here.
  • Developing theoretical guarantees for dynamically generated workflows, since the current formulation replaces a tractable bandit problem with a high-dimensional structured search once dynamic agents are included.
  • Fixing benchmark and evaluation practice. The paper explicitly calls for benchmarks validated by humans before release, and for better ways to evaluate plan correctness when no ground truth plan is available — noting that asking a model to output a format that can be tested against ground truth is itself still an open problem.

Target Audience

Researchers and practitioners working at the intersection of LLMs and automated planning, particularly those interested in neuro-symbolic pipelines, LLM-based multi-agent orchestration, and text-to-PDDL translation. It is also relevant to benchmarking and evaluation researchers, since a substantial portion of the discussion concerns flawed ground truth in widely used planning benchmarks and the reliability of LLM-as-a-judge scoring. Engineers building planning or scheduling systems that already use symbolic planners such as Fast Downward will find the planner-agnostic architecture directly applicable, though the formal optimisation sections and PDDL notation assume a technically advanced reader.

Authors’ abstract

We present an end-to-end framework for planning supported by verifiers. An orchestrator receives a human specification written in natural language and converts it into a PDDL (Planning Domain Definition Language) model, where the domain and problem are iteratively refined by sub-modules (agents) to address common planning requirements, such as time constraints and optimality, as well as ambiguities and contradictions that may exist in the human specification. We support two categories of agents: hardcoded, which are informed by logs and error traces and have a pre-defined goal (e.g., fix issues with PDDL syntax, check temporal constraints), and dynamic, which have no predefined goal but adapt to the specific domain and revise the latent planning abstraction. The validated domain and problem are then passed to an external planning engine to generate a plan. The orchestrator and agents are powered by Large Language Models (LLMs) and require no human intervention at any stage of the process. Finally, a module translates the final plan back into natural language to improve human readability while maintaining the correctness of each step. We demonstrate the flexibility and effectiveness of our framework on GPT-\{4o, 5-mini, 5.4\}, and Gemini-\{2.5, 3\}-flash across more than ten domains and tasks, including the Google NaturalPlan benchmark, Planbench, and classic planning problems like Sokoban, Blocksworld and the Tower of Hanoi, where LLMs are known to struggle even with small instances. Our framework can be integrated with any PDDL planning engine and validator (we successfully tested Fast Downward, LPG, POPF, VAL, and uVAL) and represents a significant step toward end-to-end planning aided by LLMs.

Read the original paper