Skip to content
AI.info

Research

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos Overview Research area: Computer Vision, specifically inverse vision / procedural reconstruction from instructional vide

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
arXiv
2609.00377
Published
2026-08-31
Authors
Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel

AI summary

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Overview

  • Research area: Computer Vision, specifically inverse vision / procedural reconstruction from instructional videos, intersecting with computational origami geometry.
  • Technical level: Advanced. It assumes familiarity with Vision-Language Models (VLMs), agentic tool-calling loops, and mesh/crease-pattern representations for origami.
  • Scope in one sentence: The paper introduces an agentic VLM framework that converts keyframes from instructional Pureland origami videos into an executable, parametric sequence of folding actions, and evaluates it on a newly curated benchmark called PurelandFold.

What This Paper Is About

Origami instructions are normally shared as pictures or videos, but computational tools need structured, explicit representations such as crease patterns or folding programs. This creates a gap: machines cannot ingest the intuitive visual "language of folding" that humans use. The paper's goal is to translate keyframes from an origami demonstration video into a parametric geometric state representation plus the ordered sequence of folding actions that transforms each state into the next, producing a program that can be re-rendered, edited, and analyzed.

Key Contributions

  1. A new task: Inferring procedural folding programs directly from origami demonstration sequences, rather than predicting a single static crease pattern or classifying a paper state.
  2. A parametric action space and simulator for Pureland origami, built on five composable primitives (add_vertex, fold, unfold, rotate, flip) and a state representation that extends the FOLD format with per-face orientation and an explicit layer ordering.
  3. An agentic framework that couples a pretrained VLM with a tool library (simulation actions, viewing, checkpoints, verification), a deterministic controller, and a separate visual critic model, enabling self-verification, re-planning, and rollback.
  4. PurelandFold, a curated dataset of 27 diverse Pureland origami folding sequences with keyframe-level ground-truth folding actions and geometric states, plus evaluation metrics adapted from prior work.

Main Findings

  • Validity depends on the representation, not just the model: The VLM-CP variant (predicting a standard FOLD crease pattern directly from frame pairs) completes 96% of sequences but only 8% of its outputs compile successfully with the Flat-Folder compiler.
  • Action-space representation raises validity: VLM-S, which predicts a symbolic simulator action in the paper's state and action representation, reaches 96% compilation validity on completed steps, with 85% of sequences fully completed.
  • The critic improves fine geometry: Adding the external critic in a non-agentic way (VLM-S-C) lowers Crease-Pattern Dissimilarity (CPD) from 268 ±191 (VLM-S) to 216 ±159, while other similarity scores stay roughly flat or dip slightly.
  • The full agentic framework wins across the board: FoldingAgent completes 100% of sequences and obtains the best reported values on every similarity and dissimilarity metric: TSS 0.90 ±0.1, GS 0.58 ±0.1, CS 0.77 ±0.3, FFS 0.39 ±0.3, CPD 189 ±145, with 96% compilation validity.
  • Error compounding is real, but recoverable: The paper documents a case where an incorrect fold at keyframe 5→6 is approved by the critic and propagates to frame 7; the agent then detects the divergence, rolls back to keyframe 3, and reconstructs the sequence with a corrected fold that holds to the final frame. This is attributed to the critic having an 88% success rate.
  • Performance degrades with sequence progress: All variants decline as the sequence advances, but the paper reports the full method has a steadier slope and a gap to the others that increases with complexity.
  • Human preference: In a perceptual user study of 600 judgments over a subset of 10 sequences, the method was preferred over VLM-S in 84% ±5% of judgments and over VLM-S-C in 80% ±7%. VLM-CP could not be evaluated in the user study because most of its sequences are not compilable.
  • Cost per sequence: On average, FoldingAgent issues 140 queries per sequence, consuming 4M input tokens and 68K output tokens, costing $8.9.
  • Long procedures are handled: Qualitative results include sequences of 13 and 14 keyframes involving flips, unfolds, rotations, and folds with vertex definitions at fractional positions (e.g., add_v(0.58, [2,3]), add_v(0.33, [3,14])).

Methodology in Plain English

The authors start from a manually extracted sequence of keyframes from an instructional origami video, where each keyframe shows a meaningful hand-paper interaction. Rather than asking a model to predict a whole crease pattern at once, they treat reconstruction as a step-by-step process, one keyframe transition at a time.

For the state, they describe the paper as a planar graph: vertices, faces, edges (labeled boundary, flat, mountain, or valley), plus two additions needed for vision — a per-face flag for whether the front or back of the paper faces up, and an explicit list giving the bottom-to-top stacking order of coplanar faces, grouped into planes that move together. For the actions, they restrict the task to Pureland origami, where every fold is a flat fold along a single crease, giving five composable primitives: subdivide an edge at a fractional position, fold along an edge in one of two directions, unfold, rotate by an angle, and flip about one of four axes. A simulator applies these actions and returns the updated state.

A pretrained VLM acts as the agent, zero-shot with no fine-tuning. It can call tools: the folding primitives, viewing tools that return frames, filmstrips of motion between frames, the current state as JSON, and a rendered 2D diagram of the current state, checkpoint tools to save and restore solutions, and a verification tool that queries a separate critic VLM. The critic receives a four-image grid — the real source and target photos and the rendered source and target states — and returns Match, Mismatch, or Extreme Divergence plus a written analysis. Splitting proposer from verifier is intended to stop the agent from rationalizing its own errors.

A deterministic controller parses tool calls, dispatches them, and tracks state; the agent never edits state directly. When the agent gets a Match it saves a checkpoint and moves on; on a Mismatch it retries with a different action sequence, with rollback available. If it fails three times on the same transition, it is instructed to generate an overview panel of all checkpoints against their target photos, find the earliest divergence, and roll back. To keep exploration tractable, each keyframe has a budget of C = 5 transition attempts; if the budget runs out, a lightweight selector agent picks the best explored state, and after that the resolved portion of the sequence can no longer be revisited. Total execution is capped at 300 tool calls.

Evaluation uses their new PurelandFold dataset — 27 self-captured folding sequences based on instructions from the Easy Origami category of the OrigamiWay website, which offers 60 origami models, 40 of them Pureland — annotated with ground-truth actions and geometry. They compare a ladder of ablations using the same underlying VLM (Gemini 3.1 Pro Preview for both agent and critic): direct crease-pattern prediction, action prediction in their representation, that plus a non-agentic critic, and the full agentic system. Metrics cover compilation validity, topological structure similarity, geometric similarity, constraint satisfaction, final folded state, and crease-pattern dissimilarity.

Why This Matters

  • Bridging representation gaps: The work targets the disconnect between how humans share origami knowledge (visual demonstrations) and how machines represent it (structured geometry and programs), a gap that the paper argues no prior system closes in a zero-shot, closed-loop way.
  • Agentic inverse vision as a template: The design — a VLM plus a simulator plus a separate critic plus explicit rollback — shows a pattern for turning unstructured video into executable programs, which the authors suggest generalizes beyond origami.
  • Compounding-error mitigation: Sequentially re-planned, verifiable steps offer a concrete demonstration that long multi-step visual procedures can be reconstructed without a supervised model trained on labeled procedures.

Real-world applications suggested or implied:

  • Robotic paper handling: Enabling robots to learn complex folding skills directly from the thousands of instructional videos available online.
  • Interactive instructional systems: Systems that can analyze a demonstration and provide or edit folding guidance.
  • Automated generative design: Producing editable origami programs that can be re-rendered or modified.
  • Formalizing human craft knowledge: Converting human-oriented demonstrations into formal representations for downstream computational use.

Industry relevance: The work touches domains where structured folding representations already matter — sheet-metal folding, deployable and self-folding structures, 3D/4D printing, and origami-inspired design tooling — because these pipelines need explicit geometry that this framework could in principle recover from demonstrations. The reported cost of $8.9 per sequence and an average of 140 queries per sequence is a concrete indicator of current practical feasibility.

Future Directions

  • Handling occlusion: Folds heavily occluded by the demonstrator's hands or by accumulating paper layers remain hard for the model to identify from raw keyframes; improving robustness here is a stated open problem.
  • Compound simultaneous actions: The agent struggles with actions such as rotating while flipping, which break the sequential assumptions of the parameterized action space and would need a richer action definition.
  • Extending beyond Pureland: The scope is currently Pureland origami; more advanced techniques are described as conceptually possible but would require extending the action space and possibly stronger vision-language reasoning than is currently available.
  • Reducing reliance on the VLM: The paper notes performance is highly dependent on the VLM's reasoning capability, and the critic's 88% success rate is the direct source of the error-propagation case studied, so stronger or better-calibrated verification is a natural next step.

Target Audience

Researchers in computer vision and graphics working on inverse problems, video understanding, or agentic VLM systems; computational origami and geometry researchers interested in bridging structured representations with real-world video; and robotics or design engineers exploring learning from instructional demonstrations. It is also relevant to readers interested in how tool-calling agents with external verifiers and rollback can be applied to long-horizon, physically constrained tasks.

Authors’ abstract

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.

Read the original paper