Research
FabDreamer: Exploring the Image-to-Physical Workflow Through AI-Assisted Layered Fabrication
FabDreamer: Exploring the Image-to-Physical Workflow Through AI-Assisted Layered Fabrication Overview Research area: Human-Computer Interaction (HCI) — specifically personal fabrication, creativity su
- arXiv
- 2608.13665
- Published
- 2026-08-13
- Authors
- Chenfeng Gao, Zeya Chen, Anjie Yang, Karan Ahuja, Danli Luo
AI summary
FabDreamer: Exploring the Image-to-Physical Workflow Through AI-Assisted Layered FabricationOverview
Research area: Human-Computer Interaction (HCI) — specifically personal fabrication, creativity support tools, and human-AI collaboration. The paper is published at UIST '26 (The 39th Annual ACM Symposium on User Interface Software and Technology), under the categories of human-centered computing, interactive systems and tools, empirical studies in HCI, user studies, and applied computing/media arts.
Technical level: Intermediate. The system itself combines vision-language models, segmentation, inpainting, and geometric constraint checking, but the paper's argument and findings are framed around workflow design, user studies, and human-AI initiative rather than model internals.
Scope (1 sentence): The paper presents FabDreamer, a system that carries an arbitrary image through depth-ordered layer decomposition, creative editing with real-time 3D preview, and structural verification into fabrication-ready SVGs, and evaluates it through three rounds including a study with 6 practitioners from different fabrication domains.
What This Paper Is About
Generative AI can produce a rich image in seconds, but turning that image into a physical object still requires humdrum manual work: separating the picture into physically separable parts, reconstructing content hidden behind foreground objects, and checking that every piece will survive cutting. FabDreamer's goal is to integrate these separate capabilities—segmentation, generative reconstruction, constraint checking—into a single workflow from an arbitrary image to a fabrication-ready file, for the domain of layered laser-cut art. The paper's distinctive angle is that the amount of AI initiative is deliberately varied at each stage according to how irreversible an error would be: an AI mistake in decomposition can be regenerated, but a bad cut cannot be undone.
Key Contributions
-
FabDreamer, an image-to-physical system for layered 2.5D art, combining hybrid segmentation routing, generative occlusion reconstruction, and a fabrication agent that verifies structural integrity before export. AI initiative is staged across the workflow: leading decomposition (Stage 1: Decompose), assisting creation on demand (Stage 2: Create), and advising preparation (Stage 3: Prepare). The design targets four challenges identified in the formative study: element sourcing constraints (C1), preparation bottlenecks (C2), limited spatial awareness (C3), and invisible fabrication constraints (C4).
-
Empirical findings from multi-round development, comprising a formative analysis, an early prototype evaluation with 10 novices and 3 experts, and a cross-domain practitioner study with 6 specialists. These findings include how physical awareness during design opens creative opportunities beyond error prevention, how practitioners appropriate the system's generic geometric operations for their own domains, and where the boundary lies between constraints an agent can detect from geometry and knowledge that only practitioners supply.
-
A deliberately calibrated human-AI initiative model for a workflow with irreversible physical consequences, in contrast to prior human-AI co-creation frameworks validated mainly in screen-based contexts such as writing and drawing.
-
A pipeline evaluation across 24 images (12 real photographs, 12 AI-generated) from Vecteezy in six categories, reporting which extraction route resolved each element across 238 elements and 107 layers.
Main Findings
-
Formative challenges in current practice: 76% of 25 analyzed tutorials relied on existing vector libraries rather than personal images, because preparing arbitrary images required manual tracing and prior knowledge of what would segment cleanly. Only 2 of 25 tutorials applied 3D visualization to preview designs. Material-specific limits such as minimum feature widths and elements with no physical connection to the board were rarely documented and almost never surfaced during design.
-
Early prototype study (Round 2, N=13): All 10 novices created 3 to 5 layer designs that survived laser cutting and physical assembly in approximately 25 minutes. The 3D preview received the highest satisfaction rating (mean = 4.9/5). Expert reviewers found that fabrication constraints (C4) remained entirely unaddressed in the prototype.
-
Design lesson from Round 2: Packing layer detection and segmentation into a single automated pass left users unable to intervene when the AI misread spatial relationships; users needed to iterate on the decomposition, not just accept or reject it. This directly motivated the staged, revisable design of Stage 1.
-
Practitioner study (Round 3): Six practitioners with diverse domains and laser-cut expertise participated. The abstract reports that physical awareness during design opens creative opportunities beyond error prevention, that practitioners appropriate the system's generic geometric operations for their own domains, and that the fabrication agent covers geometry-readable constraints while domain knowledge remains with the maker. Detailed quantitative outcomes for Round 3 are not reported in the excerpted content.
-
Pipeline extraction statistics: Across 238 elements in 107 layers, 77% resolved on the SAM2 first pass, 4% on a SAM2 refinement pass, 12% via generative extraction, and 7% via a bounding-box fallback. The SAM2 first pass resolved at mean confidence 0.91. The most challenging cases were diffuse natural boundaries in real photographs (distant mountains, fog gradients), which often required Stage 2 repair.
-
Timing: Each layer took approximately 39 s in the pipeline evaluation, totaling 3–4 minutes per image, with the bottleneck being API latency (VLM detection and generative inpainting) rather than local computation. During interactive use, elements appear at approximately one minute per layer.
-
Constraint thresholds are material-specific: The minimum survivable width w_min is set at plywood ≥ 1.0 mm and acrylic ≥ 1.5 mm. Changing the fabrication scale triggers re-evaluation, since features that survive at one size may fall below the threshold at another.
-
Practitioner baseline: All six participants used AI tools in other aspects of their work such as text generation or image creation, but none had integrated AI into their fabrication practice, citing a lack of tools that fit their fabrication workflows.
Methodology in Plain English
The researchers built the system iteratively alongside practitioners over three rounds, because the knowledge governing image-to-physical work is largely tacit.
Round 1 (formative study). They analyzed 25 YouTube tutorials on multilayered laser-cut design published 2019–2024, selected by topical criteria rather than popularity, and conducted a 60-minute semi-structured interview with a professional publisher with over 10 years of experience in layered laser-cut production. This produced four challenge categories (C1–C4) that became both design targets and later evaluation criteria.
Round 2 (early prototype study). They built a prototype with AI content generation, automated layer decomposition with occlusion reconstruction, real-time 3D preview, and SVG export, then had 10 novices create designs and 3 experts review the fabricated outputs. Novice failures to intervene in a single-pass decomposition and expert identification of unchecked fabrication constraints led to the three-stage design with varying AI initiative.
Round 3 (cross-domain practitioner study). Six practitioners (4 female, 2 male; ages 25–39) from the university community and local maker networks, chosen for diversity in fabrication domain and laser-cut experience, used FabDreamer in a two-task within-participant design. Task 1 (40–50 min) had all participants create a layered laser-cut souvenir from the same riverside village photograph, each required to add at least one element not in the original image. Task 2 (40–50 min) had each participant work toward a domain-relevant artifact from images they brought. Sessions lasted 90–120 minutes, conducted remotely via Zoom, screen- and audio-recorded, with a 10–15 minute interview before and after each task and interaction logs capturing AI invocations, canvas edits, and stage transitions. The contrast between each participant's behavior across the two tasks isolates domain transfer from system unfamiliarity and individual working style. All 12 artifacts (6 participants × 2 tasks) were fabricated by the research team on an xTool P2 CO2 laser cutter at participant-specified scale.
System implementation. A React and Three.js frontend with a FastAPI backend. Images are processed at 1024×1024 pixels, mapping to a 304.8 × 304.8 mm (12" × 12") physical board. Three categories of AI models are used: a VLM (Gemini 3.1 Pro) for image analysis, detection, and scene understanding; an image generation model (Gemini Flash) for stylization, inpainting, and reconstruction; and SAM2 Large for pixel-level segmentation, self-hosted on a Mac Mini (M4) with MPS acceleration. The Stage 3 fabrication agent uses Gemini 2.5 Flash with function calling.
Decomposition pipeline. Four phases: layer analysis, element extraction, background reconstruction, output preparation. Layers are processed front to back so that occlusion is handled implicitly—once front elements are extracted and inpainted away, hidden portions of deeper elements are revealed. Element extraction uses a cascade: SAM2 with bounding box and point prompts first, a refinement pass if the mask scores below a confidence threshold or is disproportionate to its bounding box, then generative extraction (redraw on a uniform background plus pixel-wise comparison) for elements with structural openings such as tree canopies or rock arches. A separate Grab pathway extracts a single described element, using chroma-key removal on a magenta background when a VLM determines the target is partially hidden. Each element yields an outline SVG (cut path) and an engrave SVG (internal detail lines via Sobel edge detection and Potrace vectorization).
Fabrication agent. It checks two failure modes detectable from SVG geometry alone. Before analysis, a union-find graph merges overlapping segments and the frame ring into connected components, inflating each polygon by 5 px to turn touching edges into measurable overlap. Floating elements are merged components not connected to the frame ring, reported with their distance to the nearest anchor. Thin elements are detected by morphological opening—eroding each polygon inward by w_min/2 and dilating back—with the erosion distance scaled as (w_min/2) × 1024/(board_mm × scale). The agent presents resolution options (auto-bridge, user-drawn bridge, ignore, or return to Stage 2), executes the fix, and re-checks the layer in a loop.
Why This Matters
Impact on research. The paper reframes "how much should AI help" as a question that depends on the physical reversibility of errors, extending human-AI initiative frameworks that have been validated mainly in screen-based contexts. It also argues that physical awareness during design is not only an error-prevention mechanism but a source of creative opportunity, and it identifies a concrete boundary: an agent can detect constraints readable from geometry, while domain knowledge stays with the maker. The work connects an upstream gap—from an arbitrary image to a fabrication-ready layered design—that prior constraint-aware tools such as Fabricaide, ScrapMap, and Towards Zero-Waste Furniture do not address because they assume a prepared design file as input. It also responds to the creativity-support concern that all-in-one pipelines disempower users, by keeping abstraction levels inspectable and exporting to standard formats.
Real-world applications (from the paper's framing and participants):
- Layered laser-cut shadow boxes and layered laser-cut art, including the running example of a Yellowstone photograph turned into an eight-layer basswood shadow box with a bear-shaped frame.
- Architectural site and stacked contour landscape models (participant P2, an architecture instructor).
- Tactile graphics such as Braille cards for blind and low-vision readers (participant P5, an accessible tech researcher, working in fabric and paper).
- Pop-up greeting cards and multi-material structures (participant P4, a makerspace manager in toy design), plus layered scene stickers for DIY decor (P3), wax-paper actuator prototypes (P1), and handmade layered fabric books (P6).
- Other mentioned targets: vinyl sticker scenes, paper-cut cards, and paper cutting traditions such as Chinese jianzhi.
Industry relevance. The export path (one SVG per layer plus PNG files in a ZIP, importable into xTool Creative Space, LightBurn, or Adobe Illustrator) is aimed at existing commercial fabrication software. The paper positions FabDreamer against commercial platforms such as xTool AIMake and Glowforge Magic Canvas, which generate single-layer designs ready for cutting but skip creative editing and offer no control over decomposition or composition, and against text-to-SVG systems like VectorFusion and SVGDreamer that are disconnected from fabrication workflows.
Future Directions
-
Extending beyond geometry-readable constraints. The fabrication agent covers floating and thin elements, but domain knowledge remains with the maker; a future question is how a system could incorporate material- or domain-specific knowledge without removing the user's authority over resolution.
-
Closing the gap between screen-ready and fabrication-ready decomposition. The pipeline's hardest cases were diffuse natural boundaries in real photographs (distant mountains, fog gradients), which often required Stage 2 repair—suggesting room for better masking or repair for low-contrast, non-object-like regions.
-
Improving latency and interactivity. Each layer took approximately 39 s, with the bottleneck reported as API latency from VLM detection and generative inpainting; reducing this would change the interactive experience of waiting roughly one minute per layer.
-
Testing broader generalization. The study covered 6 practitioners across 6 fabrication domains; the authors frame domain transfer as an open question, and the design rationale claims that layered laser cutting's core geometric constraints recur across other cutting and layered fabrication domains, which remains to be tested at scale.
-
Preserving the user's ability to work at lower levels of abstraction and interoperate with other tools, consistent with the paper's citation of vertical and horizontal movement in creativity support tool design.
Target Audience
This paper is most useful to HCI researchers working on personal fabrication, computational craft, and creativity support tools; to researchers studying human-AI collaboration and how initiative should be distributed in creative workflows; and to developers building AI-assisted design-to-fabrication pipelines. It is also relevant to practitioners and educators in laser cutting and layered art—illustrators, makerspace managers, architecture and design instructors, accessible-technology researchers producing tactile graphics, and fabric artists—because it documents where current tools break down and how participants from those domains adapted a general system to their own work. The paper's use of varied expertise levels (from a fabric artist with no prior laser-cut exposure to a makerspace manager with extensive expertise) and of two tasks (a controlled shared image and a participant-chosen domain image) makes it a useful methodological reference for cross-domain evaluation studies in HCI.
Authors’ abstract
Generative AI lets anyone create rich visual content in seconds, yet translating that content into a physically fabricable artifact still demands manual decomposition, occlusion repair, and structural verification that most tools leave entirely to the user. We present FabDreamer, an image-to-physical system that carries an image to fabrication-ready SVGs through three stages with deliberately staged AI initiative: (1) AI leads decomposition into depth-ordered layers, (2) assists on demand during creative editing with realtime 3D preview, and (3) advises on structural integrity before export. We instantiate this workflow for layered laser-cut art and evaluate it through three rounds including a formative analysis, an early prototype user evaluation (N=13), and a cross-domain practitioner study with specialists from 6 fabrication domains (N=6). Our findings show that physical awareness during design opens creative opportunities beyond error prevention, that practitioners appropriate the system's generic geometric operations for their own domains, and that the fabrication agent covers geometry-readable constraints while domain knowledge remains with the maker.