Research
MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
Overview Research area: Generative deep learning for parametric Computer-Aided Design (CAD), combining state-space sequence models with latent diffusion. Technical level: Advanced. The paper assumes f

- arXiv
- 2511.17647
- Published
- 2025-11-20
- Authors
- Liyuan Deng, Yunpeng Bai, Yongkang Dai, Xiaoshui Huang, Hongping Gan, Dongshuo Huang, Hao jiacheng, Yilei Shi
AI summary
Overview
Research area: Generative deep learning for parametric Computer-Aided Design (CAD), combining state-space sequence models with latent diffusion.
Technical level: Advanced. The paper assumes familiarity with Transformers, state-space models (Mamba), variational/latent autoencoders, and denoising diffusion probabilistic models.
Scope: The paper proposes MamTiff-CAD, a two-stage framework (a Mamba+ / Transformer autoencoder plus a multi-scale Transformer diffusion generator) and a new 13,705-model dataset (ABC-256) for generating long parametric CAD command sequences of 60 to 256 commands.
What This Paper Is About
Parametric CAD models are built from ordered command sequences (sketch, extrusion, boolean operations), and existing deep learning generators handle only short sequences — DeepCAD, for example, is limited to lengths under 60 — because the quadratic cost of Transformer self-attention does not scale to hundreds of commands. The paper's goal is to generate long, logically coherent, editable CAD command sequences of 60 to 256 commands by first compressing them into a latent space with a Mamba+ encoder and then learning that latent distribution with a multi-scale diffusion model.
Key Contributions
- A hybrid Mamba+ / Transformer autoencoder for long parametric CAD command sequences. The Mamba+ block adds a forget gate mechanism (G_f = 1 - G_b2) to a dual-branch state-space design, and a non-autoregressive Transformer decoder reconstructs the sequence from the latent vector.
- A multi-scale diffusion-based CAD sequence generator (MST-D) that fuses local and global topology. Each denoising layer runs three parallel attention branches with window sizes 64, 128, and 256, combined through a learnable gating mechanism, with time-step conditioning and sequence-aware sinusoidal positional encoding.
- A new public dataset, ABC-256, containing 13,705 CAD models with sequence lengths from 60 to 256, split into 10,964 training, 1,370 validation, and 1,371 test samples, with an average sequence length of 99 — 6.6x longer than DeepCAD's average of 15.
- Evaluations showing state-of-the-art performance on reconstruction and unconditional generation, with the authors stating (contribution text truncated in the source) that this confirms effectiveness for long-sequence generation.
Main Findings
- Reconstruction on ABC-256 (Table 2): MamTiff-CAD reaches 99.99% command accuracy, 99.93% parameter accuracy, a Median Chamfer Distance of 0.75 (reported multiplied by 10^3), an Invalid Ratio of 8.50%, and a Step Ratio of 93.93%. Baselines: DeepCAD (92.24 / 75.93 / 41.02 / 33.11% / 70.46%), MT-CAD (89.72 / 66.87 / 121.35 / 39.89% / 63.97%), MLSTM-CAD (86.09 / 65.55 / 112.89 / 42.53% / 59.85%).
- Generalization to Fusion360 (Table 3): Trained on ABC-256 and tested directly on the roughly 8,000 Fusion 360 parametric sequences, MamTiff-CAD achieves 99.99% command accuracy, 97.99% parameter accuracy, MCD 1.44, IR 5.70%, and SR 95.16%, versus DeepCAD (93.35 / 77.99 / 104.76 / 19.41% / 82.57%), MT-CAD (91.85 / 60.18 / 300.21 / 11.81% / 90.17%), and MLSTM-CAD (88.42 / 62.13 / 261.80 / 23.39% / 80.97%).
- Unconditional generation (Table 4): With 10,000 samples generated per method and converted to point clouds, MamTiff-CAD reports MMD 1.32, JSD 3.19, COV 65.31%, Unique 99.6, Novel 99.4, and Step Ratio 85.38%. For comparison: DeepCAD (2.66 / 6.49 / 56.66% / 75.8 / 88.0 / 23.96%), SkexGen (2.31 / 4.53 / 57.76% / 80.5 / 96.9 / 75.26%), HNC-CAD (1.63 / 4.25 / 62.03% / 89.2 / 91.8 / 80.86%).
- Mamba+ ablation (Table 5): Replacing the Mamba+ block with a plain Transformer block degrades reconstruction sharply (77.29% command accuracy, 63.62% parameter accuracy, MCD 64.30, IR 67.46%, SR 34.28%). Standard Mamba is close to Mamba+ (99.98 / 99.90 / 0.76 / 10.35% / 92.01%), with Mamba+ showing improvements across most metrics, which the authors attribute to the forget gate's long-range dependency capture.
- Multi-Scale Transformer ablation (Table 6): Removing MST from the diffusion generator, again over 10,000 generated models, lowers COV from 65.31% to 61.69%, worsens JSD from 3.19 to 4.92, raises MMD from 1.32 to 1.47, and drops Step Ratio from 85.38% to 77.05%.
- Latent space and sequence handling: The frozen encoder maps sequences into a latent variable Z in R^(N x 256 x 64); parameter vectors use 16 dimensions per command with unused entries padded to -1, continuous parameters normalized to a 2x2x2 cube and quantized into 256 discrete levels.
Methodology in Plain English
The framework works in two stages.
Stage 1 — compression. A parametric CAD model is written as a sequence of 256 commands, each with a command type (six types) and 16 parameters, padded with an end-of-sequence token. Each command is embedded by adding three components: a command-type embedding, a parameter embedding built from one-hot vectors of dimension 257 per parameter, and a positional embedding. These embeddings pass through four Mamba+ blocks, which split the signal into a feature-transformation branch (1D convolution plus a state-space model) and a gating branch (SiLU activation). The novel element is a forget gate computed as one minus the gating signal, which scales the transformed features so that historical information is not fully discarded; the final output adds the forget-gated features to the state-space output. A compression block reduces this to a latent vector of dimension 64. A four-block Transformer decoder then takes that latent vector together with learnable position embeddings and reconstructs the whole command sequence at once (non-autoregressive), trained with a cross-entropy loss on command types and parameters, weighted by beta = 2 for parameters, ignoring padding and unused-parameter entries.
Stage 2 — generation. The trained encoder and decoder are frozen, and a diffusion model is trained on their latent representations. Noise is added with a linear variance schedule over 1000 steps, and a multi-scale Transformer denoiser predicts that noise. Each denoising layer runs three parallel attention branches with windows of 64, 128, and 256 to capture local geometry, medium-range topology, and global structure, fusing them through a sigmoid-gated MLP. Time-step embeddings generate scale, shift, and residual-strength parameters, and sinusoidal positional encoding is added with a trainable scalar weight. At inference, the model samples a latent vector from Gaussian noise and decodes it into an executable CAD command sequence.
Implementation: PyTorch on an NVIDIA RTX 4090. The autoencoder uses AdamW with weight decay 1e-4 and learning rate 1e-3, 2000 warmup steps, gradient clipping at 1.0, batch size 32, and 300 epochs; the diffusion stage uses the Adam optimizer with beta1 = 0.9, learning rate 2e-4 decaying by a factor of 0.1 after 100,000 iterations, and batch size 64. The main text describes the second stage as 200,000 epochs while the supplementary material describes 200,000 iterations. The diffusion model uses embedding dimension 512, 6 Transformer layers, and 8 attention heads.
Why This Matters
Research impact: The work is presented as the first application of diffusion models to parametric CAD sequence generation, and it pairs that with a state-space encoder for long-range modeling. It also releases a longer-sequence benchmark (ABC-256) for a domain where prior datasets — DeepCAD at a maximum length of 60, Fusion 360 Gallery with short extrusion sequences, and ABC with 1 million B-rep models but no operation sequences — could not support long-sequence evaluation.
Real-world applications (as implied by the paper's framing):
- Industrial product design, where models containing hundreds of commands currently exceed the reach of existing generators.
- Editable 3D design tools: outputs are parametric command sequences, so users can modify them, unlike pure shape-generating approaches.
- CAD auto-completion and design assistance for long industrial workflows.
- Training and benchmarking of CAD-generation models against a released long-sequence dataset.
Industry relevance: The paper argues that existing 3D generative methods focus on geometry and neglect design logic and parametric constraints, making their outputs hard to use directly in industrial design. Preserving the command sequence keeps design intent and editability, and the reported Step Ratio (93.93% reconstruction, 85.38% generation) measures how often generated sequences convert to STEP format, a proxy for practical usability.
Future Directions
- Coverage of command types: The authors state the dataset does not cover all CAD command types used in industrial design, which limits handling of more complex topological structures.
- Boundary Representation (Brep) integration: The work parses command sequences rather than extracting features directly from Brep, and the authors call for integrating CSG and Brep representations to improve generalization and applicability.
- Model refinement: The conclusion states that future work will refine the model and explore integration with advanced design processes and interactive methods for intelligent CAD design.
- Unreported aspects: No training or inference wall-clock times, memory usage, or scaling analysis are reported, and no comparison against the multimodal CAD approaches (CAD-MLLM, FlexCAD) discussed in related work is presented, leaving open how this framework compares to those on the same long-sequence benchmark.
Target Audience
Researchers and graduate students working on generative models for 3D shapes and CAD, especially those interested in state-space models, latent diffusion, and long-sequence generation. It is also relevant to practitioners in CAD/CAE tooling and industrial design automation who need editable, parametric outputs rather than meshes or point clouds, and to anyone seeking a longer-sequence benchmark (ABC-256) for CAD command generation. Readers should be comfortable with diffusion model formulations and Transformer/SSM architecture terminology.
Authors’ abstract
Parametric Computer-Aided Design (CAD) is crucial in industrial applications, yet existing approaches often struggle to generate long sequence parametric commands due to complex CAD models' geometric and topological constraints. To address this challenge, we propose MamTiff-CAD, a novel CAD parametric command sequences generation framework that leverages a Transformer-based diffusion model for multi-scale latent representations. Specifically, we design a novel autoencoder that integrates Mamba+ and Transformer, to transfer parameterized CAD sequences into latent representations. The Mamba+ block incorporates a forget gate mechanism to effectively capture long-range dependencies. The non-autoregressive Transformer decoder reconstructs the latent representations. A diffusion model based on multi-scale Transformer is then trained on these latent embeddings to learn the distribution of long sequence commands. In addition, we also construct a dataset that consists of long parametric sequences, which is up to 256 commands for a single CAD model. Experiments demonstrate that MamTiff-CAD achieves state-of-the-art performance on both reconstruction and generation tasks, confirming its effectiveness for long sequence (60-256) CAD model generation.