Research
Brep2Shape: Boundary and Shape Representation Alignment via Self-Supervised Transformers
Overview Research area: Geometric deep learning for computer-aided design (CAD), specifically self-supervised representation learning on Boundary Representation (B-rep) models. Technical level: Advanc

- arXiv
- 2602.07429
- Published
- 2026-02-07
- Authors
- Yuanxu Sun, Yuezhou Ma, Haixu Wu, Guanyang Zeng, Muye Chen, Jianmin Wang, Mingsheng Long
AI summary
Overview
Research area: Geometric deep learning for computer-aided design (CAD), specifically self-supervised representation learning on Boundary Representation (B-rep) models.
Technical level: Advanced. The paper assumes familiarity with NURBS/Bézier parameterizations, Transformer architectures, attention mechanisms, and graph-based topological modeling.
One-sentence scope: The paper proposes Brep2Shape, a self-supervised pre-training method that learns to predict dense spatial points from parametric Bézier control points, aligning abstract boundary representations with intuitive shape representations through a Dual Transformer backbone with topology-aware attention.
What This Paper Is About
B-rep is the industry standard for encoding CAD geometry, but deep learning on it faces a representation gap: continuous methods (such as BRT) use parametric control points that are analytically precise yet visually opaque, while discrete methods (such as UV-Net and BrepNet) sample spatial points that are intuitive yet imprecise and sensitive to discretization artifacts. Brep2Shape bridges this gap by pre-training a model to map parametric control points to sampled 3D points, so the network learns the physical manifold that underlies abstract coefficients. The goal is a generalizable representation that is both mathematically rigorous and geometrically intuitive, transferable to downstream classification and segmentation tasks.
Key Contributions
-
Brep2Shape pre-training task. A self-supervised objective that predicts shape representations (dense spatial points sampled from faces and curves) directly from boundary representations (standardized Bézier primitives derived from NURBS entities), requiring no manual annotation because the target points come from analytical geometric evaluation.
-
Dual Transformer backbone. A backbone with two parallel streams that independently encode face (surface) tokens and edge (curve) tokens, preserving their distinct geometric properties rather than merging them into a single token set.
-
Topology attention. An attention module that injects a topological attention bias derived from the B-rep adjacency graph, using the face graph (faces linked by shared edges) and its topological dual, the edge graph (edges linked by shared faces), to maintain topological consistency and enable bidirectional information flow between the two streams.
-
Large-scale empirical validation. A 250,000-model pre-training corpus (Brep2Shape-250k) that demonstrates favorable data scaling and model scaling behavior, with state-of-the-art accuracy and faster convergence on four downstream benchmarks.
Main Findings
-
The representation gap is real and quantifiable. Continuous methods achieve analytical precision but decoupled spatial intuition; discrete methods are intuitive but imprecise. Digital experiments on a single control point (Figure 2) show that position and weight modifications produce opaque, spatially varying deformations, illustrating why control points are coefficients of a polynomial basis rather than points on the geometry.
-
Both pre-training data and model size scale favorably. Scaling pre-training data from 25k to 250k with a fixed 6-layer Dual Transformer reduced pre-training loss by 11.6%, while average downstream accuracy rose from 0.9379 to 0.9524. Scaling model size from 2 to 12 layers (expanding parameters from 1.3M to 3.3M) on a fixed 50k subset consistently reduced loss and improved downstream results. The authors report an average error reduction of up to 36.3% over the state-of-the-art baseline, BRT.
-
State-of-the-art downstream accuracy across four benchmarks. In Table 2, Brep2Shape achieves the best reported value in every listed column: 99.99 and 84.72 across the two classification benchmarks (versus BRT at 98.95 and 80.90, AAGNet at 96.33 and 74.72, UV-Net at 92.68 and 77.87), and 99.35, 98.02, 96.88, 83.77 across the two segmentation benchmarks (versus BRT at 95.73, 89.98, 94.48, 79.23; AAGNet at 99.29, 98.64, 82.45, 75.53; UV-Net at 98.92, 96.56, 89.03, 66.47). Fine-tuning used only 100 epochs versus 350 epochs for the baselines.
-
Strong low-label transfer. With only 100 labeled samples, Brep2Shape reaches 47.9% accuracy, compared to 41.5% for UV-Net and 44.8% for BRT.
-
Topology and edge supervision both matter. Removing the Dual Transformer (pre-training tokenizers only) raises face loss from 1.855 to 2.142 and edge loss from 1.238 to 1.287 (units of 10⁻³). Removing edge-level supervision raises face loss to 1.872, showing that edge supervision acts as a cross-entity consistency signal for adjacent faces.
-
Dual streams with topology attention outperform all tested backbone variants. In the backbone ablation, the default Dual Transformer achieves the best MFCAD++ accuracy/IoU (0.9934 / 0.9799) and lowest pre-training loss (3.093), ahead of GNN (0.9836 / 0.9522, loss 3.855), Q-Former (0.9863 / 0.9603, loss 3.158), Face Trm. (0.9907 / 0.9689, loss 3.202), Face + Edge Trm. (0.9818 / 0.9487, loss 3.251), and Dual Trm. w/ Std. Attn. (0.9853 / 0.9582, loss 3.203). The pattern indicates that classification benefits from holistic aggregation, while segmentation requires preserving entity-specific face and edge semantics.
-
Fine-tuning schedule matters less than the pre-trained representation. TMCAD accuracy rises monotonically from 82.29% at 50 epochs to 84.03% at 350 epochs, with the default 100-epoch setting (82.64%) offering a practical balance; extending to 350 epochs adds a further 1.39 percentage points.
-
Cross-domain transfer beats SSL4CAD. In-domain on Fusion360Seg, Brep2Shape-Fusion reaches 0.9201 at 1k, 0.9604 at 10k, 0.9711 at 20k, and 0.9716 at 23k (all) samples; SSL4CAD reaches 0.91, 0.95, 0.96, 0.96. On the target domain MFCAD with 100 samples, Brep2Shape-25k reaches 0.7626 and Brep2Shape-Fusion 0.7362 versus SSL4CAD's 0.66, an absolute gain of about 15%; all three converge near 0.99 or above at 10k samples.
-
Learned features are more discriminative and more linearly separable. On MFCAD++, cosine-similarity discriminability is 0.23 for Brep2Shape versus 0.21 for shape representation and 0.14 for boundary representation; frozen-representation linear probing accuracy is 0.8465 versus 0.7353 and 0.7035. The cosine analysis selects 20 random faces and compares label agreement among the top-10% most similar and bottom-10% least similar faces.
-
Qualitative behavior. Predictions are most accurate on planar coordinates, with curved surfaces harder due to higher abstraction; t-SNE of the top-3 most frequent MFCAD++ classes shows clearer class separation than boundary representations.
Methodology in Plain English
Turning NURBS into a uniform input. Raw B-rep entities are NURBS surfaces and curves with varying control-point counts, knot vectors, and weights, which cannot be fed directly into a standard deep model. The authors decompose each entity into a fixed number of Bézier primitives: n_f Bézier triangles per surface and n_e Bézier segments per curve, each primitive carrying a fixed number of control points n, where each point stores spatial coordinates plus its NURBS weight (a matrix in R^{n×4}). A Transformer encoder per entity type aggregates these primitives through a learnable [CLS] token into a single entity-level embedding, producing face tokens and edge tokens.
The self-supervised target. Instead of asking the model to reconstruct control points, the task asks it to predict where the geometry actually lies in space. Uniform sampling in the parameter domain produces 3D point sets (in R^{m×3}, padded to fixed size m) for each face and edge. The training objective minimizes the squared difference between these sampled points and the model's prediction given the control points and topology — a target that is generated analytically, so no human labels are needed.
Handling topology. Two parallel Transformer streams process face and edge tokens. Rather than standard self-attention, each stream uses topology attention, where the bias between two tokens comes from the embedding of the entity they share: two faces are connected through their shared edge in the face graph, and two edges through their shared face in the dual edge graph. This lets faces inform edges and edges inform faces while keeping the two streams separate, preserving local geometric character.
Training and evaluation. Pre-training uses AdamW with MSE loss for 100 epochs on an NVIDIA A100 GPU over the 250k-model corpus; fine-tuning switches to cross-entropy loss for 100 epochs with the same optimizer, while baselines are trained for 350 epochs. Metrics are Accuracy for classification and Accuracy plus IoU for segmentation.
Why This Matters
Impact on research. The paper reframes B-rep learning as an alignment problem rather than a choice between continuous and discrete representations, and shows that this alignment objective produces a scalable, label-free foundation model for CAD. It also demonstrates that a pure-Transformer backbone with topology-aware attention can outperform handcrafted-feature approaches (AAGNet), discretization-based approaches (UV-Net), and hierarchical recurrent-Transformer designs (BRT), while offering evidence of data and model scaling — a property the authors note B-rep learning has yet to
Authors’ abstract
Boundary representation (B-rep) is the industry standard for computer-aided design (CAD). While deep learning shows promise in processing B-rep models, existing methods suffer from a representation gap: continuous approaches offer analytical precision but are visually abstract, whereas discrete methods provide intuitive clarity at the expense of geometric precision. To bridge this gap, we introduce Brep2Shape, a novel self-supervised pre-training method designed to align abstract boundary representations with intuitive shape representations. Our method employs a geometry-aware task where the model learns to predict dense spatial points from parametric Bézier control points, enabling the network to better understand physical manifolds derived from abstract coefficients. To enhance this alignment, we propose a Dual Transformer backbone with parallel streams that independently encode surface and curve tokens to capture their distinct geometric properties. Moreover, the topology attention is integrated to model the interdependencies between surfaces and curves, thereby maintaining topological consistency. Experimental results demonstrate that Brep2Shape offers significant scalability, achieving state-of-the-art accuracy and faster convergence across various downstream tasks.Code is available at this repository: https://github.com/thuml/Brep2Shape.