Skip to content
AI.info

Research

LmPT: Conditional Point Transformer for Anatomical Landmark Detection on 3D Point Clouds

LmPT: Conditional Point Transformer for Anatomical Landmark Detection on 3D Point Clouds Overview Research area: Medical computer vision — automatic anatomical landmark detection on 3D point clouds, w

LmPT: Conditional Point Transformer for Anatomical Landmark Detection on 3D Point Clouds
arXiv
2602.02808
Published
2026-02-02
Authors
Matteo Bastico, Pierre Onghena, David Ryckelynck, Beatriz Marcotegui, Santiago Velasco-Forero, Laurent Corté, Caroline Robine--Decourcelle, Etienne Decencière

AI summary

LmPT: Conditional Point Transformer for Anatomical Landmark Detection on 3D Point Clouds

Overview

Research area: Medical computer vision — automatic anatomical landmark detection on 3D point clouds, with a focus on femoral bones across species (human and dog).

Technical level: Intermediate. The paper assumes familiarity with transformer architectures, point cloud learning, and standard keypoint-detection metrics, but the core ideas are explained at a level accessible to readers with general deep learning background.

Scope: The paper introduces LmPT, a conditional Point Transformer that uses Feature-wise Linear Modulation (FiLM) to adapt one model to different input types, and validates it on human and newly annotated dog femur point clouds for landmark detection.

What This Paper Is About

Identifying anatomical landmarks on bones such as the femur is necessary for surgical planning, kinematic modeling, and gait analysis, but manual annotation by experts is time-consuming and shows significant inter-observer variability, while rule-based automated methods tend to be tailored to specific geometries or limited landmark sets. The authors propose LmPT, a method that detects anatomical landmarks directly on 3D point clouds and can leverage homologous bones from different species for translational research. The goal is accurate, automated, cross-species femoral landmark detection that generalizes beyond a single anatomy.

Key Contributions

  1. LmPT architecture: A landmark detection method for 3D point clouds built on a Point Transformer encoder-decoder with a conditioning mechanism (FiLM) that adapts the model to different input types. Modulation is applied only to the bottleneck features of the encoder to reduce model overhead.
  2. Cross-species learning and a new dataset: The paper integrates cross-species training and introduces a newly annotated dataset of 14 dog femoral bone models, enabled by the observation that dog landmarks closely correspond to their human counterparts.
  3. Validation across species: The approach is validated on both human femurs and dog femurs, demonstrating generalization across homologous femoral structures.
  4. Public release: The code and the dog femur dataset will be made publicly available at https://github.com/Pierreoo/LandmarkPointTransformer.

Main Findings

  • Human femurs (MAE in mm, Table 1): LmPT with the PTv2 backbone achieves the lowest mean error of 2.54 mm, compared with DGCNN at 5.00 mm, LmPT-v3 at 3.87 mm, and the A&A method at 6.51 mm. For reference, the average manual error (AME) across landmarks has a mean of 3.56 mm and the maximum manual error (MME) a mean of 11.99 mm.
  • Expert-level accuracy: The paper states LmPT with the PTv2 backbone surpasses expert-level accuracy on the majority of individual landmarks to achieve a minimal MAE.
  • PTv3 versus PTv2: LmPT with the PTv3 backbone performs worse than PTv2 — the authors attribute this to PTv3's serialized mapping strategy for attention windows being more computationally efficient but less precise than PTv2's k-nearest-neighbors mechanism.
  • Dog femurs (MAE in mm, Table 2): LmPT-v2 again performs best with a mean of 1.71 mm, versus LmPT-v3 at 2.70 mm and DGCNN at 5.07 mm. LmPT-v2 also reaches the fastest convergence to perfect PCK.
  • A&A fails on dogs: The A&A method is not reported for the dog dataset because its initial atlas registration step failed for all test samples — an expected outcome given that it is designed for human femurs.
  • Cross-species training on humans: Training on both species decreases human MAE from 2.36 mm (single-species) to 2.29 mm (cross-species), indicating the model benefits from additional homologous data.
  • Cross-species training on dogs: Dog MAE rises marginally from 1.68 mm to 1.78 mm. The authors attribute this to the difference in the number of landmarks between species — the human dataset includes additional landmarks not directly applicable to dog anatomy, introducing features that confound the model.
  • PCK behavior: Both single- and cross-species training produce similar PCK curves for human and dog, but cross-species reaches a perfect score at a lower threshold.

Methodology in Plain English

The researchers represent each femur as a point cloud — a lightweight set of spatial coordinates rather than a dense volumetric scan or mesh. They build on Point Transformer architectures, which use self-attention over neighboring points and an encoder-decoder structure with pooling and unpooling to reduce and restore point cloud resolution. Enlarging or reducing the k-NN attention windows across encoder and decoder layers enables multi-scale feature learning, each layer containing Ns transformer layers.

To let one model handle more than one kind of input, they add a conditioning mechanism: a linear layer learns scale and shift parameters that apply a feature-wise affine transformation to the encoder's bottleneck features. This is the FiLM modulation, and it is applied only at the bottleneck to limit overhead. A keypoint prediction head then produces logits identifying the predicted landmark for each keypoint class.

To handle limited data, point clouds are normalized, uniformly sampled to 8192 points, and augmented with random rotation, scaling, and coronal-plane flipping to improve cross-side generalization on symmetric landmarks. All models are trained under identical conditions for 500 epochs with a batch size of 4, using a channel-wise cross-entropy loss that ignores unlabeled keypoints, an AdamW optimizer, a learning rate of 3×10⁻⁴, and a one-cycle learning rate scheduler.

Data: The human dataset from Fischer et al. contains 20 femur models (ten left, ten right) with 22 anatomical landmarks each, annotated by five experts who each processed all models and repeated the estimation four times, giving 20 annotation rounds per femur; the ground truth is the medoid of the 20 annotations. It is split into 16 train and 4 test samples. The new dog dataset contains 14 femur models from different breeds and sizes (seven left, seven right), segmented from CT scans and annotated by a veterinary expert with 11 landmarks, a subset of which also appears in the human dataset. It is split into 10 train and 4 test samples. Combined, the datasets provide 26 train and 8 test samples.

Evaluation: Localization accuracy uses mean absolute error (MAE) as the Euclidean distance between predicted and ground truth landmarks. The percentage of correct keypoints (PCK) measures whether predictions fall within a Euclidean distance threshold, evaluated across ten thresholds ranging from 1 mm to 8 mm.

Why This Matters

Impact on research: The work demonstrates that a single conditional model can learn landmark detection across homologous bones of different species, and that adding data from a second species can improve accuracy on the first. It also shows the value of point clouds as an efficient representation for medical surface analysis, and releases a new annotated dog femur dataset to support translational research. The failure of a human-specific atlas method (A&A) on all dog test samples illustrates the generalization gap that learning-based conditional models can address.

Real-world applications:

  • Surgical planning and preoperative planning based on accurate femoral landmark positions.
  • Kinematic modeling and gait analysis using the femoral mechanical and anatomical axes defined by landmarks such as the femoral head extremities, medial and lateral epicondyles, trochanters, and intercondylar notch.
  • Assessment of skeletal deformities in orthopedics.
  • Translational research using animal models, where cross-species anatomical variability is a central challenge.

Industry relevance: Automated landmarking replaces time-consuming manual annotation and reduces inter-observer variability, offering greater scalability, consistency, and efficiency for clinical and biomechanics pipelines. The paper's emphasis on 3D modeling and automated analysis pipelines points toward integration into broader orthopedic and veterinary workflows, and the public code and dataset lower the barrier to adoption.

Future Directions

  • Extending to additional species: The authors explicitly note that extending the framework to additional species in future work may further enhance generalization.
  • Improving cross-species transfer to the second species: Cross-species training slightly increased dog MAE (1.68 mm to 1.78 mm), which the authors link to the mismatch in landmark counts between datasets — resolving that asymmetry is an open problem.
  • Revisiting backbone efficiency versus precision: The PTv3 serialization scheme was less precise than PTv2's k-NN attention here, raising the question of whether efficiency-oriented point cloud backbones can be adapted for fine-grained localization tasks.
  • Broader medical integration: The authors highlight the potential of point cloud learning in the medical domain and encourage its integration into broader automated analysis pipelines.

Target Audience

Researchers and practitioners in medical image analysis, biomechanics, and orthopedic or veterinary computing who work with 3D anatomical surfaces. It is also relevant to computer vision researchers interested in point cloud transformers, keypoint detection, and conditioning mechanisms such as FiLM, as well as to engineers building automated surgical planning or gait analysis pipelines who need scalable alternatives to manual landmarking.

Authors’ abstract

Accurate identification of anatomical landmarks is crucial for various medical applications. Traditional manual landmarking is time-consuming and prone to inter-observer variability, while rule-based methods are often tailored to specific geometries or limited sets of landmarks. In recent years, anatomical surfaces have been effectively represented as point clouds, which are lightweight structures composed of spatial coordinates. Following this strategy and to overcome the limitations of existing landmarking techniques, we propose Landmark Point Transformer (LmPT), a method for automatic anatomical landmark detection on point clouds that can leverage homologous bones from different species for translational research. The LmPT model incorporates a conditioning mechanism that enables adaptability to different input types to conduct cross-species learning. We focus the evaluation of our approach on femoral landmarking using both human and newly annotated dog femurs, demonstrating its generalization and effectiveness across species. The code and dog femur dataset will be publicly available at: https://github.com/Pierreoo/LandmarkPointTransformer.

Read the original paper