Skip to content
AI.info

Research

BBoxMaskPose v2: Expanding Mutual Conditioning to 3D

Overview Research area: Computer vision — 2D/3D human pose estimation, person detection, and human instance segmentation, with a focus on heavily crowded and interacting people. Technical level: Inter

BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
arXiv
2601.15200
Published
2026-01-21
Authors
Miroslav Purkrabek, Constantin Kolomiiets, Jiri Matas

AI summary

Overview

Research area: Computer vision — 2D/3D human pose estimation, person detection, and human instance segmentation, with a focus on heavily crowded and interacting people.

Technical level: Intermediate. The paper assumes familiarity with top-down pose estimation, heatmap-based keypoint prediction, SAM-style prompted segmentation, and average precision (AP) evaluation, but the pipeline is described in accessible terms.

Scope: The paper presents BBoxMaskPose v2 (BMPv2), an upgraded iterative pipeline that mutually conditions a detector, a new pose estimator (PMPose), and a new pose-guided segmenter (SAM-pose2seg), plus a higher-cost variant (BMPv2+), an extended dataset (OCHuman-Pose), and an analysis of using the pipeline's 2D output to prompt 3D body reconstruction.

Note on completeness: The supplied paper text is truncated during Section 5.5 (robustness to domain shift). Results for that section are therefore not reported here beyond what appears before the cut-off.

What This Paper Is About

Most 2D human pose benchmarks are nearly saturated, except for crowded scenes where people's bounding boxes almost coincide. The authors build a top-down pose estimator that models keypoint probabilities and conditions on segmentation masks, and fold it into an iterative loop that has a detector, a pose estimator, and a segmenter continuously refine each other. The goal is to make both 2D pose estimation and instance segmentation reliable in crowded, interacting scenes, and then to show that this improved 2D output transfers to 3D body reconstruction.

Key Contributions

  1. PMPose, a 2D top-down pose estimation model that unifies the mask conditioning of MaskPose with the probabilistic formulation of ProbPose — predicting keypoint probability maps, presence probability, visibility, and expected OKS. It is released as a family PMPose-S/B/L/H paired with ViT-S/B/L/H backbones.
  2. SAM-pose2seg, a pose-guided human instance segmentation model built on SAM 2.1 (SAM-hiera-B+ backbone) with a frozen backbone, a human-focused fine-tuned decoder, and pose-guided prompting during training. It replaces BMPv1's complex prompt-selection strategy and reduces over-segmentation.
  3. BMPv2 and BMPv2+, upgraded versions of the BBoxMaskPose mutual-conditioning loop. BMPv2 combines the new pose and segmentation components; BMPv2+ additionally loops pose estimation and mask refinement until convergence, trading computation for accuracy.
  4. OCHuman-Pose, an extended annotation of OCHuman adding keypoints for previously unannotated instances (over 50% new instances), plus an analysis of 3D pose estimation on crowded OCHuman scenes using SAM-3D-Body prompting.

Main Findings

  • New state of the art on crowded pose estimation: BMPv2 reaches 51.3 val AP / 51.5 test AP on OCHuman, and BMPv2+ reaches 55.8 / 55.8, making it the first method to exceed 50 AP on OCHuman. On the new OCHuman-Pose, BMPv2+ reaches 85.8 val AP and 86.8 test AP. The abstract states BMPv2 surpasses the state of the art by 1.5 AP points on COCO and 6 AP points on OCHuman.
  • PMPose beats prior top-down models: PMPose-B scores 47.9 val / 48.2 test AP on OCHuman, 78.9 / 80.0 on OCHuman-Pose, and 76.9 on COCO val — improving on MaskPose-B (46.6 / 46.6, 77.1 / 78.0, 76.8). It is comparable to the iterative method BUCTD (48.3 / 47.4 on OCHuman, 74.8 on COCO) without being iterative.
  • Large segmentation gains from SAM-pose2seg: BMPv2 mask AP is 40.9 val / 41.1 test on OCHuman versus BMPv1's 33.7 / 34.0, and 69.5 on CIHP val versus 65.9. The ablation shows BMPv1 + SAM-pose2seg alone lifts mask AP from 34.0 to 40.0 on OCHuman.
  • The loop only pays off with both strong components: Ablation on OCHuman test gives BMPv1 49.2 pose AP / 34.0 mask AP; BMPv1 + SAM-pose2seg 49.5 / 40.0; BMPv1 + PMPose 50.2 / 30.8; BMPv2 51.5 / 41.1; BMPv2+ 55.8 / 41.1.
  • Mixing components naively can hurt: BMPv1 with PMPose yields worse segmentation (30.8) than standalone BMPv1 (34.0), because BMPv1's prompting was tuned for MaskPose's confidence, which PMPose does not output.
  • BMPv2+ is slightly worse on COCO than BMPv2 (78.1 vs 78.8 AP). The paper attributes this to COCO's many small instances (more than 30% of instances have a bounding box smaller than 100px), where mask refinement is unsuitable, and to annotation errors along bounding box edges.
  • CIHP bounding boxes are the one metric where BMPv2 is not SOTA (68.8 val bbox AP vs. BMPv1's 69.7), attributed to CIHP's small instances and partial-person cases; BMPv2 still leads on CIHP masks (69.5 vs. 65.9).
  • Visibility is the best cue for prompt selection: With SAM-pose2seg, prompting with visibility-selected keypoints gives 44.6 COCO val AP, 34.7 OCHu test AP, and 72.7 CIHP val AP, ahead of expected-OKS (43.9 / 34.7 / 72.0), presence probability (43.8 / 31.0 / 69.3), and confidence (38.3 / 33.6 / 66.6).
  • OCHuman's annotations are incomplete and distort evaluation: The original dataset kept only instances with maximum IoU against any other instance equal to 0 or greater than 0.5, excluding the range in between. A ViTPose*-b model scores 44.5 val AP with RTMDet bounding boxes on OCHuman but 75.3 on OCHuman-Pose, while with ground-truth boxes it scores 90.9 and 86.4 respectively — showing missing annotations were being counted as detection errors.
  • Detection is largely solved in the BMP loop: On OCHuman-Pose, BMPv2+ (86.8 test AP) slightly exceeds ViTPose*-b with ground-truth boxes (86.2 test AP) and is close on val (85.8 vs. 86.4). The authors conclude remaining errors are mostly pose estimation errors, not detection errors.
  • 3D prompting benefits from better masks, not poses: For SAM-3D-Body on OCHuman test, BMPv2 prompting yields 46.4 3D pose AP without masks and 54.8 with masks (85.2 on OCHuman-Pose with masks), versus 39.9/45.6 for BMPv1 and 38.5/40.0 for RTMDet-L. Prompting with pose helped far less than segmentation masks.
  • Domain shift: On a private dataset of close adult–infant interactions, both BMPv1 and BMPv2 significantly outperform off-the-shelf RTMDet + ViTPose. Quantitative values for this experiment are not present in the supplied text.

Methodology in Plain English

The base idea. BBoxMaskPose (BMPv1) chains three lightweight models — a detector, a pose estimator, and a segmentation model — in a self-improving loop. Each model is conditioned on the others: once an instance has been processed, its mask is blacked out so the detector does not find it again (uncovering missed people); the pose estimator receives a segmentation mask instead of just a bounding box (helping it separate overlapping bodies); and the segmenter is re-prompted using the predicted pose.

What was replaced. The two weak links were the pose estimator and the mask refinement step. BMPv2 swaps in PMPose for pose and SAM-pose2seg for segmentation.

PMPose keeps MaskPose's mask conditioning and adds ProbPose's probabilistic outputs: per-keypoint probability maps, a presence probability for whether the joint is even in the image, a visibility estimate, and an expected-OKS score used as a confidence measure. Training follows the same COCO + AIC + MPII recipe, fine-tuned with CropAugmentation so the model can learn to place keypoints outside the visible region; the semi-transparent visibility masking factor was changed to α = 0.25 from 0.2 in MaskPose.

SAM-pose2seg starts from SAM 2.1 with a frozen backbone, so generalization is preserved, and fine-tunes only the decoder to predict human instances (including clothing). Crucially, the random point prompts used in original SAM training are replaced by pose-aware prompts: the first prompt is the most visible keypoint, and later prompts are keypoints lying in the region where the previous prediction disagreed with ground truth, falling back to random points from that region when no keypoint is there. Training keypoints were produced by ProbPose, which lets the authors use datasets such as CIHP that have masks but no keypoint labels. The result is that two keypoints are sufficient to characterize an instance; the final model uses three.

The plus variant. BMPv2+ loops pose estimation and mask refinement until convergence — in practice one extra PMPose pass, since a second SAM-pose2seg step added nothing. Detection and segmentation results are identical to BMPv2.

Reaching 3D. Only after the 2D loop converges do the authors run a 3D body estimator (SAM-3D-Body). The image is encoded once, then each person is predicted by prompting with that person's bounding box and mask; predicted 3D bodies are filtered with 3D non-maximum suppression using confidences inherited from the 2D stage. Since OCHuman has no 3D ground truth, evaluation measures reprojection of 3D keypoints back to 2D.

Evaluation setup. RTMDet-L handles detection, PMPose-B handles pose, and results are reported on OCHuman, the new OCHuman-Pose, COCO, and CIHP. OCHuman-Pose was created by annotating every missing person with keypoints and reinstating annotations filtered out by the original dataset construction. It deliberately contains no masks for the new people, so it cannot be used for segmentation mAP.

Why This Matters

The work argues that crowded-scene 2D pose estimation is not solved and that fixing it is a prerequisite for reliable 3D reconstruction, because leading 3D estimators depend on accurate 2D detection, masks, or prompts. It also exposes a measurement problem: a widely used benchmark (OCHuman) has incomplete annotations that penalize correct predictions as false positives, which means reported progress on it is partly measuring dataset artifacts rather than model quality. The released OCHuman-Pose annotations, code, models, and data aim to make crowded-scene benchmarking more trustworthy.

Real-world applications:

  • Crowd monitoring and public-safety analytics, where individuals in dense groups overlap heavily and bounding-box-only detectors merge or drop people.
  • Sports and biomechanics analysis, where players frequently overlap and each athlete's skeleton must be attributed to the right body.
  • Clinical and care settings, suggested by the paper's own evaluation on adult–infant close-contact interactions, where infants are absent from standard training data and adults are only partially visible.
  • Animation, AR/VR, and motion capture, where the 3D extension could turn monocular footage of interacting people into body meshes without specialized capture hardware.

Industry relevance: The pipeline is built from lightweight, replaceable components and uses existing open models, so it maps onto practical deployment where detection, segmentation, and pose are usually separate purchase-and-integrate decisions. The BMPv2 vs. BMPv2+ split gives an explicit accuracy-versus-compute choice, and the demonstration that mask quality drives 3D body reconstruction makes the pipeline relevant to products that already prompt a 3D human model.

Future Directions

  • Closing the remaining pose error in crowds. The authors show that on OCHuman-Pose most errors are pose estimation errors rather than detection errors, so the next gains must come from better keypoint attribution, not better detectors.
  • Fixing mask refinement for small instances. BMPv2+ loses ground on COCO and CIHP bounding boxes partly because SAM-pose2seg is built for higher-resolution images and full-person instances; adapting it to small or partially visible people is an open problem.

Authors’ abstract

Most 2D human pose estimation benchmarks are nearly saturated, with the exception of crowded scenes. We introduce PMPose, a top-down 2D pose estimator that incorporates the probabilistic formulation and the mask-conditioning. PMPose improves crowded pose estimation without sacrificing performance on standard scenes. Building on this, we present BBoxMaskPose v2 (BMPv2) integrating PMPose and an enhanced SAM-based mask refinement module. BMPv2 surpasses state-of-the-art by 1.5 average precision (AP) points on COCO and 6 AP points on OCHuman, becoming the first method to exceed 50 AP on OCHuman. We demonstrate that BMP's 2D prompting of 3D model improves 3D pose estimation in crowded scenes and that advances in 2D pose quality directly benefit 3D estimation. Results on the new OCHuman-Pose dataset show that multi-person performance is more affected by pose prediction accuracy than by detection. The code, models, and data are available on https://MiraPurkrabek.github.io/BBox-Mask-Pose/.

Read the original paper