Research
BEDLAM2.0: Synthetic Humans and Cameras in Motion
BEDLAM2.0: Synthetic Humans and Cameras in Motion Overview Research area: Computer vision, specifically 3D human pose and shape (HPS) estimation from video, and synthetic training data generation. Tec

- arXiv
- 2511.14394
- Published
- 2025-11-18
- Authors
- Joachim Tesch, Giorgio Becherini, Prerana Achar, Anastasios Yiannakidis, Muhammed Kocabas, Priyanka Patel, Michael J. Black
AI summary
BEDLAM2.0: Synthetic Humans and Cameras in MotionOverview
- Research area: Computer vision, specifically 3D human pose and shape (HPS) estimation from video, and synthetic training data generation.
- Technical level: Advanced.
- Scope: The paper introduces BEDLAM2.0, a large synthetic dataset of 3D humans rendered with diverse and realistic cameras and camera motions, and demonstrates that training state-of-the-art human motion estimators on it improves accuracy over the original BEDLAM dataset.
What This Paper Is About
Estimating where a person is and how they move in world coordinates from video is hard, especially when both the human and the camera are moving, because real training data with ground-truth human and camera motion barely exists. The authors build BEDLAM2.0, a synthetic dataset that extends the widely used BEDLAM dataset with far more realistic cameras, camera motions, body shapes, clothing, hair, shoes, and 3D scenes. They then show that state-of-the-art models trained on BEDLAM2.0 are significantly more accurate than the same models trained on BEDLAM.
Key Contributions
- A major camera upgrade. BEDLAM2.0 covers focal lengths from 14mm to 400mm on a 16:9 DSLR sensor (36 x 20.25mm), compared with BEDLAM's narrow 52° or 65° horizontal field-of-view range, with 9% of videos zooming during the shot. It adds auto-generated synthetic camera motions (static, panning, tracking, dolly, orbit, zoom, and combinations, plus Perlin-noise shake) and real captured motions from phones, a tablet, and an Apple Vision Pro headset.
- Broader human diversity and realism. More body shapes with high BMIs, 4,643 motions (versus 2,311 in BEDLAM), strand-based 3D hair, graded clothing across sizes, and shoes, which were completely absent from BEDLAM.
- A released, ready-to-use dataset. Rendered videos, ground-truth SMPL-X body parameters, camera extrinsics and intrinsics, depth maps, 3D clothing/hair/shoe assets, training and test splits, plus code for training, evaluation, rendering and motion retargeting, and trained model checkpoints for CameraHMR, PromptHMR and GVHMR.
- Evidence that the dataset works. Experiments training CameraHMR, GVHMR and PromptHMR on BEDLAM (B1), BEDLAM2.0 (B2), or both, showing B2 improves accuracy over B1 and that B1+B2 gives the best world-space results.
Main Findings
- Dataset scale: 27,480 video sequences at 1280x720 and 30 frames per second, 8,048,411 PNG images, 74.52 hours of video, and 26TB of data total. Training and test splits contain 12.5M and 862K human bounding boxes respectively.
- Rendered with subframes: Images are produced from 56,338,877 rendered temporal subframes, using 7 temporal subframe images per frame and a 180-degree shutter for realistic motion blur.
- Motion pool and segments: 4,643 SMPL-X motions (compared to 2,311 in BEDLAM), sampled from AMASS (CMU, KIT, BMLmovi, BMLrub, HDM05, ACCAD, Transitions, MoSh, SOMA, PosePrior, DFaust) plus MOYO (yoga) and BEAT2 (conversational gestures). This yields 10,592 motion segments and 3,231,846 pose frames, of which 1,665,448 (51.53%) appear only once. Segments run from 4 to 16 seconds, versus a maximum of 8 seconds in BEDLAM.
- Body shapes: 1,615 sampled bodies with BMIs from 18 to 41, based on the BMI distribution of the CAESAR dataset and resampled to include more high-BMI bodies. B2 uses SMPL-X with 16 shape coefficients (B1 used 11) and a "locked" head without the hair bun. The introduction also states the dataset contains over 4K diverse body shapes.
- Hair: 40 unique strand-based hairstyles, each with between 50k and 100k 3D strand curves. Total groom vertex counts start at 1 million for straight hair and reach up to 13 million for complex curly styles such as afro, where each strand contains about 170 vertices. Nine hair material presets are used.
- Shoes: BEDLAM2.0 adds 182 unique shoes sourced from the Google Scanned Objects dataset, including 45 loafers, 6 formal shoes, 9 ballerina flats, 3 flip-flops, 18 boots, 5 football shoes, and 96 casual sport shoes. Roughly 40% of rendered images contain bodies with shoes.
- Clothing: 76 new 3D outfits were added to BEDLAM's 111, for 187 total. 50 outfits were graded into sizes XS through 6XL. Each outfit has between 6 and 28 texture variations, with a median of 10.
- Scenes and lighting: 94 preselected HDR images from PolyHaven and 15 high-quality 3D environments, up from 5 in BEDLAM. Indoor environments increase from 1 to 9, with 3 to 25 shot locations per environment.
- Camera motion mix: 86.4% of motions are synthetic and 13.6% are captured. Captured motions come from an Apple iPhone Pro 14, a Google Pixel 4a, an iPad Pro 11, and an Apple Vision Pro (90Hz head-pose capture via a custom Unity 6 app).
- Occlusion analysis: Across a random sample of 41.5k images covering 58.6k rendered bodies, 12.7% of images show more than 20% occlusion, and the top 10% most occluded bodies average 61.1% occlusion.
- Depth data: 44% of images come with 16-bit float depth maps and a corresponding center subframe render without motion blur, in EXR multilayer format.
- Test split design: 161 body shapes and 597 motions were held out, and 1,824 new sequences were rendered in 5 new environments, producing 449,061 test images with unseen pose and shape parameters.
- Single-image results (CameraHMR): Training on B2 alone beats B1 on every reported metric across 3DPW, EMDB and RICH. For example, on 3DPW PA-MPJPE drops from 43.2 (B1) to 41.1 (B2), MPJPE from 68.0 to 64.8, and PVE from 80.7 to 76.3; on RICH the values improve from 42.1/75.2/83.2 to 36.8/70.8/79.4. On EMDB, PA-MPJPE improves from 50.0 to 46.5, MPJPE from 88.7 to 74.6, and PVE from 101.6 to 86.2. Training on B2 also gives a 20% improvement in shape accuracy over B1.
- World-space results (GVHMR and PromptHMR): The combination of B1+B2 gives the best accuracy. PromptHMR trained on B1+B2 reaches WA-MPJPE100 of 72.5 and W-MPJPE100 of 116.6 on RICH, and 70.5 / 193.7 on EMDB, with Jitter of 10.2 and 11.3 respectively — better than PromptHMR trained on B1 alone (85.7 / 139.4 on RICH; 77.6 / 211.1 on EMDB) or on B2 alone (75.3 / 122.4 on RICH; 71.9 / 197.7 on EMDB). GVHMR follows the same pattern (B1+B2: 75.8 / 121.3 on RICH, 109.7 / 273.1 on EMDB).
- Synthetic-only training beats published models: GVHMR and PromptHMR trained only on synthetic data are more accurate than their originally published versions,
Authors’ abstract
Inferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This is particularly challenging when there is both human and camera motion. Progress on this topic has been limited by the lack of rich video data with ground truth human and camera movement. We address this with BEDLAM2.0, a new dataset that goes beyond the popular BEDLAM dataset in important ways. In addition to introducing more diverse and realistic cameras and camera motions, BEDLAM2.0 increases diversity and realism of body shape, motions, clothing, hair, and 3D environments. Additionally, it adds shoes, which were missing in BEDLAM. BEDLAM has become a key resource for training 3D human pose and motion regressors today and we show that BEDLAM2.0 is significantly better, particularly for training methods that estimate humans in world coordinates. We compare state-of-the art methods trained on BEDLAM and BEDLAM2.0, and find that BEDLAM2.0 significantly improves accuracy over BEDLAM. For research purposes, we provide the rendered videos, ground truth body parameters, and camera motions. We also provide the 3D assets to which we have rights and links to those from third parties.