Research
Sceniris: A Fast Procedural Scene Generation Framework
Overview Research area: Robotics — synthetic 3D scene generation for Physical AI (simulated robot environments, generative model training data, and 3D perception datasets). Technical level: Intermedia

- arXiv
- 2512.16896
- Published
- 2025-12-18
- Authors
- Jinghuan Shang, Harsh Patel, Ran Gong, Karl Schmeckpeper
AI summary
Overview
- Research area: Robotics — synthetic 3D scene generation for Physical AI (simulated robot environments, generative model training data, and 3D perception datasets).
- Technical level: Intermediate. The paper assumes familiarity with collision checking, scene graphs, pose sampling, and GPU-accelerated geometry libraries.
- Scope: Sceniris is a procedural scene generation framework that produces large batches of collision-free (and optionally robot-reachable) 3D scene variations far faster than the prior method it builds on, Scene Synthesizer.
What This Paper Is About
Procedural scene generators are used to create synthetic 3D environments, but existing methods are slow: the paper states Infinigen takes up to 10 minutes to generate a living room and Scene Synthesizer generates one table-top scene in about 10 seconds. That low throughput makes it hard to scale up datasets for robot learning, generative 3D models, and perception tasks. Sceniris is designed to remove that bottleneck by parallelizing the two dominant costs in procedural generation — pose sampling and collision checking — so that hundreds to thousands of scene variations can be produced per second from a single configuration.
Key Contributions
- A batched scene representation. Instead of storing a single 4x4 homogeneous transform per scene-graph edge, Sceniris stores a batch of transforms as a tensor (for example, an
(8, 4, 4)tensor for 8 scene instances). This enables a batched forward kinematics function that can modify poses of articulated object parts in parallel. - Batch sampling with caching. Rather than sampling one 2D point at a time from a support polygon, Sceniris samples a larger batch than required and caches it in a queue, refilling when the cache runs out. When support polygons across scene instances are related by a rigid transform, points are sampled once from a canonicalized polygon and transformed per instance.
- GPU-accelerated collision checking and an optional robot reachability check. Collision checking uses cuRobo instead of CPU-based FCL (via Trimesh), with meshes and world configurations pre-cached at initialization. Reachability uses RM4D, a reachability map using a single 4D data structure, with batched queries transferred to the GPU.
- An extended set of spatial relationships. Sceniris supports
onandinsideobject–parent relationships, single-anchor object–object relationships parameterized by anchor list, direction, distance, and distance type, aface_toorientation attribute, a multi-anchormiddlerelationship, and aratio_on_supportterm for controlling placement relative to surface boundaries.
Main Findings
- Speed-up over Scene Synthesizer: Sceniris achieves at least 234x speed-up over Scene Synthesizer by the abstract's claim, and the conclusion describes a 200+ x speed-up over the "naively multi-processed base approach." On execution time, Sceniris finished 16,384 scenes in 32s while Scene Synthesizer finished 64 scenes in 39s, which the paper describes as about 311 times the efficiency.
- Valid scene throughput: Sceniris generates about 234 times more valid scenes than Scene Synthesizer in the cold start case and about 2,936 times more in the warm start case. The caption of Figure 4 states Sceniris handles 311 and 6,321 times more environments per second than 10 multi-processed Scene Synthesizer under cold start and warm start respectively.
- Warm start performance: Warm-started Sceniris finished 16,384 scenes in 2.52s — shorter than Scene Synthesizer took to process 4 scenes.
- Component breakdown: Sceniris spent 10.17s generating 1,024 scenes, while Scene Synthesizer took 2.99s for 1 scene. The paper reports that Scene Synthesizer spends a large portion of time on overhead other than object collision checking, and identifies that this overhead is caused by repeatedly dumping the object asset to a trimesh scene; Sceniris caches this operation so it runs only once.
- Reachability cost: Running the benchmark with reachability checks turned on for the apple and the cabinet produced almost the same running time as without it. The reachability check takes a constant approximately 0.0001s per query, at most 1/20 of the collision check cost.
- Pressure test: Generating a batch of 524,288 scenes fits in an L4 GPU using about 20GB VRAM. It took 769.67s for the cold start case with 113,820 invalid scenes, resulting in 533.30 valid scenes/s — higher than the 408.77 valid scenes/s in Figure 4. Warm start achieved 5,279.84 valid scenes/s.
- Complex spatial relationships scale with a cost: Hard and Hard+ configurations increased execution time by about 20–30s over the previous level in the cold start case and about 20s in the warm start case, because some complex spatial relationships cannot be propagated to the batch by a simple rigid transformation. Sceniris still achieved about 249 valid scenes/s in Hard+ (cold start). Scene Synthesizer could not be compared here because it does not support these relationships.
- Bottleneck identified: Sceniris spends a significant amount of time initializing the collision checker, and the time for making cuRobo world configurations increases with the number of scenes, with some overhead present.
Methodology in Plain English
The authors started from Scene Synthesizer, a single-threaded procedural generator that places objects via rejection sampling: it samples one pose for an object, checks for collisions, and retries if a collision is found. They identified pose sampling and collision checking as the two main bottlenecks and attacked both with parallelization.
For sampling, they store a whole batch of scene instances in the scene graph rather than one scene at a time, sample points from support polygons in large batches, and cache the sampled points for reuse. When the support polygon in each scene instance is just a transformed copy of the same surface, they sample from one canonical polygon and apply each instance's world transform, falling back to per-polygon sampling when spatial constraints break that assumption. For collision checking, they moved from CPU-based FCL to cuRobo on the GPU, pre-caching meshes and world configurations so initialization happens once, and retrying only the scenes that actually failed rather than the whole batch. They added an optional reachability check using RM4D, batched and moved to the GPU, which rejects sampled poses that a robot cannot reach given a hypothesized robot position.
The evaluation uses a scene with three objects — an apple, a banana, and a cabinet — where the apple and cabinet are randomly placed on the plane and the banana is placed inside the upper drawer of the cabinet, with drawer joints in random states. This configuration uses only spatial relationships supported by both systems, making it a like-for-like comparison. The benchmark ran on a Google Cloud VM with 16 CPU threads (Intel Xeon CPU @ 2.20GHz) and an NVIDIA L4 GPU, with Scene Synthesizer run on 10 simultaneous processes and 10 retries allowed per object.
Why This Matters
The paper targets a practical bottleneck: procedural scene generation is the foundation for training data in physical AI, but its throughput has been too low to scale. Sceniris's stated objective is to generate 100–1000 scenes within a second on average. Because it can continuously produce randomized instances quickly, the authors argue it is promising for parallel reinforcement learning and data collection scenarios, not just one-off dataset creation.
Potential applications identified in the paper:
- Robot simulation and policy training: Diverse scenes are required in simulation to train policies and collect data, and randomizing object poses improves policy generalizability and data diversity.
- Generative simulation: Systems that generate scenes on the fly (for example GenSim, GenSim2, RoboGen, Robocasa) require a reliable approach for generating object layouts.
- Training generative 3D scene models: Learning-based scene generation requires huge datasets, which are typically produced by procedural methods, making throughput-critical upstream generators valuable.
- 3D perception tasks: Perception research can benefit from large 3D scene datasets.
- Digital twins and digital cousins: These aim to minimize the sim-to-real gap, where higher-quality 3D scenes are in greater demand.
On industry relevance, the work comes from the Robotics and AI Institute with a co-author from the University of Waterloo who did the work during an internship at the institute, and the code is released publicly at https://github.com/rai-inst/sceniris. The robot reachability feature is specifically aimed at manipulation tasks — the paper states that none of the existing scene generation frameworks supports it — which makes the framework directly usable for embodied AI pipelines rather than only for rendering or vision datasets.
Future Directions
- A better cuRobo interface for world configuration creation. The paper states that the cost of making cuRobo world configurations grows with the number of scenes and that improving this is out of scope of the work and left for future work.
- Batch sampling from heterogeneous polygons on the GPU. The paper identifies this as a potential future work that could speed up the Hard and Hard+ configurations, where valid sample polygons must be computed per scene instance.
- Larger batches on GPUs with more VRAM. The pressure test showed performance was not saturated — a larger batch and a GPU with larger VRAM could continuously improve throughput, since the cold start result of 533.30 valid scenes/s exceeded the 408.77 valid scenes/s in the main benchmark.
- Open question: scaling the comparison beyond the shared feature set. The paper could not compare Sceniris against Scene Synthesizer on the complex spatial relationship configurations because Scene Synthesizer does not support them, leaving the relative cost of the added relationships measured only against Sceniris's own prior configurations.
Target Audience
Robotics and embodied AI researchers who need large volumes of varied, collision-free simulated scenes for policy training or data collection; practitioners building procedural scene generators or working with simulation frameworks such as RoboGen, GenSim, Robocasa, or ProcTHOR; and engineers interested in GPU-accelerated geometry pipelines using cuRobo, Trimesh, and reachability maps. Readers focused on generative 3D scene models will also benefit, since procedural generators like Sceniris are the upstream source of training data for those models. Some familiarity with scene graphs, collision checking, and parallel computing will help, but the core argument — that batching and GPU acceleration remove the throughput bottleneck — is accessible to anyone who has waited on a scene generator.
Authors’ abstract
Synthetic 3D scenes are essential for developing Physical AI and generative models. Existing procedural generation methods often have low output throughput, creating a significant bottleneck in scaling up dataset creation. In this work, we introduce Sceniris, a highly efficient procedural scene generation framework for rapidly generating large-scale, collision-free scene variations. Sceniris also provides an optional robot reachability check, providing manipulation-feasible scenes for robot tasks. Sceniris is designed for maximum efficiency by addressing the primary performance limitations of the prior method, Scene Synthesizer. Leveraging batch sampling and faster collision checking in cuRobo, Sceniris achieves at least 234x speed-up over Scene Synthesizer. Sceniris also expands the object-wise spatial relationships available in prior work to support diverse scene requirements. Our code is available at https://github.com/rai-inst/sceniris