Skip to content
AI.info

Research

ShelfAware: Real-Time Semantic Localization in Quasi-Static Environments with Low-Cost Sensors

Overview Research area: Robotics — semantic visual-inertial global localization and Monte Carlo Localization (MCL), with applications to service robots and assistive navigation. Technical level: Advan

ShelfAware: Real-Time Semantic Localization in Quasi-Static Environments with Low-Cost Sensors
arXiv
2512.09065
Published
2025-12-09
Authors
Shivendra Agrawal, Jake Brawer, Ashutosh Naik, Alessandro Roncone, Bradley Hayes

AI summary

Overview

Research area: Robotics — semantic visual-inertial global localization and Monte Carlo Localization (MCL), with applications to service robots and assistive navigation.

Technical level: Advanced. The paper assumes familiarity with particle filters, observation likelihoods (beam mixture models), Jensen–Shannon divergence, visual-inertial odometry, and modern vision pipelines (YOLOv9, ResNet50, DINOv3).

Scope (one sentence): The paper presents ShelfAware, a semantic particle filter that models objects as probabilistic distributions over class counts and spatial arrangement rather than fixed landmarks, and uses an inverse semantic proposal mechanism to achieve real-time global localization on low-cost vision-only hardware in quasi-static indoor environments.

What This Paper Is About

Indoor spaces like grocery stores and warehouses are "quasi-static": walls and shelving stay put, but what sits on the shelves changes constantly. Standard particle filters such as Adaptive Monte Carlo Localization (AMCL) rely on static geometric maps and break down in long, visually repetitive aisles, producing perceptual aliasing and particle impoverishment. ShelfAware's goal is robust, start-anytime global localization — locating a robot or assistive device from an unknown initial pose and recovering after tracking loss — using only a compact RGB-D camera and a VIO camera, with no LiDAR, wheel encoders, or installed infrastructure.

Key Contributions

  1. Distributional semantic mapping and observation design. Object collections are encoded as statistical distributions over class counts and coarse spatial arrangement (count, mean range, and mean bearing sub-vectors), giving inherent robustness to semantic perturbations and inventory flux.
  2. An inverse semantic observation model inside a particle filter. Precomputed semantic viewpoints let the filter propose high-quality global pose hypotheses directly from live semantic observations, rather than relying solely on forward weighting.
  3. A real-time fusion of depth and semantic likelihoods. Depth likelihood from a synthesized 2D scan is combined with category-centric semantic similarity in a single joint observation model, running on low-cost, vision-only hardware without LiDAR or wheel encoders.
  4. A perception-agnostic evaluation across two domains. The system is validated both with a supervised closed-set pipeline (14 classes) in a motion-capture-controlled mock store and with a self-supervised open-vocabulary pipeline (60 classes) in a 3,500 sq. ft. operational grocery store.

Main Findings

  • Global localization in the mock store: ShelfAware achieved a 97% overall success rate across 100 trials, versus 15% for MCL and 10% for AMCL, with a mean time-to-convergence of 1.07 s across all setups. It reached 100% success in the Cart, Wearable, and Dynamic conditions and 88% in the Degraded/Sparse condition, versus 90% overall for the Fixed Quantity Landmark (FQL) baseline.
  • Tracking robustness: ShelfAware achieved the highest overall tracking success at 66%, ahead of FQL (59%) and far ahead of MCL and AMCL (10% each). Tracking was strongest under dynamic occlusion (80%) and weakest under sparse semantics (52%).
  • Lowest translational error in most conditions: ShelfAware obtained the lowest translational RMSE in three out of four mock-store conditions.
  • Real-world open-vocabulary deployment: In the 3,500 sq. ft. store, purely geometric methods collapsed (MCL 8%, AMCL 6% global success). ShelfAware reached 62% global localization success versus 42% for FQL under real vision, and 40% tracking success versus FQL's 38%. ShelfAware needed 6.2 s to global convergence versus FQL's 2.9 s, which the authors attribute to evaluating and injecting proposals over a vastly larger candidate pose space.
  • Upper-bound semantics ablation: With ground-truth semantic vectors and the noisy depth likelihood disabled, ShelfAware reached 98% global localization success and 92% tracking stability, versus 82% and 70% for the FQL baseline. The authors state this proves the distributional semantic representation extracts more localization signal than fixed-quantity representations even under perfect perception.
  • Depth helps with real perception but hurts with perfect semantics: The ablation showed depth observations are necessary to stabilize noisy visual features in the "real" condition, but that incorporating depth actually degrades performance when semantic perception is perfect.
  • Real-time throughput: The full pipeline ran at 9.6 Hz on a standard consumer laptop with a mid-tier GPU.

Methodology in Plain English

Map building. The team first builds a standard 2D occupancy grid of traversable space at 10×10 cm resolution. On top of it they lay a semantic map with 20×20 cm cells in (x, y) and 30 cm in z. Each voxel stores a running distribution over which object classes were seen and how many, rather than a fixed landmark identity. Detections are projected into the map using RGB-D depth and camera calibration, then refined along the camera ray using Bresenham's line algorithm until they hit the nearest occupied cell.

Two perception pipelines. In the controlled mock store, a YOLOv9 detector fine-tuned on the SKU-110K dataset proposes generic "shelf item" boxes, and a ResNet50 classifier (trained on 5,699 environment-specific frames) assigns one of 14 classes. In the real store, there is no manual annotation: the same class-agnostic detector crops boxes, a pre-trained DINOv3 Vision Transformer extracts a 768-dimensional feature per crop, and K-Means clusters them into 60 categories chosen by grid search.

Comparing what is seen to what is expected. From a live frame the system forms a semantic observation vector with three parts: what classes are present, how far they are, and in which direction. Comparison uses a weighted composite score (α=0.4, β=0.4, γ=0.2): normalized class counts compared via Jensen–Shannon divergence (robust to missed or spurious detections), distances via 1/(1+d), and angles via 1 − (d/FOV).

The particle filter. Particles are propagated using VIO motion. Every iteration, the filter checks semantic consistency: if similarity against the expected view at the current pose estimate falls below a threshold while enough class mass is visible, the inverse model kicks in. A reverse index maps each class to the poses from which it is visible, so the system unions candidates over observed classes, scores them, injects particles at the top-k poses, and reweights the whole set with both depth and semantic likelihoods. Otherwise it updates with depth alone for efficiency. Expected semantic observations are precomputed offline for poses on 10 cm cells with 36 orientation bins (a 76 MB hashmap, plus a 2.1 MB reverse index). Depth likelihoods use a standard beam mixture with z_max = 6.0 m, σ_hit = 2.0, λ = 0.1, and weights 0.85/0.05/0.05/0.05, ray-cast with CDDT acceleration on CPU to leave the GPU free for perception.

Evaluation design. Hardware was an Intel RealSense D455 RGB-D camera (~103 g) and a RealSense T265 VIO camera (~60 g) in one 3D-printed mount, connected to a Dell G15 laptop (Intel Core i7-11800H at 2.3 GHz, NVIDIA RTX 3060 with 6 GB VRAM, 32 GB RAM), tested as a chest-mounted wearable and a cart-mounted setup. The mock store held 150 products in 14 categories across nine shelves and three aisles, with OptiTrack motion capture for ground truth; RGB-D ran at 30 Hz and VIO at 200 Hz. Four conditions were tested — Cart, Wearable, Dynamic Obstacles, and Sparse Semantics (25% and 50% of products removed) — with five ~40 s trajectories each (S1–S20), 20% of shelf contents perturbed between mapping and localization, and particles initialized uniformly over free space. All methods, including MCL, AMCL, and FQL, used 1,500 particles and identical odometry and RGB-D streams, with convergence defined as within 0.7 m translation and π/4 rad rotation of ground truth.

Why This Matters

ShelfAware attacks the assumption that objects are stable landmarks — an assumption that semantic SLAM and semantic visual positioning systems typically make and that fails in any environment with restocking, occlusion, or partial observability. By treating semantics as a probability distribution over object counts and arrangements, the method stays informative as inventory churns, which is a meaningful shift for semantic particle filtering research. It also demonstrates global localization without any installed infrastructure on consumer-grade sensors, rather than server-grade compute or LiDAR.

Real-world applications include:

  • Assistive navigation for people with visual impairments, where start-anytime operation supports on-demand, shared-control assistance and recovery after lost tracking, without RFID tags or Bluetooth beacons.
  • Retail service robots that must reach the correct aisle or shelf in stores where SKU-level inventory changes constantly.
  • Warehouse and logistics robots operating among repetitive racking with frequent restocking and moving people.
  • Mobile robots in offices and laboratories, where the global geometric layout is stable but local contents and clutter are not.

Industry relevance: The system runs at 9.6 Hz on a laptop with a mid-tier GPU using two small consumer cameras, and the open-vocabulary pipeline requires zero environment-specific training or manual annotation, positioning it as a scalable, infrastructure-free localization building block for on-demand robotic assistance.

Future Directions

  • Automated map maintenance. The authors state ShelfAware currently lacks an automated update mechanism; massive inventory reorganizations outstrip the inverse proposal's capacity and require rebuilding the semantic map from scratch.
  • Handling prolonged perception failure. Extended occlusions, motion blur, and severe category sparsity reduce the information mass in the semantic vector, delaying inverse proposals and weakening forward likelihoods.
  • Managing scaling costs. Offline precomputation of the inverse semantic cache scales linearly with the number of discrete states, which is a concern for large buildings.
  • Human-factors validation. Form-factor choices were informed by prior literature and informal feedback, but no user studies with people with visual impairments were run; the authors call for measuring usability and trust with participants, along with multi-hour battery tests and interaction design.
  • Integration into a full navigation stack. Combining ShelfAware with wayfinding, obstacle avoidance, and product-retrieval interfaces such as speech or haptics is described as a natural next step.

Target Audience

Robotics researchers and engineers working on localization, SLAM, and semantic scene understanding will find the core methodological contribution most directly useful. Practitioners building service robots, retail or warehouse automation, and assistive navigation systems will benefit from the low-cost, infrastructure-free sensor configuration and the two-pipeline evaluation design. The paper is also relevant to assistive-technology and human-centered robotics researchers interested in how localization enables shared-control and on-demand assistance. Readers without a background in particle filtering or probabilistic robotics will need to work through the observation-model equations and the MCL formulation.

Authors’ abstract

Many indoor workspaces are quasi-static: their global geometric layout is stable, but local semantics change continually, producing repetitive geometry, dynamic clutter, and perceptual noise that defeat standard vision-based localization. We present ShelfAware, a semantic particle filter for robust global localization that treats scene semantics as statistical evidence over object categories rather than fixed quantity landmarks. ShelfAware fuses a depth likelihood with a category-centric semantic similarity and uses a precomputed bank of semantic viewpoints to perform inverse semantic proposals inside Monte Carlo Localization (MCL), yielding fast, targeted hypothesis generation on low-cost, vision-only hardware. To demonstrate perception-agnostic scalability, we evaluate ShelfAware across two domains. In a rigorously controlled mock retail environment, ShelfAware achieves a 97% global localization success rate, maintaining the highest tracking success (66%) across cart, wearable, and dynamic occlusion conditions. Furthermore, in a 3,500 sq. ft. operational grocery store leveraging an open-vocabulary vision pipeline, ShelfAware significantly outperforms both geometric and fixed-quantity semantic baselines. By modeling semantics distributionally and leveraging inverse proposals, ShelfAware resolves geometric aliasing, providing an infrastructure-free building block for mobile and assistive robots in dynamic real-world environments.

Read the original paper