Skip to content
AI.info

Future Horizons

The Invisible Infrastructure: How AI Is Rewriting the Physical World Through Spatial Computing and Digital Twins

Industrial digital twins are earning their keep while consumer headsets are not: BMW's virtual factory, Siemens and PepsiCo, and the 45,000 Vision Pros IDC counted in 2025 before Apple cancelled the successors.

The Invisible Infrastructure: How AI Is Rewriting the Physical World Through Spatial Computing and Digital Twins

Gabriele Masetti ·

Two Stories Wearing One Name

"Spatial computing" and "digital twins" get bundled together in press releases and conference keynotes as though they were a single wave of AI finally breaking into the physical world. They are not the same story, and conflating them obscures what is actually happening. One story — industrial digital twins built on platforms like NVIDIA Omniverse and Siemens Xcelerator — is a genuine, revenue-justified infrastructure shift that is quietly reshaping how factories, power grids, and supply chains get built and run.

The other story — consumer spatial computing, epitomized by the Apple Vision Pro — is a hardware category that has, so far, failed to find its market. Both use AI to blur the line between the physical and the digital. Only one of them is currently earning its keep at scale. The thesis of this piece is that the real transformation is happening in the unglamorous, invisible layer of enterprise infrastructure, not in the headset on your face, and that understanding why requires looking honestly at both the wins and the failures.

The Factory That Exists Twice

The clearest evidence that digital twins have moved from marketing concept to working infrastructure comes from BMW. The automaker has built what it calls a Virtual Factory using NVIDIA Omniverse and the OpenUSD (Universal Scene Description) framework, and it is not a demo — it is a production planning tool spanning over a million square meters of simulated factory floor, roughly the footprint of 140 football fields.

BMW's custom platform, FactoryExplorer, lets planning teams collaborate on layout, robotics, and logistics inside the simulated plant years before a single wall goes up. The company has publicly stated it expects the approach to cut production planning costs by up to 30 percent, and one specific workflow — collision detection between robots and equipment — has reportedly dropped from nearly four weeks of physical testing to about three days in simulation.

BMW first demonstrated this at NVIDIA's GTC conference with its future Debrecen, Hungary plant simulated more than two years before physical production began there, and the company has since said it is rolling the virtual-factory approach out globally.

BMW is not alone. NVIDIA has named Caterpillar, Foxconn, Lucid Motors, Toyota, TSMC, Wistron, and Belden among manufacturers building Omniverse-based factory digital twins, with Foxconn's Fii division citing thermal simulations run roughly 150 times faster than before through an integration with Cadence's engineering software. NVIDIA has packaged this pattern into reusable "blueprints" — pre-built reference workflows for building OpenUSD-based digital twins — with a blueprint called Mega aimed specifically at testing robot fleets in a simulated warehouse or factory before they touch the real one.

The company frames all of this under the label "physical AI": using simulation not just to visualize a facility but to generate the synthetic training data and test environment that robots and industrial AI systems need before deployment.

Siemens is running a parallel effort through Siemens Xcelerator, its digital-transformation software portfolio. At CES 2026, on 6 January, Siemens unveiled Digital Twin Composer, a tool for building large-scale "industrial metaverse" environments on NVIDIA Omniverse libraries, combining 3D visualisation, simulation and live factory data in one environment. Siemens describes it as "currently in early access with select customers" rather than generally available — a useful reminder that an unveiling and a shipping product are different events.

Siemens has also paired its digital twins with generative AI copilots aimed at the shop floor: an Industrial Copilot for Operations designed to run close to the machines themselves, giving maintenance engineers real-time decision support rather than a dashboard they have to interpret after the fact. PepsiCo has used Siemens' tools to convert selected U.S. manufacturing and warehouse facilities into high-fidelity 3D digital twins meant to model plant operations and supply-chain flow well enough to establish a measurable performance baseline before optimization begins. Siemens' CES release puts numbers on that partnership: a 20 percent increase in throughput on the initial deployment, 10 to 15 percent lower capital expenditure, and up to 90 percent of potential issues identified before any physical change was made. They are the supplier's figures, not an audited result.

Twins for Things That Cannot Be Prototyped

Factories can, at least in principle, be rebuilt. Power grids cannot. That makes utilities one of the more consequential — if less flashy — proving grounds for digital twin technology, because the alternative to simulation is often just waiting for a failure to happen and reacting to it. Xcel Energy, working with EY, built a digital twin that integrates with its core operational systems to let engineers identify overloaded transformers, monitor voltage irregularities, and improve regulatory reporting using near-real-time data rather than periodic manual inspection.

Southern California Edison has used AI-enabled digital twin modeling for vegetation and wildfire risk assessment, feeding directly into operational decisions about where to trim trees or de-energize lines rather than sitting as a standalone analytics exercise. Eversource has run an AI-driven outage-prevention pilot that the company has said avoided somewhere in the range of 40,000 customer disruptions.

Siemens' Gridscale X platform is pitched around automatically rerouting power around points of congestion, and Siemens Energy has built digital twins specifically for heat recovery steam generators aimed at predicting corrosion before it forces a shutdown.

Company Digital-twin application Reported result
BMW Virtual Factory production planning Up to 30% lower planning costs
Eversource AI-driven outage-prevention pilot ~40,000 customer disruptions avoided
Foxconn (Fii) Thermal simulation via Cadence integration ~150x faster

The common thread across the grid examples is narrower and more defensible than the industrial-metaverse rhetoric surrounding them: these are not attempts to build a photorealistic parallel world, they are attempts to fuse sensor data with a physics- and AI-informed model well enough to catch a failure a few days or weeks before it happens.

That is a real, measurable improvement over the status quo of scheduled inspection and reactive maintenance — but it is also a much more modest claim than "digital twin of the entire grid," and it is worth noticing that the utility examples that hold up best are the ones scoped to a specific asset class (a transformer fleet, a set of steam generators) rather than an entire operating company.

The Reconstruction Problem Nobody Solved Until Recently

Underneath both the factory twins and the grid twins sits a harder technical question: how do you actually get a real-world object or space into a usable 3D model in the first place? For years the honest answer was "slowly, and with specialized scanning hardware." Two AI techniques changed that calculus. Neural Radiance Fields (NeRF), introduced in 2020, showed that a neural network could learn a continuous 3D representation of a scene from a set of ordinary 2D photographs and then render novel viewpoints of that scene with startling photorealism. The catch was speed: early NeRFs could take hours to train and render.

3D Gaussian Splatting, which emerged in 2023, addressed exactly that bottleneck. Instead of querying a neural network at every point in space, it represents a scene as a large collection of 3D Gaussian "blobs" with position, shape, color, and transparency, which can be rasterized directly onto the screen — essentially treating 3D reconstruction as a much cheaper rendering problem instead of an expensive neural-inference one.

That makes it possible to go from a video walkthrough of a room or a production line to a manipulable 3D scene in a fraction of the time NeRF required, and to render it in real time rather than after a long offline render. By 2025, Gaussian Splatting had become the dominant technique for feeding AR, VR, and robotics applications with reconstructed real-world scenes, and researchers have extended it to handle moving objects and dynamic scenes, not just static ones.

That cheap reconstruction step is the unglamorous connective tissue that makes digital twins economically viable at scale: without a cheap way to turn a camera feed into a 3D model, every twin would still require the kind of manual CAD-modeling effort that made the concept a boutique exercise for the largest companies only.

The Headset That Was Supposed to Prove the Thesis

If industrial digital twins are the case for spatial-computing-as-infrastructure, Apple's Vision Pro was meant to be the case for spatial-computing-as-consumer-product — the thing that would put an AI-mediated, three-dimensional interface on every desk and in every living room. It has not worked out that way. Vision Pro's 2024 launch year sold in the range of 390,000 units, roughly in line with expectations for a first-generation, $3,499 device.

Sales then collapsed rather than built on that base. IDC's count for 2025 is 45,000 units for the whole year, down about 88 percent on 2024, and all of them shipped in the fourth quarter: Apple's manufacturing partner Luxshare halted production at the start of 2025, and the headset returned only with an M5 refresh in October. Apple cut Vision Pro advertising spend by more than 95 percent year on year across major markets, and the device was still sold directly in only 13 countries.

Then Apple stopped. On 3 June 2026 Ming-Chi Kuo reported that incoming chief executive John Ternus had cancelled both the Vision Pro successor and the lighter Vision Air, with resources moving to display-less AI glasses in 2027 and waveguide AR glasses in 2029 or later. Bloomberg's Mark Gurman gives a softer version — a Vision Pro 2 in testing, the category "on ice" rather than killed — so the two accounts should be read together. Either way, the product that was supposed to prove consumer spatial computing is no longer the plan.

Morgan Stanley analyst Erik Woodring's assessment — that cost, form factor, and a thin library of native visionOS apps are why the device never found a mainstream buyer — captures the general analyst consensus: Vision Pro impressed early adopters and reviewers, then had no second wave of demand behind them.

Period Apple Vision Pro shipments
2024 (launch year) ~390,000 units
2025 (full year, IDC) ~45,000 units
2026 Successors cancelled (Kuo, June 2026)

Meta's Quest line, the other major bet on consumer spatial computing, tells a related but distinct story. Meta still dominates the dedicated VR headset category, accounting for something like three-quarters to four-fifths of unit sales in that narrow segment even after the 2024 launch of the cheaper Quest 3S.

But Quest shipments kept falling: IDC put Meta at 1.7 million headsets over the first three quarters of 2025, 16 percent down on the same period a year earlier. What grew instead was the glasses. EssilorLuxottica's full-year results, published on 11 February 2026, reported more than 7 million smart glasses sold with Meta in 2025, more than triple the previous year. Meta's own capital allocation followed: with Reality Labs losing $4.03 billion in the first quarter of 2026, finance chief Susan Li told investors that VR investment would decrease significantly as spending moved to wearables.

That is a clear signal about where consumer appetite for "AI plus your field of view" actually sits: not in an enclosed headset that replaces your visual field, but in a lighter, cheaper, more socially normal form factor that adds a camera, a microphone and an AI assistant to glasses you would wear anyway.

The glasses have their own stumble, and it is a supply story rather than a demand one. On 6 January 2026 Meta paused shipments of the display-equipped Ray-Ban Display to the UK, France, Italy and Canada, citing extremely limited inventory and waitlists running well into the year, weeks after cutting Reality Labs spending by about 30 percent. Wanting the product and being able to buy it are not the same thing, and the category that is working is currently rationed.

Sensing the Physical World in Real Time

The third leg of this shift, alongside twins and headsets, is the sensor layer that keeps a digital twin synchronized with the object it represents — the IoT-plus-AI combination that turns a static 3D model into something that reflects current, not historical, reality. Rolls-Royce has built one of the more substantiated examples here, running continuous monitoring on more than 13,000 aircraft engines through its "IntelligentEngine" program on Microsoft Azure, using individualized models per engine rather than a single fleet-wide model, with the company citing extensions to time-between-maintenance-events as a result.

Manufacturing platforms such as PTC's ThingWorx and Siemens' Insights Hub follow a similar pattern: stream sensor data off physical equipment, run machine-learning anomaly detection against it, and feed the result back into a maintenance or operations decision rather than a static report. The pattern across these deployments is consistent — the value is not in the twin as a visualization but in the twin as a live, continuously-updated input to a decision that used to be made on a fixed schedule or after something already broke.

What the Hype Cycle Still Gets Right

None of this should be read as a case that digital twins are a solved, universally deployed technology. Gartner's own analysis is blunt about cost: building what it calls a "Level 3" digital twin — one sophisticated enough to reliably predict outcomes in a complex, dynamic system — can require an initial environment setup running into the millions of dollars, and Gartner has said outright that the expense is prohibitive for many organizations relative to what it delivers today.

Digital twins of customers, as opposed to physical assets, remain firmly in Gartner's "Innovation Trigger" phase — an admission that plenty of digital-twin marketing is running well ahead of working deployments. Market-size estimates for the sector vary by tens of billions of dollars depending on the research firm and what counts as a "digital twin," which is itself a sign of a category still being defined rather than mature.

And every twin that ingests live operational data inherits that data's security exposure: more sensors and more integration points mean more attack surface, and utilities in particular have to weigh a wildfire-risk model's value against the risk of that same model becoming a target.

The Split That Matters

The pattern across all of this research points to a specific and somewhat unfashionable conclusion. The AI-driven physical/digital merger is real, but it is happening asymmetrically. Where the return on a twin is measured in avoided downtime, reduced planning cost, or a specific engineering workflow that gets meaningfully cheaper — a collision check, a transformer inspection, an engine maintenance interval — enterprises are deploying the technology and reporting results that check out against public disclosures.

Where the return depends on a consumer strapping a headset to their face and finding enough daily value in it to justify the price and the discomfort, the technology has, twice now, underperformed its own launch expectations — and one of the two bets has now been withdrawn by the company that made it. The infrastructure BMW and Xcel Energy and Rolls-Royce are building is invisible by design — nobody buys a ticket to see a factory's OpenUSD model — and that invisibility is precisely why it is succeeding where the more visible, more hyped consumer hardware has not.

The next phase of this story is less likely to be a better headset than it is to be cheaper reconstruction (Gaussian Splatting and its successors), narrower and more disciplined twin deployments modeled on the utility examples rather than the "industrial metaverse" framing, and a growing recognition that the glasses people will actually wear look less like ski goggles and more like the ones already on their face — assuming the makers can keep them in stock.

Explore

More articles