Digital Twins & World Models: Establishing a Unified Framework for Dynamic Systems
Synopsis
This publication explores the confluence of digital twins and world models as technologies for understanding, simulating, and eventually reasoning about complex physical systems. Starting with construction, it examines how reality capture, BIM, live operational data, and fragmented project information can be reconciled into continuously updated digital representations capable of revealing relationships that conventional workflows miss. It then tracks the progression from these synchronized digital twins toward AI world models that do more than represent what is true now: they predict how environments may change, simulate counterfactual outcomes, and support planning.
Through examples ranging from early World Models research and Dreamer to Genie, JEPA, World Labs, and AMI Labs, the piece analyzes the shift toward these systems that combine simulation and prediction into increasingly capable models of the physical world, with implications for buildings, infrastructure, machinery, and the wider physical economy.
Digital Twins
A research building has a BSL-2 lab on the ninth floor. The lab holds its pressure two increments below the corridor outside it — enough that air only ever moves one direction: in.
It's a few pascals of differential, invisible to anyone standing in the hallway, and that's the whole reason the room is safe to work in.
Two floors above, an HVAC crew starts modifying ductwork for a new office buildout. It's undoubtably a small scope change. The kind that gets a single line in a project schedule.
Nobody on the crew has an obvious reason to think about a sealed lab two floors down, and nothing in their immediate scope of work tells them to.
The building does. Before a wall gets opened, the renovation's plans get run against everything the building's model knows about that laboratory: its required pressure cascade, the exhaust capacity supporting it, the relationships between shared air-handling components, and the differential-pressure readings recorded over time.
The system does more than compare the proposed design with a static threshold.
It combines the building model, equipment specifications, testing-and-balancing records, live sensor measurements, and historical operating data to estimate the laboratory's current state.
A simulation then tests how the proposed ductwork change would affect airflow under different operating conditions.
The analysis finds that the renovation would reduce the laboratory’s safety margin from comfortably stable to only a few pascals above failure. It is the one thing in the building that actually understands how two unrelated projects on two different floors are connected.
That's not fiction, and it's not quite standard practice either — not yet, not at most buildings. However it's the promise underneath one of the more prominent shifts happening in construction right now: buildings moving from static objects, documented once and then allowed to drift from that documentation, toward living systems that generate, interpret, and reconcile data continuously.
It's worth discerning precisely how that alteration works — and where it's still mostly hype, before getting into the more specific parts of this story. Because the gap between the proposal and the deployment is where the real insight derives.
Constructions Segmentation
Construction is one of the most divided project types in any industry.
Architects, structural and MEP engineers, the GC, a dozen subs; a single building passes through more disconnected hands, and more disconnected software, than almost any other asset gets built with.
Every party is tracking their own scope in their own system, on their own update cycle, and almost nothing oversees the whole.
What gets lost between those parties is specific, physical documentation. It's the redline that never makes it from the field set into the as-built, or the TAB report recording how the HVAC system actually performed at commissioning — measured airflow, operating conditions, and, where required, the room-to-room pressure differential actually achieved rather than simply designed.
The engineers who inherit the building did not witness the thousand small field decisions that shaped it, and are left reconstructing intent from documentation that may already have been outdated the day it was filed.
That's the kind of gap that let a renovation two floors up stealthily threaten a sealed lab downstairs.
Digital twins are an attempt to close that gap.
That wager is accelerating. Forrester Research has found that 55% of global software technology decision-makers are already adopting digital twins, applying them across everything from early design coordination to ongoing facilities operations.
An intriguing adoption number this early in a technology's life — and it's also, most certainly, a number that overstates how many of those deployments are doing anything close to what the term is supposed to mean. A large share of what gets called a digital twin is one half of the equation wearing the label of the whole.
Ask different people in the industry what a digital twin is and you'll get markedly different answers.
That ambiguity helps explain why the university scenario that opens this essay remains aspirational for most buildings. A static model cannot detect a pressure cascade deteriorating in real time; a live data feed, by itself, may register that a setpoint has changed without understanding why that change matters in a particular room.
Preventing the renovation from becoming a hazard requires both: a model of the building's physical and functional relationships, continuously reconciled with live operating data. That level of integration is still uncommon enough that, where it exists, it is usually the result of a deliberate owner-led strategy rather than an industry default.
A True Digital Twin
A digital twin isn't one thing, but three layers stacked on top of each other:
Layer one is reality capture.
Terrestrial LiDAR, photogrammetry, and drone scanning capture the building as it actually exists, rather than as it was originally drawn. That distinction matters because field conditions diverge from design documents throughout construction, and by handover even a carefully maintained drawing set may no longer reflect every installed condition. The industry expresses this uncertainty through standards such as the USIBD's Level of Accuracy framework, ranging from LOA 10 to LOA 50, with higher levels representing progressively tighter measurement tolerances.
If a capture is too coarse to establish a duct's true position, the digital model begins reasoning from the wrong geometry. In tightly coordinated MEP systems, even relatively small discrepancies can affect clash detection, maintenance clearances, airflow analysis, and the reliability of any simulation built on top of the model.
Layer two is BIM.
BIM elements can be developed to different Levels of Development (LOD), commonly expressed from LOD 100 through LOD 500, indicating how reliably an element's geometry, properties, and installed condition can be used at a given stage of the project.
At handover, asset information can be structured through standards such as COBie, while IFC provides an open schema for exchanging BIM data between different software platforms. The goal is interoperability: a duct modeled in one system should retain its identity, properties, and relationships when interpreted by another.
None of this is live. A BIM model, however detailed, is a photograph of the building on the day it was made — it had no way to dynamically acknowledge a renovation was underway two floors up.
Layer three is live data.
And this is the layer that begins to make the model behave like a true twin. A large research building can generate enormous streams of operational data — pressure, airflow, temperature, valve position, damper position, equipment status — but those signals rarely arrive with a common language for describing what they mean.
Building automation systems may communicate through protocols such as BACnet, Modbus, or LonWorks, yet interoperability at the network level does not guarantee interoperability at the semantic level. Different vendors and facilities use different point names, equipment hierarchies, and metadata conventions, leaving software to determine whether two differently labeled signals actually describe the same physical thing.
Standards such as Brick Schema and Project Haystack attempt to solve the operational side of that problem by giving sensors, equipment, spaces, and their relationships consistent machine-readable meaning, while IFC provides a parallel framework for exchanging BIM and built-asset information. The difficult part is connecting those worlds.
But adoption is slow, held back by legacy systems, inconsistent point naming across departments, and and the cost of mapping years of fragmented building data still make that integration slow and expensive.
Digital Twin's Hierarchy
Digital twins fall into three tiers, and almost nothing marketed as one actually reaches the top: a Level 1 system is a static BIM file, accurate only as of its last manual update; a Level 2 "digital shadow" adds automatic sensor data flowing in, but nothing flows back out — it only observes; and a Level 3 twin closes the loop, predicting, adjusting the real building, and learning from what actually happens.
Most real buildings sit at Level 1 or are inching toward Level 2 — plenty already track pressure and airflow, but a system that connects those readings to, say, a renovation permit filed two floors away, understands the implication, and raises a flag on its own is still rare.
No single standard yet defines a fully built-out twin, but two efforts matter: ISO 23247, built originally for manufacturing and increasingly borrowed by AEC, lays out a four-layer architecture: the physical asset, the sensing/control layer, the reasoning layer, and the user layer.
While the UK's Gemini Principles push a parallel idea for buildings and infrastructure specifically: a "digital thread" tying design, construction, and inspection history together so nothing gets lost.
The distinctions matter.
A digital shadow cannot tell you what would happen next if something changed.
That's a different capability entirely.
Prediction requires a model that runs forward: given a state and a hypothetical action, produce the state that follows.
It's the precondition for everything more ambitious that gets layered on afterward — simulation, forecasting, and eventually a system deciding something on its own. None of that is available to you until the model can be run forward.
That's also the capability AI research has spent the last several years trying to build, world models.
World Models Progression
A well-built digital twin can tell you what's true right now, synchronized in real time with the physical system it reflects.
What it can't do, even at its most advanced, is tell you what would become true next, if something changed.
That's a whole other league, and it's the one the AI research world has spent the better part of a decade chasing under a specific name: the world model.
Remove the branding and the concept is older than AI itself — AI as a discipline didn't formally begin until the Dartmouth workshop in 1956.
In 1943, Scottish psychologist Kenneth Craik proposed in his book The Nature of Explanation that people carry a "small-scale model" of reality in their heads specifically to test actions internally before taking them for real.
Early AI researchers tried to build the same thing in the 1960s, but those systems stayed confined to narrow, rule-bound worlds (symbolic AI), because nothing yet existed that could learn open-ended physical dynamics from experience.
Three developments changed that, and the first was the precondition for the other two. Deep learning gave systems a general-purpose way to learn rich representations directly from raw data: pixels, audio, sensor streams — rather than relying on features hand-engineered by a researcher.
That, in turn, served as a catalyst for modern reinforcement learning (which originated in the 1950s), existing as a mathematical framework for decades, only becoming capable of handling rich, high-dimensional environments once deep networks gave it something to learn from — the shift that turned "RL" into "deep RL" around the early 2010s, when DeepMind combined RL with deep neural networks to create Deep-Q-Networks.
Multimodal AI, grounding a system in how the physical world actually looks, sounds, and moves, likewise depended on deep learning's capacity to learn directly from raw sensory data instead of hand-built rules. We can largely thank the Transformer architecture for providing the unified mathematical framework capable for processing text, images, audio, and video.
The modern starting point is usually credited to Ha and Schmidhuber's 2018 paper, plainly titled World Models: an agent that compresses raw pixels into a latent representation of its environment, learns to predict how that representation evolves, and then learns to act by rolling out imagined trajectories inside its own learned model rather than the real one.
DeepMind's Dreamer family pushed the idea further, training almost entirely inside "imagination", or its very own world model. The concept here is that an intelligent agent doesn't need to experience every possibility in reality. It can learn a model of its environment and then practice inside its own internal simulator:
current state + proposed action > predicted next state + predicted reward
As a result simulating thousands or millions of experiences without executing all of them in the real environment.

Yann LeCun has since reframed this same ambition as arguably the central open problem in AI — not language modeling, but building a system that can predict the consequences of an action in the physical world competently enough to manoeuvre around it.
One models language. The other models consequence. That capacity — predicting outcomes for actions never explicitly observed, is what researchers call counterfactual reasoning, and it's the real dividing line between a world model and everything it gets casually lumped in with.
Upward counterfactual reasoning is for thinking what could've gone better, and downward counterfactual for how something could've turned out much worse.
Indeed, a world model may sound straightforward. However no one can yet agree on the details, not just what gets represented into the model, but at what fidelity?
What "understanding physics" currently looks like
A world model learns how an environment evolves by extracting regularities from experience rather than relying solely on explicitly programmed rules.
It does not need to be handed a rule stating that an unsupported object will fall or that a wall constrains movement. Instead, those relationships can emerge from training data: sequences of observations like video and, for action-conditioned models, records linking an agent's actions to the states that follow.
One test is quite simple: have an agent move an object, leave the scene, and return.
Is the object still where it was left?
Earlier generative systems often struggled with this kind of persistence, allowing objects or environmental details to drift, disappear, or reappear inconsistently once they left the visible frame. Newer world models have improved substantially. DeepMind's Genie 3, for example, can preserve previously observed environmental details across revisits and maintain largely consistent interactive worlds for several minutes. But long-horizon spatial memory and physical consistency remain unresolved problems: errors still accumulate over time, interaction remains constrained, and generated environments cannot yet be assumed to reproduce real-world dynamics reliably.
AI researcher and World Labs co-founder and CEO Fei-Fei Li, together with the World Labs team, offers a useful functional taxonomy for making sense of the field: renderers, simulators, and planners.
Renderers generate what a world looks like. Simulators represent how its geometry, physics, and dynamics behave. Planners determine what an agent should do within it. The categories are not rigidly separate — increasingly, the same systems span more than one, but they expose an important difference in maturity. Renderers are already commercially capable of producing impressive, interactive environments, while physically reliable simulation and general-purpose planning remain much harder problems.
The long-term objective is to collapse all three capabilities into a unified world model that can represent a world, predict how it will change, and reason about how to act within it.
What remains unresolved is generalization: building world models that can preserve physical consistency, predict consequences, and support useful planning across unfamiliar situations and over long time horizons. That problem is now being attacked from several directions by Google DeepMind, NVIDIA, Alibaba, Tencent, and specialist laboratories including World Labs, AMI Labs, and Odyssey. Their approaches differ; from spatial world generation, to abstract predictive representations, to real-time interactive simulation. But they converge on the same objective: giving AI systems models of the world reliable enough to reason about what happens next.
Two camps
This brings us to this: the field hasn't agreed on what should actually be inside a world model.
One camp treats world modelling as a rendering problem: if you can generate video realistic enough, you've effectively built a simulator. The previously mentioned DeepMind's Genie 3 is the clearest example.
Give it a text prompt and it generates a fully explorable 3D environment you can walk through in real time — at about 24 frames per second, holding together for a few minutes before it starts to drift. Water splashes, light falls across surfaces, objects hold their shape as you turn away and look back. None of that behavior was hand-coded. Genie 3 learned it purely from watching video.
DeepMind has been explicit about why this matters. A model that can predict how an environment evolves and how an agent's actions alter what happens next — can become a training ground for agents themselves, exposing them to an effectively unlimited curriculum of simulated situations before those situations ever have to occur in the physical world.
Waymo has already taken that idea into a specific real-world domain. Its Waymo World Model, built on top of Genie 3 and adapted for autonomous driving, generates controllable camera and LiDAR simulations of driving environments, including rare and safety-critical events that would be difficult, dangerous, or impractical to reproduce deliberately on public roads.
OpenAI approached the same frontier from a different direction. From Sora's original introduction, the company described video generation as part of a broader effort to teach models to understand and simulate the physical world in motion. Sora was still fundamentally a generative video model, not a reliable physics simulator, but the research underneath it was larger than just video creation.
Evidently, OpenAI discontinued Sora on April 26, 2026, with the API scheduled to close on September 24. Reporting at the time estimated that Sora was costing roughly $1 million per day to operate, though that figure was externally reported rather than publicly audited by the company.
What is more important is what happened to the research.
According to reporting on OpenAI's internal reorganization, the Sora research team shifted toward long-horizon world simulation for robotics. Sora research lead Bill Peebles described the longer-term prize as “automating the physical economy.”
The other camp thinks this whole approach is solving the wrong problem, and its most vocal figure is Yann LeCun — who left Meta specifically to found a new lab, AMI Labs, built around this disagreement.
His architecture, JEPA (Joint Embedding Predictive Architecture), throws the pixels away entirely. Where Genie 3 tries to predict what a scene will look like frame by frame, JEPA only tries to predict a compressed, abstract sketch of what will happen — something closer to "the ball's trajectory bends downward and slows" rather than a photorealistic rendering of that ball at frame 47.

The reasoning: a video frame is mostly clutter from a physics standpoint. Lighting, texture, the pattern on a far wall — all of it is expensive to render and almost entirely irrelevant to the causal question a planner actually needs answered.
Every bit of model capacity spent perfecting how a shadow falls is capacity that didn't go toward learning that objects fall, roll, and collide the way they do. LeCun's bet is that you get a more efficient, more genuinely predictive world model by throwing away exactly the part that makes Genie 3's demos so visually impressive.