Most of what gets sold as a digital twin is a dashboard with a 3D rendering attached, and the buyer usually works this out after the modelling budget is gone.

Sit through enough vendor presentations and you will see the same demonstration. A rotating three-dimensional model of a factory or a turbine or a building. Live sensor values overlaid on it. Temperatures updating, a gauge going amber, a red dot where something needs attention. It is genuinely useful and it looks impressive.

A digital twin is a specific thing, and the difference has engineering consequences, which is why it is worth being pedantic about.

What the term was supposed to mean

The idea predates the phrase. When Apollo 13 was in trouble, NASA had a physical duplicate of the spacecraft on the ground, and engineers used it to work out procedures before sending them up. The ground copy was not a display of telemetry. It was a system you could try things on.

Michael Grieves introduced the concept in a product lifecycle management context around 2002, and NASA’s 2010 technology roadmap gave the definition that most later ones descend from: an integrated multiphysics, multiscale simulation of a vehicle that uses the best available physical models and sensor data to mirror the life of its flying twin.

The load-bearing words are simulation and models. Not visualisation. Not monitoring.

The four questions

Here is the test I apply. It is not standard, but every part of it is checkable, which is the point.

Is there a model that produces state you did not measure?

The defining capability is inference. A thermal model that gives you the temperature at a point where no sensor exists. A structural model that gives you accumulated fatigue at a weld from load history. If every number on the screen traces back to a sensor reading, there is no model, and no amount of rendering changes that.

Does data flow both ways?

A twin ingests measurements and updates its own state or parameters. A model that was calibrated once at commissioning and has run open-loop since is a simulation with a live display bolted on. The coupling is what makes it track the specific physical asset rather than the design intent.

This is where most implementations quietly fail. Building the pipeline from sensors into a model that assimilates them, continuously, in production, is hard. Building a pipeline from sensors into a rendering is easy.

Has anyone checked whether it is right?

The question I ask that produces the longest silence.

If the twin makes predictions, someone can compare them to what happened. What is the error distribution? Under what operating conditions does it degrade? When was it last revalidated? A model that has never been scored against outcomes is an opinion with a rendering budget.

Validation is expensive and unglamorous and it is the entire difference between a model you can act on and a model you can look at.

Can you run a counterfactual?

Can you ask what happens if the load increases forty percent, or if this pump fails at 3am, and get an answer that is not simply a replay of a similar past event? The Apollo 13 ground copy was valuable because you could try things on it that had not happened.

If the system can only tell you what is happening now and what happened before, it is monitoring.

Why the confusion is commercially convenient

The word carries a premium. Procurement processes and government programmes increasingly ask for digital twin capability by name, which creates demand for the label independent of the capability.

Meanwhile the visualisation layer is where the demonstrable value is early. It is quick to build, it looks like progress in a steering committee, and it does deliver real benefit, because plenty of industrial operations genuinely lack a single live view of their asset. Nobody is defrauded exactly. The label just drifts to cover whatever was actually shipped.

There is a definitional problem underneath too. There is no widely enforced standard for what qualifies. ISO 23247 provides a reference architecture for manufacturing, and it is useful, but a reference architecture describes structure rather than setting a bar for predictive validity. So the term stays elastic.

Why it matters beyond terminology

If it were only a naming argument I would not care.

The consequence is that organisations believe they have a capability they do not have. A team told they have a digital twin of a production line will assume they can evaluate a process change before making it. When they try, they discover the system can display the line and cannot simulate it, and by then the modelling budget has been spent on integration and rendering.

The second consequence is that it damages the credibility of the real thing. Multiphysics models coupled to live data with proper uncertainty quantification are genuinely powerful and genuinely hard. When the term gets attached to dashboards, the engineering leader who was burned by a dashboard will not fund the real project.

The third is that the hard part gets underfunded. A real twin needs a validated physics model, which needs someone who understands the physics, the numerics and the asset. That person is expensive and scarce. A visualisation needs front-end developers, who are neither. Budgets flow along the path of available talent, and the modelling work gets deferred to a phase two that does not arrive.

If you are buying one

Ask for the model. Ask what governing equations it solves and what it assumes. If the answer is about the platform and the connectors and the data lake, you are buying an integration project.

Ask for the validation report. Prediction against measurement, with error bars, on an asset like yours. Not a case study. A report.

Ask what it predicts that you cannot currently measure, and what you would do differently if you had that number. If nobody can answer the second half, the project has no decision attached to it and will not survive its first budget review.

Ask what happens when the asset changes. Physical systems get modified, and a model calibrated to last year’s configuration is wrong in a way that is hard to notice, because it keeps producing plausible numbers.

None of this rules out buying the dashboard. Live visibility into an operation you could not see before is worth paying for. Buy it under its own name, at its own price, and keep the modelling budget separate.

Related: