DeploySci perspectives 6 min read

The problem: the virtual success that does not transfer

A model can look convincing in a demonstration and still respond to the wrong thing. An inspection system may associate a defect with a particular reflection. A physical simulation may favor a design because an interface was simplified. A control model may depend on timing or measurements that are more consistent in software than they are on the actual device.

This is the practical challenge of moving from simulation to reality: the environment used to develop the solution is an approximation of the environment in which it must work. More training data or a more complex model does not automatically close that gap. The development programme has to investigate which differences matter.

Why the gap matters beyond model accuracy

In industrial inspection, the consequences can include faults that escape review and acceptable parts that are rejected unnecessarily. In a process application, an inaccurate assumption can lead to an intervention that works only within a narrow operating range. For the customer, the issue is the reliability of the resulting decision and the effort needed to keep the system useful.

The impact also reaches the development team. If physical testing begins only after extensive virtual optimization, a newly discovered mismatch may require changes to the model, the data and the prototype together. An early transfer plan makes those dependencies visible while changes are still easier to make.

Use a virtual lab to test competing explanations

Simulation is valuable because it makes some variables easier to control independently. A virtual inspection scene can keep a fault fixed while changing illumination or viewpoint. A thermal model can compare geometries under the same boundary assumptions. These controlled comparisons help investigate whether a candidate responds to the mechanism that matters or to an incidental feature.

One established public approach, domain randomization, varies aspects of a simulated environment during training. Research by Tobin and colleagues explored this idea for transferring visual models from simulation to real robotic tasks. It is a useful example of designing for variation rather than relying on one idealized scene.

For a client application, variation still needs a purpose. Adding arbitrary complexity can consume effort without representing a meaningful operating condition. The virtual lab should support the specific scientific question: what needs to remain stable, what can change, and which failure would alter the design?

Public reference: Tobin et al.: Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World ↗

Build the simplest credible reference first

A new method needs something meaningful to improve upon. The reference could be the current inspection workflow, a straightforward image-processing approach, an existing scientific model or a simple statistical rule. This comparison helps establish whether additional complexity produces a useful advantage under the same conditions.

Advanced methods can then be explored where the problem justifies them. Scientific machine learning may combine observed data with physical structure. Synthetic data can expose a candidate to rare but relevant situations. Uncertainty-aware decisions can direct unfamiliar inputs to review. These are possible ingredients of a solution; their value depends on the application and the evidence.

Refine the proposition around a useful operating range

A simulation programme needs a value proposition that survives outside the virtual environment. Define the customer decision, the current alternative and the operating conditions in which an improvement would be useful. For inspection, the benefit may involve fewer missed faults and less unnecessary review together; improving one while making the other impractical may not support adoption.

When physical findings expose a mismatch, revisit both the method and the claim. A narrower first application can be valuable if its performance is better supported. The purpose of refinement is to identify a useful capability the evidence can substantiate, with a deliberate research plan for the limitations that prevent broader use.

Design the real-world evaluation before tuning the result

Identify which real observations will be used to develop the system and which will remain independent for evaluation. Samples that share a production run, physical part or acquisition setup may be more alike than their file names suggest. A useful evaluation therefore considers how the data were produced, not just how many records are available.

Compare performance across the variations that matter to the intended user. Depending on the problem, that can include different batches, material finishes, camera conditions, equipment states or operating periods. Report consequential failure modes alongside aggregate measures, and state where the evidence is too limited to support a conclusion.

The purpose is to learn the operating range of the prototype. A successful test on one setup should not silently become a claim about every setup. When a mismatch appears, use it to decide whether the model, the virtual environment, the sensing arrangement or the original requirement needs to change.

An example: reflective parts and rare surface faults

Consider a manufacturer inspecting reflective components. A surface fault may be subtle, while a harmless change in lighting produces a strong visual difference. Training only on convenient images can make the problem appear easier than it is in production.

A useful development route compares the same fault under different imaging conditions and acceptable surfaces under those conditions too. A small physical rig then checks whether the virtual assumptions reflect the actual camera, illumination and part presentation. Candidate methods can be judged against the same reference process and independent physical samples.

This illustrative case explains the role of computational discovery: create experiments that would otherwise be difficult to organize, expose fragile assumptions and guide a focused physical prototype. Each comparison should clarify what the next version needs to do better.

Transfer the whole decision into the operating environment

An edge deployment has to work with the available camera, compute, timing and operator workflow. The model is only one part of that system. Image acquisition, preprocessing, communication and the handling of exceptions can determine whether the overall decision arrives in time and can be acted on.

A staged introduction can begin with recorded evaluations and then a shadow trial alongside the existing process. The team can observe alerts and failures before those outputs control an operational action. Monitoring, a clear review route and a fallback keep the deployment connected to the evidence that justified it.

Use the physical lab to guide modernization

A focused bench prototype helps establish which parts of the real system need to change. A mismatch may originate in acquisition, material behaviour or an interface rather than in the computational model. Comparing those explanations can direct modernization towards the actual constraint and preserve equipment or software that already serves its purpose.

The realization strategy then follows the evidence: refine the prototype, evaluate it in a representative workflow and establish ownership of the resulting capability. This links laboratory learning to the deployment benefit without treating a successful virtual experiment as a complete implementation plan. Each step should reduce uncertainty about the decision the customer ultimately needs to make.

What a useful transfer package contains

The outcome should include the agreed prototype, a record of the virtual and physical evaluations, the operating conditions tested and the remaining limitations. A practical handover also identifies configuration, dependencies, responsibilities and the changes that should trigger another evaluation.

DeploySci treats the virtual lab as part of a connected development programme. The aim is to discover promising approaches quickly, build prototypes that expose the important unknowns and establish a defensible route into real-world use. The finish line is an application whose behavior can be understood and evaluated in its intended setting.

What would progress look like?

Tell us what needs to work better and where the technical uncertainty sits. We will help define a focused R&D, prototype or deployment engagement.

Get Started