Guide

Cut lab to production scale-up risk

A formulation that works in the lab can fail at plant scale, and finding out costs a qualification batch. Multi-fidelity models learn from both sources at once: plentiful small-scale data for the shape of the response, and the few expensive full-scale runs for the correction, so you predict plant behavior before committing.

Why scale-up breaks models, what multi-fidelity changes, and what it asks of data you already have.

Plot titled Not All Data Costs the Same: five expensive lab tests marked accurate and scarce against thirty cheap simulations marked noisy and biased, both compared with the true behavior curve.
Accurate and scarce, or cheap and biased. Scale-up forces the same trade.
The failure mode

Why scale-up is where projects die

Every scale you move up changes mixing, heat transfer, and residence time, so the model that fitted the bench does not describe the reactor.

The formulation did not change, but the conditions it meets did, and effects that were negligible in a flask can dominate in a vessel. That is why a result which looked settled in the lab reopens at plant scale, and why the answer usually arrives in the form of a failed qualification batch.

It is more useful to treat bench, pilot, and plant as three different systems that happen to share a recipe, and to ask what each one is actually good for.

01 Cheap, plentiful

Bench

Fast and repeatable, so you can afford many runs. It gives you the shape of the response: which inputs matter, and in which direction.

02 Costly, informative

Pilot

Fewer runs, conditions closer to production. This is where the bench picture starts to bend, and where the size of the correction becomes visible.

03 Rare, decisive

Plant

Expensive, slow to schedule, and the only scale that answers the question you actually need answered. You get very few of these.

The method

What multi-fidelity means here

Rather than discarding cheap data, the model uses it for the overall trend and uses the handful of expensive runs to learn the offset between scales.

The cheap runs are treated as informative but biased, so the model keeps their structure and corrects their level. The expensive runs are treated as accurate but scarce, so they anchor that correction instead of carrying the whole fit on their own.

What comes out is one model across both sources, with a confidence range that reflects how much full-scale evidence stands behind each prediction. Multi-Fidelity Modeling goes through the mechanics.

Plot titled Multi-Fidelity, Use Everything: a DIM-GP fit that follows the true behavior curve, anchored by the expensive points and filled in by the cheap points, with an uncertainty band around it.
One fit over both sources, with the band showing where the full-scale evidence is thin.
Proof
1.8%

Predicted before the batch was run

In a chemical scenario STOCHOS predicted scale-up within about 1.8 percent of the first qualified plant batch, before that batch was run. An example from one process setup, not a guarantee for every scale-up.

The timing is what makes the number worth anything. A prediction that exists before the batch can still change a decision, where the same number afterwards is only a post-mortem.

What a model reaches on your process depends on your data and on how far apart your scales behave.

Adoption

The training set is already in your records

It uses the bench and pilot data you already generate, so the change is which run you choose next, not how you run it.

There is no new instrumentation and no separate data programme. The runs already sitting in your records are the training set, and each new result updates the model as the campaign proceeds.

Same data

Bench and pilot records you already keep, including runs from earlier campaigns.

Different choice

The model points at the run that would tell you the most, so pilot work gets aimed rather than replaced.

Same infrastructure

Training and prediction run locally, so formulation data does not leave your systems.

The wider picture for process work sits on AI for Chemical R&D.

Common questions

01How many full-scale runs do I need?

Few, which is the point. In multi-fidelity modeling the cheap lab and pilot data carry the trend and the few expensive full-scale runs carry the correction. The exact count depends on how far apart the scales behave.

02Does multi-fidelity modeling replace pilot plant work?

No. The model tells you which pilot runs are worth doing and what to expect from them, so pilot time goes into the runs that reduce the most risk.

03Can STOCHOS combine mixed data sources?

Yes. Combining sources of different cost and accuracy is what multi-fidelity modeling is for. Lab measurements, pilot data, and simulations can feed one model of the same process.

Related pages
Method

Multi-Fidelity Modeling

How cheap and expensive data sources combine into one model.

Solution

AI for Chemical R&D

Smarter experiments and scale-up for formulation and process work.

Product

STOCHOS

The predictive engine behind the multi-fidelity fit, run on your own hardware.

News and Guides collects the rest of the guides and project updates.

Next step

Bring us your next scale-up

Send the bench and pilot data you already have. We build the multi-fidelity model with STOCHOS and show what it predicts for the next scale.

Request a Demo How STOCHOS works

Local-first. Your data stays on your infrastructure.

Partners, customers, and research collaborators
AnsysCADFEMSimuTech GroupMEScoTSNEBoschZFGEMUDLRAdler LackeMankiewiczDuluxPlixxentFraunhoferHochschule NiederrheinFUELL Lab AutomationHumotionUniversitaet HamburgRobert Bosch Stiftung