Benchmark report

STOCHOS DIM-GP: Benchmark Results Across Tabular Learning, Physics Simulation and Optimization

Probabilistic AI for engineering, evaluated on workstation hardware

PI Probaligence GmbH · STOCHOS · 7 September 2026

Request the PDF Request a Demo
14evaluated tasks and studies
12top-three placements
public and internal comparisons
203benchmark entries outranked
across 12 ranked comparisons
< 8 GBDIM-GP GPU-memory budget
≤ 1consumer GPU per run
< 1 dayreported DIM-GP model training

Overview

Engineering teams work with material measurements, design parameters, simulation fields and physical tests. Turning that information into a better design requires several connected steps: predict performance, understand which inputs matter, compare alternatives and choose the next simulation or experiment. The cost of each evaluation limits how many designs a team can explore.

STOCHOS brings probabilistic modeling, sensitivity analysis and Bayesian optimization into one engineering environment. Its DIM-GP model family covers tabular regression and classification, geometry-based field prediction and transient simulation. Predictions include uncertainty estimates that can inform both design assessment and the selection of the next evaluation.

Why we take a workstation-focused approach

STOCHOS supports foundation models and agentic AI workflows within a platform designed for affordable workstation hardware. STOCHOS Flow makes modeling and optimization accessible through visual workflows, connects to existing simulation tools and Python, and supports sharing trained models through prediction applications.

Our focus is the combination: a probabilistic modeling platform that can serve materials, simulation and design teams, incorporate their physical knowledge, and operate within their own computing environment.

Why these benchmarks matter

A benchmark gives competing methods a common problem, evaluation data and scoring rules. We evaluated DIM-GP across fourteen benchmark tasks and studies, covering tabular learning, fluid dynamics, structural mechanics and optimization. The series combines externally scored competitions with evaluations against published baselines and a virtual formulation study. We publish the results with named comparisons, evaluation conditions and computing requirements so customers can assess STOCHOS’s capabilities across a broad range of tasks.

Selected results

The strongest results span several application classes. In the PLAID snapshots of 4 September 2026 (7 September for 2D_profile), DIM-GP ranks first on four of the six boards examined when the standalone reconstructed-solver entry is excluded. This excludes the PhysicsX simulator on Tensile2d [44]. With all entries included, DIM-GP ranks first on three of six boards and second on Tensile2d. The comparison field includes NPco from NP Company and PXTransolver from PhysicsX.

On AhmedML and DrivAerML, DIM-GP has the lowest reported errors among the ten methods listed, including Emmi AI’s AB-UPT [40], across all five AhmedML channels and both DrivAerML surface channels.

On LagrangeBench, DIM-GP improves on the best of three published baselines in 19 of 21 dataset–metric comparisons. The baselines include Google DeepMind’s Graph Network-based Simulator (GNS) [34]. On DeformingPlate, DIM-GP achieves 57% lower RMSE than Google DeepMind’s MeshGraphNets [48] and ranks second among the three methods compared under the canonical protocol (§9).

In our reconstruction of the TabArena-Lite comparison, DIM-GP ranks third among 80 model entries, or sixth among 85 entries including AutoML systems, in an internal comparison using published results (§4). The field includes TabPFN-3 from Prior Labs, acquired by SAP in July 2026. On TALENT-300, it has the lowest average rank in each task type among the four models in our additional comparison. In a virtual concrete-formulation study based on public laboratory data, STOCHOS meets the modeled specification in 149 of 150 runs, compared with 137 for BoTorch’s qNEHVI.

The reported DIM-GP models train in less than one day on a local workstation, using at most one consumer GPU per run and a GPU-memory budget below 8 GB. CPU-only configurations are identified in the benchmark sections.

The scorecard

Bold type highlights DIM-GP results and selected summary findings. Error metrics are lower-is-better unless otherwise stated.

Benchmark DIM-GP result Field Interpretation
PLAID VKI-LS59 (§13) 1st 21 entries 7.9% lower aggregate error than PXTransolver.
PLAID Hyperelasticity (§14) 1st 22 entries Approximately 6% lower aggregate error than PXTransolver.
PLAID ElastoPlastoDynamics (§16) 1st 8 entries Lowest errors on both displacement channels.
PLAID Tensile2d (§15) 2nd 22 entries Second overall at an aggregate error of 0.0007.
PLAID Rotor37 (§12) 3rd 16 entries Aggregate error approximately 0.00049 versus 0.00043 for the leader: a gap of about 0.00006.
PLAID 2D_profile (§17) 3rd 22 entries Aggregate error 0.0152 versus the leader’s 0.0134: a gap of 0.0018.
TabArena-51 (§4) 3rd (internal) 80 entries Internal reconstruction; sixth of 85 with AutoML systems.
TALENT-300 (§5) 1st 4 models Lowest average rank on common completed subsets.
AhmedML (§10) 1st 10 methods Lowest error on all five channels; 6.7 h training.
DrivAerML (§11) 1st 10 methods Lowest error on both surface channels; under 15 h training.
LagrangeBench (§7) 19 of 21 cells 3 baselines Compared with the per-cell best of three methods.
DeepMind Water-3D (§8) 8.5% higher MSE GNS Compared with published GNS mean; fewer gradient updates.
DeepMind DeformingPlate (§9) 2nd 3 methods 57% lower RMSE than MeshGraphNets; M4GN reports lower RMSE.
Bayesian optimization (§18) 1st in success count 6 approaches Highest success count in the virtual study: 149/150 runs.
Swipe to explore figure →
1 2 3 5 10 20 50 80 placement within the field (log scale); the bar is the size of the field PLAID VKI-LS59 PLAID Hyperelasticity PLAID ElastoPlastoDynamics TALENT-300 LagrangeBench Bayesian optimization PLAID Tensile2d AhmedML DrivAerML DeformingPlate PLAID Rotor37 TabArena-51 (internal) DeepMind Water-3D PLAID 2D_profile 1st of 21 1st of 22 1st of 8 1st of 4 19/21 metric cells improved 1st of 6 2nd of 22 1st of 10 1st of 10 2nd of 3 3rd of 16 3rd of 80 8.5% higher mean MSE than GNS 3rd of 22

The scorecard counts each tabular suite once, each PLAID board separately, and the other studies individually. Across its 12 ranked comparisons, DIM-GP outranks 203 competing benchmark entries. Each competing entry counts once per benchmark comparison.

A transient benchmark example

This animation shows a held-out PLAID 2D ElastoPlastoDynamics test design (§16), comparing the reference simulation with DIM-GP's predicted 41-step transient response.

Reference simulationDIM-GP prediction
Perforated steel plate pulled to rupture, Held-out design 62526,738 nodes · 52,772 triangles · 41 steps of 1 ms · RRMSE U_x 0.0018, U_y 0.0021 (board metric)

PLAID positions use the archived 4 September 2026 snapshots, except for 2D_profile, which uses the 7 September snapshot. Automotive results are reported as of 7 September 2026.

Computational requirements

One DrivAerML reference simulation requires approximately 61,000 CPU core-hours according to the dataset’s reported setup (§11). Once reference data are available, a trained surrogate can evaluate additional geometries at substantially lower computational cost. Figure 2 reports the training resources for the automotive comparisons.

Swipe to explore figure →
AhmedML DrivAerML 0 20 40 60 80 training hours on one RTX 4090 6.7 h ~77 h <15 h ~67 h* Training on an RTX 4090 DIM-GP AB-UPT (see timing notes) *DrivAerML AB-UPT: schedule estimate from measured update time; surface-only timing differs from the published accuracy configuration.

Scope of the comparisons

The comparison includes named research methods and submissions from commercial vendors. On the five PLAID boards with entries from all three, DIM-GP has lower aggregate error than PXTransolver and NPco on VKI-LS59, Hyperelasticity and Tensile2d, and higher error on Rotor37 and 2D_profile. The separate PhysicsX simulator entry is reported in §15 [44].

Swipe to explore figure →
VKI-LS59 Hyperelasticity Tensile2d Rotor37 2D_profile 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 board error x 1000, lower is better 9.3 10.1 11.2 10.1 10.7 16.5 0.7 1.3 1.3 0.5 0.5 0.4 15.2 13.4 13.8 Named entries in the archived PLAID comparison DIM-GP PXTransolver (PhysicsX) NPco (NP Company)

1. STOCHOS in the engineering workflow

1.1 What DIM-GP is

DIM-GP stands for Deep Infinite Mixture of Gaussian Processes, the probabilistic model family within STOCHOS. It supports predictions from tables, particle states and unstructured meshes. An engineer can use these models to estimate a material property, predict an aerodynamic field or model a transient response, with uncertainty estimates accompanying the predictions.

The tabular evaluation uses a fixed configuration across the 51 TabArena and 300 TALENT datasets, including validation-based adaptation. Physics evaluations use task-appropriate configurations. The benchmark sections report the results of these configured workflows.

1.2 Incorporating physics into DIM-GP

Swipe to explore figure →
FROM ENGINEERING DATA TO DESIGN DECISIONS STOCHOS capabilities at a glance YOUR INPUTS YOUR APPLICATIONS Measurements & tables Observed data Geometry & simulation Meshes, fields and trajectories Engineering knowledge Optional priors and constraints STOCHOS DIM-GP Learn from your data Incorporate physical information Automate supported features Predict Properties, fields and dynamics Assess uncertainty Probabilistic predictions Guide optimization Select the next design or experiment

DIM-GP can incorporate inexpensive physical priors, post-training conditioning on known relationships, and geometric or physical features that help it learn from engineering data. STOCHOS calculates supported geometric features automatically during mesh processing; engineers can supply application-specific physical priors and constraints. Multi-fidelity modeling also allows lower-cost simulations and higher-fidelity simulations or measurements to contribute to the same modeling workflow.

1.3 From models to engineering decisions

STOCHOS Flow connects data import, model fitting, validation, sensitivity analysis and optimization in a visual workflow. Sensitivity analysis identifies influential inputs; Bayesian optimization uses predictions and uncertainty to propose the next design or experiment. Engineers can connect Ansys Workbench or Python solvers and export prediction applications for colleagues to use. Agentic AI assists with workflow creation and troubleshooting, with local model options available.

The benchmark series evaluates the prediction and optimization methods that support this workflow.

1.4 How the results are evaluated

PLAID provides externally scored evidence. Participants receive test inputs, such as geometries and operating conditions, and submit prediction files while the reference test outputs remain withheld. The platform scores the predictions and uses hidden test subsets to discourage overfitting (PLAID benchmark protocol, §4.3 and Appendix C). This allows STOCHOS to be evaluated without distributing its proprietary implementation.

The automotive and tabular sections compare our evaluations with published reference results; §4.2 explains the submission requirements behind the internal TabArena comparison. The formulation study uses a virtual test bench based on public laboratory data. Each section identifies its comparison field, metrics and evaluation conditions.


2. Tabular benchmarks and evaluation protocol

TabArena-51 TALENT-300
Datasets 51 curated, real-world 300
Protocol TabArena-Lite (single split, fold 0) single seed, fixed splits
Metrics roc_auc (30 sets), log_loss (8), rmse (13) RMSE / accuracy
Field 79 model variants + 5 AutoML systems 28–31 published methods (official) + 3 foundation models we added
Hardware 1× RTX 4090; DIM-GP <8 GB same

Hardware. The experiments used an Intel Core i9-13900KS workstation (24 cores, 32 threads), 64 GB of DDR5-4200 memory and NVIDIA GeForce RTX 4090 graphics under Windows 11 Pro. The workstation contains two cards, used for separate experiments; each DIM-GP run uses at most one GPU, with GPU memory below 8 GB. Published baseline timings identify their own hardware.

Our TabArena comparison uses fold 0 under the official TabArena-Lite protocol and the corresponding published Lite results. TALENT uses its fixed splits.

2.0 Definitions

Let be the set of methods, the set of datasets, and the error of method on dataset under that dataset’s own metric — one minus ROC-AUC on the 30 binary sets, log-loss on the 8 multiclass sets, RMSE on the 13 regression sets. Lower is better throughout. Because these three quantities share no scale, no statistic below averages across datasets directly.

Elo. TabArena fits a Bradley–Terry model [29] to pairwise wins across datasets. The model assigns method a probability of beating method :

Reported Elo is the median over 200 task-level bootstrap rounds, with confidence bounds at the 2.5% and 97.5% quantiles. Elo reflects who wins; normalized error also reflects the size of error differences. Section 4.2 reports both, together with mean rank.

Normalized error. Per dataset, the error is rescaled against the best and the median method on that dataset, then clipped:

Here, denotes the best observed method and the median or a worse result. Clipping limits the influence of a single unusually large error on the aggregate.

Headroom (§2.1) measures how much a dataset separates methods at all, using the same two quantities as that denominator:

means the median error approaches the best observed error. Larger values indicate greater separation between the best and median.

Cost. Training and inference costs normalize the complete fit and prediction pass by the number of rows, then take the median across datasets:

These measures include fixed per-dataset overhead. Section 4.5 estimates fixed and per-row contributions with a robust Theil–Sen fit [30], .

2.1 What is in TabArena-51

TabArena contains 51 curated real-world datasets [1], checked for leakage, duplication and mislabelled task types. Each dataset has a fixed evaluation metric (Figure 4).

Swipe to explore figure →
binary (roc_auc) multiclass (log_loss) regression (rmse) 0 5 10 15 20 25 30 35 datasets 30 8 13 task type and metric 10 3 10 4 10 5 rows 0 1 2 3 4 5 6 7 size: 748–150,000 median 6,497 10 0 10 1 10 2 10 3 features 0 2 4 6 8 10 12 14 features: 4–1776 median 17 TabArena-51
Task Datasets Metric Rows (median) Rows (range) Features (median)
Binary 30 roc_auc 9,911 748 – 150,000 21
Multiclass 8 log_loss 2,445 898 – 78,053 37
Regression 13 rmse 6,497 907 – 53,940 9

The suite spans 748 to 150,000 rows and 4 to 1,776 features. Binary classification accounts for 30 of the 51 datasets and uses ROC-AUC, a measure of predicted ordering.

Where the data comes from, and how hard it is

TabArena covers commercial and scientific applications; business, marketing and finance account for nearly half the datasets. Figure 5 and the table below summarize their domains and the median-to-best error gaps defined in §2.0.

Swipe to explore figure →
0 5 10 15 datasets industry & manufacturing environmental science & climate education physics & astronomy technology & internet biology & life sciences chemistry & material science medical & healthcare finance business & marketing 1 1 1 3 4 5 6 6 8 16 application field (TabArena's own labelling) 10 3 10 4 10 5 training rows 10 1 10 2 10 3 features features = rows shape of the tasks binary multiclass regression 0 20 40 60 80 headroom: how far the best method beats the median method (%) Bioresponse diabetes APSFailure how much the choice of method matters
Field Datasets Training rows Median headroom
Business & marketing 16 1,000 – 86,586 12.3 %
Finance 8 666 – 100,000 5.8 %
Chemistry & material science 6 598 – 30,486 10.5 %
Medical & healthcare 6 498 – 47,678 10.5 %
Biology & life sciences 5 604 – 2,500 15.0 %
Technology & internet 4 902 – 7,256 15.3 %
Physics & astronomy 3 1,002 – 52,035 24.9 %
Industry & manufacturing 1 50,666 51.9 %
Environmental science & climate 1 1,722 9.1 %
Education 1 2,949 6.9 %

The suite includes imbalanced industrial and credit tasks, as well as categorical inputs. These characteristics make the stated metric and consistent preprocessing important to the comparison.

2.2 What is in TALENT-300

TALENT provides 300 datasets [2]: 119 regression, 101 binary and 80 multiclass, spanning 252 to 634,460 rows (median 3,389) and 3 to 970 features (median 16), across 15 application domains — finance and economics (53 datasets), healthcare (31), image and signal (28), industrial and sensors (27), biology and genomics (21), and others.

Swipe to explore figure →
binary multiclass regression 0 20 40 60 80 100 120 140 datasets 101 80 119 task type 10 3 10 4 10 5 rows 0 10 20 30 40 size: 252–634,460 median 3,389 0 20 40 datasets Web, Social Chemistry Biology Other Industrial Image Healthcare Finance application domains (top 8 of 15) TALENT-300

TabArena emphasizes curated comparisons against an actively maintained field. TALENT extends the evaluation across more datasets and application domains.


3. Tabular model families and competitive context

3.1 The three families in the field

The comparison includes three established families: gradient-boosted trees, neural networks trained per dataset and pretrained tabular models.

Gradient-boosted decision trees — LightGBM [18], CatBoost [19] and XGBoost [20] — are widely used tabular methods that fit per dataset and support CPU execution. The comparison includes both default and tuned configurations.

Neural networks for tabular data include RealMLP [21], TabM [22], ModernNCA [23] and xRFM [24]. RealMLP uses defaults developed across multiple datasets and is the highest-ranked non-pretrained entry in this comparison, at tenth place.

Tabular foundation models learn across tasks before adapting to a new dataset. The field includes TabPFN [5, 6] and its variants, TabICL [7], TabFM [3], EXAONE-Tabular [4], TabDPT [9], LimiX [8], Mitra [10] and other published models [11–17]. TabFM and EXAONE-Tabular lead the TabArena comparison presented here.

3.2 Parameter counts and interpretation

The released configurations illustrate the range of model sizes: TabFM [3] contains 1,639,444,298 parameters (1.64 billion) and EXAONE-Tabular [4] contains 20,807,434. Their TabArena-Lite Elo ratings in this comparison are 1793 and 1749, respectively.

Swipe to explore figure →
10 7 10 8 10 9 parameters (log scale) 1600 1650 1700 1750 1800 1850 Elo TabFM 1639.4 M EXAONE-Tabular 20.8 M Published baseline model sizes and Elo

4. TabArena-51

DIM-GP ranks third among 80 model entries by Elo, mean rank and normalized error in our internal reconstruction of TabArena-Lite. Including AutoML systems places it sixth among 85 entries. The following sections describe the scoring check, comparison field and computational cost.

Hover for each entry's score and training time. The amber point is DIM-GP in the internal comparison of §4; hardware differences are discussed in §4.5.

4.1 Scoring validation

Our comparison was rebuilt from TabArena’s published per-method results and compared against the official splits_lite CSVs:

Method official (lite) ours Δ
TabFM+ 1836 1840.9 +4.9
AutoGluon 1.6 (noncommercial) 1831 1832.5 +1.5
TabFM (default) 1796 1799.8 +3.8
AutoGluon 1.6 (extreme) 1757 1757.9 +0.9
EXAONE-Tabular (default) 1746 1744.7 −1.3
AutoGluon 1.5 (extreme) 1646 1645.1 −0.9
TabPFN-3 (default) 1639 1638.5 −0.5
TabPFN-2.6 (default) 1604 1603.2 −0.8
TabICLv2 (default) 1555 1555.7 +0.7

The selected reference ratings are reproduced within 4.9 Elo points. The official comparison data were obtained from the leaderboard Space’s published results.

4.2 Standing — internal

Comparison status. TabArena’s submission protocol requires executable training and evaluation code integrated into its benchmark. STOCHOS is licensed commercial software and was not provided through that integration route. Its results are therefore reported as an internal comparison with TabArena’s published results.

Swipe to explore figure →
1300 1400 1500 1600 1700 1800 Elo — TabArena-51, models field, TabArena-Lite protocol CatBoost (tuned) LightGBM (tuned + ensembled) TabM (tuned + ensembled) TabDPT-Turbo (default) RealMLP (tuned) TabDPT (tuned + ensembled) RealMLP (tuned + ensembled) RealTabPFN-2.5 (default) RealTabPFN-2.5 (tuned) TabICLv2 (default) RealTabPFN-2.5 (tuned + ensembled) TabPFN-2.6 (default) TabPFN-3 (default) DIM-GP (ours) EXAONE-Tabular (default) TabFM (default) 1391 1396 1401 1413 1416 1423 1490 1523 1547 1555 1600 1613 1644 1652 1749 1793

Selected positions from the models field (AutoML systems excluded):

# Model Overall Class. Regr. Binary Multi. Small Medium
1 TabFM 1793 1752 2168 1780 1699 1785 1911
2 EXAONE-Tabular 1749 1747 1919 1756 1752 1698 2021
3 DIM-GP 1652 1635 1861 1654 1606 1649 1751
4 TabPFN-3 1644 1617 1897 1634 1594 1624 1789
5 TabPFN-2.6 1613 1590 1847 1585 1645 1596 1747
6 RealTabPFN-2.5 (t+e) 1600 1573 1849 1547 1745 1595 1700
7 TabICLv2 1555 1554 1696 1573 1520 1536 1691
10 RealMLP (t+e) 1490 1475 1666 1489 1451 1476 1608
16 CatBoost (tuned) 1391 1388 1494 1378 1457 1354 1561

All entries are (default) configurations unless marked (t+e) = tuned + ensembled. Rows 1–7 are foundation models, 10 a neural network, 16 tree-based.

With systems included, DIM-GP is 6th of 85 at 1647, just above AutoGluon 1.5 extreme (1646) and TabPFN-3 (1639); AutoGluon 1.6 (1757/1831) and TabFM+ (1836) rank above.

Agreement across ranking measures

Mean rank averages each method’s position across datasets; normalized error also reflects the size of error differences (§2.0). Both give DIM-GP the same placement as Elo:

# Model Elo Mean rank Normalized error Win rate
1 TabFM 1793 6.18 0.1860 0.934
2 EXAONE-Tabular 1749 7.66 0.2501 0.916
3 DIM-GP 1652 11.27 0.3902 0.870
4 TabPFN-3 1644 11.66 0.3909 0.865
5 TabPFN-2.6 1613 13.18 0.4384 0.846
6 RealTabPFN-2.5 (t+e) 1600 13.83 0.4483 0.838
7 TabICLv2 1555 16.24 0.4728 0.807

DIM-GP is third under all three measures, with the same ordering of the top nine entries. The margin to TabPFN-3 is small: approximately eight Elo points and normalized errors of 0.3902 versus 0.3909. The paired analysis below assesses the differences across datasets.

Split by task type, the picture matches §4.3:

Place Normalized error Directly ahead
Classification (38 datasets) 3 of 78 0.4234 EXAONE 0.2543
Regression, rmse (13 datasets) 5 of 77 0.3256 TabPFN-3 0.2966, Nori-30M 0.3183

Against TabPFN-3, DIM-GP has lower normalized error on classification (0.4234 versus 0.4334) and higher error on regression. TabArena contains 13 regression datasets; the complementary TALENT comparison covers 116 (§5.2).

Statistical comparisons

The paired analysis compares methods on the same 51 datasets using a Wilcoxon signed-rank test [31] on normalized-error differences and an exact sign test on wins and losses. Negative differences favour DIM-GP.

Rival W–L Mean Δ Wilcoxon Sign Conclusion
LightGBM (tuned + ens.) 44–7 −0.353 <0.001 <0.001 DIM-GP better
CatBoost (tuned + ens.) 39–12 −0.338 <0.001 0.0002 DIM-GP better
LimiX 40–10 −0.283 <0.001 <0.001 DIM-GP better
TabICLv2 30–20 −0.076 0.059 0.203 no significant difference detected
TabPFN-2.6 32–19 −0.045 0.124 0.092 no significant difference detected
TabPFN-3 27–23 +0.001 0.730 0.672 not distinguishable
EXAONE-Tabular 22–29 +0.140 0.011 0.401 mixed
TabFM 9–41 +0.204 <0.001 <0.001 TabFM better

A Friedman test over thirteen leading methods rejects equal performance ( , ). The Nemenyi post-hoc comparison [32] uses a critical distance of 2.56 average-rank points at (Figure 9).

Swipe to explore figure →
Figure 9. Critical-difference diagram (Friedman test, Nemenyi post-hoc, \alpha = 0.05) over the thirteen leading methods on all 51 datasets. Better average ranks lie to the right; bars join methods that the critical distance of 2.56 rank points cannot separate. Only TabFM is significantly better than DIM-GP. Ranks are computed within these thirteen methods and therefore differ from ranks over the full field of 80.

Within this thirteen-method comparison, DIM-GP’s average rank is 5.16 and only TabFM is significantly ahead of DIM-GP under the Nemenyi analysis. RealMLP, TabDPT-Turbo, LimiX and both tuned tree ensembles are significantly behind. Neither paired test detects a significant difference between DIM-GP and TabPFN-3.

Pairwise p-values are unadjusted across the eight comparisons; the Nemenyi procedure controls its thirteen-method comparison. EXAONE-Tabular differs under the unadjusted Wilcoxon test but falls within the Nemenyi critical distance. The paired table and model-only aggregate table use different normalization pools, so their mean-error differences are computed separately.

4.3 Per-axis reading

Swipe to explore figure →
Classi- fication Regres- sion Binary Multi- class Small datasets Medium datasets 1500 1600 1700 1800 1900 2000 2100 2200 Elo TabFM EXAONE-Tabular DIM-GP TabPFN-3 TabPFN-2.6

Selected task-wise comparisons against TabPFN-3:

Axis DIM-GP TabPFN-3 Δ
Classification 1635 1617 +18
Binary 1654 1634 +20
Small datasets 1649 1624 +25
Multiclass 1606 1594 +12
Regression 1861 1897 −36
Medium datasets 1751 1789 −38

DIM-GP has higher Elo on four of the six reported axes and lower Elo on regression and medium-sized datasets in these descriptive subgroup comparisons.

4.4 Which field DIM-GP is compared in

TabArena distinguishes individual models from AutoML systems that manage model selection and training budgets. The evaluated DIM-GP configuration is fixed across datasets. Section 4.2 reports its placement in both comparison fields.

4.5 Computational cost and hardware

Using TabArena’s cost definition (§2.0), DIM-GP records median times of 30.26 seconds per 1,000 training rows and 2.509 seconds per 1,000 prediction rows on the RTX 4090 (51 datasets, fold 0), placing 32nd of 80 by reported fit cost. The robust timing fit separates fixed per-dataset and per-row contributions:

Fixed cost / dataset Per 1,000 rows Median s/1K measured on
TabICLv2 5.7 s 0.42 s 1.96 official board
TabPFN-3 14.2 s 0.71 s 3.59 official board
EXAONE-Tabular 8.3 s 2.43 s 6.17 official board
DIM-GP 10.5 s 17.74 s 30.26 our RTX 4090
TabFM 33.6 s 26.00 s 40.79 official board

These fits describe complete training or conditioning procedures. Baseline timings come from the published leaderboard; DIM-GP timings are measured on our RTX 4090.

Hardware sensitivity. The official benchmark records H200 timings for TabPFN-3, TabICLv2, TabPFN-2.6 and LimiX; RTX PRO 6000 timings for TabFM and TabDPT-Turbo; and A100 timings for TabPFN-Wide. To examine the effect of hardware, we ran TabICLv2 through the official harness on our RTX 4090, using the same data splits, preprocessing and eight-fold bagging configuration:

Dataset Rows fit, RTX 4090 fit, H200 Factor
blood-transfusion-service-center 499 4.33 s 4.39 s 0.99
diabetes 512 4.33 s 4.09 s 1.06
qsar-biodeg 705 5.66 s 5.47 s 1.03
churn 3,400 6.24 s 5.65 s 1.10
bank-marketing 45,211 27.39 s 13.26 s 2.07
Bioresponse 3,751 173.92 s 70.21 s 2.48
APSFailure 76,000 298.29 s 107.18 s 2.78

The RTX 4090-to-H200 fit-time ratio ranges from 0.99 to 2.78 across these datasets, showing that hardware effects depend on the workload.

Applying the fitted size-dependent factor to DIM-GP gives an illustrative H200-equivalent estimate of 21.02 s/1K, versus 30.26 s/1K measured on our RTX 4090. This is not a measured DIM-GP H200 run; transferring a calibration from another model adds uncertainty.

The official entries use eight-fold bagging, while DIM-GP uses a single fit. The cost comparison reports these configurations on the shared tasks and split.


Full TabArena-51 board, 80 entries (click to expand; every column sorts; ¹ DIM-GP times are our own measurement on an RTX 4090, section 4.5)
#modelTypeElofit s/1kinfer s/1kClass.Regr.BinaryMulti.SmallMedium
1TabFM (default)Foundation Model1793.040.796.9911752.02168.01780.01699.01785.01911.0
2EXAONE-Tabular (default)Foundation Model1749.06.170.6231747.01919.01756.01752.01698.02021.0
3DIM-GPFoundation Model1651.530.26¹2.509¹1635.41860.91654.01606.11649.01751.0
4TabPFN-3 (default)Foundation Model1644.03.590.3951617.01897.01634.01594.01624.01789.0
5TabPFN-2.6 (default)Foundation Model1613.05.750.61590.01847.01585.01645.01596.01747.0
6RealTabPFN-2.5 (tuned + ensembled)Foundation Model1600.02059.949.7851573.01849.01547.01745.01595.01700.0
7TabICLv2 (default)Foundation Model1555.01.960.1461554.01696.01573.01520.01536.01691.0
8RealTabPFN-2.5 (tuned)Foundation Model1547.02059.941.031535.01727.01516.01661.01538.01654.0
9RealTabPFN-2.5 (default)Foundation Model1523.05.720.6111531.01626.01536.01542.01562.01513.0
10RealMLP (tuned + ensembled)Neural Network1490.02791.9713.8861475.01666.01489.01451.01476.01608.0
11TabDPT (tuned + ensembled)Foundation Model1423.06155.04386.1581367.01786.01369.01388.01433.01459.0
12RealMLP (tuned)Neural Network1416.02791.970.3731409.01539.01418.01404.01407.01510.0
13TabDPT-Turbo (default)Foundation Model1413.02.020.1831383.01638.01392.01377.01416.01468.0
14TabM (tuned + ensembled)Neural Network1401.02462.021.9881404.01491.01408.01414.01373.01547.0
15LightGBM (tuned + ensembled)Tree-based1396.0416.632.2361381.01542.01376.01433.01359.01570.0
16CatBoost (tuned)Tree-based1391.01347.320.0371388.01494.01378.01457.01354.01561.0
17iLTM (tuned + ensembled)Foundation Model1389.012741.09397.8821388.01481.01408.01337.01363.01524.0
18CatBoost (tuned + ensembled)Tree-based1388.01347.320.3641381.01505.01370.01457.01349.01570.0
19TabDPT (tuned)Foundation Model1377.06155.0439.4521321.01731.01329.01317.01391.01400.0
20ChimeraBoost (tuned + ensembled)Tree-based1367.0518.460.5221370.01440.01394.01308.01304.01630.0
21ModernNCA (tuned + ensembled)Neural Network1365.04618.57.7351334.01582.01341.01335.01297.01649.0
22TabM (tuned)Neural Network1364.02462.020.2311365.01448.01361.01409.01343.01482.0
23LimiX (default)Foundation Model1356.027.336.111393.01314.01375.01503.01400.01289.0
24XGBoost (tuned + ensembled)Tree-based1355.0700.961.4381359.01417.01362.01372.01312.01539.0
25ChimeraBoost (tuned)Tree-based1344.0518.460.0451361.01363.01381.01311.01286.01579.0
26LightGBM (tuned)Tree-based1341.0416.630.3811332.01451.01316.01426.01300.01518.0
27CatBoost (default)Tree-based1340.05.810.0251348.01390.01370.01291.01296.01525.0
28ModernNCA (tuned)Neural Network1337.04618.50.471349.01376.01363.01321.01288.01540.0
29XGBoost (tuned)Tree-based1332.0700.960.2131330.01414.01328.01368.01296.01491.0
30TabSwift (default)Foundation Model1332.01.190.071319.01474.01354.01205.01338.01371.0
31xRFM (tuned + ensembled)Other1328.0866.122.0071304.01495.01296.01363.01315.01414.0
32TabPFNv2 (tuned + ensembled) [35.29% IMPUTED]Foundation Model1312.02943.3917.365
33iLTM (tuned)Foundation Model1294.012741.0960.1011300.01339.01313.01269.01278.01386.0
34Mitra (default) [35.29% IMPUTED]Foundation Model1292.087.362.432
35ChimeraBoost (default)Tree-based1285.02.540.051302.01293.01319.01261.01223.01518.0
36xRFM (tuned)Other1282.0866.120.0971258.01444.01257.01284.01264.01381.0
37TabDPT (default)Foundation Model1278.045.4239.4051214.01620.01221.01202.01295.01274.0
38TabM (default)Neural Network1274.07.810.2371288.01287.01290.01305.01266.01340.0
39TabICL (default) [29.41% IMPUTED]Foundation Model1273.06.861.52
40EBM (tuned + ensembled)Tree-based1255.01877.680.141299.01168.01301.01316.01252.01310.0
41TabPFNv2 (tuned) [35.29% IMPUTED]Foundation Model1252.02943.390.262
42TorchMLP (tuned + ensembled)Neural Network1250.02832.931.8011266.01253.01275.01250.01230.01343.0
43RealMLP (default)Neural Network1243.010.441.7141257.01253.01293.01130.01246.01272.0
44SAP-RPT-OSS (default)Foundation Model1239.013.962.0811249.01272.01239.01310.01278.01163.0
45BetaTabPFN (default) [25.49% IMPUTED]Foundation Model1236.0203.01.155
46TabPFNv2 (default) [35.29% IMPUTED]Foundation Model1221.03.270.315
47ModernNCA (default)Neural Network1217.013.740.3161180.01398.01201.01111.01206.01276.0
48EBM (tuned)Tree-based1207.01877.680.0141245.01122.01248.01253.01217.01214.0
49ExtraTrees (tuned + ensembled)Tree-based1197.0257.640.8461191.01253.01167.01304.01195.01225.0
50EBM (default)Tree-based1188.06.90.0151242.01024.01250.01230.01194.01201.0
51TorchMLP (tuned)Neural Network1187.02832.930.1121199.01204.01200.01214.01181.01232.0
52XGBoost (default)Tree-based1184.02.060.1221200.01167.01204.01204.01128.01370.0
53FastaiMLP (tuned + ensembled)Neural Network1169.0594.954.6491210.01048.01215.01207.01179.01170.0
54ExtraTrees (tuned)Tree-based1168.0257.640.0681159.01225.01144.01234.01169.01178.0
55Nori-30M (default) [74.51% IMPUTED]Foundation Model1164.00.530.077
56RandomForest (tuned + ensembled)Tree-based1161.0377.140.7471167.01158.01140.01288.01141.01238.0
57Nori (default) [74.51% IMPUTED]Foundation Model1156.00.530.077
58LightGBM (default)Tree-based1154.02.20.1711150.01200.01141.01201.01134.01228.0
59RandomForest (tuned)Tree-based1124.0377.140.0911130.01117.01109.01224.01098.01206.0
60FastaiMLP (tuned)Neural Network1104.0594.950.3361134.01014.01143.01109.01127.01053.0
61PerpetualBooster (tuned + ensembled)Tree-based1082.0176.260.4991060.01177.01106.0817.01091.01060.0
62iLTM (default)Foundation Model1080.0301.065.8271135.0860.01163.01019.01058.01156.0
63TabSTAR (tuned)Foundation Model1075.039935.254.2881115.0939.01143.01000.01115.0958.0
64TabSTAR (tuned + ensembled)Foundation Model1072.039935.2520.1661105.0976.01133.0988.01114.0950.0
65OrionMSP (default) [25.49% IMPUTED]Foundation Model1050.012.42.558
66TorchMLP (default)Neural Network1040.08.960.1291045.01024.01052.01019.01023.01087.0
67PerpetualBooster (tuned)Tree-based1036.0176.260.1911009.01133.01057.0740.01049.0998.0
68xRFM (default)Other1030.03.140.741978.01203.0939.01111.01043.01006.0
69ExtraTrees (default)Tree-based1007.01.970.251987.01074.0997.0942.01036.0929.0
70RandomForest (default)Tree-based1000.00.430.0531000.01000.01000.01000.01000.01000.0
71TabFlex (default) [25.49% IMPUTED]Foundation Model982.00.80.119
72FastaiMLP (default)Neural Network979.03.120.3121003.0881.01021.0926.0988.0962.0
73TabSTAR (default)Foundation Model976.0398.794.6451018.0789.01059.0808.01019.0827.0
74KNN (tuned + ensembled)Baseline970.0129.171.6271000.0854.0995.01021.0968.0972.0
75PerpetualBooster (default)Tree-based940.022.760.025933.0952.0978.0672.0962.0876.0
76Linear (tuned + ensembled)Baseline903.0240.780.308968.0468.0981.0912.0899.0907.0
77Linear (tuned)Baseline880.0240.780.068946.0409.0958.0890.0880.0871.0
78KNN (tuned)Baseline822.0129.170.103842.0722.0846.0805.0820.0807.0
79Linear (default)Baseline821.01.230.115887.0271.0919.0718.0837.0757.0
80KNN (default)Baseline606.00.190.037566.0625.0612.0214.0623.0559.0

5. TALENT-300

We compare DIM-GP with TALENT’s published methods and with three additional tabular foundation models evaluated in our own runs.

5.1 Comparison with TALENT’s published methods

TALENT publishes results for 28 regression methods and 31 classification methods, including XGBoost, CatBoost, LightGBM, TabPFN, TabR, ModernNCA and RealMLP.

Task DIM-GP avg rank field runner-up
Regression 2.30 — 1st of 29 90 datasets CatBoost 6.74
Binary 4.14 — 1st of 32 95 datasets TabR 9.14
Multiclass 3.00 — 1st of 32 77 datasets RealMLP 6.89

DIM-GP has the lowest average rank in each task type within this published comparison field.

5.2 Additional foundation-model comparison

Our additional comparison evaluates TabPFN-3, TabICLv2 and LimiX on the TALENT datasets. Average ranks use the common set completed by all four methods, with lower values indicating better performance:

Swipe to explore figure →
Regression (n=116) Binary (n=97) Multiclass (n=62) 0.0 0.5 1.0 1.5 2.0 2.5 3.0 average rank (lower is better) TALENT-300, foundation-model field we assembled ourselves DIM-GP TabPFN-v3 TabICLv2 LimiX
Task datasets DIM-GP TabPFN-3 TabICLv2 LimiX of those, DIM-GP first
Regression (RMSE) 116 2.026 2.198 2.509 3.267 39
Binary (Accuracy) 97 1.577 2.567 2.526 2.897 60
Multiclass (Accuracy) 62 1.871 2.290 1.919 3.065 26

DIM-GP has the lowest average rank in all three task types. Section 5.3 reports completion counts for these models and for Mitra, which was excluded from the rank comparison because of its lower completion rate.

5.3 Running on one consumer GPU

The comparison uses one RTX 4090 with 24 GB of physical memory. DIM-GP’s allocation is limited to less than 8 GB. Completion counts for the tested implementations and settings are:

Model Completed Failed Failure mode
DIM-GP 300 / 300 0
TabICLv2 298 2 OOM on the largest
TabPFN-3 295 5 OOM on the largest
LimiX 276 24 OOM on the largest
Mitra 205 95 O(N_train × N_test) memory; dropped from §5.2

DIM-GP completed all 300 datasets with GPU memory below 8 GB. The baseline evaluation allows caps on the in-context sample count and ensemble size to manage resource use. The accuracy comparison in §5.2 uses the common completed datasets under those settings.

5.4 How to read these two tables

The DIM-GP and additional foundation-model evaluations use a single seed and fixed splits; TALENT’s published tables average multiple seeds. Classification is scored by accuracy, including threshold selection for binary tasks; TabPFN-3’s reported TALENT result uses AUC. Average ranks and win counts permit comparison across datasets with different target scales.

6. Beyond tabular data: physics surrogates

The physics evaluation covers particle dynamics, contact and large deformation, aerodynamic fields and transient structural response.

6.1 What is being evaluated

Two members of the STOCHOS model family are evaluated:

  • DIM-GP Particle — predicts the evolution of particle and mesh-node states (§§7–9).
  • DIM-GP — the surrogate for fields on unstructured meshes, evaluated in §§10–17.

6.2 Why these particular benchmarks

The benchmarks cover the following physical regimes and comparison fields:

Benchmark Regime Comparison field
LagrangeBench (§7) Lagrangian SPH dynamics, 7 datasets GNS, SEGNN, CoRGI
DeepMind Water-3D (§8) Lagrangian, long-horizon rollout GNS
DeepMind DeformingPlate (§9) Quasi-static solid contact, tetrahedral mesh, 400 steps MeshGraphNets and M4GN — canonical protocol
AhmedML (§10) Steady 3D automotive aerodynamics AB-UPT and eight other published methods
DrivAerML (§11) Steady 3D automotive aerodynamics, high fidelity AB-UPT and eight other published methods
PLAID Rotor37 (§12) Parametric 3D compressor CFD Public leaderboard
PLAID VKI-LS59 (§13) 2D transonic turbine cascade, RANS Public leaderboard
PLAID 2D Multiscale Hyperelasticity (§14) Finite-strain RVEs, variable topology Public leaderboard
PLAID Tensile2d (§15) 2D elastoplastic specimen, parametric Public leaderboard
PLAID 2D ElastoPlastoDynamics (§16) Transient plate rupture, 41 steps Public leaderboard
PLAID 2D_profile (§17) Transonic airfoils, shape-only input Public leaderboard

PLAID sections report the archived board state of 4 September 2026, except for 2D_profile, which uses the 7 September snapshot. Submissions from commercial vendors are identified by their entry names, so each comparison refers to a specific submission in that snapshot.

Swipe to explore figure →
Figure 12. Reference data from four physics benchmarks. Top row: (a) dam break, (b) lid-driven cavity and (c) Taylor–Green vortex from LagrangeBench [33], followed by (d) DeepMind Water-3D [34]. Particle colours show per-step displacement; 3D cutaways expose the interior. Bottom row: (e) AhmedML [39], coloured by static pressure coefficient, and (f) PLAID Rotor37 [41], coloured by surface pressure.
Swipe to explore figure →
Figure 13. Reference fields from the six PLAID benchmarks (§§12–17): (a) Tensile2d, von Mises stress; (b) 2D Multiscale Hyperelasticity, strain-energy density on a deformed volume element with exaggerated displacement; (c) VKI-LS59, turbine-cascade Mach number; (d) 2D ElastoPlastoDynamics, transverse displacement at the final rupture step; (e) Rotor37, blade surface pressure; and (f) 2D_profile, airfoil Mach number.

7. LagrangeBench — verified against the benchmark’s own evaluator

LagrangeBench [33] is a Lagrangian-simulation benchmark: seven SPH datasets (2D and 3D Taylor–Green vortex, reverse Poiseuille flow, lid-driven cavity, and 2D dam break), a fixed protocol, and published baselines. It supports comparison across three published methods: GNS [34], SEGNN [35] and CoRGI [36] all publish numbers on the same cells under the same rules.

7.1 Protocol

Scoring follows the benchmark exactly: 26-frame windows made of 6 seed frames and 20 predicted steps, with RPF-2D scored over 384 and RPF-3D over 192 test windows. Three metrics per dataset — MSE20 (position error over the 20 predicted steps, per dimension), Sinkhorn (a distributional divergence, blind to particle identity) and Ekin (squared error of total kinetic energy).

7.2 Result — lower error in 19 of 21 comparisons

Each DIM-GP result is compared with the lowest published error for that dataset and metric across GNS, SEGNN and CoRGI.

dataset metric best published (holder) DIM-GP comparison
DAM-2D MSE20 1.55e-5 (CoRGI) 3.49e-7 44× lower error
DAM-2D Sinkhorn 2.82e-6 (CoRGI) 6.48e-8 44× lower error
DAM-2D Ekin 2.18e-5 (CoRGI) 8.88e-8 245× lower error
LDC-2D MSE20 1.4e-5 (GNS) 9.34e-6 1.50× lower error
LDC-2D Sinkhorn 5.07e-7 (CoRGI) 1.15e-7 4.39× lower error
LDC-2D Ekin 3.81e-7 (CoRGI) 2.76e-7 1.38× lower error
RPF-2D MSE20 1.54e-6 (CoRGI) 1.226e-6 1.26× lower error
RPF-2D Sinkhorn 2.08e-8 (CoRGI) 1.49e-9 14× lower error
RPF-2D Ekin 2.39e-6 (CoRGI) 7.224e-6 3.02× higher error
TGV-2D MSE20 3.81e-6 (CoRGI) 2.29e-6 1.67× lower error
TGV-2D Sinkhorn 1.05e-7 (CoRGI) 3.56e-8 2.95× lower error
TGV-2D Ekin 2.90e-7 (CoRGI) 1.35e-7 2.15× lower error
LDC-3D MSE20 3.86e-5 (CoRGI) 2.44e-5 1.58× lower error
LDC-3D Sinkhorn 2.64e-7 (CoRGI/SEGNN) 1.48e-7 1.79× lower error
LDC-3D Ekin 1.55e-8 (CoRGI) 1.48e-8 1.05× lower error
RPF-3D MSE20 1.64e-5 (SEGNN) 1.330e-5 1.23× lower error
RPF-3D Sinkhorn 1.33e-7 (CoRGI) 6.41e-8 2.08× lower error
RPF-3D Ekin 1.34e-6 (SEGNN) 1.263e-6 parity
TGV-3D MSE20 5.2e-3 (SEGNN) 4.76e-3 1.10× lower error
TGV-3D Sinkhorn 6.4e-6 (SEGNN) 3.07e-6 2.09× lower error
TGV-3D Ekin 2.21e-3 (CoRGI/SEGNN) 1.86e-3 1.19× lower error

Compared with each method individually, DIM-GP has lower error than GNS on all 21 cells, lower error than SEGNN on 20 with one parity, and lower error than CoRGI on 19.

DIM-GP has the lowest error on all three metrics for DAM-2D, LDC-2D, TGV-2D, LDC-3D and TGV-3D.

Under the Neural-SPH 400-step protocol [37] on DAM-2D, evaluated over all 25 trajectories, the corresponding values are 3.12e-2 / 4.90e-4 / 1.99e-4 against a published best of 8.4e-2 / 7.5e-3 / 2.1e-3, a factor of 2.7 / 15 / 10.6. Since the 20-step protocol is comparatively short, this establishes that the margin persists at twenty times the horizon.

Swipe to explore figure →
MSE20 Sinkhorn Ekin DAM-2D LDC-2D LDC-3D RPF-2D RPF-3D TGV-2D TGV-3D lower lower lower lower lower lower lower lower lower lower lower higher lower lower parity lower lower lower lower lower lower LagrangeBench: every scored cell against the best published number

7.3 The remaining two comparisons

RPF-3D Ekin — parity. DIM-GP records 1.263e-6 against SEGNN’s 1.34e-6 and is reported as parity.

RPF-2D Ekin — higher error than CoRGI. DIM-GP has 3.9-fold lower error than GNS and 2.5-fold lower error than SEGNN on this cell, while CoRGI reports the lowest error.

7.4 Cost

Each reported model was trained on a single RTX 4090, at roughly 20–90 seconds per epoch depending on dataset and resolution.


8. DeepMind Water-3D — measured against the published GNS result

Water-3D is the 3D fluid dataset from Learning to Simulate Complex Physics with Graph Networks [34]: approximately 14,000 particles over 800 timesteps. It tests prediction over a much longer rollout than the 20-step LagrangeBench evaluation.

8.1 Result

Scored under the paper’s own protocol — 100 test sequences, full 800-step rollouts, seed frames excluded:

model rollout MSE (mean) median thickness speed MMD gradient updates
GNS, as published [34] 0.01010 ~20M
DIM-GP, accuracy setting 0.01096 0.00984 82% 87% 0.00265 ~1.9M (10% of GNS)
DIM-GP, physics setting 0.01238 0.01109 90% 106% 0.00249 ~1.6M

The accuracy setting has a mean rollout MSE 8.5% above the published GNS value, with approximately one tenth as many gradient updates. The update count describes the training schedule; wall-clock time also depends on the cost of each update.

The second setting improves the physical diagnostics, reaching 90% of reference fluid-column thickness and 106% of characteristic speed, compared with 82% and 87% for the accuracy setting. The two configurations illustrate the choice between minimizing position error and preserving bulk fluid behaviour.

8.2 Comparison scope

The comparison uses the published GNS result on Water-3D under the full 800-step protocol. LagrangeBench reports a different quantity, MSE over 20 predicted steps, while NeuralMPM [38] reports full-rollout results on 2D fluid datasets. Those results belong to their respective datasets and protocols and are evaluated separately from Water-3D.


9. DeepMind DeformingPlate — comparison under the canonical protocol

DeformingPlate is the solid-mechanics case of the MeshGraphNets suite [48]: a rigid actuator with a prescribed motion is pressed into a hyperelastic plate that is clamped along one edge, and the plate’s node positions are to be predicted over 400 quasi-static steps on a tetrahedral mesh of 700–2,200 nodes. The task includes contact and large deformation, and every trajectory has its own plate geometry, actuator shape and path. 1,200 trajectories are used for training, 100 for validation and 100 for the test score.

9.1 Protocol

The score follows the original paper: RMSE of world positions in metres over all nodes and all 399 predicted steps of the 100 test trajectories, with errors pooled before taking the square root and reported . The seed frame is excluded. The checkpoint was selected on the validation split, and the test split was scored once.

9.2 Result

method RMSE-all data and protocol
canonical: DeepMind data, official split, original error formula
MeshGraphNets [48] 15.1 original paper
DIM-GP 6.45 this work, checkpoint chosen on the validation split
M4GN [49] 2.65 reported; no code released, not reproduced
not on the canonical protocol, listed for completeness
MGN-T [50] 3.21 data regenerated with COMSOL
HCMT [51] 7.3 different error formula; its own MeshGraphNets baseline reads 7.8, DIM-GP reads 4.9 under it
ROBIN [52] 4.98 square root before the mean over trajectories; its own MeshGraphNets baseline reads 8.8
EvoMesh [53] 12.9 trained on 500 of the 1,200 trajectories
BSMS-GNN [54] 16.0 dataset re-split 1000/200/200

DIM-GP reports RMSE of 6.45 versus 15.1 for MeshGraphNets, a 57% reduction, and ranks second among the three canonical-protocol entries. M4GN reports the lowest value, 2.65. The remaining rows use different data, splits or error conventions and are listed separately.

Swipe to explore figure →
Figure 15. DeepMind DeformingPlate, test trajectory 92 at the moment of deepest indentation: the reference finite-element solution (left) and the DIM-GP rollout from the first frame alone (right), plate coloured by displacement from its rest shape, actuator in grey; the scene is turned so that the actuator approaches from above. Rollout RMSE of this trajectory 0.93\times 10^{-3}; dataset mean 6.45\times 10^{-3}.

9.3 Cost

Training completed in less than one day on a single RTX 4090, with peak GPU memory below 8 GB.


10. AhmedML — internal evaluation against published baselines

AhmedML [39] is a 500-geometry CFD dataset over the Ahmed body, a standard automotive bluff-body. The published comparison in AB-UPT [40] provides the external baseline values used below.

10.1 Protocol and metric

The metric follows that comparison: per design, the Frobenius relative error , averaged over designs, reported as a percentage. Five channels are scored — surface pressure and wall shear stress ; volume velocity , vorticity and total pressure . Evaluation uses raw full-resolution clouds from the 50 test designs in AB-UPT’s published split.

10.2 Result

relative L2 (%) training on one RTX 4090
DIM-GP [63] 2.97 3.81 1.85 6.30 1.85 6.7 h, 8 GB
AB-UPT [40] 3.01 3.88 1.90 6.52 1.98 ~77 h¹
Transformer [55] 3.41 4.03 2.09 6.76 2.16
Transolver [56] 3.45 4.00 2.05 8.22 2.16
OFormer [57] 4.12 4.60 3.63 15.06 4.08
UPT [58] 4.25 5.80 2.73 15.03 3.10
Graph U-Net [59] 6.46 7.29 4.15 53.66 5.18
GINO [60] 7.90 8.18 6.23 71.81 8.10
PointNet [61] 8.02 10.09 5.44 66.04 6.13
LNO [62] 12.95 11.50 7.59 72.49 8.48

AhmedML test geometry run_241 (official test split): surface pressure of the reference simulation and of the DIM-GP prediction, displayed on a decimated surface mesh. Drag to orbit, scroll to zoom, right-drag to pan.

¹ Measured for AB-UPT’s 200K-update schedule on the same RTX 4090 (§10.3). Both models use the same training data, including surface and volume fields. Other methods have no timing on this hardware.

DIM-GP places first of ten on all five channels in the comparison with published baselines [40]. Relative to AB-UPT, error is lower by 1.3% on surface pressure, 1.8% on wall shear, 2.6% on volume velocity, 3.4% on vorticity and 6.6% on total pressure.

The DIM-GP row reports our internal evaluation as of 7 September 2026. Baseline values come from the published comparison [40]. The ranking is based on point estimates.

10.3 Cost — measured on identical hardware

DIM-GP trains in 6.7 h on one RTX 4090 under an 8 GB cap. We executed AB-UPT’s released code on the same hardware and measured ~77 h for its 200K-update training schedule. Both training times are validated on our hardware, with the same data and split, surface and volume together. DIM-GP therefore requires approximately 91% less training time in this configuration.


11. DrivAerML — internal evaluation against published baselines

DrivAerML [43] comprises 484 usable variants of the DrivAer notchback vehicle, each solved by hybrid RANS–LES (SA- -DDES, OpenFOAM) on a ~160M-cell mesh at a cost of roughly 40 hours on 1,536 CPU cores per design — on the order of 61,000 core-hours per geometry. The learning task is steady: one geometry in, the time-averaged surface pressure and wall-shear-stress fields out (~8.8M surface cells). Evaluation follows AB-UPT [40]: its published 400/34/50 split and per-design relative L2 error, using the Frobenius norm over the three shear components and averaging over the 50 test designs.

relative L2 (%) surface pressure wall shear trained on training on one RTX 4090
DIM-GP [63] 3.74 7.15 surface only under 15 h
AB-UPT [40] 3.82 7.29 surface + volume ~67 h²
Transformer [55] 4.35 8.26 surface + volume
Transolver [56] 4.81 8.95 surface + volume
OFormer [57] 4.85 8.92 surface + volume
UPT [58] 7.44 12.93 surface + volume
GINO [60] 13.03 21.71 surface + volume
Graph U-Net [59] 16.13 27.84 surface + volume
LNO [62] 20.51 36.44 surface + volume
PointNet [61] 23.63 41.85 surface + volume

DrivAerML test car run_11: surface pressure of the reference simulation and of the DIM-GP prediction, displayed on a decimated surface mesh. Drag to orbit, scroll to zoom, right-drag to pan.

² Schedule estimate from measured throughput of the released implementation in a surface-only configuration on our RTX 4090 (§11.1). The published AB-UPT accuracy uses surface and volume data.

DIM-GP places 1st of 10 on both channels, with 2.1% less surface-pressure error and 1.9% less wall-shear error than runner-up AB-UPT. The DIM-GP values report our internal evaluation as of 7 September 2026. Baseline values come from the published comparison [40]. The ranking is based on point estimates.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 16. DrivAerML test car run_11, side view: time-averaged surface pressure of the reference simulation, the DIM-GP prediction, and their difference. Every cell of the 8.8-million-cell surface is predicted; the view shows a 400,000-cell random subset.

11.1 Computational cost on identical hardware

The released AB-UPT implementation (Noether, Emmi AI, v2026.4.0) was executed on the same RTX 4090 workstation and 400-design split in a surface-only configuration. Its measured update time of 1.21 seconds implies approximately 67 hours for 200,000 updates. This is a schedule estimate derived from measured throughput, compared with under 15 hours of DIM-GP wall-clock training, including evaluation checkpointing.

The timing was measured in our Windows environment with the available kernels. The published AB-UPT accuracy uses surface and volume data, whereas the timed configuration uses surface data alone and omits domain-branched cross-attention.

12. PLAID Rotor37 — live leaderboard, third placeboard state as of 4 September 2026

Rotor37 [41] is part of the PLAID benchmark suite (Safran, CC-BY-SA): 1,000 training and 200 test designs of a transonic axial compressor blade, 29,773 surface nodes each, with two operating scalars (rotational speed , inlet pressure ) as input and three surface fields (density, pressure, temperature) plus three global quantities (mass flow, compression ratio, isentropic efficiency) as output. All six are scored by a server-side leaderboard.

In the 4 September 2026 snapshot, DIM-GP ranks 3rd of 16 at a total error of 0.000488 (displayed 0.0005), behind NPco (~0.00043) and PXTransolver (~0.00047) and ahead of Super-MARIO (~0.00077). The gap to NPco is approximately 0.00006 in aggregate RRMSE, or 0.006 percentage points when expressed as percentages (approximately 0.049% versus 0.043%).

RRMSE (lower is better) Density Pressure Temp. Massflow Compr. ratio Efficiency total
NPco 0.0007 0.0007 0.0003 0.0003 0.0003 0.0003 0.0004
PXTransolver 0.0008 0.0008 0.0003 0.0003 0.0003 0.0003 0.0005
DIM-GP 0.0008 0.0008 0.0003 0.0003 0.0003 0.0003 0.0005
Super-MARIO 0.0013 0.0013 0.0005 0.0005 0.0005 0.0005 0.0008
MMGP 0.0031 0.003 0.0008 0.0005 0.0005 0.0005 0.0014
MeshFiLM 0.0027 0.0027 0.0007 0.0008 0.0009 0.0008 0.0014
MARIO 0.0035 0.0034 0.001 0.0008 0.0007 0.0006 0.0017
Baburu 0.0042 0.0042 0.0014 0.0007 0.0007 0.0009 0.002
Stealth 0.0029 0.0029 0.0009 0.0027 0.0027 0.0019 0.0023
ICLGS 0.0047 0.0047 0.0011 0.0021 0.0019 0.0014 0.0026
Vi-Transformer 0.0063 0.0062 0.0019 0.001 0.0011 0.0007 0.0029
Augur 0.0055 0.0053 0.0012 0.0028 0.0028 0.0019 0.0033
tp140205 0.0139 0.0136 0.0025 0.0047 0.0043 0.0042 0.0072
MGN 0.0114 0.0114 0.0024 0.0061 0.006 0.0071 0.0074
GeoFunFlow-3D 0.0319 0.0315 0.0134 0.0376 0.0383 0.0138 0.0278
FNO 0.084 0.0836 0.0086 0.0046 0.0042 0.0031 0.0313
DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 17. Rotor37, a held-out blade: surface pressure on both blade sides from the reference RANS solution, the DIM-GP prediction, and their difference.

13. PLAID VKI-LS59 — live leaderboard, first placeboard state as of 4 September 2026

VKI-LS59 [41] is the suite’s 2D transonic turbine cascade: 671 training and 168 test profiles of a linear cascade (Safran, CC-BY-SA), 36,421 nodes each, solved by compressible RANS (Spalart–Allmaras) with the passage reaching Mach 1.78. Inputs are two operating scalars (inlet angle, outlet Mach number) and the mesh; the competition scores two nodal fields (Mach number and turbulent viscosity ) and six global quantities ( , power, the pressure and temperature ratios and , the isentropic efficiency and the outlet angle).

RRMSE (lower is better) nut mach Q power Pr Tr eth_is angle_out total
DIM-GP 0.0216 0.0093 0.0025 0.0062 0.0015 0 0.0306 0.0022 0.0093
PXTransolver 0.0205 0.0084 0.0039 0.0044 0.0015 0 0.0402 0.0023 0.0101
NPco 0.0243 0.0094 0.0025 0.0067 0.0018 0 0.0433 0.0019 0.0112
MARIO 0.0259 0.0112 0.0052 0.0077 0.0018 0 0.0453 0.0023 0.0124
ICLGS 0.0344 0.0155 0.0035 0.0092 0.0013 0 0.0375 0.0037 0.0131
MeshFiLM 0.0283 0.0102 0.0077 0.0091 0.002 0 0.0486 0.0025 0.0135
CRT 0.0337 0.0185 0.0015 0.0053 0.0018 0 0.0466 0.0023 0.0137
SAIR 0.0278 0.0122 0.0015 0.0049 0.0025 0 0.0586 0.0026 0.0138
Stealth 0.035 0.0176 0.0029 0.0067 0.0027 0 0.0699 0.0032 0.0173
gantnera 0.0329 0.0131 0.0109 0.0083 0.0026 0 0.0669 0.0037 0.0173
Vi-Transformer 0.0498 0.0232 0.0052 0.0083 0.0024 0 0.0621 0.0031 0.0193
MARIO (Samy) 0.0403 0.0208 0.0054 0.008 0.0029 0 0.0898 0.0033 0.0213
FNO 0.0846 0.018 0.0047 0.0062 0.0019 0 0.0539 0.0027 0.0215
test 0.0412 0.0186 0.0034 0.0052 0.0029 0 0.0976 0.0041 0.0216
Augur 0.0424 0.0221 0.012 0.0113 0.0027 0 0.0863 0.0045 0.0227
Rrrra 0.0423 0.0182 0.0189 0.0158 0.0038 0 0.0852 0.0071 0.0239
MMGP+ 0.0822 0.0309 0.0023 0.0057 0.0026 0 0.1224 0.0033 0.0312
MGN 0.0771 0.0156 0.0716 0.0403 0.0064 0.0001 0.1625 0.0241 0.0497
transolver 0.0532 0.0139 1 1 1 1 1 1 0.7584
Naive_approach_GP 0.0669 0.0384 1 1 1 1 1 1 0.7632
akabalan_1fieldTesting 0.2849 0.0175 1 1 1 1 1 1 0.7878

The entry holds first place of 21 at 0.0093, approximately 8% lower aggregate error than PXTransolver. It leads the efficiency column and maintains competitive errors across the other scored outputs.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 18. VKI-LS59, held-out cascade design 447: Mach number in the passage from the reference RANS solution, the DIM-GP prediction, and their difference (design RRMSE 0.0113). The passage shock and the suction-side expansion are placed within a few cells.

14. PLAID 2D Multiscale Hyperelasticity — live leaderboard, first placeboard state as of 4 September 2026

This benchmark [41] is the suite’s variable-topology case: 764 training and 376 test representative volume elements of a porous hyperelastic material, each a different microstructure with five to eight holes (constant total porosity, node counts from 4,100 to 7,100), solved by FEniCS at finite strain. Inputs are three macroscopic strain components and the mesh; the competition scores seven nodal fields — the displacements , the four components of the first Piola–Kirchhoff stress, and the strain-energy density — together with the homogenised effective energy. Every design has its own mesh topology.

RRMSE (lower is better) u1 u2 P11 P12 P22 P21 psi eff. energy total
DIM-GP 0.003 0.0031 0.0093 0.0139 0.0095 0.0136 0.0263 0.002 0.0101
PXTransolver 0.0026 0.0028 0.0093 0.014 0.0095 0.0139 0.0262 0.0074 0.0107
test 0.0065 0.0064 0.0166 0.0256 0.017 0.0254 0.0267 0.0042 0.0161
NPco 0.0029 0.0028 0.0198 0.0285 0.0202 0.0283 0.0259 0.0037 0.0165
plaidtest 0.0068 0.007 0.0181 0.0278 0.0184 0.0275 0.0292 0.0057 0.0176
T&D 0.0059 0.0061 0.0201 0.0298 0.0205 0.0297 0.0266 0.0067 0.0182
Stealth 0.0055 0.0058 0.0231 0.033 0.0236 0.0329 0.0264 0.0039 0.0193
MiSe-GNN 0.012 0.014 0.0173 0.0312 0.0176 0.0314 0.0328 0.0111 0.0209
aravabt 0.006 0.0066 0.026 0.0368 0.0263 0.0364 0.0265 0.0055 0.0213
Augur 0.0109 0.0114 0.0208 0.0336 0.0212 0.033 0.0274 0.0188 0.0221
MeshFiLM 0.0091 0.0097 0.0261 0.0377 0.0267 0.0377 0.0281 0.008 0.0229
test 0.0119 0.0167 0.0262 0.0403 0.0267 0.0399 0.0295 0.0231 0.0268
FNO 0.0115 0.0117 0.0353 0.0513 0.0359 0.051 0.0329 0.012 0.0302
test1 0.0132 0.014 0.0329 0.0532 0.0326 0.0519 0.033 0.0117 0.0303
Vi-Transformer 0.0173 0.0172 0.0337 0.0581 0.0343 0.0571 0.0312 0.0113 0.0325
Unet 0.0291 0.0283 0.0349 0.0498 0.0347 0.05 0.035 0.018 0.035
Pmh 0.0161 0.0176 0.0366 0.0643 0.0367 0.0625 0.0362 0.0226 0.0366
ICLGS 0.0244 0.029 0.0473 0.0872 0.0482 0.0864 0.0433 0.0183 0.048
MARIO 0.0336 0.0377 0.0536 0.1067 0.0539 0.1053 0.0456 0.022 0.0573
GP 0.0748 0.075 0.0809 0.167 0.0815 0.1667 0.054 0.0216 0.0902
MuFi-MiSe 0.0108 0.0132 0.0305 0.0466 0.0307 0.0465 0.0315 1 0.1512
transolver 0.0129 0.0121 0.0327 0.0613 0.0332 0.0544 0.0335 1 0.155

The entry ranks first of 22 at 0.010093 (displayed 0.0101), approximately 6% below PXTransolver in aggregate error. The four stress errors are comparable to PXTransolver’s. DIM-GP has the lowest effective-energy error, 0.0020, followed by NPco at 0.0037 and PXTransolver at 0.0074.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 19. 2D Multiscale Hyperelasticity, held-out RVE 395 (five holes): first Piola–Kirchhoff stress P_{11} from the FEniCS reference, the DIM-GP prediction, and their difference (design RRMSE over the seven fields 0.0136). The stress concentrations between the holes are reproduced in place.

15. PLAID Tensile2d — live leaderboard, second place behind a declared solverboard state as of 4 September 2026

Tensile2d [41] is the suite’s entry-level structural case: a 2D plane-strain, quasi-static, small-strain elastoplastic specimen under a uniform top traction, meshed once per design with 6,000–10,000 nodes (Safran, Z-set solver). Inputs are six scalars — the load , four parameters of the isotropic hardening law and the Young’s modulus — and the mesh; the competition scores five nodal fields ( , , , , ) and three scalars (the maximum von Mises stress, and the maximum and on the top edge). 500 designs carry labels, 200 form the test set.

RRMSE (lower is better) U1 U2 sig11 sig22 sig12 max vM max U2 top max sig22 top total
PhysicsX Agent (reverse-engineered) 0 0 0 0 0 0 0 0 0
DIM-GP 0.0004 0.0005 0.0013 0.0006 0.0012 0.0004 0.001 0.0004 0.0007
NPco 0.0002 0.0003 0.0013 0.0006 0.0009 0.0044 0.0006 0.0016 0.0013
PhysicsX Transolver 0.0002 0.0002 0.0012 0.0006 0.0009 0.0046 0.0008 0.0017 0.0013
Stealth 0.0003 0.0003 0.0012 0.0006 0.0009 0.0048 0.0008 0.0016 0.0013
MeshFiLM 0.0002 0.0003 0.0012 0.0006 0.0009 0.005 0.0022 0.0015 0.0015
plaidtest 0.0013 0.0017 0.0029 0.0013 0.0018 0.0051 0.003 0.0017 0.0023
PassionTraining 0.0009 0.0011 0.0021 0.001 0.0016 0.0066 0.0044 0.0019 0.0025
MMGP 0.0015 0.0009 0.0031 0.0013 0.0021 0.005 0.0053 0.0017 0.0026
test 0.0015 0.0022 0.0036 0.0016 0.0024 0.0069 0.0025 0.0019 0.0028
ICLGS 0.0028 0.004 0.0048 0.0018 0.0032 0.0065 0.0021 0.0018 0.0034
CRT 0.0013 0.0015 0.0057 0.0022 0.0039 0.0079 0.0024 0.0026 0.0034
sangmin12312343 0.0018 0.0018 0.003 0.0019 0.0023 0.0108 0.0066 0.0016 0.0037
MARIO 0.0023 0.003 0.004 0.0017 0.0023 0.0088 0.0063 0.0023 0.0038
MMVT 0.0043 0.0051 0.0089 0.0037 0.0051 0.0068 0.0073 0.0019 0.0054
Augur 0.0037 0.0048 0.0081 0.0035 0.005 0.0101 0.014 0.0034 0.0066
MGN 0.0034 0.0043 0.0047 0.0013 0.0016 0.0169 0.0292 0.0022 0.008
Vi-Transformer 0.0086 0.0091 0.0184 0.0102 0.0146 0.009 0.0203 0.0021 0.0116
FNO 0.0174 0.011 0.025 0.0057 0.0135 0.0085 0.0152 0.0021 0.0123
JBone 0.3054 0.3891 0.4022 0.2354 0.268 0.0052 0.0017 0.0017 0.2011
MuFi-MiSe 0.0023 0.0036 0.0042 0.002 0.0024 1 1 1 0.3768
transolver 0.0029 0.0036 0.0049 0.0016 0.0022 1 1 1 0.3769

DIM-GP ranks second of 22 at a displayed aggregate error of 0.0007. In its article How an AI Agent Cracked an Industry Benchmark, PhysicsX reports that AI agents reconstructed the reference simulator by inferring the constitutive law and simulation setup from the supplied training data [44]. Its entry ranks first with zero error at the leaderboard’s displayed precision.

The DIM-GP workflow runs on the CPU in minutes.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 20. Tensile2d, a held-out specimen: von Mises stress from the reference elastoplastic solution, the DIM-GP prediction, and their difference.

16. PLAID 2D ElastoPlastoDynamics — live leaderboard, first placeboard state as of 4 September 2026

This is a transient solid-mechanics case [41]: a 200 × 100 mm steel plate with up to four holes and edge notches, clamped on the left and pulled at 500 mm/s on the right, computed by the benchmark’s authors with an explicit finite-element solver, a non-linear, non-local constitutive law and element erosion, and delivered as 41 time steps of the two displacement fields on 19,000–30,000 nodes per design. There are no input scalars: the geometry is the only variable, 1,000 designs are labelled and 18 form the test set. The competition scores the full and trajectories.

Transient field prediction substantially increases the scale of the learning problem. Each design contains a sequence of complete mesh fields rather than a single steady-state solution. Here, 41 time steps across 19,000–30,000 nodes produce approximately 0.8–1.2 million node-time states per design and output field. Memory use and training cost therefore grow rapidly with both spatial resolution and sequence length, particularly for architectures that represent mesh states as large token sequences. This makes the benchmark a demanding test of an architecture’s ability to model high-resolution spatial fields and their evolution over time.

RRMSE (lower is better) U_x U_y total
DIM-GP 0.0005 0.0167 0.0086
test 0.0008 0.017 0.0089
DAFNO 0.0025 0.0291 0.0158
FNO 0.0031 0.0399 0.0215
Vi-Transformer 0.0186 0.0269 0.0227
MGN 0.0073 0.0403 0.0238
MARIO 0.0059 0.058 0.0319
Augur 0.0264 0.0427 0.0346

The entry holds first place of eight at 0.008605, with the lowest displayed and errors in the archived table. The displayed values are 0.0005 for DIM-GP and 0.0008 for the next entry.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 21. 2D ElastoPlastoDynamics, held-out design 8 drawn in its deformed shape at the final step: vertical displacement U_y from the reference simulation, the DIM-GP prediction, and their difference (design RRMSE U_x 0.0099, U_y 0.0197).

The board metric normalises each field by its maximum over the whole test set, with constants of roughly 340 mm ( ) and 25 mm ( ).


17. PLAID 2D_profile — live leaderboard, third placeboard state as of 7 September 2026

2D_profile [41] is the suite’s shape-only aerodynamics case: 300 training and 100 test airfoil profiles in a transonic flow, each with its own triangular mesh of 35,000–39,000 nodes cut close to the profile, and no input scalars at all — the geometry is the whole input, with the effective angle of attack baked into the shape. The competition scores four nodal fields (Mach number, pressure and the two velocity components). The test cases include transonic shocks.

RRMSE (lower is better) Mach Pressure Velocity-x Velocity-y total
PXTransolver 0.0134 0.0108 0.0155 0.0139 0.0134
NPco 0.0148 0.0097 0.0172 0.0136 0.0138
DIM-GP 0.0167 0.011 0.0196 0.0134 0.0152
test-ts-dino 0.0161 0.0114 0.0184 0.0149 0.0152
Stest 0.0168 0.0111 0.0198 0.0135 0.0153
plaidtest 0.0163 0.0111 0.0197 0.0146 0.0154
Stealth 0.0166 0.011 0.0195 0.0154 0.0157
MM-GP+ 0.0208 0.0139 0.0243 0.016 0.0187
Transolver+ 0.0209 0.0131 0.0267 0.019 0.0199
test 0.0211 0.0154 0.0242 0.0192 0.02
Super MARIO 0.022 0.0158 0.0251 0.0196 0.0206
MARIO 0.023 0.0146 0.0264 0.021 0.0213
MeshFiLM 0.0265 0.0177 0.0316 0.0243 0.025
CRT 0.0278 0.0171 0.0323 0.0244 0.0254
MuFi-MiSe 0.0279 0.018 0.0328 0.0315 0.0276
MiSe-GNN 0.0309 0.0212 0.0358 0.026 0.0285
Vi-Transformer 0.036 0.0167 0.0403 0.0307 0.0309
MMGP 0.0439 0.0208 0.0471 0.0342 0.0365
Augur 0.0469 0.0248 0.0538 0.0445 0.0425
ICLGS 0.0777 0.0248 0.0885 0.0601 0.0628
FNO 0.0988 0.0785 0.1148 0.0967 0.0972
elzigomario 0.1893 0.1004 0.2295 0.1199 0.1598

DIM-GP ranks third of 22 in the 7 September 2026 snapshot, with an aggregate error of 0.0152 versus PXTransolver’s 0.0134. The numerical gap is 0.0018, equivalent to 0.18 percentage points when these relative errors are expressed as percentages (1.52% versus 1.34%). DIM-GP’s Velocity-y error of 0.0134 is the lowest displayed value on the board; its pressure error of 0.011 is joint-third.

DIM-GP predictionreference
referenceDIM-GP

Drag the slider: reference simulation on the left of the divider, DIM-GP prediction on the right (same colour scale). The three-panel figure below adds the difference map.

Swipe to explore figure →
Figure 22. 2D_profile, held-out design 272: Mach number from the elsA reference, the DIM-GP prediction, and their difference (design RRMSE 0.0192, peak Mach 2.95). The prediction reproduces the leading-edge supersonic pocket and shocks; the largest differences occur near the shock fronts.

18. Bayesian optimization — concrete mix design, six approaches, one virtual problem

Bayesian optimization uses predictions and uncertainty estimates to select the next design to evaluate. Here, STOCHOS searches for concrete recipes that combine high 28-day compressive strength with low cradle-to-gate CO and material cost. Seven ingredient masses per cubic metre are varied, subject to a total mass of 2,195–2,551 kg.

The study uses a fixed virtual test bench fitted to the 1,022 usable rows of the public Concrete Compressive Strength dataset (UCI, 1,030 laboratory recipes) [45]. It predicts strength; CO and cost follow from the ingredient quantities, sourced emission factors and an indicative price basket. Six approaches receive the same starting recipes and a budget of 60 evaluations: STOCHOS, BoTorch’s qNEHVI [46], BayBE’s desirability approach [47], a random-forest optimizer, D-optimal design of experiments and random search.

Swipe to explore figure →
HOW BAYESIAN OPTIMIZATION WORKS Each evaluation improves the information used to select the next candidate. 1 Start with observations Known designs and outcomes 2 Fit a probabilistic model Predictions and uncertainty 3 Choose the next candidate Objectives, constraints and information 4 Evaluate the candidate Simulation or experiment Add the result and repeat until the target is reached or the budget is used

Part A — reaching a specification. Five levels of increasing difficulty each demand a minimum strength together with a CO and a cost ceiling on the same recipe; difficulty is characterized by an estimated probability that a random valid recipe qualifies (1 in 9 for S1 down to 1 in 66,667 for S4.5), and an unreachable sixth level serves as a control that no method passes. Ten repeated runs per level, at round sizes of one, three and six recipes. The table counts, per level, how many of the 10 runs at one recipe per round found a qualifying recipe.

Choose the difficulty level to see how many runs met the strength, CO₂ and cost requirements as the experiment budget increased.

Runs that reached the specification (of 10) S1 S2 S3 S4 S4.5
STOCHOS 10 10 10 10 10
BoTorch (qNEHVI) 10 10 10 10 8
BayBE 10 10 10 9 2
Random forest 10 10 7 3 1
Classical DoE 10 10 6 0 0
Random search 10 7 4 1 0

Over all five levels and all three round sizes STOCHOS reaches the specification in 149 of 150 runs, against 137 for BoTorch, 111 for BayBE, 87 for the random forest, 78 for classical DoE and 66 for random search; it is the only method that solves the extreme level in every sequential run. Among runs in which both STOCHOS and BoTorch succeeded, STOCHOS reached the target sooner in 99 of 136 comparisons, later in 16 and at the same experiment in 21 (reported two-sided sign test ).

Part B — discovering the trade-off set. With no single target, a method is scored on how much of the achievable trade-off region it discovers: the hypervolume of its evaluated recipes against a frozen reference point, averaged over 30 repeated runs; the table gives the coverage after 60 experiments at round sizes of one, five and ten.

Higher coverage means a broader set of useful trade-offs. Hover for exact values; error bars show the standard error across 30 runs.

Trade-off coverage after 60 experiments (hypervolume fraction) 1 / round 5 / round 10 / round
STOCHOS 0.9225 0.9150 0.8822
BoTorch (qNEHVI) 0.9105 0.9045 0.8917
BayBE 0.8900 0.8838 0.8726
Random forest 0.7740 0.7785 0.7679
Classical DoE 0.7345 0.7345 0.7345
Random search 0.6840 0.6840 0.6840

STOCHOS leads at round sizes one and five and is second at ten, where the batch of ten commits a sixth of the budget before any result returns. Choosing one recipe at a time, it reaches 80% coverage in a median of 32 experiments and 90% in 46 (on 20 of 30 runs, the most of any method); random forest, classical DoE and random search never reach 80% inside the budget. The extended-budget study of the classical design makes the gap concrete: a D-optimal plan needs 180 experiments to reach the 70% coverage that STOCHOS passes at 21.

Choose two objectives to compare the best trade-offs each method found across 30 runs. Higher strength, lower CO₂ and lower cost are preferred: towards the lower right in the strength views and towards the lower left in the CO₂–cost view.

Compare how many experiments each method needs to reach a given level of trade-off coverage. Higher curves indicate broader coverage at the same experiment budget.

Cost of the decisions. Part B’s sequential runs have median elapsed times of 26.7 minutes for STOCHOS, 50.2 minutes for BoTorch and 147.6 minutes for BayBE over 30 runs per method, each with a 60-experiment budget. STOCHOS uses approximately 47% less elapsed time than BoTorch in this comparison, measured for model fitting and proposal generation in the virtual benchmark. The methods with the lowest computational overhead, classical DoE and random search, have lower reported coverage.

Hover to see the time spent fitting models and proposing recipes. Compare the median computational time per campaign with the trade-off coverage above.


Appendix A. Sources and authorship

Public leaderboard rows, published baseline values and internal experiments are distinguished in the relevant sections. Numerical ranks refer to the stated comparison fields and snapshot dates. PLAID scores are computed by the benchmark platform; STOCHOS training and timing measurements are reported by the authors.

This report is authored by PI Probaligence GmbH, the developer of STOCHOS. References link to the benchmark datasets, published methods and evaluation protocols used in the comparisons.


References

References identify benchmark datasets, evaluation procedures, methods and author descriptions. For tabular entries, registered TabArena references were used where available.

Benchmarks

[1] TabArena: a living benchmark for tabular machine learning. Leaderboard https://huggingface.co/spaces/TabArena/leaderboard, code https://github.com/autogluon/tabrepo. Results accessed 18 August 2026 (42 registered methods).

[2] TALENT: a tabular analytics and learning toolbox. https://github.com/LAMDA-Tabular/TALENT.

Pretrained tabular models

[3] TabFM. Google Research, released 30 June 2026. https://github.com/google-research/tabfm. Released configuration: 1,639,444,298 parameters (§3). Consult the repository for license terms.

[4] EXAONE-Tabular. LG AI Research, released 31 July 2026. https://github.com/LGAI-Research/EXAONE-Tabular. Consult the repository for code and model-license terms.

[5] Prior Labs. TabPFN-3: Technical Report. arXiv:2605.13986. https://priorlabs.ai/technical-reports/tabpfn-3. Consult the model release for license terms.

[6] TabPFN-2.6 and RealTabPFN-2.5. arXiv:2511.08667.

[7] TabICL: a tabular foundation model for in-context learning on large data. arXiv:2502.05564. TabICLv2: arXiv:2602.11139. BSD-3-Clause.

[8] LimiX. arXiv:2509.03505.

[9] TabDPT. arXiv:2410.18164.

[10] Mitra. arXiv:2510.21204.

[11] TabPFN-Wide. arXiv:2510.06162.

[12] TabSTAR. arXiv:2505.18125.

[13] SAP-RPT-OSS. arXiv:2506.10707.

[14] OrionMSP. arXiv:2511.02818.

[15] iLTM. arXiv:2511.15941.

[16] Nori / Nori-30M. https://github.com/Synthefy/synthefy-nori.

[17] TabSwift. https://github.com/LAMDA-Tabular/TabSwift.

Trees, networks and classical baselines

[18] LightGBM: a highly efficient gradient boosting decision tree. NeurIPS 2017. https://papers.nips.cc/paper_files/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html

[19] CatBoost: unbiased boosting with categorical features. arXiv:1706.09516.

[20] XGBoost: a scalable tree boosting system. arXiv:1603.02754.

[21] RealMLP — Better by default: strong pre-tuned MLPs and boosted trees on tabular data. arXiv:2407.04491.

[22] TabM. arXiv:2410.24210.

[23] ModernNCA. arXiv:2407.03257.

[24] xRFM. arXiv:2508.10053.

[25] Explainable Boosting Machine (EBM), Lou et al., KDD 2013.

[26] Random Forests, Breiman (2001), and Extremely Randomized Trees, Geurts et al. (2006).

[27] TorchMLP / FastaiMLP. arXiv:2003.06505.

[28] PerpetualBooster. https://perpetual-ml.com/.

Statistical methods used in this document

[29] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39(3/4), 324–345 (1952). The model underlying the Elo ratings defined in §2.0.

[30] H. Theil (1950); P. K. Sen. Estimates of the regression coefficient based on Kendall’s tau. JASA 63(324), 1379–1389 (1968). The robust slope estimator used in §4.5.

[31] F. Wilcoxon. Individual comparisons by ranking methods. Biometrics Bulletin 1(6), 80–83 (1945). The paired test used in §4.2.

[32] J. Demšar. Statistical comparisons of classifiers over multiple data sets. JMLR 7, 1–30 (2006). On why paired, non-parametric tests are the appropriate instrument for benchmark comparisons of this shape.

Physics benchmarks and surrogate models (§§6–17)

[33] LagrangeBench: a Lagrangian fluid mechanics benchmarking suite. A. P. Toshev, G. Galletti, F. Fritz, S. Adami, N. A. Adams. NeurIPS 2023 Datasets & Benchmarks. arXiv:2309.16342. Datasets on Zenodo, CC-BY.

[34] Learning to simulate complex physics with graph networks (GNS). A. Sanchez-Gonzalez, J. Godwin, T. Pfaff, R. Ying, J. Leskovec, P. W. Battaglia. ICML 2020. arXiv:2002.09405. Source of the Water-3D dataset and protocol used in §8.

[35] Geometric and physical quantities improve E(3) equivariant message passing (SEGNN). J. Brandstetter, R. Hesselink, E. van der Pol, E. J. Bekkers, M. Welling. ICLR 2022 (spotlight). arXiv:2110.02905. Baseline in [33].

[36] CoRGI: convolutional residual global interactions. KDD 2026. arXiv:2511.22938. The strongest published per-cell numbers on most of the board in §7.2.

[37] Neural-SPH. arXiv:2402.06275. Source of the 400-step long-horizon protocol used in the supplement to §7.2.

[38] NeuralMPM. TMLR. arXiv:2408.15753. The only surveyed method reporting full-rollout MSE in the same convention as [34] (§8.2).

[39] AhmedML. N. Ashton et al. High-fidelity CFD dataset over the Ahmed body, 500 geometries. https://huggingface.co/datasets/neashton/ahmedml.

[40] AB-UPT: anchored-branched universal physics transformers. arXiv:2502.09692. Source of the baseline tables used in §§10.2 and 11. https://arxiv.org/abs/2502.09692.

[41] PLAID: a benchmark suite of physics-learning datasets. Safran. arXiv:2505.02974. Rotor37 leaderboard hosted at https://huggingface.co/PLAIDcompetitions. CC-BY-SA 4.0.

[42] EqGINO: equivariant geometry-informed Fourier neural operators for 3D PDEs. S. Kim, J. Song, S. Shin, G. Cho, S. Kim, C. Park. arXiv:2606.03260.

[43] DrivAerML: high-fidelity computational fluid dynamics dataset for road-car external aerodynamics. N. Ashton et al. arXiv:2408.11969. 500 morphed DrivAer variants, SA-sigma-DDES, ~160M cells; source of the solver-cost figures quoted in §11.

[44] Douglas Boubert. How an AI Agent Cracked an Industry Benchmark. PhysicsX, 11 August 2026. https://www.physicsx.ai/newsroom/how-an-ai-agent-cracked-an-industry-benchmark. Author description of the solver-based Tensile2d entry (§15).

[45] I-C. Yeh. Modeling of strength of high-performance concrete using artificial neural networks. Cement and Concrete Research 28(12), 1998. Dataset: UCI Machine Learning Repository, Concrete Compressive Strength (1,030 recipes).

[46] M. Balandat et al. BoTorch: A framework for efficient Monte-Carlo Bayesian optimization. NeurIPS 2020. botorch.org — qNEHVI acquisition, version 0.18.1 in §18.

[47] BayBE — Bayesian back end, Merck KGaA. emdgroup.github.io/baybe — version 0.15.0 in §18.

[48] T. Pfaff, M. Fortunato, A. Sanchez-Gonzalez, P. W. Battaglia. Learning mesh-based simulation with graph networks (MeshGraphNets). ICLR 2021. arXiv:2010.03409. Source of the DeformingPlate data and protocol of §9.

[49] M4GN: mesh-based multi-segment hierarchical graph network. Transactions on Machine Learning Research, 2025 (openreview.net/forum?id=R3vDbqWa1v). Reported 2.65 on DeformingPlate; no code released.

[50] MGN-T: MeshGraphNet-Transformer for solid mechanics. arXiv:2601.23177. DeformingPlate data regenerated with COMSOL.

[51] Y.-Y. Yu et al. Learning Flexible Body Collision Dynamics with Hierarchical Contact Mesh Transformer. ICLR 2024. https://arxiv.org/abs/2312.12467.

[52] Diffusion-Based Hierarchical Graph Neural Networks for Simulating Nonlinear Solid Mechanics (ROBIN). NeurIPS 2025. https://arxiv.org/abs/2506.06045.

[53] EvoMesh: Adaptive Physical Simulation with Hierarchical Graph Evolutions. ICML 2025. https://hbell99.github.io/evo-mesh/.

[54] Y. Cao, M. Chai, M. Li and C. Jiang. Efficient Learning of Mesh-Based Physical Simulation with Bi-Stride Multi-Scale Graph Neural Network. ICML 2023. https://proceedings.mlr.press/v202/cao23a.html.

[55] A. Vaswani et al. Attention Is All You Need. NeurIPS 2017. https://arxiv.org/abs/1706.03762.

[56] H. Wu et al. Transolver: A Fast Transformer Solver for PDEs on General Geometries. ICML 2024. https://arxiv.org/abs/2402.02366.

[57] Z. Li et al. Transformer for Partial Differential Equations’ Operator Learning. ICLR 2023. https://arxiv.org/abs/2205.13671.

[58] B. Alkin et al. Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators. NeurIPS 2024. https://arxiv.org/abs/2402.12365.

[59] H. Gao and S. Ji. Graph U-Nets. ICML 2019. https://arxiv.org/abs/1905.05178.

[60] Z. Li et al. Geometry-Informed Neural Operator for Large-Scale 3D PDEs. NeurIPS 2023. https://arxiv.org/abs/2309.00583.

[61] C. R. Qi et al. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. CVPR 2017. https://arxiv.org/abs/1612.00593.

[62] T. Wang and C. Wang. Latent Neural Operator for Solving Forward and Inverse PDE Problems. NeurIPS 2024. https://arxiv.org/abs/2406.03923.

[63] PI Probaligence. How STOCHOS Works: The DIM-GP Algorithm. Technical description (not a peer-reviewed method paper). https://probaligence.com/how-stochos-works/.

Enlarged figure

Partners, customers, and research collaborators
AnsysCADFEMSimuTech GroupMEScoTSNENAFEMS MemberBoschZFGEMUDLRAdler LackeMankiewiczDuluxPlixxentFraunhoferHochschule NiederrheinFUELL Lab AutomationHumotionUniversitaet HamburgRobert Bosch StiftungITficient