Beyond Prediction Accuracy: Evaluating Machine Learning Generalization for Crude-Oil Wax Precipitation

Authors

  • Yaman Hamed CERDAS, Department of Applied Science, Universiti Teknologi PETRONAS, 32610 Seri Iskandar, Perak, Malaysia
  • Louis Tan Eng Hao

Keywords:

wax precipitation, machine learning, cross-crude validation, SHAP Analysis

Abstract

Wax precipitation alters crude-oil flow behaviour and provides the solid material that may later form pipeline deposits. Machine-learning models can accurately predict wax weight, but their ability to generalize beyond the training domain remains uncertain. This study evaluated artificial neural network (ANN), decision tree (DT), random forest (RF), and XGBoost models using 550 PVTSim-generated cases representing five crude-oil compositions. Model inputs included temperature, pressure, crude class, and the carbon fractions C1–C6, C7–C20, and C21+. Two validation strategies were employed: prediction under unseen pressure conditions within the represented crude compositions and leave-one-crude-out validation to assess transferability to unseen crude compositions. All models achieved R² above 0.999 for unseen pressure conditions, with ANN producing the lowest root mean square error (RMSE) of 0.0361 weight percent (wt.%) and a normalized RMSE (NRMSE) of 0.261%. Performance declined substantially when an entire crude composition was excluded from training, and no model generalized reliably across all five crude compositions. SHAP analysis identified temperature as the dominant predictor. Heavy crude oils, characterized by lower light-end (C1–C6) and higher heavy-end (C21+) fractions, consistently produced higher predicted wax weight than light crude oils, whereas pressure had only a weak influence within the simulated range.

Downloads

Published

2026-08-30