2025
Natural hydrogen (H2) has been largely overlooked in the energy transition due to limited understanding of its geological occurrence, particularly in sedimentary basins where systematic exploration is rare. Our analysis uses Ramsay-2 well from South Australia, which showed a high concentration of Hydrogen as a reference dataset. This study investigates machine learning (ML) regression techniques for predicting hydrogen-rich zones in a historical well using borehole data and derived parameters (i.e., total porosity and water saturation). Four different input feature sets were tested using three supervised ML algorithms—K-Nearest Neighbors (KNN), ExtraTrees, and CatBoost Regressors. Uncertainty bounds were quantified using conformal prediction (via MAPIE), which provides statistically rigorous confidence intervals for model outputs with guaranteed coverage probabilities. The best-performing models have been selected based on regression metrics like RMSE, MSE, R2-score, and coverage score, and then the optimum ExtraTrees Regressor model is used to predict potential hydrogen-bearing zones in the historical well. These predictions align well with known hydrogen intervals in the Ramsay-2, suggesting geological consistency. The study demonstrates the value of ML-based screening for natural hydrogen exploration, especially where labeled data is sparse and conventional methods are limited.
Natural Hydrogen, Machine learning, Borehole logging, Kulpara formation, Stansbury Basin