Classical ML β€” Lithology Sequence Models

Classical machine-learning baselines for automated lithological sequence identification from well-log data.

This repository contains trained Random Forest, XGBoost, and LightGBM models for pointwise lithology classification followed by minimum-thickness sequence segmentation and sequence-level evaluation.

The models were trained using the 400-well NoraResearchLab/Lithology-Training-Dataset and evaluated using the independent NoraResearchLab/lithology-sequence-benchmark.


Research Organization

NORA Research Lab β€” Neural Operated Reasoning Agent

Building intelligence for the real world.

GitHub Hugging Face LinkedIn X Website


Project Overview

Lithology identification from well logs is commonly approached as a pointwise classification problem. However, geological interpretation is inherently sequential: lithological units occur as intervals and transitions rather than isolated depth samples.

This project therefore evaluates classical machine-learning models through a complete workflow:

Well Logs
    ↓
Feature Preparation
    ↓
Pointwise Lithology Classification
    ↓
Minimum-Thickness Segmentation
    ↓
Lithological Sequences
    ↓
Depth-Level Evaluation
    ↓
Sequence-Level Evaluation

The objective is not simply to maximize pointwise accuracy, but to determine how well models recover realistic lithological sequences and how they generalize to previously unseen wells.


Datasets

Training Dataset

NoraResearchLab/Lithology-Training-Dataset

  • 400 wells
  • Well-log based lithology classification
  • Used for model development and training
  • Contains the log-derived features used by the classical ML models

Dataset:

https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset


Independent Benchmark

NoraResearchLab/lithology-sequence-benchmark

The independent benchmark is used to evaluate generalization to wells not used during model training.

Dataset:

https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark

The reported benchmark evaluation contains 66 wells.


Models

Three classical machine-learning approaches were trained and evaluated.

Model Saved Model Parameters
Random Forest models/random_forest.joblib models/random_forest_params.json
XGBoost models/xgboost.joblib models/xgboost_params.json
LightGBM models/lightgbm.joblib models/lightgbm_params.json

Input Features

The models use the following well-log features:

  • GR β€” Gamma Ray
  • RHOB β€” Bulk Density
  • NPHI β€” Neutron Porosity
  • PEF β€” Photoelectric Factor
  • DT β€” Compressional Sonic / Delta-T
  • log10(RT) β€” Log-transformed Resistivity
  • CALI β€” Caliper
  • Missingness indicators for the corresponding input measurements

The missingness indicators allow the models to distinguish between measured values and unavailable or missing log responses.


Evaluation Methodology

The evaluation is performed in multiple stages.

1. Pointwise Classification

Each depth sample is classified into a lithological class using the trained machine-learning model.

2. Minimum-Thickness Segmentation

The raw pointwise predictions are converted into geological intervals using minimum-thickness constraints.

This reduces unrealistic rapid class switching and produces more meaningful lithological sequences.

3. Depth-Level Evaluation

Predictions are evaluated at the individual depth-sample level using:

  • Accuracy
  • Macro F1
  • Confident Macro F1

4. Sequence-Level Evaluation

The resulting lithological sequences are evaluated using:

  • Sequence similarity

This provides a more geological interpretation of model performance than pointwise accuracy alone.


Benchmark Results

Evaluation on the independent benchmark:

Model Wells Depth Accuracy Depth Macro F1 Confident Macro F1 Sequence Similarity CV Macro F1
Random Forest 66 0.7914 0.6279 0.5900 0.5927 0.7599
XGBoost 66 0.7271 0.5757 0.5293 0.5173 0.7571
LightGBM 66 0.6936 0.5321 0.4834 0.4818 0.7266

Model Performance Visualization

Lithology Classical ML Model Evaluation

Benchmark Interpretation

Random Forest achieved the strongest performance across the reported independent benchmark metrics.

Key results:

  • 79.14% depth accuracy
  • 62.79% depth macro F1
  • 59.00% confident macro F1
  • 59.27% sequence similarity
  • 75.99% cross-validation macro F1

The results indicate that Random Forest provides the strongest overall classical baseline among the three evaluated approaches on this benchmark.


Blind-Well Generalization Test

A separate blind-well experiment was conducted on:

Well 16/2-7

The purpose was to measure how well model performance transfers to a previously unseen well.

Model Original Training Benchmark Blind Well Accuracy Approx. Performance Drop
Random Forest ~85–90% 60.17% ~25–30%
XGBoost ~88–91% 44.95% ~43–46%
LightGBM ~89–92% 40.65% ~48–51%

These results demonstrate a substantial generalization gap between conventional model-development performance and performance on an unseen well.


Key Research Findings

1. Severe Generalization Gap

All three models experienced a substantial reduction in performance when evaluated on the blind well.

This indicates that high performance on training or conventional validation splits does not necessarily translate into robust geological generalization.

The result highlights the importance of well-level and spatially independent evaluation for lithology classification.


2. Random Forest Generalized Best

Random Forest achieved the strongest blind-well accuracy:

60.17%

compared with:

  • XGBoost: 44.95%
  • LightGBM: 40.65%

Within this experiment, Random Forest demonstrated greater robustness to the distribution shift between the training wells and the blind well.


3. Sandstone Generalization Challenge

The blind-well evaluation revealed particularly difficult sandstone predictions.

Although sandstone-related performance was stronger during model development, the model showed substantially lower sandstone recall on the blind well.

This suggests that the physical/log-response characteristics associated with sandstone can vary significantly between wells and geological settings.

A model can therefore learn statistical relationships that perform well within the training distribution without learning sufficiently transferable geological representations.


Academic Implications

The experiments demonstrate several important considerations for machine-learning-based lithology interpretation:

  1. Random sample splitting can overestimate real-world performance.
  2. Well-level holdout evaluation is essential for measuring generalization.
  3. Sequence-level metrics provide additional information beyond pointwise accuracy.
  4. Tree-based ensemble models can behave differently under geological distribution shift.
  5. Lithology classes may exhibit substantially different transferability between wells.
  6. Training on more geographically and geologically diverse wells is likely to be important for robust generalization.

These results motivate further research into:

  • geographically independent validation
  • basin-level cross-validation
  • domain adaptation
  • physics-informed features
  • sequence models
  • transformer-based well-log models
  • uncertainty estimation
  • larger and more diverse lithology datasets

Reproducibility

The repository contains the trained model artifacts and associated parameter files:

models/
β”œβ”€β”€ random_forest.joblib
β”œβ”€β”€ random_forest_params.json
β”œβ”€β”€ xgboost.joblib
β”œβ”€β”€ xgboost_params.json
β”œβ”€β”€ lightgbm.joblib
└── lightgbm_params.json

The model development workflow is:

NoraResearchLab/Lithology-Training-Dataset
                    ↓
             Feature Preparation
                    ↓
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       ↓            ↓            ↓
 Random Forest    XGBoost     LightGBM
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    ↓
       Independent Benchmark
                    ↓
        Sequence-Level Evaluation

Hyperparameters used for each trained model are provided in the corresponding *_params.json files.


Intended Use

These models are intended primarily for:

  • research
  • benchmarking
  • experimentation
  • educational purposes
  • automated lithology interpretation research
  • comparison against future deep-learning approaches

They are research baselines rather than production-ready geological interpretation systems.

Predictions should not be treated as a substitute for qualified geological or petrophysical interpretation.


Limitations

Geological Distribution Shift

Performance can decline substantially when the model encounters wells with different geological or petrophysical characteristics.

Dataset Diversity

The ability of the models to generalize depends strongly on the geological diversity represented in the training wells.

Class Imbalance

Some lithology classes may be represented more heavily than others, making macro-level metrics particularly important.

Pointwise Classification

The underlying models are classical tabular classifiers and do not explicitly learn long-range geological dependencies.

Sequence Reconstruction

Minimum-thickness segmentation improves the geological structure of predictions but is still a post-processing procedure rather than a learned sequence model.


Future Research

Future research can investigate:

  • larger multi-basin training datasets
  • geographically separated train/test splits
  • cross-basin evaluation
  • sequence-aware machine learning
  • Hidden Markov Models
  • CRF-based lithology decoding
  • LSTM/GRU architectures
  • Temporal Convolutional Networks
  • Transformer-based well-log models
  • multimodal well-log learning
  • uncertainty-aware lithology prediction
  • physics-informed machine learning
  • transfer learning between geological regions

The classical ML models in this repository provide a baseline against which these approaches can be evaluated.


Citation

If you use these models, datasets, or benchmark results in research, please reference NORA Research Lab and the corresponding repositories.

NORA Research Lab

Training Dataset

https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset

Independent Benchmark

https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark


Acknowledgements

This work was developed by NORA Research Lab as part of ongoing research into machine learning, geological intelligence, and automated subsurface interpretation.

NORA Research Lab β€” Building intelligence for the real world.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train NoraResearchLab/lithology-classical-ml