Classical ML β Lithology Sequence Models
Classical machine-learning baselines for automated lithological sequence identification from well-log data.
This repository contains trained Random Forest, XGBoost, and LightGBM models for pointwise lithology classification followed by minimum-thickness sequence segmentation and sequence-level evaluation.
The models were trained using the 400-well NoraResearchLab/Lithology-Training-Dataset and evaluated using the independent NoraResearchLab/lithology-sequence-benchmark.
Research Organization
NORA Research Lab β Neural Operated Reasoning Agent
Building intelligence for the real world.
Project Overview
Lithology identification from well logs is commonly approached as a pointwise classification problem. However, geological interpretation is inherently sequential: lithological units occur as intervals and transitions rather than isolated depth samples.
This project therefore evaluates classical machine-learning models through a complete workflow:
Well Logs
β
Feature Preparation
β
Pointwise Lithology Classification
β
Minimum-Thickness Segmentation
β
Lithological Sequences
β
Depth-Level Evaluation
β
Sequence-Level Evaluation
The objective is not simply to maximize pointwise accuracy, but to determine how well models recover realistic lithological sequences and how they generalize to previously unseen wells.
Datasets
Training Dataset
NoraResearchLab/Lithology-Training-Dataset
- 400 wells
- Well-log based lithology classification
- Used for model development and training
- Contains the log-derived features used by the classical ML models
Dataset:
https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset
Independent Benchmark
NoraResearchLab/lithology-sequence-benchmark
The independent benchmark is used to evaluate generalization to wells not used during model training.
Dataset:
https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark
The reported benchmark evaluation contains 66 wells.
Models
Three classical machine-learning approaches were trained and evaluated.
| Model | Saved Model | Parameters |
|---|---|---|
| Random Forest | models/random_forest.joblib |
models/random_forest_params.json |
| XGBoost | models/xgboost.joblib |
models/xgboost_params.json |
| LightGBM | models/lightgbm.joblib |
models/lightgbm_params.json |
Input Features
The models use the following well-log features:
GRβ Gamma RayRHOBβ Bulk DensityNPHIβ Neutron PorosityPEFβ Photoelectric FactorDTβ Compressional Sonic / Delta-Tlog10(RT)β Log-transformed ResistivityCALIβ Caliper- Missingness indicators for the corresponding input measurements
The missingness indicators allow the models to distinguish between measured values and unavailable or missing log responses.
Evaluation Methodology
The evaluation is performed in multiple stages.
1. Pointwise Classification
Each depth sample is classified into a lithological class using the trained machine-learning model.
2. Minimum-Thickness Segmentation
The raw pointwise predictions are converted into geological intervals using minimum-thickness constraints.
This reduces unrealistic rapid class switching and produces more meaningful lithological sequences.
3. Depth-Level Evaluation
Predictions are evaluated at the individual depth-sample level using:
- Accuracy
- Macro F1
- Confident Macro F1
4. Sequence-Level Evaluation
The resulting lithological sequences are evaluated using:
- Sequence similarity
This provides a more geological interpretation of model performance than pointwise accuracy alone.
Benchmark Results
Evaluation on the independent benchmark:
| Model | Wells | Depth Accuracy | Depth Macro F1 | Confident Macro F1 | Sequence Similarity | CV Macro F1 |
|---|---|---|---|---|---|---|
| Random Forest | 66 | 0.7914 | 0.6279 | 0.5900 | 0.5927 | 0.7599 |
| XGBoost | 66 | 0.7271 | 0.5757 | 0.5293 | 0.5173 | 0.7571 |
| LightGBM | 66 | 0.6936 | 0.5321 | 0.4834 | 0.4818 | 0.7266 |
Model Performance Visualization
Benchmark Interpretation
Random Forest achieved the strongest performance across the reported independent benchmark metrics.
Key results:
- 79.14% depth accuracy
- 62.79% depth macro F1
- 59.00% confident macro F1
- 59.27% sequence similarity
- 75.99% cross-validation macro F1
The results indicate that Random Forest provides the strongest overall classical baseline among the three evaluated approaches on this benchmark.
Blind-Well Generalization Test
A separate blind-well experiment was conducted on:
Well 16/2-7
The purpose was to measure how well model performance transfers to a previously unseen well.
| Model | Original Training Benchmark | Blind Well Accuracy | Approx. Performance Drop |
|---|---|---|---|
| Random Forest | ~85β90% | 60.17% | ~25β30% |
| XGBoost | ~88β91% | 44.95% | ~43β46% |
| LightGBM | ~89β92% | 40.65% | ~48β51% |
These results demonstrate a substantial generalization gap between conventional model-development performance and performance on an unseen well.
Key Research Findings
1. Severe Generalization Gap
All three models experienced a substantial reduction in performance when evaluated on the blind well.
This indicates that high performance on training or conventional validation splits does not necessarily translate into robust geological generalization.
The result highlights the importance of well-level and spatially independent evaluation for lithology classification.
2. Random Forest Generalized Best
Random Forest achieved the strongest blind-well accuracy:
60.17%
compared with:
- XGBoost: 44.95%
- LightGBM: 40.65%
Within this experiment, Random Forest demonstrated greater robustness to the distribution shift between the training wells and the blind well.
3. Sandstone Generalization Challenge
The blind-well evaluation revealed particularly difficult sandstone predictions.
Although sandstone-related performance was stronger during model development, the model showed substantially lower sandstone recall on the blind well.
This suggests that the physical/log-response characteristics associated with sandstone can vary significantly between wells and geological settings.
A model can therefore learn statistical relationships that perform well within the training distribution without learning sufficiently transferable geological representations.
Academic Implications
The experiments demonstrate several important considerations for machine-learning-based lithology interpretation:
- Random sample splitting can overestimate real-world performance.
- Well-level holdout evaluation is essential for measuring generalization.
- Sequence-level metrics provide additional information beyond pointwise accuracy.
- Tree-based ensemble models can behave differently under geological distribution shift.
- Lithology classes may exhibit substantially different transferability between wells.
- Training on more geographically and geologically diverse wells is likely to be important for robust generalization.
These results motivate further research into:
- geographically independent validation
- basin-level cross-validation
- domain adaptation
- physics-informed features
- sequence models
- transformer-based well-log models
- uncertainty estimation
- larger and more diverse lithology datasets
Reproducibility
The repository contains the trained model artifacts and associated parameter files:
models/
βββ random_forest.joblib
βββ random_forest_params.json
βββ xgboost.joblib
βββ xgboost_params.json
βββ lightgbm.joblib
βββ lightgbm_params.json
The model development workflow is:
NoraResearchLab/Lithology-Training-Dataset
β
Feature Preparation
β
ββββββββββββββΌβββββββββββββ
β β β
Random Forest XGBoost LightGBM
ββββββββββββββΌβββββββββββββ
β
Independent Benchmark
β
Sequence-Level Evaluation
Hyperparameters used for each trained model are provided in the corresponding *_params.json files.
Intended Use
These models are intended primarily for:
- research
- benchmarking
- experimentation
- educational purposes
- automated lithology interpretation research
- comparison against future deep-learning approaches
They are research baselines rather than production-ready geological interpretation systems.
Predictions should not be treated as a substitute for qualified geological or petrophysical interpretation.
Limitations
Geological Distribution Shift
Performance can decline substantially when the model encounters wells with different geological or petrophysical characteristics.
Dataset Diversity
The ability of the models to generalize depends strongly on the geological diversity represented in the training wells.
Class Imbalance
Some lithology classes may be represented more heavily than others, making macro-level metrics particularly important.
Pointwise Classification
The underlying models are classical tabular classifiers and do not explicitly learn long-range geological dependencies.
Sequence Reconstruction
Minimum-thickness segmentation improves the geological structure of predictions but is still a post-processing procedure rather than a learned sequence model.
Future Research
Future research can investigate:
- larger multi-basin training datasets
- geographically separated train/test splits
- cross-basin evaluation
- sequence-aware machine learning
- Hidden Markov Models
- CRF-based lithology decoding
- LSTM/GRU architectures
- Temporal Convolutional Networks
- Transformer-based well-log models
- multimodal well-log learning
- uncertainty-aware lithology prediction
- physics-informed machine learning
- transfer learning between geological regions
The classical ML models in this repository provide a baseline against which these approaches can be evaluated.
Citation
If you use these models, datasets, or benchmark results in research, please reference NORA Research Lab and the corresponding repositories.
NORA Research Lab
- GitHub: https://github.com/Nora-Research-Lab
- Hugging Face: https://huggingface.co/NoraResearchLab
- LinkedIn: https://www.linkedin.com/company/nora-research-lab
- X: https://x.com/noraresearchlab
- Website: https://noraresearchlab.site
Training Dataset
https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset
Independent Benchmark
https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark
Acknowledgements
This work was developed by NORA Research Lab as part of ongoing research into machine learning, geological intelligence, and automated subsurface interpretation.
NORA Research Lab β Building intelligence for the real world.
