nafisatibrahim's picture
Update README.md
3e3e38e verified
|
Raw History Blame Contribute Delete
3.79 kB
---
license: mit
tags:
- tabular-classification
- credit-risk
- xgboost
- scikit-learn
pipeline_tag: tabular-classification
---
# Model Card: Credit Risk Default Champion Model
> Numeric performance metrics and operating threshold are filled in once the final-model calibration and threshold cells (`notebooks/03_modeling.ipynb`, sections XXII-XXIII) have been run.
## Model details
| Property | Value |
|---|---|
| Model type | XGBoost, `scale_pos_weight`-balanced, hyperparameter-tuned via `RandomizedSearchCV`, retrained on a reduced 49-variable feature set (chosen by cumulative SHAP importance, one-hot dummies grouped back to their source variable), Platt-scaled for calibration. Selected over Logistic Regression, Random Forest, KNN, LightGBM, CatBoost, and several XGBoost variants (resampling, cost-sensitive objective, blend/stack ensembles) compared in `notebooks/03_modeling.ipynb` |
| Version | v0.1 |
| Trained by | Nafisat Ibrahim, Marienne Dosso, Bintou Ba |
| Date trained | TBD |
| Framework | scikit-learn, XGBoost |
## Intended use
**Primary use:** Predicts the probability that a personal loan will default, using only borrower and loan characteristics available at origination, surfaced through an interactive review dashboard for credit officers at NorthBay Bank (a fictional bank case study).
**Primary users:** Credit officers
**Out-of-scope uses:** Not intended for automated approve/deny decisions without human review; not validated for loan products, lenders, or borrower populations outside the LendingClub 2007-2018 dataset it was trained on.
## Training data
See `data/DATA_CARD.md`.
- Dataset: LendingClub accepted loan data, 2007-2018 ([`BuildersLab/loan-application-dataset`](https://huggingface.co/datasets/BuildersLab/loan-application-dataset), `feature_engineered` config)
- Train / val / test split: stratified 70/15/15 (941,744 / 201,803 / 201,802 rows)
- Features: 49 of 107 original variables (selected by cumulative SHAP importance on the full model, 90%+ of total importance), re-encoded to 119 columns after one-hot expansion of the categorical ones kept (`addr_state`, `home_ownership`, `purpose`, `verification_status`)
## Performance metrics
Pending: run `notebooks/03_modeling.ipynb` cells `final_calib_score` (Brier/PR-AUC before and after calibration) and `final_test_check` (precision/recall/F1/PR-AUC/ROC-AUC at the chosen threshold, on test) and copy the printed values here.
| Metric | Value |
|---|---|
| PR-AUC | TBD, pending run |
| Recall at threshold | TBD, pending run |
| Precision at threshold | TBD, pending run |
| False positives per 1,000 negatives | TBD, pending run |
| Business impact | TBD |
## Operating threshold
**Chosen threshold:** TBD, pending run of `final_cost_optimal` in `notebooks/03_modeling.ipynb`
**Rationale:** Cost-optimal, not F1-optimal. False negatives (missed defaults) cost more than false positives, weighted 5:1 in the cost-sensitive threshold sweep, this is also where F1 itself peaks, so it isn't a tradeoff against F1, both objectives agree on this point.
## Explainability
SHAP values are computed for every prediction. See `notebooks/04_explainability.ipynb`.
The Gemini API translates SHAP output into plain English for credit officers.
Gemini output is advisory only. The model score is the authoritative decision.
## Bias evaluation
TODO: document bias audit results across demographic and geographic groups. See `data/DATA_CARD.md`'s Bias considerations for the known open items (no formal audit conducted yet).
## Limitations
See `data/DATA_CARD.md`'s Known limitations: origination-time features only, resolved-outcome loans only, U.S./LendingClub-only population, 2007-2018 vintages.
## How to reproduce
```bash
make data
make train
make evaluate
```