Model card: fraud classifier, demo model (fraud-classifier v3)
The stateless model served by the Phase 4 Gradio demo. It scores one transaction at a time from the fields on the form; it has no access to a card's earlier transactions.
| Model | xgboost, calibrated with sigmoid scaling |
| Feature pipeline | features-1.0.0, v1 (stateless: transaction, customer and geography features) |
| Decision threshold | 0.007063651457428932 (calibrated probability, picked on validation to meet precision >= 0.50) |
| Dataset | sparkov v1, dvc_data_md5 a03532b1de5c77f62c501042f7ce782e.dir |
| Split | train 2019-01-01..2020-06-30, valid 2020-07-01..2020-09-30, test 2020-10-01..2020-12-31 |
| Git commit | b4404fa |
| MLflow run | 1d54533fac6149e58ce0d8693bfae0af (registry alias demo, version 3) |
| Exported | 2026-09-28T15:46:19+00:00 |
Performance
On the validation split (Jul-Sep 2020), at the chosen threshold: precision 0.500, recall 0.921. This model has not been evaluated on the test split; the single test evaluation of Phase 3 belongs to the registry champion, a different model:
Phase 3 champion (lightgbm, v2 features), test split |
Value | 95% CI |
|---|---|---|
| Recall | 0.986 | 0.978-0.993 |
| Precision | 0.446 | 0.423-0.468 |
| PR-AUC | 0.973 | 0.965-0.980 |
The champion's precision on test misses the project's 0.50 target. Full analysis (error analysis,
fairness by age band and gender) is in the repository's docs/model_cards/v0.3-model.md.
Explanations
Each score comes with the top reasons behind it: SHAP contributions summed per input feature and shown next to the value that was submitted. They explain the model's raw score (log-odds, before calibration): sign and ranking carry over to the probability, the magnitude does not.
Known limitations
- Simulated data (Sparkov). The task is easier than real fraud detection; these numbers will not reproduce on real cardholder data.
- No card history. The best Phase 3 models use a card's recent velocity and reach ~0.99 recall
with history (
v2features); without it this model reaches 0.92 recall on validation at the same precision target. The history-aware champion needs the online history store built in Phase 5. - Gender is a model input, and the Phase 3 fairness check found recall and false-alarm-rate gaps between groups for the champion. This model has not had its own fairness evaluation.
- Not for real decisions. A portfolio demonstration, not a production fraud system.