Model card: fraud classifier, demo model (fraud-classifier v3)

The stateless model served by the Phase 4 Gradio demo. It scores one transaction at a time from the fields on the form; it has no access to a card's earlier transactions.

Model xgboost, calibrated with sigmoid scaling
Feature pipeline features-1.0.0, v1 (stateless: transaction, customer and geography features)
Decision threshold 0.007063651457428932 (calibrated probability, picked on validation to meet precision >= 0.50)
Dataset sparkov v1, dvc_data_md5 a03532b1de5c77f62c501042f7ce782e.dir
Split train 2019-01-01..2020-06-30, valid 2020-07-01..2020-09-30, test 2020-10-01..2020-12-31
Git commit b4404fa
MLflow run 1d54533fac6149e58ce0d8693bfae0af (registry alias demo, version 3)
Exported 2026-09-28T15:46:19+00:00

Performance

On the validation split (Jul-Sep 2020), at the chosen threshold: precision 0.500, recall 0.921. This model has not been evaluated on the test split; the single test evaluation of Phase 3 belongs to the registry champion, a different model:

Phase 3 champion (lightgbm, v2 features), test split Value 95% CI
Recall 0.986 0.978-0.993
Precision 0.446 0.423-0.468
PR-AUC 0.973 0.965-0.980

The champion's precision on test misses the project's 0.50 target. Full analysis (error analysis, fairness by age band and gender) is in the repository's docs/model_cards/v0.3-model.md.

Explanations

Each score comes with the top reasons behind it: SHAP contributions summed per input feature and shown next to the value that was submitted. They explain the model's raw score (log-odds, before calibration): sign and ranking carry over to the probability, the magnitude does not.

Known limitations

  • Simulated data (Sparkov). The task is easier than real fraud detection; these numbers will not reproduce on real cardholder data.
  • No card history. The best Phase 3 models use a card's recent velocity and reach ~0.99 recall with history (v2 features); without it this model reaches 0.92 recall on validation at the same precision target. The history-aware champion needs the online history store built in Phase 5.
  • Gender is a model input, and the Phase 3 fairness check found recall and false-alarm-rate gaps between groups for the champion. This model has not had its own fairness evaluation.
  • Not for real decisions. A portfolio demonstration, not a production fraud system.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support