--- tags: - fraud-detection - tabular-classification - lightgbm - shap library_name: sklearn --- # FraudLens fraud detector (lightgbm, version 2) Credit card fraud scoring model from **FraudLens** (https://github.com/YASHR2002/fraudlens): leakage-free behavioural features, a cost-based decision threshold, SHAP explanations and LLM analyst notes. This repo is what the public demo API downloads at startup. | | Test set (fraudTest, used once) | Validation | |---|---|---| | PR-AUC | 0.9745 | 0.9853 | | Recall at threshold | 95.2% | 97.6% | | Precision at threshold | 86.1% | 90.1% | | Total cost (missed fraud + $5 per false alarm) | $25,150 | $2,975 | Decision threshold: **0.4326** (minimises total cost on validation). ## Contents - `champion/model/`: scikit-learn pipeline in MLflow format, serialised with **skops** (load with the trusted types listed in `fraudlens.models.estimators.SKOPS_TRUSTED_TYPES`). - `champion/metadata.json`: version, threshold, metrics, feature importance, fairness. - `features.json`: the 18 input features with plain-English descriptions. - `state/`: per-card state snapshot (end of the training data) and demo transactions. ## Data and limitations Trained on the synthetic Sparkov dataset (CC0), so scores are higher than real fraud detection would achieve; card numbers and personal fields in the demo files are synthetic. Gender is not a model input; age is, and fairness gaps by gender and age are reported in the project's model card. Not for production use.