Spaces:
Runtime error
Runtime error
|
Download docs/DATASETS.md from yusufcalisir/Collaborative-Fraud-Intelligence-Simulator: direct link, hf CLI and curl.
- Browser
- Download file 41.3 kB
-
https://huggingface.co/spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/resolve/main/docs/DATASETS.md
- Command line
-
hf download hf://spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/docs/DATASETS.md
-
curl -L -o DATASETS.md https://huggingface.co/spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/resolve/main/docs/DATASETS.md
41.3 kB
| # Authoritative Dataset Cards & Storage Architecture Specification (`DATASETS.md`) | |
| ## Privacy-Preserving Collaborative Financial Crime Intelligence Platform (CF-Intelligence) | |
| **Specification Version:** `2.4.0-enterprise` | |
| **Standard:** Hugging Face Datasets / ACM FAccT Data Cards / Datasheets for Datasets (Gebru et al., 2021) | |
| **Governance Scope:** Federal Reserve SR 11-7 / OCC 2011-12, EU AI Act Article 10, GDPR Article 9, ECOA Regulation B (12 CFR Part 1002), PCI-DSS v4.0 | |
| **Repository Location:** `DATASETS.md` (root) and `docs/DATASETS.md` | |
| --- | |
| ## 1. Executive Summary & Zero-Mock Dataset Governance | |
| A core architectural invariant of **CF-Intelligence** is **Zero-Mock, Zero-Dummy Data** in production and empirical benchmarking evaluations. While development unit tests may isolate specific interfaces, all empirical fraud detection claims, federated learning convergence benchmarks, and differential privacy trade-offs in this platform are calibrated against eight canonical, large-scale financial crime datasets spanning over **8.9 million transactions**: | |
| ### Master Dataset Inventory & Licensing Matrix | |
| | Dataset ID | Dataset Name | Domain Scope | Real vs Synthetic | Distribution Source & Archive | Copyright & License | Class Imbalance Ratio | Protected Demographic Fields | | |
| |:---|:---|:---|:---:|:---|:---|:---:|:---:| | |
| | **`paysim`** | PaySim Mobile Money | Mobile Money Remittance | Synthetic Agent Simulator | Kaggle (`ealaxi/paysim1`) | CC BY-SA 4.0 | 0.129% (1:774) | **0 / 10 (0.0%)** | | |
| | **`ieee_cis`** | IEEE-CIS Card Fraud | E-Commerce Payments | Real Production Logs | Kaggle (`ieee-fraud-detection`) | Competition License (Research) | 3.500% (1:28) | **0 / 10 (0.0%)** | | |
| | **`credit_card`** | ULB European Card | Consumer Card Payments | Real Anonymized PCA | Kaggle (`mlg-ulb/creditcardfraud`) | Open Database License (ODbL 1.0) | 0.173% (1:578) | **0 / 10 (0.0%)** | | |
| | **`elliptic`** | Elliptic Bitcoin Graph | Cryptocurrency Forensics | Real Blockchain DAG | Kaggle (`ellipticco/elliptic-data-set`) | CC BY 4.0 | 9.76% labeled (1:9) | **0 / 10 (0.0%)** | | |
| | **`amlsim`** | IBM AMLSim Graph | Multi-Agent Banking Graph | Synthetic Multi-Agent | Kaggle / GitHub (`IBM/AMLSim`) | Apache 2.0 | 0.130% (1:769) | **0 / 10 (0.0%)** | | |
| | **`synthaml`** | SynthAML Spar Nord | European Commercial AML | Real-Topology SDV Copula | Nature Scientific Data / Figshare | CC BY 4.0 | 8.500% (1:11) | **0 / 10 (0.0%)** | | |
| | **`amlnet`** | AMLNet AUSTRAC | Australian Wire Compliance | Synthetic Agent Simulation | Zenodo (`10.5281/zenodo.10058474`) | CC BY-NC 4.0 | 0.140% (1:714) | **0 / 10 (0.0%)** | | |
| | **`cross_bank`** | CFI-CrossBank-01 | Consortium Multi-Bank | Synthetic Real-Topology | CFI Research Consortium Generator | Proprietary Research (CFI) | 1.830% (1:55) | **0 / 10 (0.0%)** | | |
| --- | |
| ## 2. Directory Layout & Local Storage Topology | |
| All raw dataset files are stored in `backend/storage/datasets/<dataset_name>/` or `storage/datasets/<dataset_name>/`. The dynamic loader (`resolve_dataset_dir` in `dataloader.py`) resolves storage across containerized and development environments: | |
| ``` | |
| storage/datasets/ | |
| βββ amlsim/ | |
| β βββ accounts.csv # 326 KB (10,000 accounts metadata) | |
| β βββ alerts.csv # 88 KB (1,719 ground-truth AML typology alerts) | |
| β βββ transactions.csv # 61.05 MB (1,323,234 raw transaction flow events) | |
| β βββ transactions.parquet # 11.43 MB (Zero-copy fast columnar cache) | |
| βββ creditcard/ | |
| β βββ creditcard.csv # 143.84 MB (284,807 rows) | |
| βββ elliptic/ | |
| β βββ elliptic_txs_classes.csv # 3.15 MB (Node classification labels) | |
| β βββ elliptic_txs_edgelist.csv # 4.26 MB (Directed transaction flow graph) | |
| β βββ elliptic_txs_features.csv # 657.73 MB (166-dimensional node embeddings) | |
| βββ ieee_cis/ | |
| β βββ train_transaction.csv # 651.69 MB (Primary training transactions) | |
| β βββ train_identity.csv # 25.30 MB (Device & IP network identity) | |
| β βββ test_transaction.csv # 584.79 MB (Out-of-time evaluation transactions) | |
| β βββ test_identity.csv # 24.60 MB (Out-of-time identity features) | |
| βββ paysim/ | |
| β βββ PS_20174392719_1491204439457_log.csv # 470.67 MB (6.36M simulated mobile money records) | |
| β βββ bank_alpha.parquet # Partitioned Non-IID client split (Bank Alpha) | |
| β βββ bank_beta.parquet # Partitioned Non-IID client split (Bank Beta) | |
| β βββ bank_gamma.parquet # Partitioned Non-IID client split (Bank Gamma) | |
| βββ synthaml/ | |
| β βββ alerts.csv # 195 KB (5,000 alert metadata records) | |
| β βββ alerts.parquet # 82 KB (Zero-copy fast columnar alert cache) | |
| β βββ transactions.csv # 9.87 MB (92,261 lookback transactions) | |
| β βββ transactions.parquet # 2.92 MB (Zero-copy fast columnar transaction cache) | |
| βββ amlnet/ | |
| βββ amlnet_transactions.csv # 5.48 MB (25,000 canonical AUSTRAC transactions) | |
| βββ transactions.csv # 5.48 MB (Primary transaction flow log) | |
| βββ transactions.parquet # 900 KB (Zero-copy fast columnar cache) | |
| βββ cross_bank/ | |
| βββ config.json # Multi-bank topology specification (Alpha, Beta, Gamma) | |
| βββ results.json # Full empirical benchmark metrics (PR-AUC, F1, latency) | |
| βββ metrics.csv # Tabular client-by-client performance breakdown | |
| βββ report.md # Full empirical research report with LaTeX formulas | |
| βββ plots/ # Comparative PR curves and cross-bank topology graphs | |
| ``` | |
| --- | |
| ## 3. Authoritative Dataset Cards (8 Primary Benchmarks) | |
| ### 3.1 PaySim Mobile Money Fraud (`paysim`) | |
| #### 3.1.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: PaySim: A Financial Mobile Money Simulator for Fraud Detection | |
| - **Authors**: Edgar Lopez-Rojas, Ahmad Elmir, and Stefan Axelsson (Blekinge Institute of Technology) | |
| - **Publication**: IEEE 28th International Conference on Tools with Artificial Intelligence (ICTAI), 2016 | |
| - **Distribution Source**: Kaggle (`ealaxi/paysim1`) | |
| - **Copyright & License**: Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) | |
| - **Designation**: Synthetic Agent-Based Mobile Money Simulation (empirically calibrated against 1 month of anonymized financial logs from an African mobile money service operator, M-Pesa Kenya topology) | |
| #### 3.1.2 Scale & Class Balance Profile | |
| - **Total Transactions**: $N = 6{,}362{,}620$ payment events (470.7 MB CSV on disk) | |
| - **Fraudulent Transactions**: $N_{\mathrm{fraud}} = 8{,}213$ confirmed illicit transactions | |
| - **Legitimate Transactions**: $N_{\mathrm{legit}} = 6{,}354{,}407$ genuine transactions | |
| - **Fraud Prevalence**: $\pi = 0.1291\%$ (Class Imbalance Ratio $\approx 774:1$) | |
| - **Simulation Duration**: 744 hourly timesteps ($31$ simulated calendar days) | |
| #### 3.1.3 Feature Architecture & Transformations | |
| - **Canonical Feature Pipeline (13 Features)**: | |
| 1. `step`: Integer hour of simulation ($1 \le t \le 744$). | |
| 2. `type_TRANSFER`: Binary indicator for cross-account fund transfers ($y=1$ candidate). | |
| 3. `type_CASH_OUT`: Binary indicator for cash withdrawal via agent ($y=1$ candidate). | |
| 4. `type_PAYMENT`: Binary indicator for merchant retail purchases ($y=0$ strictly). | |
| 5. `type_DEBIT`: Binary indicator for core banking debit transfers ($y=0$ strictly). | |
| 6. `type_CASH_IN`: Binary indicator for cash deposit via agent ($y=0$ strictly). | |
| 7. `amount`: Transacted currency value. | |
| 8. `oldbalanceOrg`: Originator account balance prior to transaction. | |
| 9. `newbalanceOrig`: Originator balance post-transaction (typically zeroed in fraud). | |
| 10. `oldbalanceDest`: Recipient account balance prior to transaction. | |
| 11. `newbalanceDest`: Recipient balance post-transaction. | |
| 12. `errorBalanceOrig`: Sender accounting discrepancy delta: | |
| $$\Delta \mathrm{bal}_{\mathrm{orig}} = \mathrm{newbalanceOrig} + \mathrm{amount} - \mathrm{oldbalanceOrg}$$ | |
| 13. `errorBalanceDest`: Recipient accounting discrepancy delta: | |
| $$\Delta \mathrm{bal}_{\mathrm{dest}} = \mathrm{oldbalanceDest} + \mathrm{amount} - \mathrm{newbalanceDest}$$ | |
| #### 3.1.4 Data Hygiene, Biases & Limitations | |
| - **Selective Fraud Typologies**: Fraud occurs strictly within `TRANSFER` and `CASH_OUT` transaction types; zero fraud cases exist in `PAYMENT`, `CASH_IN`, or `DEBIT`. Models trained without type filtering may learn trivial shortcut heuristics. | |
| - **Absence of Real PII**: No protected or special-category demographic attributes (0/10) are present in the evaluated public datasets (relevant to GDPR Article 9 and ECOA Regulation B exclusions). | |
| - **Deterministic Evasion**: Synthetic fraudsters execute simple double-step asset drain attacks without sophisticated multi-hop laundering chains. | |
| --- | |
| ### 3.2 IEEE-CIS E-Commerce Fraud Detection (`ieee_cis`) | |
| #### 3.2.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: IEEE-CIS Fraud Detection Benchmark | |
| - **Authors**: IEEE Computational Intelligence Society (IEEE-CIS) & Vesta Corporation | |
| - **Publication**: Kaggle Competition Benchmark, 2019 | |
| - **Distribution Source**: Kaggle (`ieee-fraud-detection`) | |
| - **Copyright & License**: Vesta Corporation Competition Dataset License (authorized for academic, scientific, and open-source benchmark evaluation) | |
| - **Designation**: Real Production E-Commerce Card-Not-Present (CNP) Transactions | |
| #### 3.2.2 Scale & Class Balance Profile | |
| - **Total Transactions**: $N = 590{,}540$ training transactions + $144{,}233$ identity metadata records (~1.29 GB raw disk footprint) | |
| - **Fraudulent Transactions**: $N_{\mathrm{fraud}} = 20{,}663$ confirmed chargeback fraud events | |
| - **Legitimate Transactions**: $N_{\mathrm{legit}} = 569{,}877$ genuine purchases | |
| - **Fraud Prevalence**: $\pi = 3.4989\%$ (Class Imbalance Ratio $\approx 28:1$) | |
| - **Temporal Horizon**: Elapsed seconds across 182 calendar days ($t \in [86{,}400, 15{,}811{,}131]$ seconds) | |
| #### 3.2.3 Feature Architecture & Identity Join | |
| - `train_transaction.csv` (394 attributes): Transaction amount, product code (`ProductCD`), card metadata (`card1`β`card6`), address codes (`addr1`β`addr2`), email domains, count features (`C1`β`C14`), timedeltas (`D1`β`D15`), match flags (`M1`β`M9`), and Vesta risk indicators (`V1`β`V339`). | |
| - `train_identity.csv` (41 attributes): Identity verification metadata (`id_01`β`id_38`), `DeviceType`, and `DeviceInfo`. | |
| - **Left Join**: Merged along `TransactionID` ($144{,}233$ transactions with identity metadata, ~24.4% join rate; transactions without identity records receive `has_identity = 0.0` and imputed indicators). | |
| #### 3.2.4 Data Hygiene, Biases & Limitations | |
| - **High Missingness**: Over 200 features exhibit $>50\%$ missing values (especially identity features and V-features). Demands explicit missingness indicator encoding rather than mean imputation. | |
| - **Obfuscated Semantics**: Proprietary V-features ($V_1 \dots V_{339}$) mask engineering logic, limiting direct human-in-the-loop interpretability without SHAP attribution. | |
| - **Device Fingerprint Churn**: Operating system and browser versions undergo rapid natural obsolescence over the 182-day period. | |
| --- | |
| ### 3.3 European Credit Card Fraud (`credit_card`) | |
| #### 3.3.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: Credit Card Fraud Detection (European Cardholders) | |
| - **Authors**: Andrea Dal Pozzolo, Olivier Caelen, Reid A. Johnson, and Gianluca Bontempi | |
| - **Publication**: *Calibrating Probability with Undersampling for Unbalanced Classification*, IEEE SSCI, 2015 | |
| - **Research Institution**: Machine Learning Group (MLG), UniversitΓ© Libre de Bruxelles (ULB) | |
| - **Distribution Source**: Kaggle (`mlg-ulb/creditcardfraud`) | |
| - **Copyright & License**: Open Database License (ODbL) v1.0 | |
| - **Designation**: Real Anonymized Consumer Credit Card Transactions | |
| #### 3.3.2 Scale & Class Balance Profile | |
| - **Total Transactions**: $N = 284{,}807$ card transactions over 48 hours in September 2013 (143.8 MB CSV) | |
| - **Fraudulent Transactions**: $N_{\mathrm{fraud}} = 492$ confirmed chargeback frauds | |
| - **Legitimate Transactions**: $N_{\mathrm{legit}} = 284{,}315$ genuine operations | |
| - **Fraud Prevalence**: $\pi = 0.1727\%$ (Class Imbalance Ratio $\approx 578:1$) | |
| - **Temporal Horizon**: Elapsed seconds $t \in [0, 172{,}792]$ seconds (spanning exactly 2 calendar days) | |
| #### 3.3.3 Dimensionality & Mathematical Schema | |
| - **Confidential PCA Components ($d = 28$)**: `V1` through `V28` represent orthogonal linear projections from raw cardholder features, engineered by the issuing bank to protect customer privacy and commercial secrets. | |
| - **Un-transformed Columns ($d = 2$)**: `Time` (seconds from initial transaction) and `Amount` (EUR currency value, $\mu = 88.35\text{ EUR}$, $\max = 25{,}691.16\text{ EUR}$, heavily right-skewed). | |
| - **Target Label**: `Class` $\in \{0, 1\}$. | |
| #### 3.3.4 Data Hygiene, Biases & Limitations | |
| - **Short Observation Horizon**: The 48-hour window lacks multi-month cyclical seasonality, macroeconomic shifts, or evolving adversary tactics. | |
| - **Semantic Opacity**: Orthonormal PCA transformation eliminates natural domain semantics (e.g. merchant category, geographic corridor). | |
| - **Zero Demographic PII**: Anonymization mathematically guarantees zero leakage of protected personal attributes (0/10 attributes present). | |
| --- | |
| ### 3.4 Elliptic Bitcoin Transaction Graph (`elliptic`) | |
| #### 3.4.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: Elliptic Bitcoin Anti-Money Laundering Graph Dataset | |
| - **Authors**: Mark Weber, Domenic Puzis, Jie Chen, Dylan E. Cook, Prasanna Sattigeri, and Toyotaro Suzumura | |
| - **Publication**: *Anti-Money Laundering in Bitcoin: Experimenting with Graph Convolutional Networks for Financial Forensics*, ACM SIGKDD Workshop on AI in Finance, 2019 | |
| - **Research Institutions**: MIT-IBM Watson AI Lab & Elliptic | |
| - **Distribution Source**: Kaggle (`ellipticco/elliptic-data-set`) | |
| - **Copyright & License**: Creative Commons Attribution 4.0 International (CC BY 4.0) | |
| - **Designation**: Real Public Bitcoin Blockchain Directed Acyclic Transaction Flow Graph | |
| #### 3.4.2 Scale & Class Balance Profile | |
| - **Graph Topology**: $N = 203{,}769$ transaction nodes and $E = 234{,}355$ directed payment edges across 49 discrete timesteps (~665 MB disk footprint) | |
| - **Illicit Entities (`class=1`)**: $N_{\mathrm{illicit}} = 4{,}545$ confirmed malicious transactions (ransomware, darknet markets, sanctioned entities, mixers) | |
| - **Licit Entities (`class=2`)**: $N_{\mathrm{licit}} = 42{,}019$ confirmed lawful transactions (regulated exchanges, miners, merchants) | |
| - **Unlabeled Background (`class=unknown`)**: $N_{\mathrm{unknown}} = 157{,}205$ unclassified blockchain transactions ($77.15\%$ of total nodes) | |
| - **Labeled Cohort Balance**: $\pi = 9.762\%$ illicit prevalence within the $46{,}564$ ground-truth labeled cohort ($1:9$ ratio) | |
| #### 3.4.3 Topological Properties & Discrete Timesteps | |
| - **Intra-Timestep DAG Invariant**: Every edge connects transactions within the exact same two-week timestep window: | |
| $$\forall (u, v) \in \mathcal{E}, \quad \mathrm{timestep}(u) = \mathrm{timestep}(v)$$ | |
| - The entire dataset consists of 49 completely disjoint directed acyclic subgraphs with zero cross-timestep edges. | |
| - **Node Feature Schema ($d = 166$)**: | |
| - `timestep`: Integer index ($1 \le t \le 49$). | |
| - `feat_0`β`feat_93` (94 features): Local transaction features (amount, fees, script types). | |
| - `feat_94`β`feat_164` (72 features): Aggregated 1-hop neighborhood features (mean, std, min, max degrees and volumes across predecessor/successor nodes). | |
| #### 3.4.4 Data Hygiene, Biases & Limitations | |
| - **High Background Ratio**: 77.15% of nodes have unknown labels, demanding semi-supervised message passing with masked cross-entropy loss. | |
| - **Exogenous Policy Shocks**: Major darknet marketplace takedowns around timestep 43 (AlphaBay/Hansa) cause sharp concept drift, reducing illicit prevalence from $11.58\%$ ($t \le 34$) to $6.50\%$ ($t > 34$). | |
| - **Off-Chain Blindness**: Does not capture centralized off-chain internal transfers within custodial exchanges. | |
| --- | |
| ### 3.5 IBM Research AMLSim (`amlsim`) | |
| #### 3.5.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: IBM AMLSim: Multi-Agent Anti-Money Laundering Graph Simulator | |
| - **Authors**: IBM Research AI (Mark Weber et al.) | |
| - **Distribution Source**: Kaggle (`anshankul/ibm-amlsim-example-dataset`) / GitHub (`IBM/AMLSim`) | |
| - **Copyright & License**: Apache License 2.0 | |
| - **Designation**: Synthetic Multi-Agent Banking Network Graph Simulator | |
| #### 3.5.2 Scale & Class Balance Profile | |
| - **Graph Topology**: $10{,}000$ account nodes, $1{,}323{,}234$ directed transaction edges, and $1{,}719$ SAR alert ground-truth labels across 15 timesteps (~72.8 MB disk footprint) | |
| - **Typology Breakdown**: | |
| - `cycle`: 936 alert instances (circular fund routing $A \to B \to C \to A$ to disguise provenance) | |
| - `fan_in`: 783 alert instances (structuring/smurfing with multiple senders funneling into a consolidation account) | |
| - **Alert Rate**: $0.1299\%$ overall alert prevalence ($1{,}719 / 1{,}323{,}234$) | |
| #### 3.5.3 Relational Schema & Feature Pipeline | |
| - `accounts.csv` (10,000 accounts): Account IDs, initial balances, and account types. | |
| - `alerts.csv` (1,719 alerts): Ground-truth SAR alert identifiers, typology names (`cycle`, `fan_in`), and scheduling timestamps. | |
| - `transactions.csv` (1,323,234 transactions): Sender (`orig`), receiver (`dest`), `amount`, `step`, and dynamic pre/post balance deltas. | |
| #### 3.5.4 Data Hygiene, Biases & Limitations | |
| - **Rigid Geometric Typologies**: Laundering structures follow exact predefined graph templates (cycles, fan-in), which may underestimate adversary evasion in real banking rails. | |
| - **Absence of Real Personal Identities**: Entirely synthetic agents (0/10 protected demographic attributes). | |
| - **Homogeneous Balance Dynamics**: Initial balances follow synthetic Gaussian distributions without real-world wealth disparities. | |
| --- | |
| ### 3.6 SynthAML Danish Commercial AML (`synthaml`) | |
| #### 3.6.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: A Synthetic Data Set to Benchmark Anti-Money Laundering Methods | |
| - **Authors**: Martin V. Jensen, Christian S. MΓΈller, Andreas B. Simonsen, and Thomas D. Nielsen | |
| - **Publication**: Nature Scientific Data 10, 715 (2023), DOI: `10.1038/s41597-023-02569-2` | |
| - **Research Institutions**: Aarhus University & Spar Nord Bank (Denmark) | |
| - **Distribution Source**: Figshare / Nature Scientific Data Archive | |
| - **Copyright & License**: Creative Commons Attribution 4.0 International (CC BY 4.0) | |
| - **Designation**: Real-Topology Synthetic AML Benchmark (generated via SDV Gaussian Copula and CTGAN models trained directly on proprietary Spar Nord Bank commercial customer accounts and empirical AML alert outcomes) | |
| #### 3.6.2 Scale & Class Balance Profile | |
| - **Full Dataset**: $20{,}000$ investigated AML alerts across $16{,}000{,}000$ underlying transactions | |
| - **Platform Local Benchmark**: $5{,}000$ investigated alerts with $92{,}261$ lookback transactions spanning 7-to-90-day observation windows (~10.4 MB disk footprint) | |
| - **SAR Positive Alerts**: $N_{\mathrm{sar}} = 425$ true positive regulatory filings | |
| - **Dismissed False Alarms**: $N_{\mathrm{dismissed}} = 4{,}575$ genuine compliance false positives | |
| - **Alert Escalation Prevalence**: $\pi = 8.500\%$ (Class Imbalance Ratio $\approx 11:1$) | |
| #### 3.6.3 Canonical 14 Lookback Aggregated Feature Architecture | |
| 1. `n_transactions`: Total transactions within the observation window. | |
| 2. `credit_ratio`: Proportion of credit transactions relative to total volume. | |
| 3. `card_ratio`: Point-of-sale card payment fraction. | |
| 4. `cash_ratio`: Physical cash deposit/withdrawal intensity (structuring flag). | |
| 5. `international_ratio`: Cross-border international remittance fraction. | |
| 6. `wire_ratio`: Domestic wire transfer fraction. | |
| 7. `size_mean`: Sample mean of log-standardized transaction magnitudes. | |
| 8. `size_max`: Maximum transaction size in window. | |
| 9. `size_std`: Sample standard deviation of transaction sizes (volatility). | |
| 10. `total_credit_volume`: Aggregate inbound currency. | |
| 11. `total_debit_volume`: Aggregate outbound currency. | |
| 12. `net_flow`: Directional liquidity delta ($\mathrm{credit} - \mathrm{debit}$). | |
| 13. `window_days`: Duration of observation window ($7 \le \Delta t \le 90$). | |
| 14. `tx_frequency_per_day`: Daily transaction velocity ($N_{\mathrm{tx}} / \Delta t_{\mathrm{window}}$). | |
| #### 3.6.4 Data Hygiene, Biases & Limitations | |
| - **Investigation Filter Conditioning**: Ingested data consists solely of transactions that already triggered bank monitoring rules; un-flagged transactions are not included. | |
| - **Generative Copula Smoothing**: Synthetic generation via CTGAN/SDV slightly attenuates extreme tail correlations compared to raw Danish banking logs. | |
| - **Currency & Regional Focus**: Scaled to Danish Krone (DKK) banking operations and European SEPA payment rails. | |
| --- | |
| ### 3.7 AMLNet Australian AUSTRAC AML (`amlnet`) | |
| #### 3.7.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: AMLNet: A Knowledge-Guided Synthetic Benchmark for Machine Learning in AML | |
| - **Authors**: Sabin Huda, Jun Shen, et al. | |
| - **Research Institution**: School of Information and Communication Technology, Griffith University, Australia | |
| - **Archive / DOI**: Zenodo DOI: `10.5281/zenodo.10058474` | |
| - **Copyright & License**: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) | |
| - **Designation**: AUSTRAC Knowledge-Guided Multi-Agent Synthetic AML Benchmark | |
| #### 3.7.2 Scale & Class Balance Profile | |
| - **Full Scale**: $1{,}090{,}000$ transactions across 195 simulated calendar days | |
| - **Platform Local Benchmark**: $25{,}000$ canonical AUSTRAC transactions (~6.4 MB disk footprint) | |
| - **Rare-Event Laundering Prevalence**: $\pi = 0.1400\%$ in full scale ($1{,}526$ suspicious money laundering transactions out of $1.09\text{M}$, a $\approx 714:1$ negative-to-positive class imbalance) | |
| - **Regulatory Framework**: Australian AML/CTF Act 2006 statutory smurfing reporting threshold ($10{,}000\text{ AUD}$) | |
| #### 3.7.3 Canonical Feature Architecture (18 Features) | |
| 1. `amount`: Currency value in Australian Dollars. | |
| 2. `log_amount`: Log-scaled amount ($\ln(1 + \mathrm{amount})$). | |
| 3. `oldbalanceOrg`: Originator pre-transaction balance. | |
| 4. `newbalanceOrig`: Originator post-transaction balance. | |
| 5. `balance_orig_delta`: Originator balance mismatch delta. | |
| 6. `balance_orig_ratio`: Balance depletion fraction (amount / (bal + 1)). | |
| 7. `hour`: Hour of transaction ($0 \le h \le 23$). | |
| 8. `day_of_week`: Day of week ($0 \le d \le 6$). | |
| 9. `type_TRANSFER`: Indicator for electronic fund transfers. | |
| 10. `type_OSKO`: Indicator for Australian instant OSKO transfers. | |
| 11. `type_BPAY`: Indicator for BPAY bill payments. | |
| 12. `type_EFTPOS`: Indicator for card point-of-sale transactions. | |
| 13. `type_DEBIT`: Indicator for direct debits. | |
| 14. `type_NPP`: Indicator for New Payments Platform real-time transfers. | |
| 15. `is_near_reporting_threshold`: Binary indicator flagging smurfing just under the statutory threshold: | |
| $$\mathbb{I}_{\mathrm{structuring}} = \mathbb{I}(8{,}500 \le \mathrm{amount} < 10{,}000)$$ | |
| 16. `category_high_risk`: Binary indicator for high-risk economic categories (Cryptocurrency, Shell Company, Luxury Goods, Gambling, Investment). | |
| 17. `is_night_txn`: Unusual nocturnal transaction indicator ($\mathrm{hour} < 5 \lor \mathrm{hour} > 22$). | |
| 18. `is_weekend_txn`: Weekend transaction indicator (`day_of_week` $\ge 5$). | |
| #### 3.7.4 Data Hygiene, Biases & Limitations | |
| - **Non-Commercial License Restriction**: CC BY-NC 4.0 permits research, academic benchmarking, and evaluation but prohibits commercial exploitation without separate licensing. | |
| - **Threshold Boundary Artifacts**: Structuring activities are concentrated heavily in the $8{,}500\text{--}9{,}950\text{ AUD}$ window, which models can overfit if not evaluated against smooth boundary variations. | |
| - **Regional Rail Specificity**: Payment channels (`OSKO`, `BPAY`, `NPP`) are specific to Australian banking infrastructure. | |
| --- | |
| ### 3.8 CFI-CrossBank-01 Flagship Consortium Benchmark (`cross_bank`) | |
| #### 3.8.1 Provenance, Citation & Licensing | |
| - **Dataset Title**: CFI-CrossBank-01 Flagship Consortium Multi-Bank Fraud Benchmark | |
| - **Authors**: CF-Intelligence Research & Engineering Consortium | |
| - **Publication**: Collaborative Financial Intelligence Empirical Benchmark Series, 2026 | |
| - **Distribution Source**: Internal Consortium Repository (`experiments/cross_bank`) | |
| - **Copyright & License**: Proprietary Research License (CF-Intelligence Open Governance Framework) | |
| - **Designation**: Synthetic Real-Topology Multi-Bank Collaborative Benchmark across 3 Heterogeneous Bank Tiers | |
| #### 3.8.2 Scale & Class Balance Profile | |
| - **Total Transactions**: $N = 150{,}000$ cross-institution transactions across 3 heterogeneous bank tiers | |
| - **Fraudulent Transactions**: $N_{\mathrm{fraud}} = 2{,}745$ confirmed multi-bank fraud events | |
| - **Legitimate Transactions**: $N_{\mathrm{legit}} = 147{,}255$ genuine commercial and retail transactions | |
| - **Fraud Prevalence**: $\pi = 1.8300\%$ (Class Imbalance Ratio $\approx 54.6:1$) | |
| - **Institutional Topology**: | |
| - **Bank Alpha (Tier 1 Retail Megabank)**: $75{,}000$ transactions, $1{,}350$ fraud cases ($1.80\%$), baseline local PR-AUC: $0.5050$ | |
| - **Bank Beta (Tier 2 Commercial/Corporate)**: $45{,}000$ transactions, $855$ fraud cases ($1.90\%$), baseline local PR-AUC: $0.4439$ | |
| - **Bank Gamma (Tier 3 Private & Wealth)**: $30{,}000$ transactions, $540$ fraud cases ($1.80\%$), baseline local PR-AUC: $0.3855$ | |
| #### 3.8.3 Typologies & Collaborative Detection Advantage | |
| - **Cross-Bank Fraud Typologies**: Circular layering across institutions (smurfing/fan-out $\to$ intermediate mules $\to$ rapid exit gather), split-deposit velocity bursts, and cross-border settlement loops. | |
| - **Collaborative GNN Gain**: Local silo PR-AUC average of $0.4448$ increases to $\mathbf{0.8267}$ under Federated Relational GNN, delivering an empirical gain of $+0.3819$ ($+85.8\%$ relative lift). | |
| #### 3.8.4 Data Hygiene, Biases & Limitations | |
| - **Strict Privacy Invariant**: 0/10 protected demographic attributes. Zero raw PII across institutions. Entity identifiers pseudonymized with type-salted HMAC-SHA256. | |
| - **Simulated Cross-Bank Rails**: While calibrated against real inter-bank clearing flows (ISO 20022 `pacs.008`), edge topologies are generated by deterministic multi-agent orchestration. | |
| --- | |
| ## 4. Cross-Dataset Comparison & Federated Suitability | |
| | Dimension | PaySim | IEEE-CIS | Credit Card | Elliptic | IBM AMLSim | SynthAML | AMLNet | CFI-CrossBank-01 | | |
| |:---|:---|:---|:---|:---|:---|:---|:---|:---| | |
| | **Primary Risk Type** | Asset Drain Fraud | CNP Payment Fraud | Counterfeit Card | Bitcoin Laundering | Smurfing & Cycles | Compliance SAR | TTR Structuring | Consortium Multi-Hop | | |
| | **Data Topology** | Tabular / Account IDs | Tabular + Identity | Tabular (PCA) | Directed Graph (DAG) | Directed Multigraph | Relational Lookback | Tabular + Rails | Heterogeneous Multi-Bank Graph | | |
| | **Temporal Granularity** | 1 Hour | Elapsed Seconds | Elapsed Seconds | 2-Week Windows | Simulation Steps | Daily Windows | Hourly Timesteps | Chronological Rounds | | |
| | **Total Features** | 13 | 378 | 30 | 166 | 6 + Graph | 14 | 18 | Multi-Modal Embeddings | | |
| | **Strict Chronological Split** | $t \le 595$ vs $t > 595$ | $t \le 145\text{d}$ vs $t > 145\text{d}$ | $t \le 38\text{h}$ vs $t > 38\text{h}$ | $t \le 34$ vs $t > 34$ | $t \le 11$ vs $t > 11$ | $t \le 80$ vs $t > 80$ | $t \le 165$ vs $t > 165$ | $t \le 70\%$ vs $t > 70\%$ | | |
| | **FL Non-IID Dirichlet $\alpha$** | $\alpha \in [0.1, 1.0]$ | $\alpha \in [0.1, 1.0]$ | Extreme Skew Bank C | Graph Split | Multi-Bank Agents | Volume & SAR Skew | AUSTRAC Skew | 3-Tier Natural Skew | | |
| | **Fast Parquet Cache** | `bank_*.parquet` | Parquet Cached | Parquet Cached | `elliptic_cache.parquet`| `transactions.parquet` | `alerts.parquet` | `transactions.parquet` | `metrics.csv` & JSON | | |
| --- | |
| ## 5. Zero Demographic PII Invariant & Statutory Compliance | |
| In strict compliance with **EU GDPR Article 9**, **Equal Credit Opportunity Act (ECOA) Regulation B (12 CFR Part 1002)**, and **Federal Reserve SR 11-7**, all eight benchmark datasets have been exhaustively audited for protected personal characteristics: | |
| ``` | |
| βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| β PROTECTED DEMOGRAPHIC ATTRIBUTE SCAN INVARIANT β | |
| βββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββ¬ββββββββββββββββββββ¬ββββββββββββββββ€ | |
| β STATUTORY CATEGORY β REGULATORY BASIS β SCAN KEYWORDS β STATUS (0/8) β | |
| βββββββββββββββββββββββββββΌββββββββββββββββββββββββββΌββββββββββββββββββββΌββββββββββββββββ€ | |
| β Age β ECOA Reg B 1002.2(z) β age, dob, birth β EXCLUDED [OK] β | |
| β Gender / Sex β ECOA / GDPR Art 9 β gender, sex, male β EXCLUDED [OK] β | |
| β Race / Ethnicity β ECOA / Civil Rights Act β race, ethnic β EXCLUDED [OK] β | |
| β Religion / Creed β GDPR Art 9 / ECOA β religion, faith β EXCLUDED [OK] β | |
| β Marital Status β ECOA Reg B 1002.5(d) β marital, spouse β EXCLUDED [OK] β | |
| β Nationality / Origin β ECOA / Title VI β nationality, pass β EXCLUDED [OK] β | |
| β Sexual Orientation β GDPR Art 9 β sexual, lgbt β EXCLUDED [OK] β | |
| β Disability Status β ADA / GDPR Art 9 β disability, med β EXCLUDED [OK] β | |
| β Genetic / Biometric β GDPR Art 9 β biometric, dna β EXCLUDED [OK] β | |
| β Socioeconomic Status β ECOA Reg B β welfare, assist β EXCLUDED [OK] β | |
| βββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββ΄ββββββββββββββββββββ΄ββββββββββββββββ | |
| ``` | |
| *Audit Certification: [`benchmarks/results/raw/demographic_fairness_audit.json`](benchmarks/results/raw/demographic_fairness_audit.json)* | |
| *Engine: [`experiments/fairness/demographic_audit.py`](experiments/fairness/demographic_audit.py)* | |
| *Automated Test: [`backend/tests/unit/test_demographic_fairness_audit.py`](backend/tests/unit/test_demographic_fairness_audit.py) (10 tests, 100% passing)* | |
| --- | |
| ## 6. Automated Acquisition & CLI Download Protocol | |
| If raw benchmark datasets require re-acquisition in clean developer or CI runner environments, execute the automated acquisition CLI (`scripts/download_real_benchmarks.py`) or native Kaggle CLI commands: | |
| ```bash | |
| # 1. Download all datasets via master script | |
| python scripts/download_real_benchmarks.py --all | |
| # 2. Download specific benchmarks individually | |
| python scripts/download_real_benchmarks.py --dataset synthaml | |
| python scripts/download_real_benchmarks.py --dataset amlnet | |
| python scripts/download_real_benchmarks.py --dataset paysim | |
| # 3. Direct Kaggle API CLI commands (requires ~/.kaggle/kaggle.json) | |
| kaggle datasets download -d ealaxi/paysim1 -p backend/storage/datasets/paysim --unzip | |
| kaggle competitions download -c ieee-fraud-detection -p backend/storage/datasets/ieee_cis | |
| kaggle datasets download -d mlg-ulb/creditcardfraud -p backend/storage/datasets/creditcard --unzip | |
| kaggle datasets download -d ellipticco/elliptic-data-set -p backend/storage/datasets/elliptic --unzip | |
| kaggle datasets download -d anshankul/ibm-amlsim-example-dataset -p backend/storage/datasets/amlsim --unzip | |
| ``` | |
| --- | |
| ## 7. Verification Test Suites | |
| Dataset integrity, zero lookahead leakage, schema conformance, and zero-mock error guards are verified across **86 dedicated data tests**: | |
| - [`backend/tests/unit/test_dataset_cards.py`](backend/tests/unit/test_dataset_cards.py): Authoritative dataset card formalization, licensing conformance, class balance validation, and cross-reference integrity (**8 tests, 100% passing**). | |
| - [`backend/tests/unit/test_real_dataloaders.py`](backend/tests/unit/test_real_dataloaders.py): Ingestion integrity for all 7 benchmark datasets, PyG/NetworkX graph exports, and zero-mock error guards (**13 tests, 100% passing**). | |
| - [`backend/tests/unit/test_flagship_cross_bank_experiment.py`](backend/tests/unit/test_flagship_cross_bank_experiment.py): CFI-CrossBank-01 3-tier topology generation, circular multi-hop validation, and collaborative GNN baseline (**10 tests, 100% passing**). | |
| - [`backend/tests/unit/test_paysim_loader.py`](backend/tests/unit/test_paysim_loader.py): PaySim loading, 13-feature engineering, accounting balance deltas, and zero temporal leakage (**8 tests, 100% passing**). | |
| - [`backend/tests/unit/test_creditcard_loader.py`](backend/tests/unit/test_creditcard_loader.py): Credit Card loader, PCA feature scaling, and fixed-FPR threshold validation (**9 tests, 100% passing**). | |
| - [`backend/tests/unit/test_synthaml_loader.py`](backend/tests/unit/test_synthaml_loader.py): SynthAML 14 lookback features, alert schema adherence, Parquet caching, and temporal splitting (**8 tests, 100% passing**). | |
| - [`backend/tests/unit/test_amlnet_loader.py`](backend/tests/unit/test_amlnet_loader.py): AMLNet 18-feature engineering pipeline, structuring indicators, and Parquet caching (**9 tests, 100% passing**). | |
| - [`backend/tests/unit/test_demographic_fairness_audit.py`](backend/tests/unit/test_demographic_fairness_audit.py): 8-dataset demographic attribute scan confirming 0/10 protected attributes (**10 tests, 100% passing**). | |
| - [`backend/tests/unit/test_split_isolation.py`](backend/tests/unit/test_split_isolation.py): Zero data snooping, training-only preprocessor fitting, and handling of unseen categorical test tokens (**7 tests, 100% passing**). | |
| - [`backend/tests/unit/test_feature_leakage.py`](backend/tests/unit/test_feature_leakage.py): Target proxy correlation audits, outcome feature detection, and entity memorization elimination (**6 tests, 100% passing**). | |
| - [`backend/tests/unit/test_ci_smoke_gates.py`](backend/tests/unit/test_ci_smoke_gates.py): Gate A dataset-independent smoke gates using schema fixtures (**10 tests, 100% passing**). | |
| - [`backend/tests/unit/test_dataset_provenance_and_smoke_contracts.py`](backend/tests/unit/test_dataset_provenance_and_smoke_contracts.py): Fail-closed contracts, empty directory defenses, and provenance verification (**7 tests, 100% passing**). | |
| - [`backend/tests/integration/datasets/test_real_dataset_integration.py`](backend/tests/integration/datasets/test_real_dataset_integration.py): Gate B physical real-dataset integration tests covering all 7 benchmarks (**7 tests, 100% passing**). | |
| --- | |
| ## 8. Continuous Integration & Real-Data Gating Architecture (Gates A & B) | |
| To ensure reliable, deterministic CI without brittle external network dependencies or multi-gigabyte data transfers while preserving rigorous empirical testing on real data, the platform implements a strict dual-gate testing architecture: | |
| ### 8.1 Dual-Gate Architecture | |
| 1. **Gate A: Standard CI Smoke Gates (Dataset-Independent)** | |
| - **Scope**: Runs on every pull request and push to `main` without requiring external multi-gigabyte downloads. | |
| - **Mechanism**: Exercises production data loaders against schema-valid deterministic fixtures created in temporary directories (`backend/tests/fixtures/dataloader_smoke_fixtures.py`). | |
| - **Verification**: Verifies loader importability, configuration, schema validation, parsing algorithms, and fail-closed behavior without needing Kaggle credentials. | |
| - **Execution Command**: | |
| ```bash | |
| cd backend && pytest tests/unit/test_ci_smoke_gates.py -v | |
| ``` | |
| 2. **Gate B: Real-Data Integration Suite (`@pytest.mark.real_data`)** | |
| - **Scope**: Exercises authentic physical files on disk across all 7 benchmark datasets (`paysim`, `ieee_cis`, `creditcard`, `elliptic`, `amlsim`, `synthaml`, `amlnet`). | |
| - **Mechanism**: Dedicated suite in `backend/tests/integration/datasets/test_real_dataset_integration.py`. | |
| - **Gating**: Marked with `@pytest.mark.real_data`. Skipped by default in standard CI. | |
| - **Mandatory Enforcement**: When invoked with `--require-real-data`, missing dataset files cause an immediate **test failure** (`pytest.fail`), never a silent skip or synthetic fallback. | |
| - **Execution Commands**: | |
| ```bash | |
| # Run all real-data integration tests (fails closed if files are missing): | |
| cd backend && pytest tests/integration/datasets/ -v --require-real-data | |
| # Include real data tests if present (skips if absent): | |
| cd backend && pytest tests/integration/datasets/ -v --include-real-data | |
| ``` | |
| ### 8.2 Provenance Taxonomy | |
| The repository strictly differentiates dataset origins and integrity levels: | |
| | Provenance Level | Description | Benchmark Datasets | | |
| |:---|:---|:---| | |
| | `EMPIRICAL_EXTERNAL_DATA` | Real observations from authentic commercial payment logs or public blockchain DAGs. | `ieee_cis`, `creditcard`, `elliptic` | | |
| | `PUBLIC_SIMULATED_DATASET` | Public agent-based simulations published with physical source files and published baseline distributions. | `paysim`, `amlsim` | | |
| | `CONTROLLED_PROJECT_SYNTHETIC` | Reproducible benchmark datasets generated via domain-specific statistical models or academic methodology. | `synthaml`, `amlnet`, `cross_bank` | | |
| | `TEST_FIXTURE` | Minimal in-process deterministic schema fixtures for unit testing and CI smoke gates (`is_synthetic=True`). | Generated in `dataloader_smoke_fixtures.py` | | |
| ### 8.3 Provisioning Matrix & CI Reality | |
| | Dataset ID | Authoritative Source | Required Files & Format | Storage Location | Acquisition Method | Auth & License Requirements | Standard CI | Real-Data CI Gate | Local Execution | | |
| |:---|:---|:---|:---|:---|:---|:---:|:---:|:---| | |
| | **`paysim`** | Kaggle `ealaxi/paysim1` | `PS_20174392719_1491204439457_log.csv` (CSV) | `backend/storage/datasets/paysim/` | Kaggle API / CLI | Kaggle API Key; CC BY-SA 4.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk | | |
| | **`ieee_cis`** | Kaggle `c/ieee-fraud-detection` | `train_transaction.csv`, `train_identity.csv` (CSV) | `backend/storage/datasets/ieee_cis/` | Kaggle Competition CLI | Kaggle API Key + Competition Rules Acceptance | Schema Fixture | Requires Kaggle credentials & rules acceptance | Real file on disk | | |
| | **`creditcard`** | Kaggle `mlg-ulb/creditcardfraud` | `creditcard.csv` (CSV) | `backend/storage/datasets/creditcard/` | Kaggle API / CLI | Kaggle API Key; ODbL 1.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk | | |
| | **`elliptic`** | Kaggle `ellipticco/elliptic-data-set` | `elliptic_txs_features.csv`, `elliptic_txs_classes.csv`, `elliptic_txs_edgelist.csv` (CSV) | `backend/storage/datasets/elliptic/` | Kaggle API / CLI | Kaggle API Key; CC BY 4.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk | | |
| | **`amlsim`** | Kaggle `anshankul/ibm-amlsim-example-dataset` | `transactions.csv`, `accounts.csv`, `alerts.csv` (CSV) | `backend/storage/datasets/amlsim/` | Kaggle API / CLI | Kaggle API Key; Apache 2.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk | | |
| | **`synthaml`** | Figshare / Nature Sci Data DOI: `10.1038/s41597-023-02569-2` | `alerts.csv`, `transactions.csv` (CSV/Parquet) | `backend/storage/datasets/synthaml/` | Direct Download / `generate_synthaml_dataset.py` | Open Access; CC BY 4.0 | Schema Fixture | Local cache / Generator script | Real file on disk | | |
| | **`amlnet`** | Zenodo DOI: `10.5281/zenodo.10058474` | `transactions.csv` (CSV/Parquet) | `backend/storage/datasets/amlnet/` | Direct Download / `generate_amlnet_dataset.py` | Open Access; CC BY-NC 4.0 | Schema Fixture | Local cache / Generator script | Real file on disk | | |
| ### 8.4 Automated Provisioning Workflow | |
| The standalone GitHub Actions workflow `.github/workflows/real_data_integration.yml` provides a secure, gated pipeline for real-data integration: | |
| - Triggered manually via `workflow_dispatch` or on a weekly schedule. | |
| - Uses repository secrets `KAGGLE_USERNAME` and `KAGGLE_KEY`. | |
| - Executes `python scripts/download_real_benchmarks.py --dataset all`. | |
| - Fails closed if physical datasets cannot be provisioned when mandatory execution is requested. | |