Spaces:
Runtime error
Download docs/DATASETS.md from yusufcalisir/Collaborative-Fraud-Intelligence-Simulator: direct link, hf CLI and curl.
- Browser
- Download file 41.3 kB
-
https://huggingface.co/spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/resolve/main/docs/DATASETS.md
- Command line
-
hf download hf://spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/docs/DATASETS.md
-
curl -L -o DATASETS.md https://huggingface.co/spaces/yusufcalisir/Collaborative-Fraud-Intelligence-Simulator/resolve/main/docs/DATASETS.md
Authoritative Dataset Cards & Storage Architecture Specification (DATASETS.md)
Privacy-Preserving Collaborative Financial Crime Intelligence Platform (CF-Intelligence)
Specification Version: 2.4.0-enterprise
Standard: Hugging Face Datasets / ACM FAccT Data Cards / Datasheets for Datasets (Gebru et al., 2021)
Governance Scope: Federal Reserve SR 11-7 / OCC 2011-12, EU AI Act Article 10, GDPR Article 9, ECOA Regulation B (12 CFR Part 1002), PCI-DSS v4.0
Repository Location: DATASETS.md (root) and docs/DATASETS.md
1. Executive Summary & Zero-Mock Dataset Governance
A core architectural invariant of CF-Intelligence is Zero-Mock, Zero-Dummy Data in production and empirical benchmarking evaluations. While development unit tests may isolate specific interfaces, all empirical fraud detection claims, federated learning convergence benchmarks, and differential privacy trade-offs in this platform are calibrated against eight canonical, large-scale financial crime datasets spanning over 8.9 million transactions:
Master Dataset Inventory & Licensing Matrix
| Dataset ID | Dataset Name | Domain Scope | Real vs Synthetic | Distribution Source & Archive | Copyright & License | Class Imbalance Ratio | Protected Demographic Fields |
|---|---|---|---|---|---|---|---|
paysim |
PaySim Mobile Money | Mobile Money Remittance | Synthetic Agent Simulator | Kaggle (ealaxi/paysim1) |
CC BY-SA 4.0 | 0.129% (1:774) | 0 / 10 (0.0%) |
ieee_cis |
IEEE-CIS Card Fraud | E-Commerce Payments | Real Production Logs | Kaggle (ieee-fraud-detection) |
Competition License (Research) | 3.500% (1:28) | 0 / 10 (0.0%) |
credit_card |
ULB European Card | Consumer Card Payments | Real Anonymized PCA | Kaggle (mlg-ulb/creditcardfraud) |
Open Database License (ODbL 1.0) | 0.173% (1:578) | 0 / 10 (0.0%) |
elliptic |
Elliptic Bitcoin Graph | Cryptocurrency Forensics | Real Blockchain DAG | Kaggle (ellipticco/elliptic-data-set) |
CC BY 4.0 | 9.76% labeled (1:9) | 0 / 10 (0.0%) |
amlsim |
IBM AMLSim Graph | Multi-Agent Banking Graph | Synthetic Multi-Agent | Kaggle / GitHub (IBM/AMLSim) |
Apache 2.0 | 0.130% (1:769) | 0 / 10 (0.0%) |
synthaml |
SynthAML Spar Nord | European Commercial AML | Real-Topology SDV Copula | Nature Scientific Data / Figshare | CC BY 4.0 | 8.500% (1:11) | 0 / 10 (0.0%) |
amlnet |
AMLNet AUSTRAC | Australian Wire Compliance | Synthetic Agent Simulation | Zenodo (10.5281/zenodo.10058474) |
CC BY-NC 4.0 | 0.140% (1:714) | 0 / 10 (0.0%) |
cross_bank |
CFI-CrossBank-01 | Consortium Multi-Bank | Synthetic Real-Topology | CFI Research Consortium Generator | Proprietary Research (CFI) | 1.830% (1:55) | 0 / 10 (0.0%) |
2. Directory Layout & Local Storage Topology
All raw dataset files are stored in backend/storage/datasets/<dataset_name>/ or storage/datasets/<dataset_name>/. The dynamic loader (resolve_dataset_dir in dataloader.py) resolves storage across containerized and development environments:
storage/datasets/
βββ amlsim/
β βββ accounts.csv # 326 KB (10,000 accounts metadata)
β βββ alerts.csv # 88 KB (1,719 ground-truth AML typology alerts)
β βββ transactions.csv # 61.05 MB (1,323,234 raw transaction flow events)
β βββ transactions.parquet # 11.43 MB (Zero-copy fast columnar cache)
βββ creditcard/
β βββ creditcard.csv # 143.84 MB (284,807 rows)
βββ elliptic/
β βββ elliptic_txs_classes.csv # 3.15 MB (Node classification labels)
β βββ elliptic_txs_edgelist.csv # 4.26 MB (Directed transaction flow graph)
β βββ elliptic_txs_features.csv # 657.73 MB (166-dimensional node embeddings)
βββ ieee_cis/
β βββ train_transaction.csv # 651.69 MB (Primary training transactions)
β βββ train_identity.csv # 25.30 MB (Device & IP network identity)
β βββ test_transaction.csv # 584.79 MB (Out-of-time evaluation transactions)
β βββ test_identity.csv # 24.60 MB (Out-of-time identity features)
βββ paysim/
β βββ PS_20174392719_1491204439457_log.csv # 470.67 MB (6.36M simulated mobile money records)
β βββ bank_alpha.parquet # Partitioned Non-IID client split (Bank Alpha)
β βββ bank_beta.parquet # Partitioned Non-IID client split (Bank Beta)
β βββ bank_gamma.parquet # Partitioned Non-IID client split (Bank Gamma)
βββ synthaml/
β βββ alerts.csv # 195 KB (5,000 alert metadata records)
β βββ alerts.parquet # 82 KB (Zero-copy fast columnar alert cache)
β βββ transactions.csv # 9.87 MB (92,261 lookback transactions)
β βββ transactions.parquet # 2.92 MB (Zero-copy fast columnar transaction cache)
βββ amlnet/
βββ amlnet_transactions.csv # 5.48 MB (25,000 canonical AUSTRAC transactions)
βββ transactions.csv # 5.48 MB (Primary transaction flow log)
βββ transactions.parquet # 900 KB (Zero-copy fast columnar cache)
βββ cross_bank/
βββ config.json # Multi-bank topology specification (Alpha, Beta, Gamma)
βββ results.json # Full empirical benchmark metrics (PR-AUC, F1, latency)
βββ metrics.csv # Tabular client-by-client performance breakdown
βββ report.md # Full empirical research report with LaTeX formulas
βββ plots/ # Comparative PR curves and cross-bank topology graphs
3. Authoritative Dataset Cards (8 Primary Benchmarks)
3.1 PaySim Mobile Money Fraud (paysim)
3.1.1 Provenance, Citation & Licensing
- Dataset Title: PaySim: A Financial Mobile Money Simulator for Fraud Detection
- Authors: Edgar Lopez-Rojas, Ahmad Elmir, and Stefan Axelsson (Blekinge Institute of Technology)
- Publication: IEEE 28th International Conference on Tools with Artificial Intelligence (ICTAI), 2016
- Distribution Source: Kaggle (
ealaxi/paysim1) - Copyright & License: Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
- Designation: Synthetic Agent-Based Mobile Money Simulation (empirically calibrated against 1 month of anonymized financial logs from an African mobile money service operator, M-Pesa Kenya topology)
3.1.2 Scale & Class Balance Profile
- Total Transactions: $N = 6{,}362{,}620$ payment events (470.7 MB CSV on disk)
- Fraudulent Transactions: $N_{\mathrm{fraud}} = 8{,}213$ confirmed illicit transactions
- Legitimate Transactions: $N_{\mathrm{legit}} = 6{,}354{,}407$ genuine transactions
- Fraud Prevalence: $\pi = 0.1291%$ (Class Imbalance Ratio $\approx 774:1$)
- Simulation Duration: 744 hourly timesteps ($31$ simulated calendar days)
3.1.3 Feature Architecture & Transformations
- Canonical Feature Pipeline (13 Features):
step: Integer hour of simulation ($1 \le t \le 744$).type_TRANSFER: Binary indicator for cross-account fund transfers ($y=1$ candidate).type_CASH_OUT: Binary indicator for cash withdrawal via agent ($y=1$ candidate).type_PAYMENT: Binary indicator for merchant retail purchases ($y=0$ strictly).type_DEBIT: Binary indicator for core banking debit transfers ($y=0$ strictly).type_CASH_IN: Binary indicator for cash deposit via agent ($y=0$ strictly).amount: Transacted currency value.oldbalanceOrg: Originator account balance prior to transaction.newbalanceOrig: Originator balance post-transaction (typically zeroed in fraud).oldbalanceDest: Recipient account balance prior to transaction.newbalanceDest: Recipient balance post-transaction.errorBalanceOrig: Sender accounting discrepancy delta:
errorBalanceDest: Recipient accounting discrepancy delta:
3.1.4 Data Hygiene, Biases & Limitations
- Selective Fraud Typologies: Fraud occurs strictly within
TRANSFERandCASH_OUTtransaction types; zero fraud cases exist inPAYMENT,CASH_IN, orDEBIT. Models trained without type filtering may learn trivial shortcut heuristics. - Absence of Real PII: No protected or special-category demographic attributes (0/10) are present in the evaluated public datasets (relevant to GDPR Article 9 and ECOA Regulation B exclusions).
- Deterministic Evasion: Synthetic fraudsters execute simple double-step asset drain attacks without sophisticated multi-hop laundering chains.
3.2 IEEE-CIS E-Commerce Fraud Detection (ieee_cis)
3.2.1 Provenance, Citation & Licensing
- Dataset Title: IEEE-CIS Fraud Detection Benchmark
- Authors: IEEE Computational Intelligence Society (IEEE-CIS) & Vesta Corporation
- Publication: Kaggle Competition Benchmark, 2019
- Distribution Source: Kaggle (
ieee-fraud-detection) - Copyright & License: Vesta Corporation Competition Dataset License (authorized for academic, scientific, and open-source benchmark evaluation)
- Designation: Real Production E-Commerce Card-Not-Present (CNP) Transactions
3.2.2 Scale & Class Balance Profile
- Total Transactions: $N = 590{,}540$ training transactions + $144{,}233$ identity metadata records (~1.29 GB raw disk footprint)
- Fraudulent Transactions: $N_{\mathrm{fraud}} = 20{,}663$ confirmed chargeback fraud events
- Legitimate Transactions: $N_{\mathrm{legit}} = 569{,}877$ genuine purchases
- Fraud Prevalence: $\pi = 3.4989%$ (Class Imbalance Ratio $\approx 28:1$)
- Temporal Horizon: Elapsed seconds across 182 calendar days ($t \in [86{,}400, 15{,}811{,}131]$ seconds)
3.2.3 Feature Architecture & Identity Join
train_transaction.csv(394 attributes): Transaction amount, product code (ProductCD), card metadata (card1βcard6), address codes (addr1βaddr2), email domains, count features (C1βC14), timedeltas (D1βD15), match flags (M1βM9), and Vesta risk indicators (V1βV339).train_identity.csv(41 attributes): Identity verification metadata (id_01βid_38),DeviceType, andDeviceInfo.- Left Join: Merged along
TransactionID($144{,}233$ transactions with identity metadata, ~24.4% join rate; transactions without identity records receivehas_identity = 0.0and imputed indicators).
3.2.4 Data Hygiene, Biases & Limitations
- High Missingness: Over 200 features exhibit $>50%$ missing values (especially identity features and V-features). Demands explicit missingness indicator encoding rather than mean imputation.
- Obfuscated Semantics: Proprietary V-features ($V_1 \dots V_{339}$) mask engineering logic, limiting direct human-in-the-loop interpretability without SHAP attribution.
- Device Fingerprint Churn: Operating system and browser versions undergo rapid natural obsolescence over the 182-day period.
3.3 European Credit Card Fraud (credit_card)
3.3.1 Provenance, Citation & Licensing
- Dataset Title: Credit Card Fraud Detection (European Cardholders)
- Authors: Andrea Dal Pozzolo, Olivier Caelen, Reid A. Johnson, and Gianluca Bontempi
- Publication: Calibrating Probability with Undersampling for Unbalanced Classification, IEEE SSCI, 2015
- Research Institution: Machine Learning Group (MLG), UniversitΓ© Libre de Bruxelles (ULB)
- Distribution Source: Kaggle (
mlg-ulb/creditcardfraud) - Copyright & License: Open Database License (ODbL) v1.0
- Designation: Real Anonymized Consumer Credit Card Transactions
3.3.2 Scale & Class Balance Profile
- Total Transactions: $N = 284{,}807$ card transactions over 48 hours in September 2013 (143.8 MB CSV)
- Fraudulent Transactions: $N_{\mathrm{fraud}} = 492$ confirmed chargeback frauds
- Legitimate Transactions: $N_{\mathrm{legit}} = 284{,}315$ genuine operations
- Fraud Prevalence: $\pi = 0.1727%$ (Class Imbalance Ratio $\approx 578:1$)
- Temporal Horizon: Elapsed seconds $t \in [0, 172{,}792]$ seconds (spanning exactly 2 calendar days)
3.3.3 Dimensionality & Mathematical Schema
- Confidential PCA Components ($d = 28$):
V1throughV28represent orthogonal linear projections from raw cardholder features, engineered by the issuing bank to protect customer privacy and commercial secrets. - Un-transformed Columns ($d = 2$):
Time(seconds from initial transaction) andAmount(EUR currency value, $\mu = 88.35\text{ EUR}$, $\max = 25{,}691.16\text{ EUR}$, heavily right-skewed). - Target Label:
Class$\in {0, 1}$.
3.3.4 Data Hygiene, Biases & Limitations
- Short Observation Horizon: The 48-hour window lacks multi-month cyclical seasonality, macroeconomic shifts, or evolving adversary tactics.
- Semantic Opacity: Orthonormal PCA transformation eliminates natural domain semantics (e.g. merchant category, geographic corridor).
- Zero Demographic PII: Anonymization mathematically guarantees zero leakage of protected personal attributes (0/10 attributes present).
3.4 Elliptic Bitcoin Transaction Graph (elliptic)
3.4.1 Provenance, Citation & Licensing
- Dataset Title: Elliptic Bitcoin Anti-Money Laundering Graph Dataset
- Authors: Mark Weber, Domenic Puzis, Jie Chen, Dylan E. Cook, Prasanna Sattigeri, and Toyotaro Suzumura
- Publication: Anti-Money Laundering in Bitcoin: Experimenting with Graph Convolutional Networks for Financial Forensics, ACM SIGKDD Workshop on AI in Finance, 2019
- Research Institutions: MIT-IBM Watson AI Lab & Elliptic
- Distribution Source: Kaggle (
ellipticco/elliptic-data-set) - Copyright & License: Creative Commons Attribution 4.0 International (CC BY 4.0)
- Designation: Real Public Bitcoin Blockchain Directed Acyclic Transaction Flow Graph
3.4.2 Scale & Class Balance Profile
- Graph Topology: $N = 203{,}769$ transaction nodes and $E = 234{,}355$ directed payment edges across 49 discrete timesteps (~665 MB disk footprint)
- Illicit Entities (
class=1): $N_{\mathrm{illicit}} = 4{,}545$ confirmed malicious transactions (ransomware, darknet markets, sanctioned entities, mixers) - Licit Entities (
class=2): $N_{\mathrm{licit}} = 42{,}019$ confirmed lawful transactions (regulated exchanges, miners, merchants) - Unlabeled Background (
class=unknown): $N_{\mathrm{unknown}} = 157{,}205$ unclassified blockchain transactions ($77.15%$ of total nodes) - Labeled Cohort Balance: $\pi = 9.762%$ illicit prevalence within the $46{,}564$ ground-truth labeled cohort ($1:9$ ratio)
3.4.3 Topological Properties & Discrete Timesteps
- Intra-Timestep DAG Invariant: Every edge connects transactions within the exact same two-week timestep window:
- The entire dataset consists of 49 completely disjoint directed acyclic subgraphs with zero cross-timestep edges.
- Node Feature Schema ($d = 166$):
timestep: Integer index ($1 \le t \le 49$).feat_0βfeat_93(94 features): Local transaction features (amount, fees, script types).feat_94βfeat_164(72 features): Aggregated 1-hop neighborhood features (mean, std, min, max degrees and volumes across predecessor/successor nodes).
3.4.4 Data Hygiene, Biases & Limitations
- High Background Ratio: 77.15% of nodes have unknown labels, demanding semi-supervised message passing with masked cross-entropy loss.
- Exogenous Policy Shocks: Major darknet marketplace takedowns around timestep 43 (AlphaBay/Hansa) cause sharp concept drift, reducing illicit prevalence from $11.58%$ ($t \le 34$) to $6.50%$ ($t > 34$).
- Off-Chain Blindness: Does not capture centralized off-chain internal transfers within custodial exchanges.
3.5 IBM Research AMLSim (amlsim)
3.5.1 Provenance, Citation & Licensing
- Dataset Title: IBM AMLSim: Multi-Agent Anti-Money Laundering Graph Simulator
- Authors: IBM Research AI (Mark Weber et al.)
- Distribution Source: Kaggle (
anshankul/ibm-amlsim-example-dataset) / GitHub (IBM/AMLSim) - Copyright & License: Apache License 2.0
- Designation: Synthetic Multi-Agent Banking Network Graph Simulator
3.5.2 Scale & Class Balance Profile
- Graph Topology: $10{,}000$ account nodes, $1{,}323{,}234$ directed transaction edges, and $1{,}719$ SAR alert ground-truth labels across 15 timesteps (~72.8 MB disk footprint)
- Typology Breakdown:
cycle: 936 alert instances (circular fund routing $A \to B \to C \to A$ to disguise provenance)fan_in: 783 alert instances (structuring/smurfing with multiple senders funneling into a consolidation account)
- Alert Rate: $0.1299%$ overall alert prevalence ($1{,}719 / 1{,}323{,}234$)
3.5.3 Relational Schema & Feature Pipeline
accounts.csv(10,000 accounts): Account IDs, initial balances, and account types.alerts.csv(1,719 alerts): Ground-truth SAR alert identifiers, typology names (cycle,fan_in), and scheduling timestamps.transactions.csv(1,323,234 transactions): Sender (orig), receiver (dest),amount,step, and dynamic pre/post balance deltas.
3.5.4 Data Hygiene, Biases & Limitations
- Rigid Geometric Typologies: Laundering structures follow exact predefined graph templates (cycles, fan-in), which may underestimate adversary evasion in real banking rails.
- Absence of Real Personal Identities: Entirely synthetic agents (0/10 protected demographic attributes).
- Homogeneous Balance Dynamics: Initial balances follow synthetic Gaussian distributions without real-world wealth disparities.
3.6 SynthAML Danish Commercial AML (synthaml)
3.6.1 Provenance, Citation & Licensing
- Dataset Title: A Synthetic Data Set to Benchmark Anti-Money Laundering Methods
- Authors: Martin V. Jensen, Christian S. MΓΈller, Andreas B. Simonsen, and Thomas D. Nielsen
- Publication: Nature Scientific Data 10, 715 (2023), DOI:
10.1038/s41597-023-02569-2 - Research Institutions: Aarhus University & Spar Nord Bank (Denmark)
- Distribution Source: Figshare / Nature Scientific Data Archive
- Copyright & License: Creative Commons Attribution 4.0 International (CC BY 4.0)
- Designation: Real-Topology Synthetic AML Benchmark (generated via SDV Gaussian Copula and CTGAN models trained directly on proprietary Spar Nord Bank commercial customer accounts and empirical AML alert outcomes)
3.6.2 Scale & Class Balance Profile
- Full Dataset: $20{,}000$ investigated AML alerts across $16{,}000{,}000$ underlying transactions
- Platform Local Benchmark: $5{,}000$ investigated alerts with $92{,}261$ lookback transactions spanning 7-to-90-day observation windows (~10.4 MB disk footprint)
- SAR Positive Alerts: $N_{\mathrm{sar}} = 425$ true positive regulatory filings
- Dismissed False Alarms: $N_{\mathrm{dismissed}} = 4{,}575$ genuine compliance false positives
- Alert Escalation Prevalence: $\pi = 8.500%$ (Class Imbalance Ratio $\approx 11:1$)
3.6.3 Canonical 14 Lookback Aggregated Feature Architecture
n_transactions: Total transactions within the observation window.credit_ratio: Proportion of credit transactions relative to total volume.card_ratio: Point-of-sale card payment fraction.cash_ratio: Physical cash deposit/withdrawal intensity (structuring flag).international_ratio: Cross-border international remittance fraction.wire_ratio: Domestic wire transfer fraction.size_mean: Sample mean of log-standardized transaction magnitudes.size_max: Maximum transaction size in window.size_std: Sample standard deviation of transaction sizes (volatility).total_credit_volume: Aggregate inbound currency.total_debit_volume: Aggregate outbound currency.net_flow: Directional liquidity delta ($\mathrm{credit} - \mathrm{debit}$).window_days: Duration of observation window ($7 \le \Delta t \le 90$).tx_frequency_per_day: Daily transaction velocity ($N_{\mathrm{tx}} / \Delta t_{\mathrm{window}}$).
3.6.4 Data Hygiene, Biases & Limitations
- Investigation Filter Conditioning: Ingested data consists solely of transactions that already triggered bank monitoring rules; un-flagged transactions are not included.
- Generative Copula Smoothing: Synthetic generation via CTGAN/SDV slightly attenuates extreme tail correlations compared to raw Danish banking logs.
- Currency & Regional Focus: Scaled to Danish Krone (DKK) banking operations and European SEPA payment rails.
3.7 AMLNet Australian AUSTRAC AML (amlnet)
3.7.1 Provenance, Citation & Licensing
- Dataset Title: AMLNet: A Knowledge-Guided Synthetic Benchmark for Machine Learning in AML
- Authors: Sabin Huda, Jun Shen, et al.
- Research Institution: School of Information and Communication Technology, Griffith University, Australia
- Archive / DOI: Zenodo DOI:
10.5281/zenodo.10058474 - Copyright & License: Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
- Designation: AUSTRAC Knowledge-Guided Multi-Agent Synthetic AML Benchmark
3.7.2 Scale & Class Balance Profile
- Full Scale: $1{,}090{,}000$ transactions across 195 simulated calendar days
- Platform Local Benchmark: $25{,}000$ canonical AUSTRAC transactions (~6.4 MB disk footprint)
- Rare-Event Laundering Prevalence: $\pi = 0.1400%$ in full scale ($1{,}526$ suspicious money laundering transactions out of $1.09\text{M}$, a $\approx 714:1$ negative-to-positive class imbalance)
- Regulatory Framework: Australian AML/CTF Act 2006 statutory smurfing reporting threshold ($10{,}000\text{ AUD}$)
3.7.3 Canonical Feature Architecture (18 Features)
amount: Currency value in Australian Dollars.log_amount: Log-scaled amount ($\ln(1 + \mathrm{amount})$).oldbalanceOrg: Originator pre-transaction balance.newbalanceOrig: Originator post-transaction balance.balance_orig_delta: Originator balance mismatch delta.balance_orig_ratio: Balance depletion fraction (amount / (bal + 1)).hour: Hour of transaction ($0 \le h \le 23$).day_of_week: Day of week ($0 \le d \le 6$).type_TRANSFER: Indicator for electronic fund transfers.type_OSKO: Indicator for Australian instant OSKO transfers.type_BPAY: Indicator for BPAY bill payments.type_EFTPOS: Indicator for card point-of-sale transactions.type_DEBIT: Indicator for direct debits.type_NPP: Indicator for New Payments Platform real-time transfers.is_near_reporting_threshold: Binary indicator flagging smurfing just under the statutory threshold:
category_high_risk: Binary indicator for high-risk economic categories (Cryptocurrency, Shell Company, Luxury Goods, Gambling, Investment).is_night_txn: Unusual nocturnal transaction indicator ($\mathrm{hour} < 5 \lor \mathrm{hour} > 22$).is_weekend_txn: Weekend transaction indicator (day_of_week$\ge 5$).
3.7.4 Data Hygiene, Biases & Limitations
- Non-Commercial License Restriction: CC BY-NC 4.0 permits research, academic benchmarking, and evaluation but prohibits commercial exploitation without separate licensing.
- Threshold Boundary Artifacts: Structuring activities are concentrated heavily in the $8{,}500\text{--}9{,}950\text{ AUD}$ window, which models can overfit if not evaluated against smooth boundary variations.
- Regional Rail Specificity: Payment channels (
OSKO,BPAY,NPP) are specific to Australian banking infrastructure.
3.8 CFI-CrossBank-01 Flagship Consortium Benchmark (cross_bank)
3.8.1 Provenance, Citation & Licensing
- Dataset Title: CFI-CrossBank-01 Flagship Consortium Multi-Bank Fraud Benchmark
- Authors: CF-Intelligence Research & Engineering Consortium
- Publication: Collaborative Financial Intelligence Empirical Benchmark Series, 2026
- Distribution Source: Internal Consortium Repository (
experiments/cross_bank) - Copyright & License: Proprietary Research License (CF-Intelligence Open Governance Framework)
- Designation: Synthetic Real-Topology Multi-Bank Collaborative Benchmark across 3 Heterogeneous Bank Tiers
3.8.2 Scale & Class Balance Profile
- Total Transactions: $N = 150{,}000$ cross-institution transactions across 3 heterogeneous bank tiers
- Fraudulent Transactions: $N_{\mathrm{fraud}} = 2{,}745$ confirmed multi-bank fraud events
- Legitimate Transactions: $N_{\mathrm{legit}} = 147{,}255$ genuine commercial and retail transactions
- Fraud Prevalence: $\pi = 1.8300%$ (Class Imbalance Ratio $\approx 54.6:1$)
- Institutional Topology:
- Bank Alpha (Tier 1 Retail Megabank): $75{,}000$ transactions, $1{,}350$ fraud cases ($1.80%$), baseline local PR-AUC: $0.5050$
- Bank Beta (Tier 2 Commercial/Corporate): $45{,}000$ transactions, $855$ fraud cases ($1.90%$), baseline local PR-AUC: $0.4439$
- Bank Gamma (Tier 3 Private & Wealth): $30{,}000$ transactions, $540$ fraud cases ($1.80%$), baseline local PR-AUC: $0.3855$
3.8.3 Typologies & Collaborative Detection Advantage
- Cross-Bank Fraud Typologies: Circular layering across institutions (smurfing/fan-out $\to$ intermediate mules $\to$ rapid exit gather), split-deposit velocity bursts, and cross-border settlement loops.
- Collaborative GNN Gain: Local silo PR-AUC average of $0.4448$ increases to $\mathbf{0.8267}$ under Federated Relational GNN, delivering an empirical gain of $+0.3819$ ($+85.8%$ relative lift).
3.8.4 Data Hygiene, Biases & Limitations
- Strict Privacy Invariant: 0/10 protected demographic attributes. Zero raw PII across institutions. Entity identifiers pseudonymized with type-salted HMAC-SHA256.
- Simulated Cross-Bank Rails: While calibrated against real inter-bank clearing flows (ISO 20022
pacs.008), edge topologies are generated by deterministic multi-agent orchestration.
4. Cross-Dataset Comparison & Federated Suitability
| Dimension | PaySim | IEEE-CIS | Credit Card | Elliptic | IBM AMLSim | SynthAML | AMLNet | CFI-CrossBank-01 |
|---|---|---|---|---|---|---|---|---|
| Primary Risk Type | Asset Drain Fraud | CNP Payment Fraud | Counterfeit Card | Bitcoin Laundering | Smurfing & Cycles | Compliance SAR | TTR Structuring | Consortium Multi-Hop |
| Data Topology | Tabular / Account IDs | Tabular + Identity | Tabular (PCA) | Directed Graph (DAG) | Directed Multigraph | Relational Lookback | Tabular + Rails | Heterogeneous Multi-Bank Graph |
| Temporal Granularity | 1 Hour | Elapsed Seconds | Elapsed Seconds | 2-Week Windows | Simulation Steps | Daily Windows | Hourly Timesteps | Chronological Rounds |
| Total Features | 13 | 378 | 30 | 166 | 6 + Graph | 14 | 18 | Multi-Modal Embeddings |
| Strict Chronological Split | $t \le 595$ vs $t > 595$ | $t \le 145\text{d}$ vs $t > 145\text{d}$ | $t \le 38\text{h}$ vs $t > 38\text{h}$ | $t \le 34$ vs $t > 34$ | $t \le 11$ vs $t > 11$ | $t \le 80$ vs $t > 80$ | $t \le 165$ vs $t > 165$ | $t \le 70%$ vs $t > 70%$ |
| FL Non-IID Dirichlet $\alpha$ | $\alpha \in [0.1, 1.0]$ | $\alpha \in [0.1, 1.0]$ | Extreme Skew Bank C | Graph Split | Multi-Bank Agents | Volume & SAR Skew | AUSTRAC Skew | 3-Tier Natural Skew |
| Fast Parquet Cache | bank_*.parquet |
Parquet Cached | Parquet Cached | elliptic_cache.parquet |
transactions.parquet |
alerts.parquet |
transactions.parquet |
metrics.csv & JSON |
5. Zero Demographic PII Invariant & Statutory Compliance
In strict compliance with EU GDPR Article 9, Equal Credit Opportunity Act (ECOA) Regulation B (12 CFR Part 1002), and Federal Reserve SR 11-7, all eight benchmark datasets have been exhaustively audited for protected personal characteristics:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PROTECTED DEMOGRAPHIC ATTRIBUTE SCAN INVARIANT β
βββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββ¬ββββββββββββββββββββ¬ββββββββββββββββ€
β STATUTORY CATEGORY β REGULATORY BASIS β SCAN KEYWORDS β STATUS (0/8) β
βββββββββββββββββββββββββββΌββββββββββββββββββββββββββΌββββββββββββββββββββΌββββββββββββββββ€
β Age β ECOA Reg B 1002.2(z) β age, dob, birth β EXCLUDED [OK] β
β Gender / Sex β ECOA / GDPR Art 9 β gender, sex, male β EXCLUDED [OK] β
β Race / Ethnicity β ECOA / Civil Rights Act β race, ethnic β EXCLUDED [OK] β
β Religion / Creed β GDPR Art 9 / ECOA β religion, faith β EXCLUDED [OK] β
β Marital Status β ECOA Reg B 1002.5(d) β marital, spouse β EXCLUDED [OK] β
β Nationality / Origin β ECOA / Title VI β nationality, pass β EXCLUDED [OK] β
β Sexual Orientation β GDPR Art 9 β sexual, lgbt β EXCLUDED [OK] β
β Disability Status β ADA / GDPR Art 9 β disability, med β EXCLUDED [OK] β
β Genetic / Biometric β GDPR Art 9 β biometric, dna β EXCLUDED [OK] β
β Socioeconomic Status β ECOA Reg B β welfare, assist β EXCLUDED [OK] β
βββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββ΄ββββββββββββββββββββ΄ββββββββββββββββ
Audit Certification: benchmarks/results/raw/demographic_fairness_audit.json
Engine: experiments/fairness/demographic_audit.py
Automated Test: backend/tests/unit/test_demographic_fairness_audit.py (10 tests, 100% passing)
6. Automated Acquisition & CLI Download Protocol
If raw benchmark datasets require re-acquisition in clean developer or CI runner environments, execute the automated acquisition CLI (scripts/download_real_benchmarks.py) or native Kaggle CLI commands:
# 1. Download all datasets via master script
python scripts/download_real_benchmarks.py --all
# 2. Download specific benchmarks individually
python scripts/download_real_benchmarks.py --dataset synthaml
python scripts/download_real_benchmarks.py --dataset amlnet
python scripts/download_real_benchmarks.py --dataset paysim
# 3. Direct Kaggle API CLI commands (requires ~/.kaggle/kaggle.json)
kaggle datasets download -d ealaxi/paysim1 -p backend/storage/datasets/paysim --unzip
kaggle competitions download -c ieee-fraud-detection -p backend/storage/datasets/ieee_cis
kaggle datasets download -d mlg-ulb/creditcardfraud -p backend/storage/datasets/creditcard --unzip
kaggle datasets download -d ellipticco/elliptic-data-set -p backend/storage/datasets/elliptic --unzip
kaggle datasets download -d anshankul/ibm-amlsim-example-dataset -p backend/storage/datasets/amlsim --unzip
7. Verification Test Suites
Dataset integrity, zero lookahead leakage, schema conformance, and zero-mock error guards are verified across 86 dedicated data tests:
backend/tests/unit/test_dataset_cards.py: Authoritative dataset card formalization, licensing conformance, class balance validation, and cross-reference integrity (8 tests, 100% passing).backend/tests/unit/test_real_dataloaders.py: Ingestion integrity for all 7 benchmark datasets, PyG/NetworkX graph exports, and zero-mock error guards (13 tests, 100% passing).backend/tests/unit/test_flagship_cross_bank_experiment.py: CFI-CrossBank-01 3-tier topology generation, circular multi-hop validation, and collaborative GNN baseline (10 tests, 100% passing).backend/tests/unit/test_paysim_loader.py: PaySim loading, 13-feature engineering, accounting balance deltas, and zero temporal leakage (8 tests, 100% passing).backend/tests/unit/test_creditcard_loader.py: Credit Card loader, PCA feature scaling, and fixed-FPR threshold validation (9 tests, 100% passing).backend/tests/unit/test_synthaml_loader.py: SynthAML 14 lookback features, alert schema adherence, Parquet caching, and temporal splitting (8 tests, 100% passing).backend/tests/unit/test_amlnet_loader.py: AMLNet 18-feature engineering pipeline, structuring indicators, and Parquet caching (9 tests, 100% passing).backend/tests/unit/test_demographic_fairness_audit.py: 8-dataset demographic attribute scan confirming 0/10 protected attributes (10 tests, 100% passing).backend/tests/unit/test_split_isolation.py: Zero data snooping, training-only preprocessor fitting, and handling of unseen categorical test tokens (7 tests, 100% passing).backend/tests/unit/test_feature_leakage.py: Target proxy correlation audits, outcome feature detection, and entity memorization elimination (6 tests, 100% passing).backend/tests/unit/test_ci_smoke_gates.py: Gate A dataset-independent smoke gates using schema fixtures (10 tests, 100% passing).backend/tests/unit/test_dataset_provenance_and_smoke_contracts.py: Fail-closed contracts, empty directory defenses, and provenance verification (7 tests, 100% passing).backend/tests/integration/datasets/test_real_dataset_integration.py: Gate B physical real-dataset integration tests covering all 7 benchmarks (7 tests, 100% passing).
8. Continuous Integration & Real-Data Gating Architecture (Gates A & B)
To ensure reliable, deterministic CI without brittle external network dependencies or multi-gigabyte data transfers while preserving rigorous empirical testing on real data, the platform implements a strict dual-gate testing architecture:
8.1 Dual-Gate Architecture
Gate A: Standard CI Smoke Gates (Dataset-Independent)
- Scope: Runs on every pull request and push to
mainwithout requiring external multi-gigabyte downloads. - Mechanism: Exercises production data loaders against schema-valid deterministic fixtures created in temporary directories (
backend/tests/fixtures/dataloader_smoke_fixtures.py). - Verification: Verifies loader importability, configuration, schema validation, parsing algorithms, and fail-closed behavior without needing Kaggle credentials.
- Execution Command:
cd backend && pytest tests/unit/test_ci_smoke_gates.py -v
- Scope: Runs on every pull request and push to
Gate B: Real-Data Integration Suite (
@pytest.mark.real_data)- Scope: Exercises authentic physical files on disk across all 7 benchmark datasets (
paysim,ieee_cis,creditcard,elliptic,amlsim,synthaml,amlnet). - Mechanism: Dedicated suite in
backend/tests/integration/datasets/test_real_dataset_integration.py. - Gating: Marked with
@pytest.mark.real_data. Skipped by default in standard CI. - Mandatory Enforcement: When invoked with
--require-real-data, missing dataset files cause an immediate test failure (pytest.fail), never a silent skip or synthetic fallback. - Execution Commands:
# Run all real-data integration tests (fails closed if files are missing): cd backend && pytest tests/integration/datasets/ -v --require-real-data # Include real data tests if present (skips if absent): cd backend && pytest tests/integration/datasets/ -v --include-real-data
- Scope: Exercises authentic physical files on disk across all 7 benchmark datasets (
8.2 Provenance Taxonomy
The repository strictly differentiates dataset origins and integrity levels:
| Provenance Level | Description | Benchmark Datasets |
|---|---|---|
EMPIRICAL_EXTERNAL_DATA |
Real observations from authentic commercial payment logs or public blockchain DAGs. | ieee_cis, creditcard, elliptic |
PUBLIC_SIMULATED_DATASET |
Public agent-based simulations published with physical source files and published baseline distributions. | paysim, amlsim |
CONTROLLED_PROJECT_SYNTHETIC |
Reproducible benchmark datasets generated via domain-specific statistical models or academic methodology. | synthaml, amlnet, cross_bank |
TEST_FIXTURE |
Minimal in-process deterministic schema fixtures for unit testing and CI smoke gates (is_synthetic=True). |
Generated in dataloader_smoke_fixtures.py |
8.3 Provisioning Matrix & CI Reality
| Dataset ID | Authoritative Source | Required Files & Format | Storage Location | Acquisition Method | Auth & License Requirements | Standard CI | Real-Data CI Gate | Local Execution |
|---|---|---|---|---|---|---|---|---|
paysim |
Kaggle ealaxi/paysim1 |
PS_20174392719_1491204439457_log.csv (CSV) |
backend/storage/datasets/paysim/ |
Kaggle API / CLI | Kaggle API Key; CC BY-SA 4.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk |
ieee_cis |
Kaggle c/ieee-fraud-detection |
train_transaction.csv, train_identity.csv (CSV) |
backend/storage/datasets/ieee_cis/ |
Kaggle Competition CLI | Kaggle API Key + Competition Rules Acceptance | Schema Fixture | Requires Kaggle credentials & rules acceptance | Real file on disk |
creditcard |
Kaggle mlg-ulb/creditcardfraud |
creditcard.csv (CSV) |
backend/storage/datasets/creditcard/ |
Kaggle API / CLI | Kaggle API Key; ODbL 1.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk |
elliptic |
Kaggle ellipticco/elliptic-data-set |
elliptic_txs_features.csv, elliptic_txs_classes.csv, elliptic_txs_edgelist.csv (CSV) |
backend/storage/datasets/elliptic/ |
Kaggle API / CLI | Kaggle API Key; CC BY 4.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk |
amlsim |
Kaggle anshankul/ibm-amlsim-example-dataset |
transactions.csv, accounts.csv, alerts.csv (CSV) |
backend/storage/datasets/amlsim/ |
Kaggle API / CLI | Kaggle API Key; Apache 2.0 | Schema Fixture | Provisions via Kaggle API | Real file on disk |
synthaml |
Figshare / Nature Sci Data DOI: 10.1038/s41597-023-02569-2 |
alerts.csv, transactions.csv (CSV/Parquet) |
backend/storage/datasets/synthaml/ |
Direct Download / generate_synthaml_dataset.py |
Open Access; CC BY 4.0 | Schema Fixture | Local cache / Generator script | Real file on disk |
amlnet |
Zenodo DOI: 10.5281/zenodo.10058474 |
transactions.csv (CSV/Parquet) |
backend/storage/datasets/amlnet/ |
Direct Download / generate_amlnet_dataset.py |
Open Access; CC BY-NC 4.0 | Schema Fixture | Local cache / Generator script | Real file on disk |
8.4 Automated Provisioning Workflow
The standalone GitHub Actions workflow .github/workflows/real_data_integration.yml provides a secure, gated pipeline for real-data integration:
- Triggered manually via
workflow_dispatchor on a weekly schedule. - Uses repository secrets
KAGGLE_USERNAMEandKAGGLE_KEY. - Executes
python scripts/download_real_benchmarks.py --dataset all. - Fails closed if physical datasets cannot be provisioned when mandatory execution is requested.