GCC Fintech & Digital Banking AI: 10-Model Suite

This repository contains ten small scikit-learn models for fintech and digital-banking questions in the Gulf Cooperation Council (GCC) region, with a focus on Kuwait. They cover payment fraud, consumer credit scoring, buy-now-pay-later (BNPL) default, crypto portfolio risk, open-banking adoption, churn at Islamic digital banks, remittance flows, KYC risk, anomalies in central bank digital currency (CBDC) transactions, and neobank profitability. A simple AutoML loop picked each model by comparing five scikit-learn algorithms.

All training data was generated synthetically by the training script (see Training data). The suite is a prototype and teaching reference for fintech analysts, students and developers. It is not a validated credit, fraud or compliance system.

Author: AgenThink, Kuwait City

Models

"Selected algorithm" is the candidate with the best 3-fold cross-validation score on the training split. Test metrics come from a 20% hold-out split. All numbers are copied from automl_results.json.

# Model file Task (target) Type Selected algorithm CV score Test score Test MAE
1 model_payment_fraud.pkl Digital Payment Fraud Detection Classification (2 classes) GradientBoosting Acc 0.9828 Acc 0.9854 –
2 model_credit_score.pkl Consumer Credit Score (300–850) Regression GradientBoosting R² 0.7558 R² 0.7620 22.3323
3 model_bnpl_default.pkl BNPL Default Risk Classification (2 classes) GradientBoosting Acc 0.9654 Acc 0.9692 –
4 model_crypto_risk.pkl Crypto Portfolio Risk (VaR %, 1–95) Regression Ridge R² 0.8648 R² 0.8679 4.0182
5 model_open_banking.pkl Open Banking Adoption Classification (2 classes) LogisticRegression Acc 0.9416 Acc 0.9454 –
6 model_islamic_churn.pkl Islamic Digital Banking Churn Classification (2 classes) GradientBoosting Acc 0.9553 Acc 0.9646 –
7 model_remittance_flow.pkl Remittance Flow (million KWD, 0.1–500) Regression GradientBoosting R² 0.8937 R² 0.8829 1.4441
8 model_kyc_risk.pkl KYC Identity Verification Risk Classification (2 classes) GradientBoosting Acc 0.9857 Acc 0.9896 –
9 model_cbdc_anomaly.pkl CBDC Transaction Anomaly Detection Classification (2 classes) GradientBoosting Acc 0.9898 Acc 0.9883 –
10 model_neobank_profit.pkl (see note) Neobank Monthly Profit (KWD) Regression ExtraTrees (per JSON) R² 0.1853 R² 0.2219 3996.5126

Note: model/metadata mismatch (model 10). The uploaded model_neobank_profit.pkl does not match the training script or automl_results.json:

  • What the pickle contains: a GradientBoostingRegressor with n_estimators=100 and max_depth=5, trained on 10 unnamed input columns (n_features_in_=10, no feature_names_in_).
  • What the metadata describes: the JSON reports an ExtraTrees model over 20 named features.

The pickle was evidently produced by a separate run that is not included here. Its inputs, column order and accuracy are unknown, and the metrics in the table do not describe it. Treat the file as unusable until you regenerate it with gcc_fintech_ai.py. Even the scripted version scores poorly (test R² 0.22).

The other nine pickles match the script: same estimator class, same hyperparameters, and the same 19–20 named features in the same order. They were saved with scikit-learn 1.8.0.

The script fixes the candidate algorithms and their hyperparameters:

  • Regression: RandomForestRegressor(n_estimators=80, max_depth=10), GradientBoostingRegressor(n_estimators=80, max_depth=4, learning_rate=0.1), ExtraTreesRegressor(n_estimators=80, max_depth=10), Ridge(alpha=10.0), DecisionTreeRegressor(max_depth=8).
  • Classification: the matching classifiers, plus LogisticRegression(max_iter=300).

All candidates use random_state=42. The CV score of every candidate is stored under all_scores in the JSON.

The classification labels are all binary, with 1 as the positive class. Each label is created by thresholding a synthetic score at a fixed quantile, which sets the approximate positive rate:

Task Positive class (1) Approx. positive rate
Payment Fraud fraud ~5%
BNPL Default default ~18%
Open Banking adopts ~50%
Islamic Churn churns ~22%
KYC high risk ~12%
CBDC anomaly ~7%

Dashboard

GCC Fintech AI dashboard

The training script generates this dashboard. It shows per-model scores, selected feature importances, views of the synthetic data, and a heatmap comparing the candidate algorithms.

Repository files

File Size Description
gcc_fintech_ai.py 37 KB Full pipeline: synthetic data generation, AutoML selection, evaluation, dashboard, model export
automl_results.json 16 KB For each model: task type, selected algorithm, CV and test metrics, all candidate scores, feature list, feature importances, short domain note
gcc_fintech_ai_dashboard.png 763 KB Results dashboard
model_payment_fraud.pkl 202 KB GradientBoostingClassifier
model_credit_score.pkl 204 KB GradientBoostingRegressor
model_bnpl_default.pkl 198 KB GradientBoostingClassifier
model_crypto_risk.pkl 1 KB Ridge regressor
model_open_banking.pkl 2 KB LogisticRegression
model_islamic_churn.pkl 205 KB GradientBoostingClassifier
model_remittance_flow.pkl 198 KB GradientBoostingRegressor
model_kyc_risk.pkl 203 KB GradientBoostingClassifier
model_cbdc_anomaly.pkl 179 KB GradientBoostingClassifier
model_neobank_profit.pkl 440 KB GradientBoostingRegressor with 10 unnamed inputs. Does not match the script or JSON (see note above)
README.md – This model card

Training data

No real-world data was used. For each task, gcc_fintech_ai.py generates 12,000 synthetic rows from NumPy random distributions, with seed 42. The ranges were chosen to look plausible for the GCC: amounts in KWD, Kuwait remittance corridors, and Islamic-banking attributes such as Zakat services and Shariah-board trust.

How targets and labels are built:

  • Regression targets are hand-written formulas plus Gaussian noise. For example, the credit score is a weighted sum of payment history, utilisation, credit age, national status, government or oil-sector employment, income, delinquencies, bankruptcies, property ownership and savings, clipped to 300–850.
  • Neobank profit has noise with standard deviation 5,000 KWD and no clipping, which explains the low R².
  • Classification labels come from thresholding a hand-written risk score at a fixed quantile.
  • Unused inputs: several inputs do not appear in the generating formulas. Remittance Flow generates harvest_season_dest, but the model does not use it as an input.

Evaluation protocol:

  • 80/20 train/test split, with random_state=42.
  • 3-fold cross-validation on the training portion for model selection.
  • The selected model is refit on that 80% and scored on the 20% hold-out.

The gcc_note strings in the JSON, such as "Kuwait remits $15B/yr", are contextual remarks by the author. The code does not derive them, and this card does not verify them.

Features / inputs

Columns must match the names and order below. The nine consistent pickles store feature_names_in_.

  1. Payment Fraud (20): amount_kwd, hour, day_of_week, is_ramadan, merchant_category, cross_border, new_device, new_merchant, velocity_1hr, velocity_24hr, distance_from_home_km, card_present, customer_age, account_age_months, avg_txn_amount, declined_prev_24hr, weekend, ip_country_mismatch, unusual_time, high_risk_merchant
  2. Credit Score (20): age, is_national, monthly_income_kwd, employment_type, years_employed, existing_loans, loan_to_income, payment_history_score, credit_utilization, num_credit_cards, credit_age_months, delinquencies_2yr, bankruptcies, savings_kwd, property_owner, has_guarantor, gcc_residence_years, family_size, oil_sector_employee, govt_employee
  3. BNPL Default (20): purchase_amount_kwd, installments, customer_age, is_national, monthly_income_kwd, existing_bnpl_count, credit_score, payment_history, employment_stable, merchant_category, is_luxury, ramadan_purchase, eid_purchase, first_bnpl, device_type, app_engagement, time_of_purchase, weekend_purchase, referral_source, gcc_expat
  4. Crypto Risk (20): portfolio_value_kwd, btc_pct, eth_pct, altcoin_pct, stablecoin_pct, defi_exposure, nft_exposure, leverage_ratio, holding_period_days, shariah_screened, exchange_risk, custody_type, market_correlation, volatility_30d, liquidity_score, investor_experience, regulatory_jurisdiction, tax_reporting, stop_loss_set, diversification_score
  5. Open Banking (20): age, income_tier, tech_savviness, current_bank_satisfaction, num_banking_apps, fintech_usage, data_privacy_concern, smartphone_usage_hrs, is_national, education_level, trust_in_fintech, peer_adoption, incentive_offered, employer_supports, islamic_compliance, num_financial_products, previous_fintech_issue, cbk_endorsement, social_influence, convenience_score
  6. Islamic Churn (20): account_age_months, products_held, monthly_transactions, app_logins_mo, avg_balance_kwd, zakat_service_used, halal_invest_used, profit_rate_satisfaction, shariah_board_trust, customer_service_score, app_rating, complaint_count, competitor_offer, salary_transferred, family_members_same_bank, is_national, age, digital_only, ramadan_engagement, fee_complaint
  7. Remittance Flow (19): corridor_enc, month, is_ramadan, is_eid, sender_count_k, avg_amount_kwd, oil_price_usd, kwd_to_dest_rate, transfer_fee_pct, mobile_adoption, destination_inflation, employment_rate_gcc, visa_restrictions, bank_vs_app, day_of_month, academic_season, school_fees_season, gcc_salary_day, digital_platform
    • corridor_enc (LabelEncoder, alphabetical): 0=Kuwait-Bangladesh, 1=Kuwait-Egypt, 2=Kuwait-India, 3=Kuwait-Jordan, 4=Kuwait-Lebanon, 5=Kuwait-Nepal, 6=Kuwait-Pakistan, 7=Kuwait-Philippines, 8=Kuwait-Sri Lanka, 9=Kuwait-UK
  8. KYC Risk (20): doc_enc, nationality_risk, doc_age_months, doc_expiry_days, pep_status, adverse_media_hits, address_verified, biometric_match, liveness_score, document_authenticity, previous_rejections, cross_border_applicant, income_source_clear, beneficial_owner_clear, sanctions_check_clear, source_of_funds_clear, digital_footprint, third_party_data_match, application_consistency, processing_channel
    • doc_enc: 0=Civil ID, 1=Other, 2=Passport, 3=Power of Attorney, 4=Residence Permit, 5=Trade License
  9. CBDC Anomaly (20): amount_kwd, transaction_type, sender_wallet_age_days, receiver_wallet_age_days, hour, velocity_5min, velocity_1hr, cross_border_cbdc, merchant_type, is_programmable, smart_contract_call, gas_fee_anomaly, wallet_cluster_risk, amount_round_number, mixing_pattern, unusual_recipient_count, cbdc_to_crypto_bridge, govt_sanctioned_merchant, kyc_tier, zakat_payment
  10. Neobank Profit:
    • Script (20 columns): active_users_k, monthly_active_rate, avg_balance_kwd, transaction_revenue_kwd, interchange_revenue, subscription_revenue, cac_kwd, ltv_kwd, churn_rate_mo, products_per_user, nps_score, regulatory_cost_kwd, tech_cost_kwd, marketing_kwd, islamic_product_pct, b2b_revenue_pct, months_operating, gcc_countries, license_type, breakeven_achieved
    • The uploaded pickle expects 10 unnamed inputs, and which columns they are is unknown.

Inputs such as merchant_category, employment_type, custody_type, regulatory_jurisdiction and transaction_type are integer codes that the script draws at random. They have no documented category mapping. Ordinal scores are integers from 1 to 5, and flags are 0 or 1.

How to use

The models were saved with joblib.dump. Load them with joblib, using scikit-learn 1.8.0.

import joblib
import pandas as pd
from huggingface_hub import hf_hub_download

path = hf_hub_download("agenthinkmesh/gcc_fintech_ai", "model_payment_fraud.pkl")
model = joblib.load(path)

X = pd.DataFrame([{
    "amount_kwd": 7500.0,
    "hour": 3,
    "day_of_week": 4,
    "is_ramadan": 0,
    "merchant_category": 12,
    "cross_border": 1,
    "new_device": 1,
    "new_merchant": 1,
    "velocity_1hr": 6,
    "velocity_24hr": 15,
    "distance_from_home_km": 1200.0,
    "card_present": 0,
    "customer_age": 34,
    "account_age_months": 8.0,
    "avg_txn_amount": 45.0,
    "declined_prev_24hr": 1,
    "weekend": 0,
    "ip_country_mismatch": 1,
    "unusual_time": 1,
    "high_risk_merchant": 1,
}])[list(model.feature_names_in_)]

print(model.predict(X))        # 1 = flagged as fraud
print(model.predict_proba(X))  # [P(legit), P(fraud)]

The expected columns for each model are also listed in automl_results.json under <model_key>.features. This does not apply to model_neobank_profit.pkl, as noted above.

Rerunning the training pipeline (this also regenerates a consistent model_neobank_profit.pkl):

pip install numpy pandas matplotlib scikit-learn==1.8.0 joblib
# The script writes to OUT="/home/claude/gcc_fintech_ai" (line 29); edit that path first.
python gcc_fintech_ai.py

The script writes all ten .pkl files, automl_results.json, the dashboard PNG and a short README.md into OUT. If OUT points at a clone of this repository, the generated README.md replaces this card. Exact numbers can vary slightly across NumPy and scikit-learn versions.

Intended use

  • Demonstrating and teaching a tabular AutoML workflow on GCC fintech themes.
  • Prototyping dashboards, scoring APIs or what-if tools before real data is available.
  • Serving as a template for retraining on real, properly governed transaction and customer data.

Limitations and risks

  • Synthetic data only. Every relationship was hand-written by the script author. No model here has been validated on real customer, transaction, credit-bureau, KYC or market data. Do not use these models to make credit, lending, fraud-blocking, AML/KYC onboarding, investment or regulatory-compliance decisions about real people or transactions.
  • Accuracy is misleading for imbalanced tasks. A classifier that always predicts the negative class would already score about 0.95 on Payment Fraud, 0.93 on CBDC Anomaly, 0.88 on KYC, 0.82 on BNPL Default and 0.78 on Islamic Churn. The reported accuracies of 0.96–0.99 are only modestly above these baselines. The script reports no precision, recall, ROC-AUC or calibration.
  • Low-performing models:
    • Neobank Profitability has a test R² of 0.22 and an MAE of about 4,000 KWD. Its uploaded pickle is also not the model the metrics describe.
    • Consumer Credit Scoring has a test R² of 0.76 and an MAE of about 22 points.
  • Sensitive attributes. Credit, BNPL, open-banking and churn models use nationality or residency proxies (is_national, gcc_expat). KYC uses nationality_risk. Several synthetic labels depend on these directly: the credit score adds points for is_national, and BNPL risk adds weight for gcc_expat. Using them in real credit or KYC decisions may be discriminatory or unlawful, and requires fairness and legal review.
  • Pickle security. .pkl files can execute code when loaded. Only load files you trust.

License

MIT

Citation

@misc{agenthink2026gccfintech,
  author       = {AgenThink},
  title        = {GCC Fintech \& Digital Banking AI: 10-Model Suite},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/agenthinkmesh/gcc_fintech_ai}}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support