GCC Retail & Consumer AI

Ten scikit-learn models for retail and consumer analytics framed around Kuwait and the GCC: customer lifetime value, loyalty churn, Ramadan sales, basket size, mall footfall, price sensitivity, product returns, luxury demand, flash-sale conversion and omnichannel revenue attribution. One script, gcc_retail_ai.py, generates a synthetic dataset for each task (12,000 rows each). It then compares five candidate algorithms with 3-fold cross-validation and saves the best one per task as a joblib pickle. The models are trained only on simulated data. They are demonstrations and templates, not validated business tools.

Author: AgenThink, Kuwait City

Models

Test metrics are on a 20% held-out split of the same synthetic data (random_state=42). CV is the mean 3-fold cross-validation score on the 80% training split. All numbers come from automl_results.json. Money values are in KWD.

# File Task Type Selected algorithm CV score Test score Test MAE
1 model_clv.pkl Customer Lifetime Value Regression RandomForest R² 0.9976 R² 0.9983 460.53 KWD
2 model_loyalty_churn.pkl Loyalty Program Churn Prediction Classification GradientBoosting Acc 0.9524 Acc 0.9554 –
3 model_ramadan_sales.pkl Ramadan Sales Forecasting Regression Ridge R² 0.8549 R² 0.8788 847.57 KWD
4 model_basket_size.pkl Basket Size Prediction Regression Ridge R² 0.9800 R² 0.9787 7.91 KWD
5 model_mall_footfall.pkl Mall Footfall Forecasting Regression GradientBoosting R² 0.9184 R² 0.9187 1,680.01 visitors
6 model_price_sensitivity.pkl Price Sensitivity Scoring Regression Ridge R² 0.7822 R² 0.778 0.082 (elasticity)
7 model_product_return.pkl Product Return Prediction Classification LogisticRegression Acc 0.9611 Acc 0.9571 –
8 model_luxury_demand.pkl Luxury Goods Demand Forecasting Regression GradientBoosting R² 0.8256 R² 0.8289 6.19 (index)
9 model_flash_sale.pkl Flash Sale Conversion Rate Regression Ridge R² 0.9303 R² 0.9356 0.024 (rate)
10 model_attribution.pkl Omnichannel Revenue Attribution Regression GradientBoosting R² 0.9354 R² 0.935 17.35 KWD

The candidates were RandomForest, GradientBoosting, ExtraTrees, DecisionTree, and either Ridge (for regression) or LogisticRegression (for classification). The CV score of every candidate is stored under all_scores in automl_results.json. The saved estimator classes match the best_algo values in the JSON. The script's docstring lists Flash Sale Conversion as a classification task, but it is trained and saved as a regressor that predicts a conversion rate.

Dashboard

GCC Retail & Consumer AI dashboard

The script generates this dashboard. It shows the test score for each model, plots of the synthetic data, feature importances and a heatmap of the algorithm comparison.

Repository files

File Description
README.md This model card
gcc_retail_ai.py Full pipeline: synthetic data generation, AutoML selection, evaluation, dashboard, and saving of models and results
automl_results.json For each model: task, type, selected algorithm, CV and test metrics, CV scores of all candidates, ordered feature list, feature importances, and a free-text gcc_note
gcc_retail_ai_dashboard.png Performance and data dashboard (about 1 MB)
model_*.pkl (10 files) Fitted scikit-learn estimators saved with joblib. model_clv.pkl is about 10 MB (a random forest); the others are small. The pickles record scikit-learn 1.8.0 as the training version

Training data

All training data is synthetic. The script generates it with NumPy (np.random.seed(42)) from hand-chosen distributions. It does not ship the data; running the script recreates it. Each task has 12,000 independent rows and 20 input features.

  • Regression targets are hand-written formulas over a subset of the features, plus Gaussian noise, clipped to a range. For example, CLV is roughly purchase frequency × order value × 30, plus tier, tenure, referral and subscription terms.
  • Classification targets come from a hand-weighted risk score. A row is labelled 1 if its score is at or above a quantile: 0.78 for churn, giving about 22% positives, and 0.82 for returns, about 18%.
  • Several features do not appear in the target formula and are pure noise.
  • Categorical variables such as category, store_type, customer_segment, channel and payment_type are integer codes (for example 1–8). Their mapping to real categories is not documented.
  • The gcc_note strings in the JSON are the author's contextual remarks, for example "Post-Iftar 8-11pm is GCC retail's peak window". They are not derived from or validated by data.

Features and targets

Feature order matters. The order below matches features in automl_results.json and model.feature_names_in_.

  1. CLV: target clv_kwd (50–50,000). Features: age, is_national, monthly_income_kwd, tenure_months, purchase_freq_mo, avg_order_kwd, category_diversity, loyalty_tier, online_pct, ramadan_spend_boost, eid_spend_boost, referred_friends, app_engagement, returns_rate, credit_card_user, family_size, nps_score, social_media_influence, subscription_holder, years_in_gcc
  2. Loyalty Churn: target churns (0/1). Features: tenure_months, loyalty_tier, points_balance, points_expiring_soon, months_inactive, purchase_freq_drop, competitor_offer, app_logins_mo, email_open_rate, complaint_count, last_redemption_days, is_national, age, family_account, ramadan_engagement, nps_score, avg_order_kwd, channel_shift, price_sensitivity, vip_event_attended
  3. Ramadan Sales: target sales_kwd (200–300,000). Features: day_of_ramadan, hour, category, store_type, year, is_weekend, is_last_10_days, prev_year_sales_kwd, temperature_c, promotions_active, discount_depth, competitor_promo, staff_level, inventory_level, online_share, footfall_idx, expat_pct, iftar_time_hr, suhoor_spike, charity_giving_spike
  4. Basket Size: target basket_kwd (5–1,500). Features: customer_segment, store_format, hour, day_of_week, is_ramadan, is_eid, is_weekend, loyalty_tier, income_tier, family_size, promotion_active, discount_kwd, prev_basket_kwd, items_count, fresh_food_pct, private_label_pct, time_in_store_min, uses_list, payment_type, temperature_c
  5. Mall Footfall: target footfall (500–120,000). Features: mall_tier, hour, day_of_week, month, is_ramadan, is_eid, is_summer, temperature_c, is_weekend, event_happening, school_holiday, prayer_time, rain, sandstorm, anchor_store_promo, parking_capacity_pct, new_store_opened, competitor_mall_event, public_holiday, prev_week_footfall
  6. Price Sensitivity: target price_elasticity (−1.5 to 0). Features: income_tier, is_national, age, category, brand_loyalty_score, price_comparison_freq, coupon_usage, private_label_affinity, oil_economy_phase, household_size, has_children, monthly_spend_kwd, deals_app_user, impulse_buyer, bulk_buyer, luxury_affinity, online_shopper, subscription_user, expat_remittance, recession_worried
  7. Product Return: target returned (0/1). Features: category, price_kwd, channel, customer_tier, days_to_return, size_issue, quality_complaint, gift_purchase, impulse_purchase, description_mismatch, is_ramadan, is_eid, prev_return_rate, free_return_policy, influencer_purchase, size_chart_viewed, reviews_read, payment_type, first_purchase, luxury_brand
  8. Luxury Demand: target luxury_demand_idx (10–200). Features: brand_tier, category, month, is_eid, is_ramadan, oil_price_usd, gcc_gdp_growth, tourist_arrivals_k, national_income_idx, social_media_buzz, celebrity_endorsement, new_collection, gifting_season, gold_price_idx, youth_bulge, female_consumer_pct, duty_free_competition, counterfeit_pressure, store_experience, loyalty_vip_count
  9. Flash Sale Conversion: target conversion_rate (0.01–0.95). Features: discount_depth, sale_duration_hr, category, original_price_kwd, stock_units, customer_segment, notification_sent, email_push, sms_push, social_promotion, hour_launched, is_ramadan, is_eid, day_of_week, prev_sale_conversion, influencer_amplified, loyalty_exclusive, app_only, bundle_offer, free_shipping
  10. Omnichannel Attribution: target online_revenue_attributed (KWD, 1–5,000). Features: touchpoints_count, instagram_touch, snapchat_touch, tiktok_touch, whatsapp_touch, email_touch, sms_touch, store_visit, app_browse, web_browse, influencer_touch, paid_search, days_to_purchase, is_ramadan, is_eid, customer_age, order_value_kwd, first_touch_channel, last_touch_channel, device_type

For the Ridge and LogisticRegression models, the feature_importance values in the JSON are normalized absolute coefficients on unscaled inputs. They depend on feature scale.

How to use

The models are plain scikit-learn estimators saved with joblib. Use scikit-learn 1.8.x to avoid unpickling problems.

# pip install scikit-learn==1.8.* pandas joblib huggingface_hub
import json, joblib, pandas as pd
from huggingface_hub import hf_hub_download

repo = "agenthinkmesh/gcc_retail_ai"
results = json.load(open(hf_hub_download(repo, "automl_results.json"), encoding="utf-8"))
model = joblib.load(hf_hub_download(repo, "model_loyalty_churn.pkl"))

features = results["loyalty_churn"]["features"]   # same order as model.feature_names_in_
row = {
    "tenure_months": 14.0, "loyalty_tier": 2, "points_balance": 1200.0, "points_expiring_soon": 1,
    "months_inactive": 6.0, "purchase_freq_drop": 0.5, "competitor_offer": 1, "app_logins_mo": 0,
    "email_open_rate": 0.2, "complaint_count": 2, "last_redemption_days": 240.0, "is_national": 0,
    "age": 34, "family_account": 0, "ramadan_engagement": 0, "nps_score": 4.0,
    "avg_order_kwd": 45.0, "channel_shift": 1, "price_sensitivity": 4, "vip_event_attended": 0,
}
X = pd.DataFrame([row], columns=features)
print("churn probability:", model.predict_proba(X)[0, 1])

Regression models work the same way: build a one-row DataFrame with that model's features and call model.predict(X). Loading a pickle can execute arbitrary code, so only load files from sources you trust.

Re-running training

pip install numpy pandas matplotlib scikit-learn joblib
# The output directory is hard-coded: edit OUT = "/home/claude/gcc_retail_ai" near the top of the script first
python gcc_retail_ai.py

The script regenerates the data, all 10 models, automl_results.json, the dashboard and a short README.md. It will overwrite a README in the output directory.

Intended use

  • Examples and starting templates for tabular AutoML on retail and CRM-style problems with GCC seasonal features (Ramadan, Eid, summer heat, prayer times).
  • Teaching, prototyping and schema design before retraining on real, consented customer data.

Out of scope: real pricing, credit, marketing-targeting or customer-treatment decisions without retraining and validation on real data.

Limitations and risks

  • Synthetic data only. High R² and accuracy values show that the models recover the hand-written formulas in the generator. They say nothing about real retail behaviour, and performance on real data is unknown.
  • Evaluation is in-distribution. Rows are i.i.d. draws from the generator. "Forecasting" tasks such as Ramadan sales and mall footfall are not evaluated on future time periods. Only accuracy is reported for the classifiers, with no precision, recall or calibration, even though positive rates are about 18–22%.
  • Sensitive attributes. Some models take nationality- or origin-related inputs: is_national in CLV, churn and price sensitivity; expat_remittance in price sensitivity; expat_pct in Ramadan sales. Using such scores to set prices, offers or service levels for individuals could be discriminatory or unlawful. Any real use needs a legal and fairness review and human oversight.
  • Encoded categories are undocumented. Integer category and segment codes have no published mapping, so real-world inputs cannot be encoded faithfully without redefining them.
  • Unscaled linear models. Coefficient-based importances depend on feature scale.

License

MIT

Citation

@misc{agenthink2026gccretail,
  author       = {AgenThink},
  title        = {GCC Retail \& Consumer AI},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/agenthinkmesh/gcc_retail_ai}}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support