Instructions to use agenthinkmesh/gcc_retail_ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use agenthinkmesh/gcc_retail_ai with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("agenthinkmesh/gcc_retail_ai", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
GCC Retail & Consumer AI
Ten scikit-learn models for retail and consumer analytics framed around Kuwait and the GCC: customer lifetime value, loyalty churn, Ramadan sales, basket size, mall footfall, price sensitivity, product returns, luxury demand, flash-sale conversion and omnichannel revenue attribution. One script, gcc_retail_ai.py, generates a synthetic dataset for each task (12,000 rows each). It then compares five candidate algorithms with 3-fold cross-validation and saves the best one per task as a joblib pickle. The models are trained only on simulated data. They are demonstrations and templates, not validated business tools.
Author: AgenThink, Kuwait City
Models
Test metrics are on a 20% held-out split of the same synthetic data (random_state=42). CV is the mean 3-fold cross-validation score on the 80% training split. All numbers come from automl_results.json. Money values are in KWD.
| # | File | Task | Type | Selected algorithm | CV score | Test score | Test MAE |
|---|---|---|---|---|---|---|---|
| 1 | model_clv.pkl |
Customer Lifetime Value | Regression | RandomForest | R² 0.9976 | R² 0.9983 | 460.53 KWD |
| 2 | model_loyalty_churn.pkl |
Loyalty Program Churn Prediction | Classification | GradientBoosting | Acc 0.9524 | Acc 0.9554 | – |
| 3 | model_ramadan_sales.pkl |
Ramadan Sales Forecasting | Regression | Ridge | R² 0.8549 | R² 0.8788 | 847.57 KWD |
| 4 | model_basket_size.pkl |
Basket Size Prediction | Regression | Ridge | R² 0.9800 | R² 0.9787 | 7.91 KWD |
| 5 | model_mall_footfall.pkl |
Mall Footfall Forecasting | Regression | GradientBoosting | R² 0.9184 | R² 0.9187 | 1,680.01 visitors |
| 6 | model_price_sensitivity.pkl |
Price Sensitivity Scoring | Regression | Ridge | R² 0.7822 | R² 0.778 | 0.082 (elasticity) |
| 7 | model_product_return.pkl |
Product Return Prediction | Classification | LogisticRegression | Acc 0.9611 | Acc 0.9571 | – |
| 8 | model_luxury_demand.pkl |
Luxury Goods Demand Forecasting | Regression | GradientBoosting | R² 0.8256 | R² 0.8289 | 6.19 (index) |
| 9 | model_flash_sale.pkl |
Flash Sale Conversion Rate | Regression | Ridge | R² 0.9303 | R² 0.9356 | 0.024 (rate) |
| 10 | model_attribution.pkl |
Omnichannel Revenue Attribution | Regression | GradientBoosting | R² 0.9354 | R² 0.935 | 17.35 KWD |
The candidates were RandomForest, GradientBoosting, ExtraTrees, DecisionTree, and either Ridge (for regression) or LogisticRegression (for classification). The CV score of every candidate is stored under all_scores in automl_results.json. The saved estimator classes match the best_algo values in the JSON. The script's docstring lists Flash Sale Conversion as a classification task, but it is trained and saved as a regressor that predicts a conversion rate.
Dashboard
The script generates this dashboard. It shows the test score for each model, plots of the synthetic data, feature importances and a heatmap of the algorithm comparison.
Repository files
| File | Description |
|---|---|
README.md |
This model card |
gcc_retail_ai.py |
Full pipeline: synthetic data generation, AutoML selection, evaluation, dashboard, and saving of models and results |
automl_results.json |
For each model: task, type, selected algorithm, CV and test metrics, CV scores of all candidates, ordered feature list, feature importances, and a free-text gcc_note |
gcc_retail_ai_dashboard.png |
Performance and data dashboard (about 1 MB) |
model_*.pkl (10 files) |
Fitted scikit-learn estimators saved with joblib. model_clv.pkl is about 10 MB (a random forest); the others are small. The pickles record scikit-learn 1.8.0 as the training version |
Training data
All training data is synthetic. The script generates it with NumPy (np.random.seed(42)) from hand-chosen distributions. It does not ship the data; running the script recreates it. Each task has 12,000 independent rows and 20 input features.
- Regression targets are hand-written formulas over a subset of the features, plus Gaussian noise, clipped to a range. For example, CLV is roughly purchase frequency × order value × 30, plus tier, tenure, referral and subscription terms.
- Classification targets come from a hand-weighted risk score. A row is labelled
1if its score is at or above a quantile: 0.78 for churn, giving about 22% positives, and 0.82 for returns, about 18%. - Several features do not appear in the target formula and are pure noise.
- Categorical variables such as
category,store_type,customer_segment,channelandpayment_typeare integer codes (for example 1–8). Their mapping to real categories is not documented. - The
gcc_notestrings in the JSON are the author's contextual remarks, for example "Post-Iftar 8-11pm is GCC retail's peak window". They are not derived from or validated by data.
Features and targets
Feature order matters. The order below matches features in automl_results.json and model.feature_names_in_.
- CLV: target
clv_kwd(50–50,000). Features:age, is_national, monthly_income_kwd, tenure_months, purchase_freq_mo, avg_order_kwd, category_diversity, loyalty_tier, online_pct, ramadan_spend_boost, eid_spend_boost, referred_friends, app_engagement, returns_rate, credit_card_user, family_size, nps_score, social_media_influence, subscription_holder, years_in_gcc - Loyalty Churn: target
churns(0/1). Features:tenure_months, loyalty_tier, points_balance, points_expiring_soon, months_inactive, purchase_freq_drop, competitor_offer, app_logins_mo, email_open_rate, complaint_count, last_redemption_days, is_national, age, family_account, ramadan_engagement, nps_score, avg_order_kwd, channel_shift, price_sensitivity, vip_event_attended - Ramadan Sales: target
sales_kwd(200–300,000). Features:day_of_ramadan, hour, category, store_type, year, is_weekend, is_last_10_days, prev_year_sales_kwd, temperature_c, promotions_active, discount_depth, competitor_promo, staff_level, inventory_level, online_share, footfall_idx, expat_pct, iftar_time_hr, suhoor_spike, charity_giving_spike - Basket Size: target
basket_kwd(5–1,500). Features:customer_segment, store_format, hour, day_of_week, is_ramadan, is_eid, is_weekend, loyalty_tier, income_tier, family_size, promotion_active, discount_kwd, prev_basket_kwd, items_count, fresh_food_pct, private_label_pct, time_in_store_min, uses_list, payment_type, temperature_c - Mall Footfall: target
footfall(500–120,000). Features:mall_tier, hour, day_of_week, month, is_ramadan, is_eid, is_summer, temperature_c, is_weekend, event_happening, school_holiday, prayer_time, rain, sandstorm, anchor_store_promo, parking_capacity_pct, new_store_opened, competitor_mall_event, public_holiday, prev_week_footfall - Price Sensitivity: target
price_elasticity(−1.5 to 0). Features:income_tier, is_national, age, category, brand_loyalty_score, price_comparison_freq, coupon_usage, private_label_affinity, oil_economy_phase, household_size, has_children, monthly_spend_kwd, deals_app_user, impulse_buyer, bulk_buyer, luxury_affinity, online_shopper, subscription_user, expat_remittance, recession_worried - Product Return: target
returned(0/1). Features:category, price_kwd, channel, customer_tier, days_to_return, size_issue, quality_complaint, gift_purchase, impulse_purchase, description_mismatch, is_ramadan, is_eid, prev_return_rate, free_return_policy, influencer_purchase, size_chart_viewed, reviews_read, payment_type, first_purchase, luxury_brand - Luxury Demand: target
luxury_demand_idx(10–200). Features:brand_tier, category, month, is_eid, is_ramadan, oil_price_usd, gcc_gdp_growth, tourist_arrivals_k, national_income_idx, social_media_buzz, celebrity_endorsement, new_collection, gifting_season, gold_price_idx, youth_bulge, female_consumer_pct, duty_free_competition, counterfeit_pressure, store_experience, loyalty_vip_count - Flash Sale Conversion: target
conversion_rate(0.01–0.95). Features:discount_depth, sale_duration_hr, category, original_price_kwd, stock_units, customer_segment, notification_sent, email_push, sms_push, social_promotion, hour_launched, is_ramadan, is_eid, day_of_week, prev_sale_conversion, influencer_amplified, loyalty_exclusive, app_only, bundle_offer, free_shipping - Omnichannel Attribution: target
online_revenue_attributed(KWD, 1–5,000). Features:touchpoints_count, instagram_touch, snapchat_touch, tiktok_touch, whatsapp_touch, email_touch, sms_touch, store_visit, app_browse, web_browse, influencer_touch, paid_search, days_to_purchase, is_ramadan, is_eid, customer_age, order_value_kwd, first_touch_channel, last_touch_channel, device_type
For the Ridge and LogisticRegression models, the feature_importance values in the JSON are normalized absolute coefficients on unscaled inputs. They depend on feature scale.
How to use
The models are plain scikit-learn estimators saved with joblib. Use scikit-learn 1.8.x to avoid unpickling problems.
# pip install scikit-learn==1.8.* pandas joblib huggingface_hub
import json, joblib, pandas as pd
from huggingface_hub import hf_hub_download
repo = "agenthinkmesh/gcc_retail_ai"
results = json.load(open(hf_hub_download(repo, "automl_results.json"), encoding="utf-8"))
model = joblib.load(hf_hub_download(repo, "model_loyalty_churn.pkl"))
features = results["loyalty_churn"]["features"] # same order as model.feature_names_in_
row = {
"tenure_months": 14.0, "loyalty_tier": 2, "points_balance": 1200.0, "points_expiring_soon": 1,
"months_inactive": 6.0, "purchase_freq_drop": 0.5, "competitor_offer": 1, "app_logins_mo": 0,
"email_open_rate": 0.2, "complaint_count": 2, "last_redemption_days": 240.0, "is_national": 0,
"age": 34, "family_account": 0, "ramadan_engagement": 0, "nps_score": 4.0,
"avg_order_kwd": 45.0, "channel_shift": 1, "price_sensitivity": 4, "vip_event_attended": 0,
}
X = pd.DataFrame([row], columns=features)
print("churn probability:", model.predict_proba(X)[0, 1])
Regression models work the same way: build a one-row DataFrame with that model's features and call model.predict(X). Loading a pickle can execute arbitrary code, so only load files from sources you trust.
Re-running training
pip install numpy pandas matplotlib scikit-learn joblib
# The output directory is hard-coded: edit OUT = "/home/claude/gcc_retail_ai" near the top of the script first
python gcc_retail_ai.py
The script regenerates the data, all 10 models, automl_results.json, the dashboard and a short README.md. It will overwrite a README in the output directory.
Intended use
- Examples and starting templates for tabular AutoML on retail and CRM-style problems with GCC seasonal features (Ramadan, Eid, summer heat, prayer times).
- Teaching, prototyping and schema design before retraining on real, consented customer data.
Out of scope: real pricing, credit, marketing-targeting or customer-treatment decisions without retraining and validation on real data.
Limitations and risks
- Synthetic data only. High R² and accuracy values show that the models recover the hand-written formulas in the generator. They say nothing about real retail behaviour, and performance on real data is unknown.
- Evaluation is in-distribution. Rows are i.i.d. draws from the generator. "Forecasting" tasks such as Ramadan sales and mall footfall are not evaluated on future time periods. Only accuracy is reported for the classifiers, with no precision, recall or calibration, even though positive rates are about 18–22%.
- Sensitive attributes. Some models take nationality- or origin-related inputs:
is_nationalin CLV, churn and price sensitivity;expat_remittancein price sensitivity;expat_pctin Ramadan sales. Using such scores to set prices, offers or service levels for individuals could be discriminatory or unlawful. Any real use needs a legal and fairness review and human oversight. - Encoded categories are undocumented. Integer category and segment codes have no published mapping, so real-world inputs cannot be encoded faithfully without redefining them.
- Unscaled linear models. Coefficient-based importances depend on feature scale.
License
MIT
Citation
@misc{agenthink2026gccretail,
author = {AgenThink},
title = {GCC Retail \& Consumer AI},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/agenthinkmesh/gcc_retail_ai}}
}
- Downloads last month
- -
