GCC Logistics & Supply Chain AI: 10-Model Suite

This repository contains ten scikit-learn models for logistics and supply-chain questions in the Gulf Cooperation Council (GCC), with a Kuwait focus:

  • Risk classifiers: port congestion, cold-chain breaches, fleet maintenance needs, driver fatigue, supply-chain disruption and halal-logistics non-compliance.
  • Regressors: last-mile delivery time, customs clearance time, warehouse occupancy and cross-border freight cost.

One training script (gcc_logistics_ai.py) builds all ten. For each task it generates a synthetic dataset, runs a small AutoML search over five candidate algorithms, and saves the winning model. The suite is a prototyping baseline for logistics-analytics developers and operations analysts. It is not a production planning system.

Author: AgenThink, Kuwait City

Models

All scores come from automl_results.json. "CV" is the mean 3-fold cross-validation score on the 80% training split, and that score was used to pick the algorithm. "Test" is the score on the 20% hold-out split. All data is synthetic (see Training data). The JSON does not store class names or target descriptions for this suite; the labels and targets below are taken from the script.

# File Task Type / target Selected algorithm CV score Test score Test MAE
1 model_port_congestion.pkl Port Congestion Risk Classification (congested 0/1) GradientBoosting Acc 0.942 Acc 0.938 —
2 model_delivery_eta.pkl Last-Mile Delivery ETA Regression (delivery_minutes, 5–240) GradientBoosting R² 0.959 R² 0.959 7.09 min
3 model_cold_chain.pkl Cold Chain Breach Risk Classification (breach 0/1) GradientBoosting Acc 0.951 Acc 0.952 —
4 model_customs_clearance.pkl Customs Clearance Duration Regression (clearance_hours, 1–120) GradientBoosting (per JSON; see warning) R² 0.712 R² 0.731 2.81 h
5 model_fleet_maintenance.pkl Fleet Maintenance Prediction Classification (needs_maintenance 0/1) GradientBoosting Acc 0.942 Acc 0.945 —
6 model_warehouse_demand.pkl Warehouse Demand Forecasting Regression (occupancy_pct, 0.10–1.00) ExtraTrees (per JSON; see warning) R² 0.478 R² 0.430 0.0013
7 model_freight_cost.pkl Cross-Border Freight Cost Regression (log1p(freight_cost_kwd)) GradientBoosting R² 0.901 R² 0.906 0.1055 (log units)
8 model_driver_fatigue.pkl Driver Fatigue Risk Detection Classification (fatigued 0/1) GradientBoosting Acc 0.957 Acc 0.970 —
9 model_supply_chain.pkl Supply Chain Disruption Risk Classification (disrupted 0/1) GradientBoosting Acc 0.947 Acc 0.953 —
10 model_halal_compliance.pkl Halal Logistics Compliance Classification (non_compliant 0/1) GradientBoosting Acc 0.964 Acc 0.968 —

Candidate algorithms, with fixed hyperparameters from the script:

  • Regression: RandomForest, GradientBoosting, ExtraTrees (80 trees; max depth 10 for the forests and 4 for boosting), Ridge (alpha = 10) and DecisionTree (max depth 8).
  • Classification: the same set, with LogisticRegression in place of Ridge.

No hyperparameter tuning was done. Per-candidate CV scores are stored under all_scores in automl_results.json.

Warning: two committed model files do not match the script or the reported scores.

  • model_customs_clearance.pkl is a GradientBoostingRegressor with 100 estimators that expects 12 unnamed features. The script trains 80-estimator models on the 20 named features listed below.
  • model_warehouse_demand.pkl is a GradientBoostingRegressor with 100 estimators that expects 7 unnamed features. The JSON reports an ExtraTrees model trained on 20 features.

This repository does not document where these two files came from, what their feature order is, or how they score. The metrics in the table above describe the script's models, not these files. Treat both files as unusable until they are regenerated.

The warehouse task is also weak as defined in the script. Replaying the seeded generator shows about 97.6% of occupancy_pct values at the 1.00 cap. That explains the near-zero MAE (0.0013) alongside a low R² (0.430).

Accuracy vs. class balance

For every classifier, 1 is the risk class. Labels come from thresholding a synthetic risk score at a quantile. The shares below were computed by replaying the seeded generator.

Model Share labeled 1 (risk) Majority baseline Test accuracy
Port congestion 28% 0.72 0.938
Cold chain breach 28% 0.72 0.952
Fleet maintenance 30% 0.70 0.945
Driver fatigue 25% 0.75 0.970
Supply chain disruption 28% 0.72 0.953
Halal non-compliance 20% 0.80 0.968

Only accuracy was recorded. No precision or recall figures are available for the risk class.

Dashboard

GCC Logistics & Supply Chain AI dashboard

The training script generates this dashboard. It shows the model scores, feature importances, target distributions and an AutoML algorithm-comparison heatmap.

Repository files

File Description
gcc_logistics_ai.py Training script: synthetic data generation, AutoML selection, evaluation, dashboard, model export
automl_results.json For each model: task, selected algorithm, CV and test scores, all candidate scores, feature list, feature importances, GCC note
model_*.pkl (10 files) Fitted scikit-learn estimators serialized with joblib (pickles record scikit-learn 1.8.0). Two files do not match the script (see the warning above)
gcc_logistics_ai_dashboard.png Summary dashboard image
README.md This model card

Training data

All training data is synthetic. The script generates it at run time with NumPy (np.random.seed(42)). No real shipment, port, fleet, telematics, customs or driver data was used.

  • Size and inputs: each dataset has 12,000 rows and 20 input features drawn from hand-chosen distributions.
  • Labels: classification labels come from a weighted rule plus noise, thresholded at a quantile.
  • Regression targets: hand-written formulas plus noise, clipped to a range.
  • Places: port names (for example Shuwaikh, Shuaiba and Jebel Ali) and route names (for example Kuwait-UAE) are category labels only. No real operational statistics were used.

Each dataset is split 80/20 (random_state=42). The models learn the script's rules, not real logistics behavior.

Features / inputs

For the eight script-consistent models, pass a pandas.DataFrame with exactly these columns, in this order. Those pickles store feature_names_in_.

  1. Port congestion: port_enc, month, day_of_week, vessels_in_port, vessels_waiting, crane_availability, berth_utilization, weather_score, sandstorm, is_ramadan, is_eid, customs_staffing, import_volume_teu, export_volume_teu, reefer_pct, hazmat_pct, oil_price_usd, shipping_index, truck_availability, it_system_status
  2. Delivery ETA: distance_km, hour, day_of_week, is_ramadan, is_eid, temperature_c, sandstorm, traffic_index, district, building_type, package_size, weight_kg, requires_signature, cod, customer_instructions, driver_experience, vehicle_type, route_stops, gps_accuracy, address_quality
  3. Cold chain: product_type, target_temp_c, ambient_temp_c, journey_duration_hr, vehicle_age_yr, refrigeration_quality, loading_time_min, door_openings, is_ramadan, sandstorm, power_backup, iot_monitoring, driver_training, route_length_km, customs_hold_hr, halal_requirement, humidity_target_pct, last_maintenance_days, multiple_stops, insurance_grade
  4. Customs clearance (per script): shipment_type, origin_country_risk, value_kwd, num_line_items, weight_kg, hazmat, food_item, halal_cert_present, docs_complete, broker_quality, is_ramadan, is_eid, weekend, customs_officer_workload, pre_clearance_submitted, inspection_triggered, govt_entity_importer, e_customs, prev_violations, perishable. The committed pickle does not accept this schema.
  5. Fleet maintenance: vtype_enc, age_years, mileage_k, engine_temp_c, oil_pressure, brake_wear_pct, tire_wear_pct, fuel_efficiency, ambient_temp_c, days_since_service, fault_codes, driver_harshness, load_weight_pct, route_roughness, sandstorm_exposure, idle_hours_pct, ac_usage_pct, maintenance_quality, telematics, prev_breakdowns_yr
  6. Warehouse demand (per script): warehouse_type, district, month, week_of_month, is_ramadan, is_eid, is_national_day, ecommerce_season, oil_price_usd, gdp_growth, expat_population_k, import_volume_prev_mo, retail_sales_idx, new_retailer_onboarding, warehouse_capacity_pct, automation_level, staff_availability, current_tenants, competitor_pricing, year. The committed pickle does not accept this schema.
  7. Freight cost: route_enc, weight_kg, volume_cbm, shipment_type, hazmat, reefer, urgent, mode, fuel_price_idx, oil_price_usd, carrier_competition, is_ramadan, is_eid, border_wait_hr, insurance_required, customs_complexity, distance_km, halal_cert_required, consolidation, booking_lead_days
  8. Driver fatigue: hours_driven_today, hours_driven_week, shift_start_hour, sleep_hours_last, temperature_c, is_night_shift, is_ramadan, driver_age, experience_years, route_monotony, vehicle_type, last_break_hr, heart_rate_avg, eye_blink_rate, caffeine_intake, medical_condition, sandstorm, lane_deviation_score, telematics_alert_hr, employer_compliance
  9. Supply chain disruption: supplier_country_risk, supplier_concentration, single_source_pct, inventory_days, lead_time_days, demand_volatility, geopolitical_risk, oil_price_usd, shipping_disruption_idx, port_congestion, customs_risk, forex_volatility, climate_risk, supplier_financial_health, num_backup_suppliers, digital_visibility, buffer_stock_days, local_content_pct, gcc_strategic_reserve, pandemic_resilience
  10. Halal compliance: product_category, halal_cert_type, cert_issuing_body, origin_country, storage_segregated, transport_dedicated, no_cross_contamination, documentation_complete, last_audit_days, audit_score, staff_trained, temp_monitoring, cleaning_protocol, alcohol_exposure_risk, pork_exposure_risk, supplier_halal_verified, traceability_system, customer_complaints_yr, third_party_verified, renewal_days_remaining

Categorical encodings. The *_enc columns were produced with LabelEncoder, which assigns integer codes in alphabetical order. The encoders are not saved, so apply these mappings yourself:

  • port_enc: 0=Hamad, 1=Jebel Ali, 2=King Abdulaziz, 3=Salalah, 4=Shuaiba, 5=Shuwaikh, 6=Sohar
  • vtype_enc: 0=Bus, 1=Crane, 2=Forklift, 3=Reefer, 4=Tanker, 5=Truck-Heavy, 6=Truck-Medium, 7=Van
  • route_enc: 0=Kuwait-Bahrain, 1=Kuwait-China, 2=Kuwait-Egypt, 3=Kuwait-Europe, 4=Kuwait-India, 5=Kuwait-Jordan, 6=Kuwait-Oman, 7=Kuwait-Qatar, 8=Kuwait-Saudi, 9=Kuwait-UAE

Other coded inputs, for example product_type, shipment_type, mode, district, vehicle_type in models 2 and 8, origin_country, cert_issuing_body and the 1–5 ratings, are integer codes that the script never defines. Check their ranges in gcc_logistics_ai.py.

How to use

pip install "scikit-learn==1.8.0" joblib pandas huggingface_hub

The pickles record scikit-learn 1.8.0. Other versions may fail to load them or may print warnings.

import joblib
import pandas as pd
from huggingface_hub import hf_hub_download

path = hf_hub_download("agenthinkmesh/gcc_logistics_ai", "model_delivery_eta.pkl")
model = joblib.load(path)  # GradientBoostingRegressor

features = ["distance_km", "hour", "day_of_week", "is_ramadan", "is_eid", "temperature_c",
            "sandstorm", "traffic_index", "district", "building_type", "package_size",
            "weight_kg", "requires_signature", "cod", "customer_instructions",
            "driver_experience", "vehicle_type", "route_stops", "gps_accuracy", "address_quality"]

X = pd.DataFrame([{
    "distance_km": 12.0, "hour": 17, "day_of_week": 3, "is_ramadan": 0, "is_eid": 0,
    "temperature_c": 44, "sandstorm": 0, "traffic_index": 4, "district": 2, "building_type": 1,
    "package_size": 2, "weight_kg": 1.8, "requires_signature": 0, "cod": 1,
    "customer_instructions": 0, "driver_experience": 2.5, "vehicle_type": 1, "route_stops": 9,
    "gps_accuracy": 0.85, "address_quality": 3,
}])[features]

print(model.predict(X))  # estimated delivery time in minutes

For the freight-cost model, convert the prediction to KWD with numpy.expm1(model.predict(X)). Use predict_proba for class probabilities from the classifiers.

Retraining

Requirements: numpy, pandas, matplotlib, scikit-learn, joblib.

  1. The output directory is hard-coded as OUT="/home/claude/gcc_logistics_ai". Change it to a local path first.
  2. Run python gcc_logistics_ai.py.

The script regenerates the synthetic data, reruns model selection, and writes all ten .pkl files, automl_results.json and the dashboard PNG. This also replaces the two mismatched pickles with script-consistent ones. The script overwrites README.md in the output directory, so keep a copy of this card.

Intended use

  • Education and demos of tabular ML for logistics, fleet and supply-chain analytics.
  • A template to retrain on real operational data, followed by independent validation.
  • Exploring GCC-specific operational factors such as heat, sandstorms, Ramadan and Eid, cash on delivery and halal handling.

Out-of-scope use

  • Real dispatch, routing, customs, maintenance-safety or procurement decisions.
  • Assessing or disciplining individual drivers based on fatigue or telematics predictions.
  • Halal certification or compliance determinations for real products or facilities.

Limitations and risks

  • Synthetic data only. No model has been validated on real logistics data. Scores measure recovery of the script's own formulas.
  • Mismatched artifacts. model_customs_clearance.pkl and model_warehouse_demand.pkl do not match the script or JSON (see the warning above).
  • Weak and moderate models: warehouse demand (R² 0.430, with a nearly constant target) and customs clearance (R² 0.731, MAE about 2.8 hours).
  • Safety-relevant tasks. For driver fatigue and fleet maintenance, missing a risk case can have safety consequences. Only accuracy was measured; recall on the risk class is unknown.
  • Driver privacy. The driver-fatigue model uses physiological and behavioral inputs (heart rate, blink rate, medical condition). Real deployment would raise privacy and labor-rights issues.
  • Unsaved encoders. Categorical encoders are not saved and many integer codes are undefined, so inputs are easy to mis-specify.
  • Pickle security. Loading a pickle can execute arbitrary code, so load only files you trust.
  • Expert review required. Outputs must not drive decisions about individual drivers, workers, suppliers or shipments without review by qualified operations, safety or compliance staff.

License

MIT, as stated in the original repository README. The repository contains no separate LICENSE file.

Citation

@misc{agenthink2026gcc_logistics_ai,
  author       = {AgenThink},
  title        = {GCC Logistics \& Supply Chain AI: 10-Model Suite},
  year         = {2026},
  howpublished = {Hugging Face model repository},
  url          = {https://huggingface.co/agenthinkmesh/gcc_logistics_ai}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support