Instructions to use agenthinkmesh/gcc_logistics_ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use agenthinkmesh/gcc_logistics_ai with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("agenthinkmesh/gcc_logistics_ai", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
GCC Logistics & Supply Chain AI: 10-Model Suite
This repository contains ten scikit-learn models for logistics and supply-chain questions in the Gulf Cooperation Council (GCC), with a Kuwait focus:
- Risk classifiers: port congestion, cold-chain breaches, fleet maintenance needs, driver fatigue, supply-chain disruption and halal-logistics non-compliance.
- Regressors: last-mile delivery time, customs clearance time, warehouse occupancy and cross-border freight cost.
One training script (gcc_logistics_ai.py) builds all ten. For each task it generates a synthetic dataset, runs a small AutoML search over five candidate algorithms, and saves the winning model. The suite is a prototyping baseline for logistics-analytics developers and operations analysts. It is not a production planning system.
Author: AgenThink, Kuwait City
Models
All scores come from automl_results.json. "CV" is the mean 3-fold cross-validation score on the 80% training split, and that score was used to pick the algorithm. "Test" is the score on the 20% hold-out split. All data is synthetic (see Training data). The JSON does not store class names or target descriptions for this suite; the labels and targets below are taken from the script.
| # | File | Task | Type / target | Selected algorithm | CV score | Test score | Test MAE |
|---|---|---|---|---|---|---|---|
| 1 | model_port_congestion.pkl |
Port Congestion Risk | Classification (congested 0/1) |
GradientBoosting | Acc 0.942 | Acc 0.938 | — |
| 2 | model_delivery_eta.pkl |
Last-Mile Delivery ETA | Regression (delivery_minutes, 5–240) |
GradientBoosting | R² 0.959 | R² 0.959 | 7.09 min |
| 3 | model_cold_chain.pkl |
Cold Chain Breach Risk | Classification (breach 0/1) |
GradientBoosting | Acc 0.951 | Acc 0.952 | — |
| 4 | model_customs_clearance.pkl |
Customs Clearance Duration | Regression (clearance_hours, 1–120) |
GradientBoosting (per JSON; see warning) | R² 0.712 | R² 0.731 | 2.81 h |
| 5 | model_fleet_maintenance.pkl |
Fleet Maintenance Prediction | Classification (needs_maintenance 0/1) |
GradientBoosting | Acc 0.942 | Acc 0.945 | — |
| 6 | model_warehouse_demand.pkl |
Warehouse Demand Forecasting | Regression (occupancy_pct, 0.10–1.00) |
ExtraTrees (per JSON; see warning) | R² 0.478 | R² 0.430 | 0.0013 |
| 7 | model_freight_cost.pkl |
Cross-Border Freight Cost | Regression (log1p(freight_cost_kwd)) |
GradientBoosting | R² 0.901 | R² 0.906 | 0.1055 (log units) |
| 8 | model_driver_fatigue.pkl |
Driver Fatigue Risk Detection | Classification (fatigued 0/1) |
GradientBoosting | Acc 0.957 | Acc 0.970 | — |
| 9 | model_supply_chain.pkl |
Supply Chain Disruption Risk | Classification (disrupted 0/1) |
GradientBoosting | Acc 0.947 | Acc 0.953 | — |
| 10 | model_halal_compliance.pkl |
Halal Logistics Compliance | Classification (non_compliant 0/1) |
GradientBoosting | Acc 0.964 | Acc 0.968 | — |
Candidate algorithms, with fixed hyperparameters from the script:
- Regression: RandomForest, GradientBoosting, ExtraTrees (80 trees; max depth 10 for the forests and 4 for boosting), Ridge (alpha = 10) and DecisionTree (max depth 8).
- Classification: the same set, with LogisticRegression in place of Ridge.
No hyperparameter tuning was done. Per-candidate CV scores are stored under all_scores in automl_results.json.
Warning: two committed model files do not match the script or the reported scores.
model_customs_clearance.pklis aGradientBoostingRegressorwith 100 estimators that expects 12 unnamed features. The script trains 80-estimator models on the 20 named features listed below.model_warehouse_demand.pklis aGradientBoostingRegressorwith 100 estimators that expects 7 unnamed features. The JSON reports anExtraTreesmodel trained on 20 features.This repository does not document where these two files came from, what their feature order is, or how they score. The metrics in the table above describe the script's models, not these files. Treat both files as unusable until they are regenerated.
The warehouse task is also weak as defined in the script. Replaying the seeded generator shows about 97.6% of
occupancy_pctvalues at the 1.00 cap. That explains the near-zero MAE (0.0013) alongside a low R² (0.430).
Accuracy vs. class balance
For every classifier, 1 is the risk class. Labels come from thresholding a synthetic risk score at a quantile. The shares below were computed by replaying the seeded generator.
| Model | Share labeled 1 (risk) | Majority baseline | Test accuracy |
|---|---|---|---|
| Port congestion | 28% | 0.72 | 0.938 |
| Cold chain breach | 28% | 0.72 | 0.952 |
| Fleet maintenance | 30% | 0.70 | 0.945 |
| Driver fatigue | 25% | 0.75 | 0.970 |
| Supply chain disruption | 28% | 0.72 | 0.953 |
| Halal non-compliance | 20% | 0.80 | 0.968 |
Only accuracy was recorded. No precision or recall figures are available for the risk class.
Dashboard
The training script generates this dashboard. It shows the model scores, feature importances, target distributions and an AutoML algorithm-comparison heatmap.
Repository files
| File | Description |
|---|---|
gcc_logistics_ai.py |
Training script: synthetic data generation, AutoML selection, evaluation, dashboard, model export |
automl_results.json |
For each model: task, selected algorithm, CV and test scores, all candidate scores, feature list, feature importances, GCC note |
model_*.pkl (10 files) |
Fitted scikit-learn estimators serialized with joblib (pickles record scikit-learn 1.8.0). Two files do not match the script (see the warning above) |
gcc_logistics_ai_dashboard.png |
Summary dashboard image |
README.md |
This model card |
Training data
All training data is synthetic. The script generates it at run time with NumPy (np.random.seed(42)). No real shipment, port, fleet, telematics, customs or driver data was used.
- Size and inputs: each dataset has 12,000 rows and 20 input features drawn from hand-chosen distributions.
- Labels: classification labels come from a weighted rule plus noise, thresholded at a quantile.
- Regression targets: hand-written formulas plus noise, clipped to a range.
- Places: port names (for example Shuwaikh, Shuaiba and Jebel Ali) and route names (for example Kuwait-UAE) are category labels only. No real operational statistics were used.
Each dataset is split 80/20 (random_state=42). The models learn the script's rules, not real logistics behavior.
Features / inputs
For the eight script-consistent models, pass a pandas.DataFrame with exactly these columns, in this order. Those pickles store feature_names_in_.
- Port congestion:
port_enc, month, day_of_week, vessels_in_port, vessels_waiting, crane_availability, berth_utilization, weather_score, sandstorm, is_ramadan, is_eid, customs_staffing, import_volume_teu, export_volume_teu, reefer_pct, hazmat_pct, oil_price_usd, shipping_index, truck_availability, it_system_status - Delivery ETA:
distance_km, hour, day_of_week, is_ramadan, is_eid, temperature_c, sandstorm, traffic_index, district, building_type, package_size, weight_kg, requires_signature, cod, customer_instructions, driver_experience, vehicle_type, route_stops, gps_accuracy, address_quality - Cold chain:
product_type, target_temp_c, ambient_temp_c, journey_duration_hr, vehicle_age_yr, refrigeration_quality, loading_time_min, door_openings, is_ramadan, sandstorm, power_backup, iot_monitoring, driver_training, route_length_km, customs_hold_hr, halal_requirement, humidity_target_pct, last_maintenance_days, multiple_stops, insurance_grade - Customs clearance (per script):
shipment_type, origin_country_risk, value_kwd, num_line_items, weight_kg, hazmat, food_item, halal_cert_present, docs_complete, broker_quality, is_ramadan, is_eid, weekend, customs_officer_workload, pre_clearance_submitted, inspection_triggered, govt_entity_importer, e_customs, prev_violations, perishable. The committed pickle does not accept this schema. - Fleet maintenance:
vtype_enc, age_years, mileage_k, engine_temp_c, oil_pressure, brake_wear_pct, tire_wear_pct, fuel_efficiency, ambient_temp_c, days_since_service, fault_codes, driver_harshness, load_weight_pct, route_roughness, sandstorm_exposure, idle_hours_pct, ac_usage_pct, maintenance_quality, telematics, prev_breakdowns_yr - Warehouse demand (per script):
warehouse_type, district, month, week_of_month, is_ramadan, is_eid, is_national_day, ecommerce_season, oil_price_usd, gdp_growth, expat_population_k, import_volume_prev_mo, retail_sales_idx, new_retailer_onboarding, warehouse_capacity_pct, automation_level, staff_availability, current_tenants, competitor_pricing, year. The committed pickle does not accept this schema. - Freight cost:
route_enc, weight_kg, volume_cbm, shipment_type, hazmat, reefer, urgent, mode, fuel_price_idx, oil_price_usd, carrier_competition, is_ramadan, is_eid, border_wait_hr, insurance_required, customs_complexity, distance_km, halal_cert_required, consolidation, booking_lead_days - Driver fatigue:
hours_driven_today, hours_driven_week, shift_start_hour, sleep_hours_last, temperature_c, is_night_shift, is_ramadan, driver_age, experience_years, route_monotony, vehicle_type, last_break_hr, heart_rate_avg, eye_blink_rate, caffeine_intake, medical_condition, sandstorm, lane_deviation_score, telematics_alert_hr, employer_compliance - Supply chain disruption:
supplier_country_risk, supplier_concentration, single_source_pct, inventory_days, lead_time_days, demand_volatility, geopolitical_risk, oil_price_usd, shipping_disruption_idx, port_congestion, customs_risk, forex_volatility, climate_risk, supplier_financial_health, num_backup_suppliers, digital_visibility, buffer_stock_days, local_content_pct, gcc_strategic_reserve, pandemic_resilience - Halal compliance:
product_category, halal_cert_type, cert_issuing_body, origin_country, storage_segregated, transport_dedicated, no_cross_contamination, documentation_complete, last_audit_days, audit_score, staff_trained, temp_monitoring, cleaning_protocol, alcohol_exposure_risk, pork_exposure_risk, supplier_halal_verified, traceability_system, customer_complaints_yr, third_party_verified, renewal_days_remaining
Categorical encodings. The *_enc columns were produced with LabelEncoder, which assigns integer codes in alphabetical order. The encoders are not saved, so apply these mappings yourself:
port_enc: 0=Hamad, 1=Jebel Ali, 2=King Abdulaziz, 3=Salalah, 4=Shuaiba, 5=Shuwaikh, 6=Soharvtype_enc: 0=Bus, 1=Crane, 2=Forklift, 3=Reefer, 4=Tanker, 5=Truck-Heavy, 6=Truck-Medium, 7=Vanroute_enc: 0=Kuwait-Bahrain, 1=Kuwait-China, 2=Kuwait-Egypt, 3=Kuwait-Europe, 4=Kuwait-India, 5=Kuwait-Jordan, 6=Kuwait-Oman, 7=Kuwait-Qatar, 8=Kuwait-Saudi, 9=Kuwait-UAE
Other coded inputs, for example product_type, shipment_type, mode, district, vehicle_type in models 2 and 8, origin_country, cert_issuing_body and the 1–5 ratings, are integer codes that the script never defines. Check their ranges in gcc_logistics_ai.py.
How to use
pip install "scikit-learn==1.8.0" joblib pandas huggingface_hub
The pickles record scikit-learn 1.8.0. Other versions may fail to load them or may print warnings.
import joblib
import pandas as pd
from huggingface_hub import hf_hub_download
path = hf_hub_download("agenthinkmesh/gcc_logistics_ai", "model_delivery_eta.pkl")
model = joblib.load(path) # GradientBoostingRegressor
features = ["distance_km", "hour", "day_of_week", "is_ramadan", "is_eid", "temperature_c",
"sandstorm", "traffic_index", "district", "building_type", "package_size",
"weight_kg", "requires_signature", "cod", "customer_instructions",
"driver_experience", "vehicle_type", "route_stops", "gps_accuracy", "address_quality"]
X = pd.DataFrame([{
"distance_km": 12.0, "hour": 17, "day_of_week": 3, "is_ramadan": 0, "is_eid": 0,
"temperature_c": 44, "sandstorm": 0, "traffic_index": 4, "district": 2, "building_type": 1,
"package_size": 2, "weight_kg": 1.8, "requires_signature": 0, "cod": 1,
"customer_instructions": 0, "driver_experience": 2.5, "vehicle_type": 1, "route_stops": 9,
"gps_accuracy": 0.85, "address_quality": 3,
}])[features]
print(model.predict(X)) # estimated delivery time in minutes
For the freight-cost model, convert the prediction to KWD with numpy.expm1(model.predict(X)). Use predict_proba for class probabilities from the classifiers.
Retraining
Requirements: numpy, pandas, matplotlib, scikit-learn, joblib.
- The output directory is hard-coded as
OUT="/home/claude/gcc_logistics_ai". Change it to a local path first. - Run
python gcc_logistics_ai.py.
The script regenerates the synthetic data, reruns model selection, and writes all ten .pkl files, automl_results.json and the dashboard PNG. This also replaces the two mismatched pickles with script-consistent ones. The script overwrites README.md in the output directory, so keep a copy of this card.
Intended use
- Education and demos of tabular ML for logistics, fleet and supply-chain analytics.
- A template to retrain on real operational data, followed by independent validation.
- Exploring GCC-specific operational factors such as heat, sandstorms, Ramadan and Eid, cash on delivery and halal handling.
Out-of-scope use
- Real dispatch, routing, customs, maintenance-safety or procurement decisions.
- Assessing or disciplining individual drivers based on fatigue or telematics predictions.
- Halal certification or compliance determinations for real products or facilities.
Limitations and risks
- Synthetic data only. No model has been validated on real logistics data. Scores measure recovery of the script's own formulas.
- Mismatched artifacts.
model_customs_clearance.pklandmodel_warehouse_demand.pkldo not match the script or JSON (see the warning above). - Weak and moderate models: warehouse demand (R² 0.430, with a nearly constant target) and customs clearance (R² 0.731, MAE about 2.8 hours).
- Safety-relevant tasks. For driver fatigue and fleet maintenance, missing a risk case can have safety consequences. Only accuracy was measured; recall on the risk class is unknown.
- Driver privacy. The driver-fatigue model uses physiological and behavioral inputs (heart rate, blink rate, medical condition). Real deployment would raise privacy and labor-rights issues.
- Unsaved encoders. Categorical encoders are not saved and many integer codes are undefined, so inputs are easy to mis-specify.
- Pickle security. Loading a pickle can execute arbitrary code, so load only files you trust.
- Expert review required. Outputs must not drive decisions about individual drivers, workers, suppliers or shipments without review by qualified operations, safety or compliance staff.
License
MIT, as stated in the original repository README. The repository contains no separate LICENSE file.
Citation
@misc{agenthink2026gcc_logistics_ai,
author = {AgenThink},
title = {GCC Logistics \& Supply Chain AI: 10-Model Suite},
year = {2026},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/agenthinkmesh/gcc_logistics_ai}
}
- Downloads last month
- -
