GCC Education AI: 10-Model Suite

This repository contains ten small scikit-learn models for education-sector questions in the Gulf Cooperation Council (GCC) region, with a Kuwait focus. They cover:

  • student performance
  • university dropout
  • school enrollment demand
  • teacher retention
  • Arabic language proficiency
  • STEM readiness
  • special-needs support demand
  • private tutoring demand
  • scholarship eligibility
  • EdTech adoption

A simple AutoML loop picked each model by comparing five scikit-learn algorithms. All training data was generated synthetically by the training script (see Training data). The suite is a prototype and teaching reference for education analysts, researchers and developers. It is not a validated tool for decisions about individual students or teachers.

Author: AgenThink, Kuwait City

Models

"Selected algorithm" means the candidate with the best 3-fold cross-validation score on the training split. Test metrics come from a 20% hold-out split. All numbers are copied from automl_results.json.

# Model file Task Type Selected algorithm CV score Test score Test MAE
1 model_student_performance.pkl Student Academic Performance (GPA, 30–100) Regression Ridge R² 0.8486 R² 0.8524 3.2315
2 model_dropout_risk.pkl University Dropout Risk Classification (2 classes) GradientBoosting Acc 0.9692 Acc 0.9700 –
3 model_enrollment_demand.pkl School Enrollment Demand (students) Regression Ridge R² 0.9328 R² 0.9303 80.6795
4 model_teacher_retention.pkl Teacher Retention (leaves / stays) Classification (2 classes) GradientBoosting Acc 0.9526 Acc 0.9592 –
5 model_arabic_proficiency.pkl (see note) Arabic Language Proficiency (20–100) Regression GradientBoosting R² 0.6984 R² 0.7126 4.4354
6 model_stem_readiness.pkl STEM Career Readiness Classification (2 classes) LogisticRegression Acc 0.8640 Acc 0.8688 –
7 model_special_needs.pkl (see note) Special Needs Support Demand (5–800) Regression Ridge (per JSON) R² 0.2853 R² 0.2568 15.0056
8 model_tutoring_demand.pkl Private Tutoring Demand Index (5–150) Regression GradientBoosting R² 0.8337 R² 0.8370 7.9619
9 model_scholarship.pkl Scholarship Eligibility Classification (2 classes) GradientBoosting Acc 0.9499 Acc 0.9425 –
10 model_edtech_adoption.pkl EdTech Adoption Classification (2 classes) GradientBoosting Acc 0.9445 Acc 0.9458 –

Note: Model/metadata mismatch (models 5 and 7). The uploaded pickles for these two models do not match the training script or automl_results.json:

File Pickled estimator Hyperparameters Inputs Mismatch with JSON
model_arabic_proficiency.pkl GradientBoostingRegressor n_estimators=100, max_depth=5 10 unnamed columns (n_features_in_=10, no feature_names_in_) JSON describes an 80-tree, depth-4 model over 20 named features
model_special_needs.pkl GradientBoostingRegressor n_estimators=100, max_depth=5 8 unnamed columns (n_features_in_=8, no feature_names_in_) JSON reports a Ridge model over 20 named features

These two files were evidently produced by a separate run that is not included in this repository. Their input columns, column order and accuracy are unknown, and the metrics in the table do not describe them. Treat them as unusable until they are regenerated with gcc_education_ai.py. The other eight pickles match the script: same estimator class, same hyperparameters, and the same 18–20 named features in the same order.

Candidate algorithms, with hyperparameters fixed in the script:

  • Regression: RandomForestRegressor(n_estimators=80, max_depth=10), GradientBoostingRegressor(n_estimators=80, max_depth=4, learning_rate=0.1), ExtraTreesRegressor(n_estimators=80, max_depth=10), Ridge(alpha=10.0), DecisionTreeRegressor(max_depth=8).
  • Classification: the matching classifiers, plus LogisticRegression(max_iter=300).

All candidates use random_state=42. automl_results.json records every candidate's CV score under all_scores. The pickles were saved with scikit-learn 1.8.0.

Classification labels, all binary with 1 as the positive class. The positive share follows from the quantile threshold used to build each label:

  • Dropout Risk: 1 = drops out (about 18% positive)
  • Teacher Retention: 1 = teacher leaves (about 22% positive)
  • STEM Readiness: 1 = STEM-ready (about 35% positive)
  • Scholarship: 1 = awarded (about 35% positive)
  • EdTech Adoption: 1 = adopts (about 45% positive)

Dashboard

GCC Education AI dashboard

The training script generates this dashboard. It shows per-model scores, selected feature importances, views of the synthetic data and a heatmap comparing the candidate algorithms.

Repository files

File Size Description
gcc_education_ai.py 44 KB Full pipeline: synthetic data generation, AutoML selection, evaluation, dashboard, model export
automl_results.json 16 KB Per-model task type, selected algorithm, CV/test metrics, all candidate scores, feature list, feature importances, short domain note
gcc_education_ai_dashboard.png 952 KB Results dashboard
model_student_performance.pkl 1 KB Ridge regressor
model_dropout_risk.pkl 203 KB GradientBoostingClassifier
model_enrollment_demand.pkl 1 KB Ridge regressor
model_teacher_retention.pkl 202 KB GradientBoostingClassifier
model_arabic_proficiency.pkl 465 KB GradientBoostingRegressor, 10 unnamed inputs. Does not match the script or JSON (see above)
model_stem_readiness.pkl 2 KB LogisticRegression
model_special_needs.pkl 467 KB GradientBoostingRegressor, 8 unnamed inputs. Does not match the script or JSON (see above)
model_tutoring_demand.pkl 205 KB GradientBoostingRegressor
model_scholarship.pkl 204 KB GradientBoostingClassifier
model_edtech_adoption.pkl 205 KB GradientBoostingClassifier
README.md – This model card

Training data

No real-world data was used. For each task, gcc_education_ai.py generates 12,000 synthetic rows from NumPy random distributions with seed 42. The distributions (uniform, normal, lognormal, beta, Poisson, binomial and categorical draws) use ranges chosen to look plausible for Kuwait and the GCC. For example, the tutoring rate is set to 55% and the Kuwaiti/expat shares are set by hand.

  • Regression targets are hand-written formulas over a few input columns, plus Gaussian noise, then clipped. For example, GPA is a weighted sum of previous GPA, study hours, attendance, tutoring, parental education, teacher experience, screen time, health absences and a Ramadan effect.
  • Classification labels come from thresholding a hand-written risk or eligibility score at a fixed quantile.
  • Unused columns: several inputs do not appear in any generating formula. In Enrollment Demand, covid_effect and tutor_substitute are generated but are not model inputs.
  • Train/test protocol: each task uses an 80/20 train/test split (random_state=42). Model selection uses 3-fold CV on the 80% portion. The selected model is refit on that 80% and scored on the 20% hold-out.

The gcc_note strings in automl_results.json are contextual remarks by the author, for example "Private tutoring rate 55%+ in Kuwait". The code does not derive them from data, and this card does not verify them.

Features / inputs

Columns must match the names and order below. The eight consistent pickles store feature_names_in_. Categorical codes come from LabelEncoder, so they follow alphabetical order.

  1. Student Performance (19 features): school_enc, grade_level, age, is_female, is_kuwaiti, class_size, teacher_experience, parental_education, family_income_tier, study_hours_daily, tutoring, screen_time_hours, attendance_pct, arabic_at_home, prev_gpa, extracurricular, health_absences, summer_school, digital_resources
    • school_enc: 0=International, 1=Private-Arabic, 2=Private-English, 3=Public-Arabic, 4=Public-English
  2. Dropout Risk (20 features): uni_enc, faculty, year_of_study, entry_gpa, current_gpa, is_female, is_kuwaiti, scholarship, works_part_time, commute_hours, attendance_pct, failed_courses, family_pressure, financial_stress, social_integration, mental_health_score, language_barrier, wrong_major, advisor_meetings, oil_job_pull
    • uni_enc: 0=ACK, 1=AOU, 2=AUK, 3=Abroad, 4=GUST, 5=Gulf University, 6=Kuwait University
  3. Enrollment Demand (18 features): school_type_enc, district, year, month, population_growth, expat_inflow, oil_price_usd, birth_rate_lag6yr, private_school_fees, public_capacity, new_residential_dev, expat_policy_strict, summer_exodus, online_alt_pct, nationalization_push, gdp_per_capita, fertility_rate, teachers_available
  4. Teacher Retention (20 features): subject_enc, school_type_enc, years_teaching, age, is_kuwaiti, is_female, salary_kwd, class_load, students_per_class, admin_burden, student_behaviour, parental_pressure, professional_dev, management_support, visa_renewals_left, family_in_gcc, better_offer, housing_provided, commute_km, ramadan_workload
    • subject_enc: 0=Arabic, 1=Arts, 2=English, 3=IT, 4=Islamic, 5=Math, 6=PE, 7=Science, 8=Social, 9=Special
  5. Arabic Proficiency: the script defines 20 features (dialect_enc, age, grade_level, arabic_at_home_hrs, school_arabic_hrs, is_native_arabic, quran_study, arabic_media_hrs, english_dominance, reading_habits, teacher_quality, parental_arabic_edu, digital_arabic_use, private_arabic_tutor, msa_exposure, literature_exposure, bilingual_school, travel_arab_countries, social_media_arabic, family_size). The uploaded pickle expects 10 unnamed inputs, whose identity is unknown.
  6. STEM Readiness (20 features): math_score, science_score, it_score, english_score, grade_level, is_female, stem_interest, stem_extracurricular, coding_experience, robotics_club, parental_stem_bg, stem_teacher_quality, lab_access, national_competition, tutoring_stem, critical_thinking, problem_solving, oil_sector_aspiration, university_stem_intent, scholarship_available
  7. Special Needs Demand: the script defines 20 features (district, school_count, total_students_k, consanguinity_rate, diagnosis_autism, diagnosis_adhd, diagnosis_learning, diagnosis_physical, diagnosis_speech, specialist_count, resource_rooms, budget_kwd_k, awareness_score, parental_stigma, year, early_intervention, expat_student_pct, technology_support, govt_initiative, ngo_support). The uploaded pickle expects 8 unnamed inputs, whose identity is unknown.
  8. Tutoring Demand (20 features): district, grade_level, subject_enc, exam_season, ramadan_period, school_type_enc, family_income_tier, avg_class_performance, parent_education, both_parents_working, prev_year_results, online_tutoring_avail, tutor_hourly_rate, school_quality_score, cultural_pressure, num_siblings, oil_price_env, summer_program, university_prep, population_k
  9. Scholarship (20 features): schol_enc, gpa, is_kuwaiti, is_female, family_income_kwd, parent_govt_employee, target_country, intended_major, extracurricular_score, leadership_roles, volunteer_hours, language_test_score, recommendation_score, interview_score, disability_status, orphan_status, siblings_on_scholarship, prev_scholarship, stem_major, oil_sector_major
    • schol_enc: 0=Corporate, 1=Government, 2=International, 3=Merit, 4=Need-Based, 5=University
  10. EdTech Adoption (20 features): school_type_enc, grade_level, teacher_age, teacher_tech_comfort, device_per_student, internet_quality, it_support_staff, admin_mandate, training_received, student_engagement, parent_digital_lit, platform_usability, content_arabic, post_covid, govt_initiative, budget_edtech_kwd, pilot_success, ai_tools_awareness, oil_price_env, region_connectivity

Some integer-coded inputs are random integers with no defined category mapping: school_type_enc in models 3, 4, 8 and 10, plus faculty, district and subject_enc in model 8. Ordinal scores are integers from 1 to 5, and flags are 0 or 1.

How to use

The script saved the models with joblib.dump. Load them with joblib, using scikit-learn 1.8.0.

import joblib
import pandas as pd
from huggingface_hub import hf_hub_download

path = hf_hub_download("agenthinkmesh/gcc_education_ai", "model_dropout_risk.pkl")
model = joblib.load(path)

X = pd.DataFrame([{
    "uni_enc": 6,              # Kuwait University
    "faculty": 3,
    "year_of_study": 2,
    "entry_gpa": 80.0,
    "current_gpa": 58.0,
    "is_female": 1,
    "is_kuwaiti": 1,
    "scholarship": 0,
    "works_part_time": 0,
    "commute_hours": 0.5,
    "attendance_pct": 0.55,
    "failed_courses": 3,
    "family_pressure": 3,
    "financial_stress": 4,
    "social_integration": 2,
    "mental_health_score": 45.0,
    "language_barrier": 0,
    "wrong_major": 1,
    "advisor_meetings": 1,
    "oil_job_pull": 0,
}])[list(model.feature_names_in_)]

print(model.predict(X))        # 1 = predicted dropout
print(model.predict_proba(X))  # [P(stay), P(dropout)]

automl_results.json lists the expected columns for each model under <model_key>.features. That list does not apply to model_arabic_proficiency.pkl or model_special_needs.pkl, as noted above.

Rerunning the training pipeline (this also regenerates consistent versions of the two mismatched pickles):

pip install numpy pandas matplotlib scikit-learn==1.8.0 joblib
# The script writes to OUT = "/home/claude/gcc_education_ai" (line 39); edit that path first.
python gcc_education_ai.py

The script writes all ten .pkl files, automl_results.json, the dashboard PNG and a short README.md to OUT. If you point it at a clone of this repo, that README.md replaces this card. Exact numbers can vary slightly across NumPy and scikit-learn versions.

Intended use

  • Demonstrating and teaching a tabular AutoML workflow on GCC education themes.
  • Prototyping dashboards and what-if tools before real, governed data is available.
  • A template for retraining on real institutional data, with proper ethics review.

Limitations and risks

  • Synthetic data only. The relationships are hand-written by the script author. They have not been validated against any real student, teacher, school or ministry data.

  • Do not use for decisions about individuals. Admissions, scholarships, dropout interventions, staffing and special-needs provision are all out of scope.

  • Sensitive attributes. Several inputs are protected or sensitive:

    • gender (is_female)
    • nationality (is_kuwaiti)
    • disability or orphan status
    • consanguinity_rate
    • mental-health scores

    Some synthetic labels depend on these directly. The scholarship label, for example, adds weight for is_kuwaiti. Any real-world version needs fairness and legal review.

  • Low-performing model. Special Needs Support Demand reaches only R² 0.257 on the test set, according to the JSON. Its uploaded pickle is also not the model those metrics describe.

  • Mismatched artifacts. model_arabic_proficiency.pkl and model_special_needs.pkl do not match the code or metadata. Their inputs are undocumented.

  • Accuracy compared with baselines. Positive rates are low for several tasks, so compare accuracy with the majority-class baseline:

    • Dropout Risk: about 0.82
    • Teacher Retention: about 0.78
    • STEM Readiness and Scholarship: about 0.65
    • EdTech Adoption: about 0.55

    Only accuracy is reported. There is no precision, recall or calibration.

  • Pickle security. .pkl files can execute code when loaded. Only load files you trust.

License

MIT

Citation

@misc{agenthink2026gcceducation,
  author       = {AgenThink},
  title        = {GCC Education AI: 10-Model Suite},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/agenthinkmesh/gcc_education_ai}}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support