mojo-tabular / README.md
lee101's picture
README: registry status
66bcf02 verified
|
Raw History Blame Contribute Delete
3.06 kB
metadata
license: mit
library_name: scikit-learn
pipeline_tag: tabular-classification
tags:
  - tabular
  - cpu
  - tiny
  - cog
  - app-nz

mojo-tabular

Tiny CPU tabular classifiers (one per dataset), total 12 MB, no GPU. Trained with scikit-learn; mojo- is the planned optimization path, not a Mojo implementation yet.

Results (ROC-AUC, seed 42, 60/20/20 stratified split, selection on validation only)

dataset (OpenML) val test untuned best baseline test MB
adult 0.929 0.9282 0.9272 2.64
credit-g 0.8294 0.7986 0.7958 0.00
bank-marketing 0.9354 0.9384 0.9375 2.67
phoneme 0.9523 0.9546 0.9554 5.50
diabetes 0.8528 0.8001 0.8031 1.09

Search: 16 random HistGradientBoosting configs + logistic regression + extra trees; top-3 averaged when validation improves. Single split, single seed: small sets (credit-g n=1000, diabetes n=768) have large validation/test gaps and are noisy. No comparison against XGBoost/CatBoost/TabPFN has been run, so no frontier claim is made. Adding tuned XGBoost 3.4.1 and LightGBM 4.7.0 (6 configs each) to the pool changed test AUC by at most +0.001 (adult 0.929, bank-marketing 0.9386), i.e. these results sit at the plateau of strong boosted trees; the shipped models use sklearn only to keep the image small. CatBoost and TabPFN were not compared (CatBoost hung under the 4 GiB address-space limit). Data: public OpenML sets (license listed as Public). Reproduce: train/fetch_tabular.py, train/tabular_tune.py (--gbdt needs xgboost/lightgbm), artifact hashes in results/tuned.json.

Use

import joblib, pandas as pd
b = joblib.load("models/adult.joblib")  # needs scikit-learn==1.8.0
p = sum(m.predict_proba(df[b["columns"]])[:, 1] for m in b["members"]) / len(b["members"])

Deploy on app.nz (Cog, CPU)

cd cog && python3 -m unittest discover -s tests          # contract test, no network
docker build -t ghcr.io/lee101/mojo-tabular:latest .     # 381 MB (python:3.12-slim), weights baked
docker run --rm -p 5000:5000 --memory 512m ghcr.io/lee101/mojo-tabular:latest   # ~80 MB RSS, ready in <10 s

Deploy to app.nz

POST /predictions {"input": {"dataset": "adult", "rows": "[{...}]"}} returns a JSON list of probabilities (max 1000 rows). Do not use the default cog build image: its apt layer alone is 1.78 GB (2.33 GB total). Links: GitHub, Hugging Face. Status: no registry image is published yet. The GHCR push needs a token with write:packages, and app.nz's remote builder currently runs in simulated mode (APPNZ_BUILDER unset), so the one-click badge will not resolve until an image is pushed. Build locally with the command above.

Limitations

Trained on public benchmark data only; not validated for credit, hiring or medical decisions.