# Australian Gallops Trifecta Predictor A production-quality, transparent trifecta prediction system for **Australian gallops**, powered by the **FormFav API**. It discovers tomorrow's meetings, fetches fields + form for every runner, scores each runner with a configurable weighted model, ranks them, and emits 1st/2nd/3rd trifecta combinations with confidence and explainable reasoning — never a guarantee. > Predictions are statistical estimates based on available data. They are **not** > guarantees of any racing outcome or betting success. ## Architecture ``` trifecta_bro/ api/ FormFavClient (X-API-Key, retry/backoff, cache), response validator data/ pydantic models, normalizer (scratch/abandon/missing), SQLite storage model/ feature_engine (form-quality + x/spell), scoring (configurable), pace_analysis, probability (trifecta + confidence) reporting/ Markdown dashboard + JSON exporter jobs/ daily_prediction orchestrator (python -m ...) evaluation/ backtester + metrics (no data leakage) tools.py Hermes/Agent reusable tools config/ settings.py (env-driven), weights.yaml (configurable weights) tests/ unittest suite (mocked FormFav, no network) ``` The **LLM/agent is never the prediction engine**. All scoring, ranking, probabilities, confidence and trifecta generation are deterministic statistics so results are reproducible. The agent layer is for orchestration, explanation, and querying. ## Setup ```bash cd trifecta-bro-hf-space python3 -m venv .venv && . .venv/bin/activate pip install -r requirements.txt # httpx, pydantic, pyyaml, reportlab, ... cp .env.example .env # then edit .env and paste your key ``` Configure the key (server-side only — never in source, logs, or reports): ``` FORMFAV_API_KEY=your_key_here ``` All other settings (country, race_code, timezone, cache TTL, DB path, weights path, model version) have sensible defaults and are overridable via env vars. ## Run tomorrow's analysis (live) ```bash python -m trifecta_bro.jobs.daily_prediction # or a specific date: python -m trifecta_bro.jobs.daily_prediction --date 2026-08-11 ``` This will: 1. Resolve tomorrow's date in `Australia/Sydney`. 2. `GET /form/meetings?date=...&country=au&race_code=gallops`. 3. For each AU, non-abandoned meeting, fetch every race via `GET /form`. 4. Normalise, drop scratched runners, skip abandoned races. 5. Score + rank every runner, generate trifecta + savers + confidence. 6. Persist to SQLite (`data/trifecta_bro.db`). 7. Write `data/reports/predictions-.json` and `report-.md`. ## Run without a key (offline / CI) The prediction model runs on any race-form payload — no API needed for analysis, tests, or backtesting. Try the end-to-end demo (uses bundled mock data): ```bash python run_e2e_demo.py ``` ## Backtesting ```bash python -m trifecta_bro.evaluation.backtester --date 2026-08-10 \ --actuals actuals.json ``` `actuals.json` maps `"track_slug:RACE"` -> `"1,7,9"`. Metrics: Top-1 accuracy, place accuracy, trifecta box hit rate, exact trifecta hit rate, average partial coverage, average confidence, and per-race records. The model is evaluated only on pre-race fields — the actual result is never fed back into scoring, so there is **no data leakage**. ## Tests ```bash python -m unittest discover -s tests -v ``` Covers: date logic, meetings retrieval, race retrieval, API-error handling, scratched runners, abandoned races, missing fields, form-string parsing, x/spell parsing, runner scoring, pace scoring, barrier scoring, ranking, trifecta generation, confidence, DB persistence, backtesting, and a full mocked end-to-end pipeline. ## Hermes / Agent tools `trifecta_bro/tools.py` exposes: `get_tomorrow_meetings`, `get_race_form`, `analyse_race`, `rank_runners`, `generate_trifecta`, `save_prediction`, `run_daily_analysis`, `why_ranked_first`. An agent can execute a request like "Run tomorrow's Australian gallops analysis" or "Why did you rank runner 4 first?" entirely through these deterministic wrappers. ## Tuning the model Edit `trifecta_bro/config/weights.yaml` — weights are normalised at load time, so the sum need not be exact. Weights are deliberately a **starting point**, not a claimed optimum; tune them against historical prediction/result data via the backtester, and add ML models (logistic regression, random forest, GBM) behind the same `analyse_race` interface for out-of-sample comparison. ## Limitations (FormFav tier / data) - `/predictions` (premium) returns 403 on basic keys; the system gracefully falls back to the local transparent model only. - Base-tier form data lacks an explicit jockey/trainer *strike-rate* field; the model uses career win%/place% as a transparent proxy and flags it as lower-confidence input. - Non-AU meetings (e.g. NZ jumps) appear in `/form/meetings`; the code filters strictly on `country == au`. - Where FormFav supplies no pace scenario, a derived pressure proxy is used; the pace label is therefore an estimate, not a measured value.