Trifecta-Lab / README.md
Brettapps's picture
Upload folder using huggingface_hub (part 20)
013874c verified
|
Raw History Blame Contribute Delete
5.13 kB
# Australian Gallops Trifecta Predictor
A production-quality, transparent trifecta prediction system for **Australian
gallops**, powered by the **FormFav API**. It discovers tomorrow's meetings,
fetches fields + form for every runner, scores each runner with a configurable
weighted model, ranks them, and emits 1st/2nd/3rd trifecta combinations with
confidence and explainable reasoning β€” never a guarantee.
> Predictions are statistical estimates based on available data. They are **not**
> guarantees of any racing outcome or betting success.
## Architecture
```
trifecta_bro/
api/ FormFavClient (X-API-Key, retry/backoff, cache), response validator
data/ pydantic models, normalizer (scratch/abandon/missing), SQLite storage
model/ feature_engine (form-quality + x/spell), scoring (configurable),
pace_analysis, probability (trifecta + confidence)
reporting/ Markdown dashboard + JSON exporter
jobs/ daily_prediction orchestrator (python -m ...)
evaluation/ backtester + metrics (no data leakage)
tools.py Hermes/Agent reusable tools
config/ settings.py (env-driven), weights.yaml (configurable weights)
tests/ unittest suite (mocked FormFav, no network)
```
The **LLM/agent is never the prediction engine**. All scoring, ranking,
probabilities, confidence and trifecta generation are deterministic statistics
so results are reproducible. The agent layer is for orchestration, explanation,
and querying.
## Setup
```bash
cd trifecta-bro-hf-space
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt # httpx, pydantic, pyyaml, reportlab, ...
cp .env.example .env # then edit .env and paste your key
```
Configure the key (server-side only β€” never in source, logs, or reports):
```
FORMFAV_API_KEY=your_key_here
```
All other settings (country, race_code, timezone, cache TTL, DB path, weights
path, model version) have sensible defaults and are overridable via env vars.
## Run tomorrow's analysis (live)
```bash
python -m trifecta_bro.jobs.daily_prediction
# or a specific date:
python -m trifecta_bro.jobs.daily_prediction --date 2026-08-11
```
This will:
1. Resolve tomorrow's date in `Australia/Sydney`.
2. `GET /form/meetings?date=...&country=au&race_code=gallops`.
3. For each AU, non-abandoned meeting, fetch every race via `GET /form`.
4. Normalise, drop scratched runners, skip abandoned races.
5. Score + rank every runner, generate trifecta + savers + confidence.
6. Persist to SQLite (`data/trifecta_bro.db`).
7. Write `data/reports/predictions-<date>.json` and `report-<date>.md`.
## Run without a key (offline / CI)
The prediction model runs on any race-form payload β€” no API needed for analysis,
tests, or backtesting. Try the end-to-end demo (uses bundled mock data):
```bash
python run_e2e_demo.py
```
## Backtesting
```bash
python -m trifecta_bro.evaluation.backtester --date 2026-08-10 \
--actuals actuals.json
```
`actuals.json` maps `"track_slug:RACE"` -> `"1,7,9"`. Metrics: Top-1 accuracy,
place accuracy, trifecta box hit rate, exact trifecta hit rate, average partial
coverage, average confidence, and per-race records. The model is evaluated only
on pre-race fields β€” the actual result is never fed back into scoring, so there
is **no data leakage**.
## Tests
```bash
python -m unittest discover -s tests -v
```
Covers: date logic, meetings retrieval, race retrieval, API-error handling,
scratched runners, abandoned races, missing fields, form-string parsing,
x/spell parsing, runner scoring, pace scoring, barrier scoring, ranking,
trifecta generation, confidence, DB persistence, backtesting, and a full
mocked end-to-end pipeline.
## Hermes / Agent tools
`trifecta_bro/tools.py` exposes: `get_tomorrow_meetings`, `get_race_form`,
`analyse_race`, `rank_runners`, `generate_trifecta`, `save_prediction`,
`run_daily_analysis`, `why_ranked_first`. An agent can execute a request like
"Run tomorrow's Australian gallops analysis" or "Why did you rank runner 4
first?" entirely through these deterministic wrappers.
## Tuning the model
Edit `trifecta_bro/config/weights.yaml` β€” weights are normalised at load time, so
the sum need not be exact. Weights are deliberately a **starting point**, not a
claimed optimum; tune them against historical prediction/result data via the
backtester, and add ML models (logistic regression, random forest, GBM) behind
the same `analyse_race` interface for out-of-sample comparison.
## Limitations (FormFav tier / data)
- `/predictions` (premium) returns 403 on basic keys; the system gracefully
falls back to the local transparent model only.
- Base-tier form data lacks an explicit jockey/trainer *strike-rate* field; the
model uses career win%/place% as a transparent proxy and flags it as
lower-confidence input.
- Non-AU meetings (e.g. NZ jumps) appear in `/form/meetings`; the code filters
strictly on `country == au`.
- Where FormFav supplies no pace scenario, a derived pressure proxy is used; the
pace label is therefore an estimate, not a measured value.