fathom-code / assets /DEMO_SCRIPT.md
23f2002275
feat(C): demo materials - architecture diagram, demo script, blog draft, viz reward-pie panel, README polish (preflight green)
fb74a9b
|
Raw
History Blame Contribute Delete
3.3 kB
# 90-second demo video script β€” FATHOM
Target length: 90 seconds. Recording specs: 1080p / 30 fps / OBS Studio or Windows Game Bar (Win + G). Save as `assets/demo.mp4` LOCALLY ONLY (do NOT commit large video files; upload to YouTube unlisted, link from README).
---
## Beat sheet
### [0:00 – 0:10] TITLE CARD
**On screen:** Big text β€” "FATHOM β€” the first RL-trained Recursive Language Model. A 1.5B model that reads documents 50Γ— larger than its context window."
**Voiceover:**
> "FATHOM β€” the first RL-trained Recursive Language Model. A 1.5B model that reads documents 50Γ— larger than its context window."
---
### [0:10 – 0:25] THE ENV
**On screen:** Browser β†’ `https://Pratham-math-fathom-env.hf.space/openapi.json`. Scroll the OpenAPI page so the `/reset`, `/step`, `/healthz` endpoints are visible.
**Voiceover:**
> "Our OpenEnv server gives the agent two tools β€” a sandboxed Python REPL and a recursive `llm()` call. It's a public Hugging Face Space; anyone can hit it."
---
### [0:25 – 0:45] THE REWARD
**On screen:** Cursor / VS Code with `REWARD_AUDIT.md` open. Scroll the table that lists the 5 attacks (masked-context, format-only, length gaming, recursion-spam, copy-question).
**Voiceover:**
> "We hardened the verifier against five reward-hacking attacks before training. Every component is grep-verifiable. `pytest -m reward_audit` catches masked-context exploits, format-only attacks, and length gaming."
---
### [0:45 – 1:10] THE TRAINING
**On screen:** Browser β†’ W&B run page β†’ reward curve panel (composite reward over steps). Pause briefly on the rising curve.
**Voiceover:**
> "Here's GRPO training Qwen 2.5 Coder 1.5B against the FATHOM env. Composite reward rises from baseline to a clean trained value across the run. The dashed line is an untrained Qwen baseline."
---
### [1:10 – 1:25] THE OUTCOME
**On screen:** Streamlit running at `localhost:8501` OR a terminal showing `python -m env.client --doc 200k.txt --question "..."`. Show the recursion tree visualization rendering the model's tool calls.
**Voiceover:**
> "The trained model decomposes the long document, calls itself recursively, and answers correctly using only its 4K native context."
---
### [1:25 – 1:30] CLOSE
**On screen:** README.md with the Reproduce section visible β€” Colab link badge.
**Voiceover:**
> "Full training reproducer in our Colab notebook. Code public on Hugging Face and GitHub. FATHOM."
---
## Recording checklist
- [ ] OBS or Game Bar set to 1080p / 30 fps / mic on
- [ ] Browser tabs pre-loaded so no waiting on page-loads during the take
- [ ] `outputs/plots/reward_curve.png` rendered + W&B run public BEFORE recording
- [ ] Streamlit `viz/app.py` already running on `localhost:8501`
- [ ] Single take preferred; if more, edit ruthlessly to ≀90s
- [ ] Upload to YouTube as **Unlisted**, copy URL into README + submission form
- [ ] Save the URL into `assets/DEMO_URL.txt` (one line) so the preflight can grep it
## Plan-B if recording fails
The hackathon accepts mini-blog OR slides OR video. If recording falls through:
- Use `assets/BLOG_DRAFT.md` as the writeup and post on huggingface.co/blog
- OR build a 5-slide PDF deck (Canva) and save as `assets/pitch.pdf`
- Either route satisfies the storytelling requirement