Spaces:
Sleeping
Sleeping
KevinIsInCoding
Claude Opus 4.8
feat(trials): RAG-grounded mechanism summary per clinical trial (#34)
4750fb9 unverified |
Download docs/ui-verification-todo.md from KevinIsCoding/candle-fire: direct link, hf CLI and curl.
- Browser
- Download file 5.99 kB
-
https://huggingface.co/spaces/KevinIsCoding/candle-fire/resolve/main/docs/ui-verification-todo.md
- Command line
-
hf download hf://spaces/KevinIsCoding/candle-fire/docs/ui-verification-todo.md
-
curl -L -o ui-verification-todo.md https://huggingface.co/spaces/KevinIsCoding/candle-fire/resolve/main/docs/ui-verification-todo.md
5.99 kB
| # TODO β UI designβmockβimplementβverify loop | |
| > **Status (in progress):** Phases 0β3 + 6 built and green. The launch/drive/assert engine was | |
| > extracted into an **independent, generic package `gradio-ui-verify`** (its own repo); this repo | |
| > is now a *consumer* β `.claude/skills/run-ui/candle_fire_spec.py` is the project spec, run via | |
| > `python -m gradio_ui_verify β¦`. All checks PASS. Building it caught a real bug: the combobox | |
| > placeholder + 3-char gate **never ran** (Gradio ignored `js=`/`demo.load(js=)`) β fixed via | |
| > `gr.Blocks(head=...)`. Still to do: Phase 1 (<5s smoke boot β currently ~25s), Phases 4β5 | |
| > (design-mock front half + visual-QA agent). | |
| **Gap:** UI changes can't be verified from the agent's environment (no browser tooling), so | |
| fixes ship "blind" and rely on the user restarting the app and eyeballing. We also want the | |
| front half: describe a change β see a visual mock β iterate β implement β auto-verify against | |
| that mock. | |
| **Two reframes that make this real (read first):** | |
| 1. **The "mock" is HTML, not a hand-painted PNG.** Claude can't render a raster image from a | |
| description, but it can generate an HTML mock that renders visually and *exports* to PNG. | |
| The PNG is the **reference snapshot** the automated test diffs the real app against. | |
| 2. **Gradio owns the DOM.** The Clinical Trials tab is `gr.Dropdown`/`gr.Row`/etc.; Gradio | |
| generates the markup. So "mock β code" here means **mock β Gradio components + targeted | |
| CSS/JS hooks**, not arbitrary DOM/JS. Fidelity is "close," not pixel-perfect. (Pixel control | |
| would require moving the UI to a raw-HTML/React custom component β a much bigger change, | |
| out of scope unless we decide otherwise.) | |
| Sequence deliberately: prove the risky/automated half (Phases 0β3) before investing in the | |
| design front half (Phases 4β5). Stop and report if Phase 0 fails. | |
| --- | |
| ## Phase 0 β Feasibility spike (DO FIRST; may kill the automated half) | |
| - [ ] Confirm Playwright + Chromium can install in the agent environment | |
| (`uv pip install playwright && playwright install chromium`); note if network-restricted. | |
| - [ ] Boot the app on a test port and confirm a headless Chromium can reach it | |
| (navigate to `http://127.0.0.1:<port>`, read the page title). | |
| - [ ] **If blocked:** stop; switch to the fallback (user runs the browser, agent drives via a | |
| script the user executes). Record the decision here. | |
| ## Phase 1 β Fast "UI smoke" launch mode | |
| - [ ] Add a lightweight launch path that renders the Blocks WITHOUT the heavy startup loads | |
| (cross-encoder, ~31k-chunk ChromaDB, graph) β layout checks don't need them. | |
| e.g. `CANDLE_UI_SMOKE=1` stubs `_collection`/`_graph`/`_cross_encoder` and loads a tiny | |
| trials subset for the combobox vocab. | |
| - [ ] Target cold-start < ~5s in smoke mode (vs ~60s full). | |
| - [ ] Verify the Clinical Trials tab renders with the smoke subset. | |
| ## Phase 2 β Playwright driver | |
| - [ ] Script: launch app (smoke mode, test port, background) β wait for ready β open the | |
| π₯ Clinical Trials tab. | |
| - [ ] Drive the combobox: type into facility/city, read the suggestion list, pick an option, | |
| run a search. | |
| - [ ] Capture a screenshot of the tab to a known path. | |
| ## Phase 3 β Deterministic assertions (the reliable gate; the things that actually broke) | |
| - [ ] `Recruitment status` label renders on ONE line (measure label box height / line count). | |
| - [ ] Dropdown chevron does NOT overlap the value text (compare bounding boxes). | |
| - [ ] 3-char gate: 2 chars β option list hidden; 3 chars β visible. | |
| - [ ] Placeholder shows when empty and clears on typing. | |
| - [ ] Picking a suggestion fills the box with the CLEAN value (no count suffix). | |
| - [ ] Defaults are Interventional + Recruiting; empty-result hint appears when a location has | |
| only trials outside the active filters. | |
| - [ ] Eligibility section shows for recruiting trials, hidden for completed. | |
| - [ ] Headline status-bar counts are present and non-hardcoded (match loaded data). | |
| - [ ] Assertions exit non-zero on failure (CI-usable). | |
| ## Phase 4 β Design front half: describe β HTML mock β iterate β reference PNG | |
| - [ ] Use the `design` skill (multi-artboard canvas) or `frontend-design` to turn a text | |
| description into an HTML mock of the target tab/component. | |
| - [ ] Iterate on the HTML mock (edit + re-render) until approved. | |
| - [ ] Export the approved mock to a **reference PNG** stored alongside the test | |
| (e.g. `docs/ui-mocks/clinical-trials.png`). | |
| ## Phase 5 β Close the loop: implement in Gradio + visual-QA agent | |
| - [ ] Translate the approved mock into Gradio components + CSS/JS hooks. | |
| - [ ] A `ui-reviewer` subagent takes the Phase-2 screenshot + the Phase-4 reference PNG + a | |
| requirements checklist and reports pass/fail with reasons (fidelity, spacing, alignment). | |
| - [ ] Loop: implement β screenshot β compare to reference β adjust, until it matches. | |
| - [ ] (Optional) pixel-diff the screenshot vs reference for regression alarms. | |
| ## Phase 6 β Package as a project `run-ui` skill | |
| - [ ] Commit as `.claude/skills/run-ui/` (SKILL.md + driver + smoke-mode notes) so any agent | |
| invokes it the same way; document how to add an assertion and where artifacts land. | |
| - [ ] Wire into the PR/deploy flow: no UI change ships without this passing. | |
| ## Notes / open questions | |
| - Deterministic DOM assertions (Phase 3) are the trustworthy regression gate; the visual-QA | |
| agent (Phase 5) is for subjective design judgment. Use both; don't gate on the fuzzy one. | |
| - Playwright install feasibility (Phase 0) is the single biggest unknown β everything | |
| automated depends on it. | |
| ## Definition of done | |
| - [ ] An agent can run one command to launch + assert + screenshot the Clinical Trials tab. | |
| - [ ] The regressions in Phase 3 are covered by assertions that fail loudly. | |
| - [ ] A described UI change can be mocked, approved as a reference PNG, implemented, and | |
| auto-checked against that reference. | |