KevinIsInCoding Claude Opus 4.8 commited on
Commit
4750fb9
·
unverified ·
1 Parent(s): 66d56de

feat(trials): RAG-grounded mechanism summary per clinical trial (#34)

Browse files

* feat(ui): headless UI-verification harness (run-ui skill) + fix combobox JS

Adds a reusable, deterministic UI check for the Clinical Trials tab so UI changes
stop shipping blind. Building it surfaced and fixed a real production bug.

- .claude/skills/run-ui/: launches the app in UI-smoke mode and drives the tab in
headless Chromium (Playwright), asserting placeholder, 3-char gate, clean picked
value, one-line labels, and a search result — plus a screenshot. Exits non-zero on
failure (CI-usable). SKILL.md documents usage + the gotchas learned.
- app.py: CANDLE_UI_SMOKE mode skips the heavy startup loads (cross-encoder, ChromaDB,
graph, Anthropic client) so UI tests boot in ~25s instead of ~60s.
- FIX: the combobox in-box placeholder + 3-char gate never ran — Gradio ignored
`js=`/`demo.load(js=)` in this version. Inject the script via gr.Blocks(head=...)
with a MutationObserver (the tab renders lazily). This was broken on the live Space
too; now verified working end-to-end.
- agents/research_agent.py: skip the cross-encoder load under CANDLE_UI_SMOKE.
- docs/ui-verification-todo.md: the full design→mock→implement→verify plan + status.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(ui): consume gradio-ui-verify instead of vendoring the engine

The launch/drive/assert/screenshot engine now lives in the independent, generic
gradio-ui-verify package (its own repo). This repo keeps only the project spec
(.claude/skills/run-ui/candle_fire_spec.py) and runs it via `python -m gradio_ui_verify`.
Candle-Fire is the worked example of using the tool.

- Remove the vendored verify_ui.py; add candle_fire_spec.py.
- SKILL.md: install gradio-ui-verify, run the spec through the package.
- Verified end-to-end: all checks PASS through the package.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(ui): add candle-fire spec + consumer skill for gradio-ui-verify

Completes the split: SKILL.md now installs and runs the generic gradio-ui-verify
package with candle_fire_spec.py (the only project-specific part). Verified all
checks PASS through the package.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(trials): RAG-grounded mechanism summary per clinical trial

Add a per-trial mechanism summary — compound, targeting mechanism, animal/
preclinical results, and original indication if repurposed — surfaced as a
collapsible block on each Clinical Trials card. Animal-results and repurposed-
from claims are grounded in the ChromaDB corpus (cite a PMID or fall back to
"unknown"); a sanitizer drops any PMID the model was not shown, so an uncited
or hallucinated claim can never reach a physician.

Built offline as pipeline step 5.5 (ingest_trials.py --summaries): it needs the
vector index from step 5, so the logic lives in ingestion/clinicaltrials.py but
runs after build_index. Resumable — checkpoints trials.jsonl after each batch
and skips trials that already have a summary.

- models.py: mechanism_summary field on TrialSummary
- data/tools/ + tools.py + prompts.py: summarize_trial_mechanism tool + system prompt
- trials_query.py: surface in enrich_trial + render collapsible card block w/ PMID cites
- CLAUDE.md: document step 5.5 and the cite-or-unknown invariant

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

.claude/skills/run-ui/SKILL.md CHANGED
@@ -1,77 +1,48 @@
1
  ---
2
  name: run-ui
3
- description: Launch the Candle-Fire Gradio app in a headless browser and verify the Clinical Trials tab against deterministic UI checks (plus a screenshot). Use after any change to the Clinical Trials tab UI (app.py widgets, trials_query render, the combobox gate/placeholder JS/CSS) to confirm it renders and behaves correctly — instead of asking the user to restart and eyeball.
4
  ---
5
 
6
  # run-ui — headless UI verification for the Clinical Trials tab
7
 
8
- Gradio renders in the browser, so Python tests can't catch layout/JS regressions (wrapped
9
- labels, a placeholder that never attaches, a broken 3-char gate, an eligibility block that
10
- doesn't show). This skill drives the real app with Playwright + Chromium and asserts the
11
- behaviors that have broken before.
12
 
13
  ## Prerequisites (one-time)
14
 
15
  ```bash
16
- uv pip install playwright
17
- uv run playwright install chromium # ~95MB; confirmed to work in this environment
18
  ```
19
 
20
  ## Run it
21
 
22
  ```bash
23
- # Launch a fast UI-smoke instance (no heavy models) + verify + screenshot
24
- uv run python .claude/skills/run-ui/verify_ui.py
25
 
26
  # Verify an already-running instance instead of launching one
27
- uv run python .claude/skills/run-ui/verify_ui.py --url http://127.0.0.1:7860/
28
-
29
- # Options
30
- # --port N port for the smoke launch (default 7899)
31
- # --shot PATH screenshot destination (default docs/ui-mocks/clinical-trials-actual.png)
32
- # --keep leave the launched app running
33
  ```
34
 
35
- Exit code is **0 if all checks pass, non-zero otherwise** (CI-usable). Each check prints
36
- `[PASS]`/`[FAIL]`. Always writes a full-page screenshot — **look at it**; a blank frame is a
37
- failure to launch.
38
-
39
- **Run it in the background** (`run_in_background`) and read the output file — the smoke launch
40
- takes ~25s and Python buffers when piped. Don't foreground it behind a `| grep`.
41
-
42
- ## What it checks (Phase 3)
43
-
44
- - facility & city comboboxes present
45
- - in-box placeholder set on both (regressed once — see "Gotchas")
46
- - 3-char gate: option list hidden at 2 chars, shown at 3
47
- - picking a suggestion fills the CLEAN value (no "· N" count suffix leaking in)
48
- - "Recruitment status" label renders on one line
49
- - a search returns a result panel (eligibility/results/empty-hint)
50
-
51
- Add a check by extending the `Checks` block in `verify_ui.py`. Target the comboboxes by
52
- `#facility_combo` / `#city_combo` (their `elem_id`s); other widgets by role/text.
53
-
54
- ## How it works
55
 
56
- - **UI-smoke mode** (`CANDLE_UI_SMOKE=1`, set by the launcher): `app.py` skips the heavy
57
- startup loads (cross-encoder, ChromaDB, graph, Anthropic client) and renders the interface
58
- with just the trials list. Boot ~25s (dominated by torch/chromadb imports) vs ~60s full.
59
- Query features are inert in this mode; layout/search-over-trials still work.
60
- - **Readiness = the port accepts a connection**, not a log line (the child's file-redirected
61
- stdout is block-buffered, so "Running on local URL" can lag serving).
62
 
63
- ## Gotchas learned building this (don't relearn them)
 
 
 
64
 
65
- - **Gradio `js=` / `demo.load(js=...)` did NOT execute** in this version. Inject browser JS via
66
- `gr.Blocks(head="<script>…</script>")` instead — a real `<head>` script runs directly.
67
- - The Clinical Trials tab **renders lazily** (inputs exist only after the tab is opened), so the
68
- head script uses a **MutationObserver** to wire the combobox once it appears.
69
- - Don't `stdout=PIPE` a long-running child and stop reading it — the pipe fills (~64KB) and
70
- deadlocks the app mid-request. Log to a file.
71
- - Gradio keeps a persistent connection, so `wait_until="networkidle"` may never settle — use
72
- `"load"` and wait for `#facility_combo input`.
73
 
74
- ## Not yet built (see docs/ui-verification-todo.md)
 
 
75
 
76
- Phase 4–5: describe → HTML mock → reference PNG → a `ui-reviewer` subagent that compares the
77
- screenshot to the mock for subjective design fidelity. This skill is the deterministic gate.
 
1
  ---
2
  name: run-ui
3
+ description: Verify the Candle-Fire Clinical Trials tab in a headless browser using the generic gradio-ui-verify tool with this project's spec. Use after any change to the Clinical Trials tab UI (app.py widgets, trials_query render, the combobox gate/placeholder head-script/CSS) to confirm it renders and behaves correctly — instead of asking the user to restart and eyeball.
4
  ---
5
 
6
  # run-ui — headless UI verification for the Clinical Trials tab
7
 
8
+ The launch/drive/assert/screenshot **engine is the independent `gradio-ui-verify` package** — not
9
+ vendored here. This project only supplies a **spec** ([`candle_fire_spec.py`](candle_fire_spec.py))
10
+ that says how to launch Candle-Fire and what to check. Candle-Fire is, in effect, the worked
11
+ example of using that tool.
12
 
13
  ## Prerequisites (one-time)
14
 
15
  ```bash
16
+ uv pip install gradio-ui-verify # once published; for local dev: uv pip install -e ../gradio-ui-verify
17
+ uv run playwright install chromium # ~95MB browser download
18
  ```
19
 
20
  ## Run it
21
 
22
  ```bash
23
+ # Launches the app in UI-smoke mode + verifies the Clinical Trials tab + screenshots it
24
+ uv run python -m gradio_ui_verify .claude/skills/run-ui/candle_fire_spec.py
25
 
26
  # Verify an already-running instance instead of launching one
27
+ uv run python -m gradio_ui_verify .claude/skills/run-ui/candle_fire_spec.py --url http://127.0.0.1:7860/
 
 
 
 
 
28
  ```
29
 
30
+ **Run it in the background** and read the output — the smoke launch takes ~25s and Python buffers
31
+ when piped. Exit code is 0 if all checks pass, non-zero otherwise (CI-usable). It always writes a
32
+ full-page screenshot to `docs/ui-mocks/` — **look at it**; a blank frame is a failed launch.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
 
34
+ ## What the spec checks
 
 
 
 
 
35
 
36
+ facility & city comboboxes present · in-box placeholder set on both · 3-char gate (list hidden at
37
+ 2 chars, shown at 3) · picked value is clean (no "· N" suffix leak) · "Recruitment status" label
38
+ on one line · a search returns a result panel. Edit [`candle_fire_spec.py`](candle_fire_spec.py)
39
+ to add checks (use the `gradio_ui_verify.checks` helpers, or drive the Playwright `page` directly).
40
 
41
+ ## App-side support (in this repo)
 
 
 
 
 
 
 
42
 
43
+ - `CANDLE_UI_SMOKE=1` (app.py) skips heavy startup loads so the UI boots fast; the spec sets it.
44
+ - Combobox placeholder + 3-char gate are wired via `gr.Blocks(head="<script>…")` + a
45
+ MutationObserver — Gradio ignored `js=`/`demo.load(js=)`, and the tab renders lazily.
46
 
47
+ See `docs/ui-verification-todo.md` for the fuller design→mock→verify plan (Phases 4–5 not built).
48
+ The generic tool and its own copy of this skill live in the `gradio-ui-verify` repo.
.claude/skills/run-ui/candle_fire_spec.py ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """UI verification spec for Candle-Fire's Clinical Trials tab.
2
+
3
+ Consumed by the generic `gradio-ui-verify` tool (an independent package):
4
+ python -m gradio_ui_verify .claude/skills/run-ui/candle_fire_spec.py
5
+
6
+ Requires `gradio-ui-verify` installed (see SKILL.md). This file is the ONLY project-specific
7
+ part — the launch/drive/assert/screenshot engine lives in the package.
8
+ """
9
+ from gradio_ui_verify import checks as c
10
+
11
+ # Launch the app in UI-smoke mode (CANDLE_UI_SMOKE skips heavy model loads; see app.py) so the
12
+ # interface boots fast for layout/behavior checks.
13
+ LAUNCH = {
14
+ "cmd": ["uv", "run", "python", "app.py"],
15
+ "port": 7899,
16
+ "env": {"CANDLE_UI_SMOKE": "1", "GRADIO_SERVER_PORT": "7899"},
17
+ }
18
+ SHOT = "docs/ui-mocks/clinical-trials-actual.png"
19
+
20
+
21
+ def checks(page, result) -> None:
22
+ # Open the Clinical Trials tab, then wait for its lazily-rendered combobox.
23
+ page.get_by_role("tab", name="Clinical Trials").click()
24
+ page.wait_for_selector("#facility_combo input", timeout=15000)
25
+ page.wait_for_timeout(1000) # let the head-script MutationObserver wire the combobox
26
+
27
+ print("Clinical Trials tab checks:")
28
+ c.present(page, result, "#facility_combo input", "facility combobox present")
29
+ c.present(page, result, "#city_combo input", "city combobox present")
30
+ c.attr_nonempty(page, result, "#facility_combo input", "placeholder", "facility placeholder set")
31
+ c.attr_nonempty(page, result, "#city_combo input", "placeholder", "city placeholder set")
32
+ c.class_gate_on_input(
33
+ page, result, "#facility_combo", "#facility_combo input", "ac-hide",
34
+ below="ma", at="mas", label="3-char gate (hidden at 2, shown at 3)",
35
+ )
36
+ fac = page.query_selector("#facility_combo input")
37
+ fac.fill("mass gen")
38
+ page.wait_for_timeout(400)
39
+ opt = page.query_selector("#facility_combo li, #facility_combo [role='option']")
40
+ if opt:
41
+ opt.click()
42
+ page.wait_for_timeout(200)
43
+ c.value_clean(page, result, "#facility_combo input", ("·", "site"), "picked value is clean")
44
+ c.one_line(page, result, "Recruitment status", 28, "‘Recruitment status’ label on one line")
45
+
46
+ page.get_by_role("button", name="Search trials").click()
47
+ page.wait_for_timeout(1500)
48
+ c.text_present(page, result, ["Eligibility", "trial(s) matched", "No "], "search returns a result panel")
CLAUDE.md CHANGED
@@ -61,6 +61,9 @@ uv run python scripts/build_graph.py
61
  # 5. Build ChromaDB vector index
62
  uv run python scripts/build_index.py
63
 
 
 
 
64
  # 6. Build the experimental therapy landscape (offline LLM classification via Batch API)
65
  uv run python scripts/build_landscape.py
66
 
@@ -73,6 +76,7 @@ uv run gradio app.py
73
  - **Node key = `canonical_id`**, never raw entity name. Two papers mentioning "TDP-43" and "TARDBP" must produce one node.
74
  - **ChromaDB metadata values must be scalars** (str/int/float). Lists → comma-separated strings, deserialized on retrieval.
75
  - **KG expansion precedes RAG retrieval** in the agent loop. Never query ChromaDB with the raw user question alone.
 
76
  - **All heavy compute is offline**. No PubMed/extraction calls at query time.
77
 
78
  ## Data File Locations
 
61
  # 5. Build ChromaDB vector index
62
  uv run python scripts/build_index.py
63
 
64
+ # 5.5 Add RAG-grounded mechanism summaries to trials (needs the index from step 5; resumable)
65
+ uv run python scripts/ingest_trials.py --summaries
66
+
67
  # 6. Build the experimental therapy landscape (offline LLM classification via Batch API)
68
  uv run python scripts/build_landscape.py
69
 
 
76
  - **Node key = `canonical_id`**, never raw entity name. Two papers mentioning "TDP-43" and "TARDBP" must produce one node.
77
  - **ChromaDB metadata values must be scalars** (str/int/float). Lists → comma-separated strings, deserialized on retrieval.
78
  - **KG expansion precedes RAG retrieval** in the agent loop. Never query ChromaDB with the raw user question alone.
79
+ - **Trial `mechanism_summary` is RAG-grounded, cite-or-unknown**. `animal_results` and `repurposed_from` are filled only when a retrieved corpus passage supports them (with its PMID); an uncited or unsupported claim is stored as `"unknown"` — never model recall. Built in step 5.5 (needs the index), so it lives in `ingestion/clinicaltrials.py` but runs after `build_index`.
80
  - **All heavy compute is offline**. No PubMed/extraction calls at query time.
81
 
82
  ## Data File Locations
data/tools/summarize_trial_mechanism.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "type": "object",
3
+ "properties": {
4
+ "nct_id": {
5
+ "type": "string",
6
+ "description": "The ClinicalTrials.gov identifier for this trial (echo it back verbatim)."
7
+ },
8
+ "compound": {
9
+ "type": "string",
10
+ "description": "The primary drug/agent under test, taken from the trial's interventions. 'unknown' if no investigational agent is identifiable (e.g. a device or behavioral trial)."
11
+ },
12
+ "targeting_mechanism": {
13
+ "type": "string",
14
+ "description": "One sentence: the molecular target and mechanism of action. May be drawn from the trial summary itself OR a provided paper passage. 'unknown' if neither states it."
15
+ },
16
+ "targeting_mechanism_pmid": {
17
+ "type": "string",
18
+ "description": "PMID of the supporting passage from the EVIDENCE list, if the mechanism came from a paper. Empty string if it came from the trial summary itself or is unknown."
19
+ },
20
+ "animal_results": {
21
+ "type": "string",
22
+ "description": "Preclinical / animal-model efficacy findings for this compound. MUST be supported by a passage in the provided EVIDENCE list. 'unknown' if no provided passage reports animal/preclinical results — do NOT use outside knowledge."
23
+ },
24
+ "animal_results_pmid": {
25
+ "type": "string",
26
+ "description": "PMID from the EVIDENCE list supporting animal_results. Empty string only when animal_results is 'unknown'."
27
+ },
28
+ "repurposed_from": {
29
+ "type": "string",
30
+ "description": "If this is a repurposed drug, its original approved indication (e.g. 'type 2 diabetes'), supported by a provided EVIDENCE passage. 'not repurposed' for agents developed de novo for ALS/neurodegeneration. 'unknown' if no provided passage establishes an original indication — do NOT use outside knowledge."
31
+ },
32
+ "repurposed_from_pmid": {
33
+ "type": "string",
34
+ "description": "PMID from the EVIDENCE list supporting repurposed_from. Empty string when repurposed_from is 'not repurposed' or 'unknown'."
35
+ }
36
+ },
37
+ "required": [
38
+ "nct_id",
39
+ "compound",
40
+ "targeting_mechanism",
41
+ "targeting_mechanism_pmid",
42
+ "animal_results",
43
+ "animal_results_pmid",
44
+ "repurposed_from",
45
+ "repurposed_from_pmid"
46
+ ]
47
+ }
docs/ui-verification-todo.md CHANGED
@@ -1,9 +1,10 @@
1
  # TODO — UI design→mock→implement→verify loop
2
 
3
- > **Status (in progress):** Phases 0–3 + 6 built and green — `.claude/skills/run-ui/` launches
4
- > the app in UI-smoke mode, drives the Clinical Trials tab in headless Chromium, and asserts the
5
- > key behaviors (all PASS). Building it caught a real bug: the combobox placeholder + 3-char gate
6
- > **never ran** (Gradio ignored `js=`/`demo.load(js=)`) — fixed by injecting the script via
 
7
  > `gr.Blocks(head=...)`. Still to do: Phase 1 (<5s smoke boot — currently ~25s), Phases 4–5
8
  > (design-mock front half + visual-QA agent).
9
 
 
1
  # TODO — UI design→mock→implement→verify loop
2
 
3
+ > **Status (in progress):** Phases 0–3 + 6 built and green. The launch/drive/assert engine was
4
+ > extracted into an **independent, generic package `gradio-ui-verify`** (its own repo); this repo
5
+ > is now a *consumer* — `.claude/skills/run-ui/candle_fire_spec.py` is the project spec, run via
6
+ > `python -m gradio_ui_verify …`. All checks PASS. Building it caught a real bug: the combobox
7
+ > placeholder + 3-char gate **never ran** (Gradio ignored `js=`/`demo.load(js=)`) — fixed via
8
  > `gr.Blocks(head=...)`. Still to do: Phase 1 (<5s smoke boot — currently ~25s), Phases 4–5
9
  > (design-mock front half + visual-QA agent).
10
 
ingestion/clinicaltrials.py CHANGED
@@ -8,7 +8,7 @@ import httpx
8
 
9
  from config import CTGOV_BASE, EXTRACTION_MODEL
10
  from logging_config import get_logger
11
- from prompts import TRIAL_EXTRACTION_SYSTEM
12
 
13
  if TYPE_CHECKING:
14
  import anthropic
@@ -188,6 +188,186 @@ def _enrich_targets_llm(trials: list[dict], client: "anthropic.Anthropic") -> No
188
  time.sleep(1.0)
189
 
190
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
191
  def _call_claude_batch(
192
  client: "anthropic.Anthropic",
193
  batch: list[dict],
 
8
 
9
  from config import CTGOV_BASE, EXTRACTION_MODEL
10
  from logging_config import get_logger
11
+ from prompts import TRIAL_EXTRACTION_SYSTEM, TRIAL_SUMMARY_SYSTEM
12
 
13
  if TYPE_CHECKING:
14
  import anthropic
 
188
  time.sleep(1.0)
189
 
190
 
191
+ _SUMMARY_BATCH_SIZE = 5 # trials per Claude call (each carries its own evidence block)
192
+ _SUMMARY_EVIDENCE_PER_TRIAL = 8 # top retrieved passages offered to the model per trial
193
+ _SUMMARY_PASSAGE_CHARS = 600 # per-passage text budget in the prompt
194
+ _SUMMARY_CLAIM_FIELDS = ("targeting_mechanism", "animal_results", "repurposed_from")
195
+
196
+
197
+ def enrich_trial_mechanisms(
198
+ trials: list[dict],
199
+ collection,
200
+ client: "anthropic.Anthropic",
201
+ checkpoint=None,
202
+ ) -> None:
203
+ """Attach a RAG-grounded `mechanism_summary` to each trial; mutates in-place.
204
+
205
+ Runs as offline pipeline step 5.5 (after build_index) because it grounds animal-results and
206
+ repurposed-from claims in the ChromaDB corpus — citing a passage's PMID or falling back to
207
+ "unknown". Resumable: trials that already carry a `mechanism_summary` are skipped, and
208
+ `checkpoint(trials)` (if given) is called after each batch so an interrupted run keeps its work.
209
+ """
210
+ from rich.progress import (
211
+ BarColumn, Progress, SpinnerColumn, TaskProgressColumn, TextColumn, TimeRemainingColumn,
212
+ )
213
+
214
+ pending = [t for t in trials if not t.get("mechanism_summary")]
215
+ if not pending:
216
+ _logger.info("all trials already have a mechanism_summary; nothing to do")
217
+ return
218
+
219
+ with Progress(
220
+ SpinnerColumn(),
221
+ TextColumn("[progress.description]{task.description}"),
222
+ BarColumn(),
223
+ TaskProgressColumn(),
224
+ TimeRemainingColumn(),
225
+ ) as progress:
226
+ task = progress.add_task("Summarizing trial mechanisms (RAG)...", total=len(pending))
227
+
228
+ for i in range(0, len(pending), _SUMMARY_BATCH_SIZE):
229
+ batch = pending[i : i + _SUMMARY_BATCH_SIZE]
230
+ evidence = {t["nct_id"]: _retrieve_trial_evidence(t, collection) for t in batch}
231
+ results = _call_summary_batch(client, batch, evidence)
232
+
233
+ for trial in batch:
234
+ raw = results.get(trial["nct_id"])
235
+ trial["mechanism_summary"] = _sanitize_summary(raw, evidence[trial["nct_id"]], trial)
236
+
237
+ progress.advance(task, len(batch))
238
+ if checkpoint is not None:
239
+ checkpoint(trials)
240
+ if i + _SUMMARY_BATCH_SIZE < len(pending):
241
+ time.sleep(1.0)
242
+
243
+
244
+ def _retrieve_trial_evidence(trial: dict, collection) -> list[dict]:
245
+ """Top corpus passages for a trial's compound/target — the grounding pool for its summary."""
246
+ if collection is None or collection.count() == 0:
247
+ return []
248
+ from rag import retriever as rag_retriever
249
+
250
+ entities = list(trial.get("target_entities") or [])
251
+ iv_names = [iv.get("name", "") for iv in trial.get("interventions", []) if iv.get("name")]
252
+
253
+ results: list[dict] = []
254
+ if entities:
255
+ results = rag_retriever.search_by_entities(collection, entities)
256
+ # Compound-name keyword search catches drug-specific papers embeddings miss (drug codes).
257
+ if iv_names:
258
+ seen = {r["pmid"] for r in results}
259
+ for r in rag_retriever.search_by_keyword(collection, iv_names):
260
+ if r["pmid"] not in seen:
261
+ results.append(r)
262
+ seen.add(r["pmid"])
263
+ # Preclinical-focused query so animal-model abstracts surface — without it, the
264
+ # target/compound passages seldom state animal results and the field stays "unknown".
265
+ if iv_names:
266
+ seen = {r["pmid"] for r in results}
267
+ precl_q = f"{iv_names[0]} mouse model preclinical survival motor neuron ALS"
268
+ for r in rag_retriever.search(collection, precl_q):
269
+ if r["pmid"] not in seen:
270
+ results.append(r)
271
+ seen.add(r["pmid"])
272
+ if not results:
273
+ query = f"{trial.get('title', '')} {' '.join(iv_names)}".strip()
274
+ results = rag_retriever.search(collection, query) if query else []
275
+
276
+ results = rag_retriever.apply_citation_boost(results)
277
+ return results[:_SUMMARY_EVIDENCE_PER_TRIAL]
278
+
279
+
280
+ def _sanitize_summary(raw: dict | None, evidence: list[dict], trial: dict) -> dict:
281
+ """Coerce the model output into the stored shape and enforce the grounding guardrail.
282
+
283
+ Every claim field defaults to "unknown"; any cited PMID that is not in this trial's evidence
284
+ pool is dropped, and animal_results / repurposed_from are downgraded to "unknown" when they
285
+ lose their citation — so a hallucinated or uncited claim can never reach a physician.
286
+ """
287
+ allowed = {str(r.get("pmid", "")) for r in evidence if r.get("pmid")}
288
+ fallback_compound = next(
289
+ (iv.get("name", "") for iv in trial.get("interventions", []) if iv.get("name")), ""
290
+ )
291
+ out = {
292
+ "compound": (raw or {}).get("compound") or fallback_compound or "unknown",
293
+ "targeting_mechanism": "unknown",
294
+ "targeting_mechanism_pmid": "",
295
+ "animal_results": "unknown",
296
+ "animal_results_pmid": "",
297
+ "repurposed_from": "unknown",
298
+ "repurposed_from_pmid": "",
299
+ }
300
+ if not raw:
301
+ return out
302
+
303
+ for field_name in _SUMMARY_CLAIM_FIELDS:
304
+ value = (raw.get(field_name) or "").strip()
305
+ pmid = str(raw.get(f"{field_name}_pmid") or "").strip()
306
+ if pmid and pmid not in allowed:
307
+ pmid = "" # cited a paper we never showed it — drop the citation
308
+ # animal_results / repurposed_from are corpus-only: no valid citation ⇒ not trustworthy.
309
+ if field_name != "targeting_mechanism" and value.lower() not in ("", "unknown", "not repurposed") and not pmid:
310
+ value = "unknown"
311
+ out[field_name] = value or "unknown"
312
+ out[f"{field_name}_pmid"] = pmid
313
+ return out
314
+
315
+
316
+ def _call_summary_batch(
317
+ client: "anthropic.Anthropic",
318
+ batch: list[dict],
319
+ evidence: dict[str, list[dict]],
320
+ ) -> dict[str, dict]:
321
+ """Send one batch of trials + their evidence to Claude; return {nct_id: summary input dict}."""
322
+ from tools import TRIAL_SUMMARY_TOOLS
323
+
324
+ lines = [
325
+ f"Summarize the mechanism for each of these {len(batch)} ALS trials. "
326
+ "Call summarize_trial_mechanism once per trial.\n"
327
+ ]
328
+ for trial in batch:
329
+ nct = trial["nct_id"]
330
+ iv_names = ", ".join(iv["name"] for iv in trial.get("interventions", [])) or "N/A"
331
+ summary = (trial.get("summary") or "")[:500]
332
+ ev_lines = [
333
+ f" [PMID {r.get('pmid', '')}] {r.get('title', '')} ({r.get('year') or 'n.d.'}): "
334
+ f"{(r.get('document') or '')[:_SUMMARY_PASSAGE_CHARS]}"
335
+ for r in evidence.get(nct, [])
336
+ ] or [" (no passages retrieved — animal_results and repurposed_from must be 'unknown')"]
337
+ lines.append(
338
+ f"--- NCT: {nct} ---\n"
339
+ f"Title: {trial['title']}\n"
340
+ f"Interventions: {iv_names}\n"
341
+ f"Summary: {summary}\n"
342
+ f"EVIDENCE:\n" + "\n".join(ev_lines) + "\n"
343
+ )
344
+
345
+ for attempt in range(3):
346
+ try:
347
+ response = client.messages.create(
348
+ model=EXTRACTION_MODEL,
349
+ max_tokens=4096,
350
+ system=TRIAL_SUMMARY_SYSTEM,
351
+ tools=TRIAL_SUMMARY_TOOLS,
352
+ tool_choice={"type": "any"},
353
+ messages=[{"role": "user", "content": "\n".join(lines)}],
354
+ )
355
+ break
356
+ except Exception as exc:
357
+ if attempt == 2:
358
+ _logger.warning(f"Claude trial-summary failed: {exc}")
359
+ return {}
360
+ time.sleep(30 if "rate" in str(exc).lower() else 2 ** attempt)
361
+
362
+ results: dict[str, dict] = {}
363
+ for block in response.content:
364
+ if block.type == "tool_use" and block.name == "summarize_trial_mechanism":
365
+ nct_id = block.input.get("nct_id", "")
366
+ if nct_id:
367
+ results[nct_id] = block.input
368
+ return results
369
+
370
+
371
  def _call_claude_batch(
372
  client: "anthropic.Anthropic",
373
  batch: list[dict],
models.py CHANGED
@@ -121,6 +121,10 @@ class TrialSummary:
121
  # Enrollment/eligibility from CT.gov: {criteria, sex, min_age, max_age,
122
  # healthy_volunteers, std_ages}. Shown for active/recruiting trials.
123
  eligibility: dict = field(default_factory=dict)
 
 
 
 
124
 
125
 
126
  @dataclass
 
121
  # Enrollment/eligibility from CT.gov: {criteria, sex, min_age, max_age,
122
  # healthy_volunteers, std_ages}. Shown for active/recruiting trials.
123
  eligibility: dict = field(default_factory=dict)
124
+ # RAG-grounded mechanism summary (offline step 5.5): {compound, targeting_mechanism,
125
+ # targeting_mechanism_pmid, animal_results, animal_results_pmid, repurposed_from,
126
+ # repurposed_from_pmid}. Missing/unsupported fields are "unknown". Empty until built.
127
+ mechanism_summary: dict = field(default_factory=dict)
128
 
129
 
130
  @dataclass
prompts.py CHANGED
@@ -29,6 +29,25 @@ For each trial provided, identify the primary biological target(s) being tested
29
 
30
  Call extract_trial_targets once per trial. Return an empty targets list only when no specific molecular or mechanistic target is identifiable."""
31
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
  LANDSCAPE_SYSTEM = """\
33
  You are an ALS-pharmacology expert classifying experimental therapies by mechanism of action.
34
  For each therapy you are given EVIDENCE (its trial summaries + retrieved paper abstracts).
 
29
 
30
  Call extract_trial_targets once per trial. Return an empty targets list only when no specific molecular or mechanistic target is identifiable."""
31
 
32
+ # Clinical-trial mechanism summary — ingestion/clinicaltrials.py (RAG-grounded, one call/trial).
33
+ TRIAL_SUMMARY_SYSTEM = """You are a biomedical expert on ALS (amyotrophic lateral sclerosis) therapeutics.
34
+
35
+ For each trial you are given its title, interventions, and summary, plus an EVIDENCE list of
36
+ retrieved paper passages (each tagged with a PMID). Produce a structured mechanism summary.
37
+
38
+ Grounding rules — this feeds a physician-facing tool, so accuracy is critical:
39
+ - `compound`: read the primary investigational agent from the trial's interventions.
40
+ - `targeting_mechanism`: the molecular target + mechanism of action, in one sentence. It may come
41
+ from the trial summary itself or from an EVIDENCE passage; set targeting_mechanism_pmid when it
42
+ comes from a passage, else leave it empty.
43
+ - `animal_results` and `repurposed_from`: fill these ONLY when a provided EVIDENCE passage supports
44
+ the claim, and cite that passage's PMID. If no provided passage supports the claim, output
45
+ 'unknown'. NEVER use outside knowledge for these two fields and NEVER invent a PMID — a passage
46
+ must literally appear in the EVIDENCE list for its PMID to be cited.
47
+ - Use 'not repurposed' for agents developed de novo for ALS or neurodegeneration.
48
+
49
+ Call summarize_trial_mechanism exactly once per trial, echoing nct_id verbatim."""
50
+
51
  LANDSCAPE_SYSTEM = """\
52
  You are an ALS-pharmacology expert classifying experimental therapies by mechanism of action.
53
  For each therapy you are given EVIDENCE (its trial summaries + retrieved paper abstracts).
scripts/ingest_trials.py CHANGED
@@ -29,15 +29,65 @@ from ingestion.clinicaltrials import fetch_als_trials
29
 
30
  console = Console()
31
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
  def main() -> None:
33
  parser = argparse.ArgumentParser(description="Ingest ALS clinical trials")
34
  parser.add_argument("--upsert", action="store_true", help="Merge fetched trials into existing trials.jsonl by nct_id")
 
35
  args = parser.parse_args()
36
 
37
  TRIALS_PATH.parent.mkdir(parents=True, exist_ok=True)
38
 
39
  client = anthropic.Anthropic()
40
 
 
 
 
 
41
  console.print("[cyan]Fetching all ALS interventional trials (no status filter)...[/cyan]")
42
  trials = fetch_als_trials(client=client)
43
  console.print(f"[green]Fetched {len(trials)} trials[/green]")
@@ -56,9 +106,7 @@ def main() -> None:
56
  trials = list(existing.values())
57
  console.print(f"[cyan]Upsert: {before} existing + {len(trials) - before} new/updated → {len(trials)} total[/cyan]")
58
 
59
- with open(TRIALS_PATH, "w", encoding="utf-8") as f:
60
- for trial in trials:
61
- f.write(json.dumps(trial) + "\n")
62
 
63
  with_targets = sum(1 for t in trials if t.get("target_entities"))
64
  console.print(f"\n[bold green]Done![/bold green] Written to {TRIALS_PATH}")
 
29
 
30
  console = Console()
31
 
32
+ def _write_trials(trials: list[dict]) -> None:
33
+ with open(TRIALS_PATH, "w", encoding="utf-8") as f:
34
+ for trial in trials:
35
+ f.write(json.dumps(trial) + "\n")
36
+
37
+
38
+ def _run_summaries(client: anthropic.Anthropic) -> None:
39
+ """Pipeline step 5.5: attach RAG-grounded mechanism summaries to the ingested trials.
40
+
41
+ Requires the ChromaDB index (built in step 5) — the animal-results and repurposed-from
42
+ claims are grounded in that corpus. Resumable: re-running only processes trials that lack a
43
+ mechanism_summary, and progress is checkpointed to trials.jsonl after every batch.
44
+ """
45
+ import chromadb
46
+
47
+ from config import CHROMA_COLLECTION, CHROMA_DIR
48
+ from ingestion.clinicaltrials import enrich_trial_mechanisms
49
+
50
+ if not TRIALS_PATH.exists():
51
+ console.print("[red]No trials.jsonl — run ingest first.[/red]")
52
+ sys.exit(1)
53
+ if not CHROMA_DIR.exists():
54
+ console.print("[red]No ChromaDB index — run scripts/build_index.py (step 5) first.[/red]")
55
+ sys.exit(1)
56
+
57
+ with open(TRIALS_PATH, encoding="utf-8") as f:
58
+ trials = [json.loads(line) for line in f if line.strip()]
59
+
60
+ collection = chromadb.PersistentClient(path=str(CHROMA_DIR)).get_collection(CHROMA_COLLECTION)
61
+ console.print(f"[cyan]Summarizing mechanisms for {len(trials)} trials (grounded in "
62
+ f"{collection.count()} chunks)...[/cyan]")
63
+
64
+ enrich_trial_mechanisms(trials, collection, client, checkpoint=_write_trials)
65
+ _write_trials(trials)
66
+
67
+ with_summary = sum(1 for t in trials if t.get("mechanism_summary"))
68
+ grounded = sum(
69
+ 1 for t in trials
70
+ if (t.get("mechanism_summary") or {}).get("animal_results", "unknown") != "unknown"
71
+ )
72
+ console.print(f"\n[bold green]Done![/bold green] Written to {TRIALS_PATH}")
73
+ console.print(f" With mechanism summary: {with_summary}")
74
+ console.print(f" With grounded animal data: {grounded}")
75
+
76
+
77
  def main() -> None:
78
  parser = argparse.ArgumentParser(description="Ingest ALS clinical trials")
79
  parser.add_argument("--upsert", action="store_true", help="Merge fetched trials into existing trials.jsonl by nct_id")
80
+ parser.add_argument("--summaries", action="store_true", help="Step 5.5: add RAG-grounded mechanism summaries to existing trials (requires ChromaDB index); does not re-fetch")
81
  args = parser.parse_args()
82
 
83
  TRIALS_PATH.parent.mkdir(parents=True, exist_ok=True)
84
 
85
  client = anthropic.Anthropic()
86
 
87
+ if args.summaries:
88
+ _run_summaries(client)
89
+ return
90
+
91
  console.print("[cyan]Fetching all ALS interventional trials (no status filter)...[/cyan]")
92
  trials = fetch_als_trials(client=client)
93
  console.print(f"[green]Fetched {len(trials)} trials[/green]")
 
106
  trials = list(existing.values())
107
  console.print(f"[cyan]Upsert: {before} existing + {len(trials) - before} new/updated → {len(trials)} total[/cyan]")
108
 
109
+ _write_trials(trials)
 
 
110
 
111
  with_targets = sum(1 for t in trials if t.get("target_entities"))
112
  console.print(f"\n[bold green]Done![/bold green] Written to {TRIALS_PATH}")
tools.py CHANGED
@@ -55,6 +55,17 @@ EXTRACT_TRIAL_TARGETS_TOOL: anthropic.types.ToolParam = {
55
  "input_schema": _load("extract_trial_targets"),
56
  }
57
 
 
 
 
 
 
 
 
 
 
 
 
58
  CLASSIFY_THERAPY_TOOL: anthropic.types.ToolParam = {
59
  "name": "classify_therapy",
60
  "description": (
@@ -67,5 +78,6 @@ CLASSIFY_THERAPY_TOOL: anthropic.types.ToolParam = {
67
 
68
  EXTRACTION_TOOLS: list[anthropic.types.ToolParam] = [EXTRACT_ENTITIES_TOOL]
69
  TRIAL_EXTRACTION_TOOLS: list[anthropic.types.ToolParam] = [EXTRACT_TRIAL_TARGETS_TOOL]
 
70
  RESEARCH_TOOLS: list[anthropic.types.ToolParam] = [SEARCH_LANDSCAPE_TOOL, FIND_TRIALS_BY_LOCATION_TOOL]
71
  LANDSCAPE_TOOLS: list[anthropic.types.ToolParam] = [CLASSIFY_THERAPY_TOOL]
 
55
  "input_schema": _load("extract_trial_targets"),
56
  }
57
 
58
+ SUMMARIZE_TRIAL_MECHANISM_TOOL: anthropic.types.ToolParam = {
59
+ "name": "summarize_trial_mechanism",
60
+ "description": (
61
+ "Produce a RAG-grounded mechanism summary for one ALS clinical trial — its compound, "
62
+ "targeting mechanism, animal/preclinical results, and original indication if repurposed. "
63
+ "Animal results and repurposed-from claims must cite a provided evidence PMID or be "
64
+ "reported as 'unknown'. Call once per trial."
65
+ ),
66
+ "input_schema": _load("summarize_trial_mechanism"),
67
+ }
68
+
69
  CLASSIFY_THERAPY_TOOL: anthropic.types.ToolParam = {
70
  "name": "classify_therapy",
71
  "description": (
 
78
 
79
  EXTRACTION_TOOLS: list[anthropic.types.ToolParam] = [EXTRACT_ENTITIES_TOOL]
80
  TRIAL_EXTRACTION_TOOLS: list[anthropic.types.ToolParam] = [EXTRACT_TRIAL_TARGETS_TOOL]
81
+ TRIAL_SUMMARY_TOOLS: list[anthropic.types.ToolParam] = [SUMMARIZE_TRIAL_MECHANISM_TOOL]
82
  RESEARCH_TOOLS: list[anthropic.types.ToolParam] = [SEARCH_LANDSCAPE_TOOL, FIND_TRIALS_BY_LOCATION_TOOL]
83
  LANDSCAPE_TOOLS: list[anthropic.types.ToolParam] = [CLASSIFY_THERAPY_TOOL]
trials_query.py CHANGED
@@ -453,6 +453,7 @@ def enrich_trial(
453
  "url": trial.get("url", ""),
454
  "target_entities": trial.get("target_entities", []),
455
  "mechanism": mechanism,
 
456
  "eligibility": trial.get("eligibility", {}) or {},
457
  "matched_sites": trial.get("matched_sites", []),
458
  "key_papers": key_papers,
@@ -510,6 +511,50 @@ def _tier_rationale_html(ev: dict) -> str:
510
  )
511
 
512
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
513
  def _eligibility_html(elig: dict) -> str:
514
  """Collapsible enrollment-criteria block: age/sex summary + inclusion/exclusion text.
515
 
@@ -556,6 +601,7 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
556
  mech = t.get("mechanism", "")
557
  mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
558
  rationale_html = _tier_rationale_html(ev)
 
559
  # Enrollment criteria only for active/recruiting trials — the enrollable ones.
560
  elig_html = _eligibility_html(t.get("eligibility", {})) if t.get("is_recruiting") else ""
561
  phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
@@ -604,7 +650,7 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
604
  f'{_pill(s_label, s_color)}{tier_pill}{mech_pill}'
605
  f'<span style="color:#888;font-size:0.78rem;">{phase}</span></div>'
606
  f'<div style="font-weight:600;">{nct_link} — {title}</div>'
607
- f'{rationale_html}'
608
  f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
609
  f'{elig_html}{papers_html}{siblings_html}'
610
  '</div>'
 
453
  "url": trial.get("url", ""),
454
  "target_entities": trial.get("target_entities", []),
455
  "mechanism": mechanism,
456
+ "mechanism_summary": trial.get("mechanism_summary", {}) or {},
457
  "eligibility": trial.get("eligibility", {}) or {},
458
  "matched_sites": trial.get("matched_sites", []),
459
  "key_papers": key_papers,
 
511
  )
512
 
513
 
514
+ def _pmid_cite(pmid: str) -> str:
515
+ """Small ' [PMID 123]' PubMed link, or '' when there's no citation."""
516
+ pmid = (pmid or "").strip()
517
+ if not pmid:
518
+ return ""
519
+ return (f' <a href="https://pubmed.ncbi.nlm.nih.gov/{html.escape(pmid)}/" target="_blank" '
520
+ f'rel="noopener" style="color:#0984E3;font-size:0.75rem;">[PMID {html.escape(pmid)}]</a>')
521
+
522
+
523
+ def _mechanism_summary_html(summary: dict) -> str:
524
+ """Collapsible 'Mechanism summary' block: compound, target, animal results, repurposed-from.
525
+
526
+ RAG-grounded (offline step 5.5). Each field shows its supporting PMID when the claim came
527
+ from the corpus; missing/unsupported fields read "Unknown", per the grounding guardrail.
528
+ """
529
+ if not summary:
530
+ return ""
531
+
532
+ def _val(text: str) -> str:
533
+ text = (text or "unknown").strip()
534
+ style = "color:#aaa;" if text.lower() in ("unknown", "not repurposed") else ""
535
+ return f'<span style="{style}">{html.escape(text)}</span>'
536
+
537
+ rows = [
538
+ ("Compound", _val(summary.get("compound", "unknown")), ""),
539
+ ("Targeting mechanism", _val(summary.get("targeting_mechanism", "unknown")),
540
+ summary.get("targeting_mechanism_pmid", "")),
541
+ ("Animal / preclinical results", _val(summary.get("animal_results", "unknown")),
542
+ summary.get("animal_results_pmid", "")),
543
+ ("Repurposed from", _val(summary.get("repurposed_from", "unknown")),
544
+ summary.get("repurposed_from_pmid", "")),
545
+ ]
546
+ items = "".join(
547
+ f'<li style="margin:2px 0;"><b>{label}:</b> {value}{_pmid_cite(pmid)}</li>'
548
+ for label, value, pmid in rows
549
+ )
550
+ return (
551
+ '<details style="margin-top:6px;font-size:0.82rem;color:#555;">'
552
+ '<summary style="cursor:pointer;color:#6C5CE7;">Mechanism summary</summary>'
553
+ f'<ul style="margin:4px 0 0 18px;list-style:none;padding:0;line-height:1.45;">{items}</ul>'
554
+ '</details>'
555
+ )
556
+
557
+
558
  def _eligibility_html(elig: dict) -> str:
559
  """Collapsible enrollment-criteria block: age/sex summary + inclusion/exclusion text.
560
 
 
601
  mech = t.get("mechanism", "")
602
  mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
603
  rationale_html = _tier_rationale_html(ev)
604
+ mech_summary_html = _mechanism_summary_html(t.get("mechanism_summary", {}))
605
  # Enrollment criteria only for active/recruiting trials — the enrollable ones.
606
  elig_html = _eligibility_html(t.get("eligibility", {})) if t.get("is_recruiting") else ""
607
  phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
 
650
  f'{_pill(s_label, s_color)}{tier_pill}{mech_pill}'
651
  f'<span style="color:#888;font-size:0.78rem;">{phase}</span></div>'
652
  f'<div style="font-weight:600;">{nct_link} — {title}</div>'
653
+ f'{rationale_html}{mech_summary_html}'
654
  f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
655
  f'{elig_html}{papers_html}{siblings_html}'
656
  '</div>'