Spaces:
Running
Running
KevinIsInCoding Claude Opus 4.8 commited on
feat(trials): filters, polish & eligibility on main (recover #29 + #30) (#32)
Browse files#28 (autocomplete) reached main, but stacked PRs #29 (filters + polish) and #30
(eligibility) merged into their base branches, not up to main. This lands their
content on main so the full Clinical Trials tab work ships: study-type/status
filters + defaults, layout & number-consistency polish, and enrollment/eligibility
(ingestion capture, backfill script, active-trial display).
Data change: run scripts/upload_data.py --only trials before deploying.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
- README.md +5 -5
- app.py +49 -13
- ingestion/clinicaltrials.py +18 -0
- models.py +3 -0
- scripts/backfill_eligibility.py +104 -0
- trials_query.py +68 -15
README.md
CHANGED
|
@@ -13,11 +13,11 @@ pinned: false
|
|
| 13 |
|
| 14 |
**ALS Research Intelligence for Physicians**
|
| 15 |
|
| 16 |
-
Candle-fire is a physician-facing tool that synthesizes evidence from ~
|
| 17 |
|
| 18 |
## What It Does
|
| 19 |
|
| 20 |
-
- **Two-layer retrieval**: Knowledge graph expansion (BioLORD-2023-C embeddings + NetworkX) → RAG over ~
|
| 21 |
- **Citation-weighted ranking**: Highly-cited papers surface first
|
| 22 |
- **Structured synthesis**: Claude Sonnet produces mechanism summaries, entity tables, evidence strength assessments, and trial links
|
| 23 |
- **Biomedical synonyms**: BioLORD understands that "TDP-43" = "TARDBP" = "TAR DNA-binding protein 43"
|
|
@@ -47,7 +47,7 @@ cp .env.example .env
|
|
| 47 |
Build the knowledge assets before launching the app. Each step is resumable.
|
| 48 |
|
| 49 |
```bash
|
| 50 |
-
# 1. Ingest ~
|
| 51 |
uv run python scripts/ingest_papers.py
|
| 52 |
|
| 53 |
# 2. Ingest ALS clinical trials from ClinicalTrials.gov (< 1 min, run in parallel)
|
|
@@ -131,10 +131,10 @@ Physician query
|
|
| 131 |
|
| 132 |
| Source | Content | Volume |
|
| 133 |
|---|---|---|
|
| 134 |
-
| PubMed Entrez | ALS paper abstracts + metadata | ~
|
| 135 |
| PubMed Central | Full text for Open Access papers | ~50% coverage |
|
| 136 |
| Semantic Scholar | Citation counts per paper | All papers |
|
| 137 |
-
| ClinicalTrials.gov v2 |
|
| 138 |
|
| 139 |
## Disclaimer
|
| 140 |
|
|
|
|
| 13 |
|
| 14 |
**ALS Research Intelligence for Physicians**
|
| 15 |
|
| 16 |
+
Candle-fire is a physician-facing tool that synthesizes evidence from ~10,000 curated ALS research papers and a biomedical knowledge graph. Ask a free-text question about ALS biology, drug targets, or clinical trials — get a structured, cited answer in under 30 seconds.
|
| 17 |
|
| 18 |
## What It Does
|
| 19 |
|
| 20 |
+
- **Two-layer retrieval**: Knowledge graph expansion (BioLORD-2023-C embeddings + NetworkX) → RAG over ~10,000 ALS papers
|
| 21 |
- **Citation-weighted ranking**: Highly-cited papers surface first
|
| 22 |
- **Structured synthesis**: Claude Sonnet produces mechanism summaries, entity tables, evidence strength assessments, and trial links
|
| 23 |
- **Biomedical synonyms**: BioLORD understands that "TDP-43" = "TARDBP" = "TAR DNA-binding protein 43"
|
|
|
|
| 47 |
Build the knowledge assets before launching the app. Each step is resumable.
|
| 48 |
|
| 49 |
```bash
|
| 50 |
+
# 1. Ingest ~10,000 ALS papers from PubMed + PMC full text + citation counts (~15 min)
|
| 51 |
uv run python scripts/ingest_papers.py
|
| 52 |
|
| 53 |
# 2. Ingest ALS clinical trials from ClinicalTrials.gov (< 1 min, run in parallel)
|
|
|
|
| 131 |
|
| 132 |
| Source | Content | Volume |
|
| 133 |
|---|---|---|
|
| 134 |
+
| PubMed Entrez | ALS paper abstracts + metadata | ~10,000 papers |
|
| 135 |
| PubMed Central | Full text for Open Access papers | ~50% coverage |
|
| 136 |
| Semantic Scholar | Citation counts per paper | All papers |
|
| 137 |
+
| ClinicalTrials.gov v2 | Interventional + expanded-access ALS trials | ~720 trials |
|
| 138 |
|
| 139 |
## Disclaimer
|
| 140 |
|
app.py
CHANGED
|
@@ -1,6 +1,7 @@
|
|
| 1 |
"""Gradio web UI for candle-fire — physician-facing ALS research intelligence."""
|
| 2 |
from __future__ import annotations
|
| 3 |
|
|
|
|
| 4 |
import json
|
| 5 |
from pathlib import Path
|
| 6 |
|
|
@@ -103,7 +104,13 @@ _trials = _load_trials()
|
|
| 103 |
_client = anthropic.Anthropic()
|
| 104 |
|
| 105 |
_n_chunks = _collection.count() if _collection else 0
|
| 106 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 107 |
_kg_nodes = _graph.number_of_nodes() if _graph else 0
|
| 108 |
|
| 109 |
# Experimental therapy landscape (offline-built artifact; loaded once)
|
|
@@ -171,12 +178,15 @@ footer { display: none !important; }
|
|
| 171 |
/* Autocomplete gate: hide a combobox's attached option list until ≥3 chars (see
|
| 172 |
_AUTOCOMPLETE_GATE_JS). The script toggles .ac-hide on the input's wrapper by length. */
|
| 173 |
#facility_combo.ac-hide ul, #city_combo.ac-hide ul { display: none !important; }
|
|
|
|
|
|
|
|
|
|
| 174 |
"""
|
| 175 |
|
| 176 |
_TITLE_MD = """# 🕯️ Candle-Fire
|
| 177 |
### ALS Research Intelligence for Physicians
|
| 178 |
Ask a free-text question about ALS biology, drug targets, or clinical trials.
|
| 179 |
-
Answers are synthesized from
|
| 180 |
"""
|
| 181 |
|
| 182 |
_DISCLAIMER_MD = """<div class="disclaimer">
|
|
@@ -223,23 +233,24 @@ _LOC_CHOICES = trials_query.location_choices(_LOC_INDEX)
|
|
| 223 |
_AUTOCOMPLETE_GATE_JS = f"""
|
| 224 |
() => {{
|
| 225 |
const MIN = {trials_query.MIN_AUTOCOMPLETE_CHARS};
|
| 226 |
-
const gate = (id) => {{
|
| 227 |
const root = document.getElementById(id);
|
| 228 |
if (!root) return;
|
| 229 |
const input = root.querySelector('input');
|
| 230 |
if (!input) return;
|
|
|
|
| 231 |
const apply = () => root.classList.toggle('ac-hide', input.value.trim().length < MIN);
|
| 232 |
input.addEventListener('input', apply);
|
| 233 |
input.addEventListener('focus', apply);
|
| 234 |
apply();
|
| 235 |
}};
|
| 236 |
-
gate('facility_combo');
|
| 237 |
-
gate('city_combo');
|
| 238 |
}}
|
| 239 |
"""
|
| 240 |
|
| 241 |
|
| 242 |
-
def _search_trials(facility: str, state: str, city: str, status: str) -> str:
|
| 243 |
facility = (facility or "").strip() or None
|
| 244 |
city = (city or "").strip() or None
|
| 245 |
state = None if (not state or state == "All") else state
|
|
@@ -248,8 +259,28 @@ def _search_trials(facility: str, state: str, city: str, status: str) -> str:
|
|
| 248 |
return '<div style="color:#888;padding:12px 0;">Enter a facility, state, or city to search.</div>'
|
| 249 |
|
| 250 |
matches = trials_query.search_trials_by_location(
|
| 251 |
-
_trials, facility=facility, city=city, state=state,
|
|
|
|
| 252 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 253 |
enriched = [
|
| 254 |
trials_query.enrich_trial(t, _collection, _graph, _trials)
|
| 255 |
for t in matches[:_TRIAL_ENRICH_CAP]
|
|
@@ -269,7 +300,7 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
|
|
| 269 |
gr.HTML(
|
| 270 |
f'<div class="status-bar">'
|
| 271 |
f'{_n_chunks} paper chunks · '
|
| 272 |
-
f'{_n_trials}
|
| 273 |
f'{_kg_nodes} knowledge graph nodes'
|
| 274 |
f'</div>'
|
| 275 |
)
|
|
@@ -380,12 +411,11 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
|
|
| 380 |
"**location** (state / city). Each result is enriched with recruiting status, an "
|
| 381 |
"evidence-strength tier, key supporting papers, and related trials for the same compound."
|
| 382 |
)
|
| 383 |
-
with gr.Row():
|
| 384 |
facility_tb = gr.Dropdown(
|
| 385 |
choices=_LOC_CHOICES["facilities"], value=None,
|
| 386 |
label="Facility / institution", scale=2,
|
| 387 |
filterable=True, allow_custom_value=True, elem_id="facility_combo",
|
| 388 |
-
info="Type ≥3 letters and pick a match (e.g. Mass General).",
|
| 389 |
)
|
| 390 |
state_dd = gr.Dropdown(
|
| 391 |
choices=_US_STATES, value="All", label="State", scale=1,
|
|
@@ -394,11 +424,17 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
|
|
| 394 |
choices=_LOC_CHOICES["cities"], value=None,
|
| 395 |
label="City", scale=1,
|
| 396 |
filterable=True, allow_custom_value=True, elem_id="city_combo",
|
| 397 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 398 |
)
|
| 399 |
trial_status_dd = gr.Dropdown(
|
| 400 |
choices=["All", "Recruiting", "Not recruiting"],
|
| 401 |
-
value="
|
| 402 |
)
|
| 403 |
search_btn = gr.Button("Search trials", variant="primary")
|
| 404 |
trial_results = gr.HTML(
|
|
@@ -408,7 +444,7 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
|
|
| 408 |
# Facility/city are typeable comboboxes (filterable Dropdowns) — the physician
|
| 409 |
# types and picks from the attached, busiest-first list. No per-keystroke server
|
| 410 |
# event needed; the search reads the selected/typed value directly.
|
| 411 |
-
_trial_search_inputs = [facility_tb, state_dd, city_tb, trial_status_dd]
|
| 412 |
search_btn.click(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
|
| 413 |
# Picking a facility/city from its list also runs the search immediately.
|
| 414 |
facility_tb.select(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
|
|
|
|
| 1 |
"""Gradio web UI for candle-fire — physician-facing ALS research intelligence."""
|
| 2 |
from __future__ import annotations
|
| 3 |
|
| 4 |
+
import html
|
| 5 |
import json
|
| 6 |
from pathlib import Path
|
| 7 |
|
|
|
|
| 104 |
_client = anthropic.Anthropic()
|
| 105 |
|
| 106 |
_n_chunks = _collection.count() if _collection else 0
|
| 107 |
+
# Headline trial count = recruiting, interventional trials only (the actionable set), not the
|
| 108 |
+
# full corpus (which includes completed/terminated studies and expanded-access programs).
|
| 109 |
+
_RECRUITING = {"RECRUITING", "NOT_YET_RECRUITING", "ENROLLING_BY_INVITATION", "AVAILABLE"}
|
| 110 |
+
_n_trials = sum(
|
| 111 |
+
1 for t in _trials
|
| 112 |
+
if t.get("study_type") == "INTERVENTIONAL" and t.get("status") in _RECRUITING
|
| 113 |
+
)
|
| 114 |
_kg_nodes = _graph.number_of_nodes() if _graph else 0
|
| 115 |
|
| 116 |
# Experimental therapy landscape (offline-built artifact; loaded once)
|
|
|
|
| 178 |
/* Autocomplete gate: hide a combobox's attached option list until ≥3 chars (see
|
| 179 |
_AUTOCOMPLETE_GATE_JS). The script toggles .ac-hide on the input's wrapper by length. */
|
| 180 |
#facility_combo.ac-hide ul, #city_combo.ac-hide ul { display: none !important; }
|
| 181 |
+
/* Smaller filter labels so long ones (e.g. "Recruitment status") stay on one line and the
|
| 182 |
+
dropdown chevron doesn't overlap the text. */
|
| 183 |
+
.trial-filters label span { font-size: 0.78rem !important; white-space: nowrap; }
|
| 184 |
"""
|
| 185 |
|
| 186 |
_TITLE_MD = """# 🕯️ Candle-Fire
|
| 187 |
### ALS Research Intelligence for Physicians
|
| 188 |
Ask a free-text question about ALS biology, drug targets, or clinical trials.
|
| 189 |
+
Answers are synthesized from a curated ALS research corpus and enriched by a biomedical knowledge graph.
|
| 190 |
"""
|
| 191 |
|
| 192 |
_DISCLAIMER_MD = """<div class="disclaimer">
|
|
|
|
| 233 |
_AUTOCOMPLETE_GATE_JS = f"""
|
| 234 |
() => {{
|
| 235 |
const MIN = {trials_query.MIN_AUTOCOMPLETE_CHARS};
|
| 236 |
+
const gate = (id, hint) => {{
|
| 237 |
const root = document.getElementById(id);
|
| 238 |
if (!root) return;
|
| 239 |
const input = root.querySelector('input');
|
| 240 |
if (!input) return;
|
| 241 |
+
if (hint) input.setAttribute('placeholder', hint); // in-box hint; hides once they type
|
| 242 |
const apply = () => root.classList.toggle('ac-hide', input.value.trim().length < MIN);
|
| 243 |
input.addEventListener('input', apply);
|
| 244 |
input.addEventListener('focus', apply);
|
| 245 |
apply();
|
| 246 |
}};
|
| 247 |
+
gate('facility_combo', 'Type ≥3 letters, e.g. Mass General');
|
| 248 |
+
gate('city_combo', 'Type ≥3 letters, busiest cities first');
|
| 249 |
}}
|
| 250 |
"""
|
| 251 |
|
| 252 |
|
| 253 |
+
def _search_trials(facility: str, state: str, city: str, study_type: str, status: str) -> str:
|
| 254 |
facility = (facility or "").strip() or None
|
| 255 |
city = (city or "").strip() or None
|
| 256 |
state = None if (not state or state == "All") else state
|
|
|
|
| 259 |
return '<div style="color:#888;padding:12px 0;">Enter a facility, state, or city to search.</div>'
|
| 260 |
|
| 261 |
matches = trials_query.search_trials_by_location(
|
| 262 |
+
_trials, facility=facility, city=city, state=state,
|
| 263 |
+
status=status, study_type=study_type,
|
| 264 |
)
|
| 265 |
+
|
| 266 |
+
# If the active filters hide everything, say whether broader filters would find trials —
|
| 267 |
+
# e.g. a facility with only completed studies under the default Recruiting + Interventional.
|
| 268 |
+
if not matches and (status != "All" or study_type != "All"):
|
| 269 |
+
broad = trials_query.search_trials_by_location(
|
| 270 |
+
_trials, facility=facility, city=city, state=state, status="All", study_type="All",
|
| 271 |
+
)
|
| 272 |
+
if broad:
|
| 273 |
+
where = ", ".join(p for p in (facility, city, state) if p)
|
| 274 |
+
return (
|
| 275 |
+
'<div style="background:#fff6e5;border:1px solid #ffe0a3;border-radius:8px;'
|
| 276 |
+
'padding:10px 12px;margin:6px 0;color:#7a5b00;font-size:0.9rem;">'
|
| 277 |
+
f'No <b>{html.escape((study_type or "").lower())}</b> trials that are '
|
| 278 |
+
f'<b>{html.escape((status or "").lower())}</b> at {html.escape(where)}. '
|
| 279 |
+
f'{len(broad)} trial(s) exist there under broader filters — set '
|
| 280 |
+
'<b>Study type</b> and <b>Recruitment status</b> to <b>All</b> to see them.'
|
| 281 |
+
'</div>'
|
| 282 |
+
)
|
| 283 |
+
|
| 284 |
enriched = [
|
| 285 |
trials_query.enrich_trial(t, _collection, _graph, _trials)
|
| 286 |
for t in matches[:_TRIAL_ENRICH_CAP]
|
|
|
|
| 300 |
gr.HTML(
|
| 301 |
f'<div class="status-bar">'
|
| 302 |
f'{_n_chunks} paper chunks · '
|
| 303 |
+
f'{_n_trials} recruiting interventional trials · '
|
| 304 |
f'{_kg_nodes} knowledge graph nodes'
|
| 305 |
f'</div>'
|
| 306 |
)
|
|
|
|
| 411 |
"**location** (state / city). Each result is enriched with recruiting status, an "
|
| 412 |
"evidence-strength tier, key supporting papers, and related trials for the same compound."
|
| 413 |
)
|
| 414 |
+
with gr.Row(elem_classes="trial-filters"):
|
| 415 |
facility_tb = gr.Dropdown(
|
| 416 |
choices=_LOC_CHOICES["facilities"], value=None,
|
| 417 |
label="Facility / institution", scale=2,
|
| 418 |
filterable=True, allow_custom_value=True, elem_id="facility_combo",
|
|
|
|
| 419 |
)
|
| 420 |
state_dd = gr.Dropdown(
|
| 421 |
choices=_US_STATES, value="All", label="State", scale=1,
|
|
|
|
| 424 |
choices=_LOC_CHOICES["cities"], value=None,
|
| 425 |
label="City", scale=1,
|
| 426 |
filterable=True, allow_custom_value=True, elem_id="city_combo",
|
| 427 |
+
)
|
| 428 |
+
# Filters on their own row so the labels/values have full width — no wrapping,
|
| 429 |
+
# no value running under the chevron.
|
| 430 |
+
with gr.Row(elem_classes="trial-filters"):
|
| 431 |
+
study_type_dd = gr.Dropdown(
|
| 432 |
+
choices=["Interventional", "Expanded Access", "All"],
|
| 433 |
+
value="Interventional", label="Study type", scale=1,
|
| 434 |
)
|
| 435 |
trial_status_dd = gr.Dropdown(
|
| 436 |
choices=["All", "Recruiting", "Not recruiting"],
|
| 437 |
+
value="Recruiting", label="Recruitment status", scale=1,
|
| 438 |
)
|
| 439 |
search_btn = gr.Button("Search trials", variant="primary")
|
| 440 |
trial_results = gr.HTML(
|
|
|
|
| 444 |
# Facility/city are typeable comboboxes (filterable Dropdowns) — the physician
|
| 445 |
# types and picks from the attached, busiest-first list. No per-keystroke server
|
| 446 |
# event needed; the search reads the selected/typed value directly.
|
| 447 |
+
_trial_search_inputs = [facility_tb, state_dd, city_tb, study_type_dd, trial_status_dd]
|
| 448 |
search_btn.click(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
|
| 449 |
# Picking a facility/city from its list also runs the search immediately.
|
| 450 |
facility_tb.select(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
|
ingestion/clinicaltrials.py
CHANGED
|
@@ -70,6 +70,22 @@ def fetch_als_trials(
|
|
| 70 |
return trials
|
| 71 |
|
| 72 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
def _flatten_trial(study: dict) -> dict:
|
| 74 |
proto = study.get("protocolSection", {})
|
| 75 |
id_mod = proto.get("identificationModule", {})
|
|
@@ -79,6 +95,7 @@ def _flatten_trial(study: dict) -> dict:
|
|
| 79 |
arms_mod = proto.get("armsInterventionsModule", {})
|
| 80 |
status_mod = proto.get("statusModule", {})
|
| 81 |
contacts_mod = proto.get("contactsLocationsModule", {})
|
|
|
|
| 82 |
|
| 83 |
nct_id = id_mod.get("nctId", "")
|
| 84 |
interventions = [
|
|
@@ -126,6 +143,7 @@ def _flatten_trial(study: dict) -> dict:
|
|
| 126 |
"locations": locations,
|
| 127 |
"contact_phone": contact_phone,
|
| 128 |
"contact_email": contact_email,
|
|
|
|
| 129 |
}
|
| 130 |
|
| 131 |
|
|
|
|
| 70 |
return trials
|
| 71 |
|
| 72 |
|
| 73 |
+
def _extract_eligibility(elig_mod: dict) -> dict:
|
| 74 |
+
"""Enrollment/eligibility fields from a CT.gov v2 eligibilityModule.
|
| 75 |
+
|
| 76 |
+
Shared by _flatten_trial (ingestion) and scripts/backfill_eligibility.py so the stored
|
| 77 |
+
shape is identical whichever path populated it.
|
| 78 |
+
"""
|
| 79 |
+
return {
|
| 80 |
+
"criteria": elig_mod.get("eligibilityCriteria", ""), # free text: Inclusion/Exclusion
|
| 81 |
+
"sex": elig_mod.get("sex", ""), # ALL / MALE / FEMALE
|
| 82 |
+
"min_age": elig_mod.get("minimumAge", ""), # e.g. "18 Years"
|
| 83 |
+
"max_age": elig_mod.get("maximumAge", ""),
|
| 84 |
+
"healthy_volunteers": elig_mod.get("healthyVolunteers"),
|
| 85 |
+
"std_ages": elig_mod.get("stdAges", []), # e.g. ["ADULT", "OLDER_ADULT"]
|
| 86 |
+
}
|
| 87 |
+
|
| 88 |
+
|
| 89 |
def _flatten_trial(study: dict) -> dict:
|
| 90 |
proto = study.get("protocolSection", {})
|
| 91 |
id_mod = proto.get("identificationModule", {})
|
|
|
|
| 95 |
arms_mod = proto.get("armsInterventionsModule", {})
|
| 96 |
status_mod = proto.get("statusModule", {})
|
| 97 |
contacts_mod = proto.get("contactsLocationsModule", {})
|
| 98 |
+
elig_mod = proto.get("eligibilityModule", {})
|
| 99 |
|
| 100 |
nct_id = id_mod.get("nctId", "")
|
| 101 |
interventions = [
|
|
|
|
| 143 |
"locations": locations,
|
| 144 |
"contact_phone": contact_phone,
|
| 145 |
"contact_email": contact_email,
|
| 146 |
+
"eligibility": _extract_eligibility(elig_mod),
|
| 147 |
}
|
| 148 |
|
| 149 |
|
models.py
CHANGED
|
@@ -118,6 +118,9 @@ class TrialSummary:
|
|
| 118 |
locations: list[dict] = field(default_factory=list)
|
| 119 |
contact_phone: str = ""
|
| 120 |
contact_email: str = ""
|
|
|
|
|
|
|
|
|
|
| 121 |
|
| 122 |
|
| 123 |
@dataclass
|
|
|
|
| 118 |
locations: list[dict] = field(default_factory=list)
|
| 119 |
contact_phone: str = ""
|
| 120 |
contact_email: str = ""
|
| 121 |
+
# Enrollment/eligibility from CT.gov: {criteria, sex, min_age, max_age,
|
| 122 |
+
# healthy_volunteers, std_ages}. Shown for active/recruiting trials.
|
| 123 |
+
eligibility: dict = field(default_factory=dict)
|
| 124 |
|
| 125 |
|
| 126 |
@dataclass
|
scripts/backfill_eligibility.py
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Backfill enrollment/eligibility criteria into an existing trials.jsonl.
|
| 2 |
+
|
| 3 |
+
Adds the `eligibility` field to trials that don't have it by fetching each trial's
|
| 4 |
+
eligibilityModule from ClinicalTrials.gov v2 — WITHOUT re-running the expensive LLM target
|
| 5 |
+
extraction that a full `ingest_trials.py` would. Fetches in batches via `filter.ids`, merges
|
| 6 |
+
in place, and rewrites trials.jsonl. Idempotent: re-running only fills trials still missing it.
|
| 7 |
+
|
| 8 |
+
Usage:
|
| 9 |
+
uv run python scripts/backfill_eligibility.py # backfill missing eligibility
|
| 10 |
+
uv run python scripts/backfill_eligibility.py --force # refetch for ALL trials
|
| 11 |
+
uv run python scripts/backfill_eligibility.py --dry-run # report only, no write
|
| 12 |
+
|
| 13 |
+
After it writes trials.jsonl, deploy per the runbook:
|
| 14 |
+
uv run python scripts/upload_data.py --only trials # push to the HF dataset repo
|
| 15 |
+
git push hf main
|
| 16 |
+
"""
|
| 17 |
+
from __future__ import annotations
|
| 18 |
+
|
| 19 |
+
import argparse
|
| 20 |
+
import json
|
| 21 |
+
import sys
|
| 22 |
+
import time
|
| 23 |
+
from pathlib import Path
|
| 24 |
+
|
| 25 |
+
import httpx
|
| 26 |
+
|
| 27 |
+
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
| 28 |
+
|
| 29 |
+
from config import CTGOV_BASE, TRIALS_PATH # noqa: E402
|
| 30 |
+
from ingestion.clinicaltrials import _extract_eligibility # noqa: E402
|
| 31 |
+
|
| 32 |
+
_BATCH = 50 # NCT ids per request (filter.ids)
|
| 33 |
+
_PAUSE_S = 0.34 # ~3 req/s, polite to CT.gov
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def _fetch_eligibility(nct_ids: list[str]) -> dict[str, dict]:
|
| 37 |
+
"""{nct_id: eligibility dict} for a batch of NCT ids."""
|
| 38 |
+
resp = httpx.get(
|
| 39 |
+
CTGOV_BASE,
|
| 40 |
+
params={
|
| 41 |
+
"filter.ids": ",".join(nct_ids),
|
| 42 |
+
"fields": "NCTId,EligibilityModule",
|
| 43 |
+
"pageSize": len(nct_ids),
|
| 44 |
+
},
|
| 45 |
+
timeout=30,
|
| 46 |
+
)
|
| 47 |
+
resp.raise_for_status()
|
| 48 |
+
out: dict[str, dict] = {}
|
| 49 |
+
for study in resp.json().get("studies", []):
|
| 50 |
+
proto = study.get("protocolSection", {})
|
| 51 |
+
nct = proto.get("identificationModule", {}).get("nctId", "")
|
| 52 |
+
if nct:
|
| 53 |
+
out[nct] = _extract_eligibility(proto.get("eligibilityModule", {}))
|
| 54 |
+
return out
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def main() -> int:
|
| 58 |
+
ap = argparse.ArgumentParser(description="Backfill eligibility criteria into trials.jsonl.")
|
| 59 |
+
ap.add_argument("--force", action="store_true", help="Refetch for all trials, not just missing.")
|
| 60 |
+
ap.add_argument("--dry-run", action="store_true", help="Report what would change; write nothing.")
|
| 61 |
+
args = ap.parse_args()
|
| 62 |
+
|
| 63 |
+
if not TRIALS_PATH.exists():
|
| 64 |
+
print(f"ERROR — {TRIALS_PATH} not found (run ingest_trials.py first).", file=sys.stderr)
|
| 65 |
+
return 1
|
| 66 |
+
|
| 67 |
+
trials = [json.loads(line) for line in TRIALS_PATH.read_text().splitlines() if line.strip()]
|
| 68 |
+
todo = [
|
| 69 |
+
t for t in trials
|
| 70 |
+
if t.get("nct_id") and (args.force or not (t.get("eligibility") or {}).get("criteria"))
|
| 71 |
+
]
|
| 72 |
+
print(f"{len(trials)} trials; {len(todo)} to fetch eligibility for"
|
| 73 |
+
f"{' (force)' if args.force else ''}.")
|
| 74 |
+
if not todo:
|
| 75 |
+
print("Nothing to do.")
|
| 76 |
+
return 0
|
| 77 |
+
if args.dry_run:
|
| 78 |
+
print("--dry-run: no fetch, no write.")
|
| 79 |
+
return 0
|
| 80 |
+
|
| 81 |
+
by_nct = {t["nct_id"]: t for t in trials}
|
| 82 |
+
fetched = 0
|
| 83 |
+
for i in range(0, len(todo), _BATCH):
|
| 84 |
+
ids = [t["nct_id"] for t in todo[i : i + _BATCH]]
|
| 85 |
+
try:
|
| 86 |
+
for nct, elig in _fetch_eligibility(ids).items():
|
| 87 |
+
if nct in by_nct:
|
| 88 |
+
by_nct[nct]["eligibility"] = elig
|
| 89 |
+
fetched += 1
|
| 90 |
+
except httpx.HTTPError as exc:
|
| 91 |
+
print(f" batch {i // _BATCH} failed: {exc}", file=sys.stderr)
|
| 92 |
+
print(f" {min(i + _BATCH, len(todo))}/{len(todo)}")
|
| 93 |
+
if i + _BATCH < len(todo):
|
| 94 |
+
time.sleep(_PAUSE_S)
|
| 95 |
+
|
| 96 |
+
with_crit = sum(1 for t in trials if (t.get("eligibility") or {}).get("criteria"))
|
| 97 |
+
TRIALS_PATH.write_text("".join(json.dumps(t, ensure_ascii=False) + "\n" for t in trials))
|
| 98 |
+
print(f"Fetched {fetched}; {with_crit}/{len(trials)} trials now have eligibility criteria.")
|
| 99 |
+
print(f"Wrote {TRIALS_PATH}. Next: upload_data.py --only trials, then git push hf main.")
|
| 100 |
+
return 0
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
if __name__ == "__main__":
|
| 104 |
+
raise SystemExit(main())
|
trials_query.py
CHANGED
|
@@ -112,19 +112,29 @@ def search_trials_by_location(
|
|
| 112 |
state: str | None = None,
|
| 113 |
country: str | None = None,
|
| 114 |
status: str | None = None,
|
|
|
|
| 115 |
) -> list[dict]:
|
| 116 |
"""Return trials with at least one site matching the location filters.
|
| 117 |
|
| 118 |
`status`: "Recruiting" keeps only trials with an open overall status; "Not recruiting"
|
| 119 |
-
keeps only closed ones; anything else (None / "All") keeps all.
|
| 120 |
-
|
|
|
|
|
|
|
| 121 |
"""
|
| 122 |
if not any([facility, city, state, country]):
|
| 123 |
return []
|
| 124 |
|
| 125 |
status_filter = (status or "").strip().lower()
|
|
|
|
| 126 |
results: list[dict] = []
|
| 127 |
for trial in trials:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
matched_sites = [
|
| 129 |
s for s in trial.get("locations", [])
|
| 130 |
if _site_matches(s, facility, city, state, country)
|
|
@@ -158,22 +168,31 @@ MIN_AUTOCOMPLETE_CHARS = 3 # client-side gate: the attached list stays hidden u
|
|
| 158 |
|
| 159 |
|
| 160 |
def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]:
|
| 161 |
-
"""Distinct facility + city names with their
|
| 162 |
|
| 163 |
-
|
|
|
|
|
|
|
|
|
|
| 164 |
"""
|
| 165 |
from collections import Counter
|
| 166 |
|
| 167 |
city_counts: Counter = Counter()
|
| 168 |
facility_counts: Counter = Counter()
|
| 169 |
for t in trials:
|
|
|
|
|
|
|
| 170 |
for s in t.get("locations", []):
|
| 171 |
city = (s.get("city") or "").strip()
|
| 172 |
facility = (s.get("facility") or "").strip()
|
| 173 |
if city:
|
| 174 |
-
|
| 175 |
if facility:
|
| 176 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 177 |
|
| 178 |
def _ranked(counter: Counter) -> list[tuple[str, int]]:
|
| 179 |
# busiest first, then alphabetical for stable ties
|
|
@@ -182,16 +201,17 @@ def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]
|
|
| 182 |
return {"cities": _ranked(city_counts), "facilities": _ranked(facility_counts)}
|
| 183 |
|
| 184 |
|
| 185 |
-
def location_choices(index: dict) -> dict[str, list[
|
| 186 |
-
"""Gradio combobox
|
| 187 |
|
| 188 |
-
|
| 189 |
-
|
|
|
|
| 190 |
"""
|
| 191 |
-
|
| 192 |
-
|
| 193 |
-
|
| 194 |
-
|
| 195 |
|
| 196 |
|
| 197 |
def _find_supporting_papers(
|
|
@@ -433,6 +453,7 @@ def enrich_trial(
|
|
| 433 |
"url": trial.get("url", ""),
|
| 434 |
"target_entities": trial.get("target_entities", []),
|
| 435 |
"mechanism": mechanism,
|
|
|
|
| 436 |
"matched_sites": trial.get("matched_sites", []),
|
| 437 |
"key_papers": key_papers,
|
| 438 |
"sibling_trials": sibling_trials,
|
|
@@ -489,6 +510,36 @@ def _tier_rationale_html(ev: dict) -> str:
|
|
| 489 |
)
|
| 490 |
|
| 491 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 492 |
def render_trials_html(enriched: list[dict], match_count: int) -> str:
|
| 493 |
"""Render enriched location-search results as an HTML card list."""
|
| 494 |
if not enriched:
|
|
@@ -505,6 +556,8 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
|
|
| 505 |
mech = t.get("mechanism", "")
|
| 506 |
mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
|
| 507 |
rationale_html = _tier_rationale_html(ev)
|
|
|
|
|
|
|
| 508 |
phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
|
| 509 |
title = html.escape(t.get("title", "")[:140])
|
| 510 |
nct = html.escape(t.get("nct_id", ""))
|
|
@@ -553,7 +606,7 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
|
|
| 553 |
f'<div style="font-weight:600;">{nct_link} — {title}</div>'
|
| 554 |
f'{rationale_html}'
|
| 555 |
f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
|
| 556 |
-
f'{papers_html}{siblings_html}'
|
| 557 |
'</div>'
|
| 558 |
)
|
| 559 |
|
|
|
|
| 112 |
state: str | None = None,
|
| 113 |
country: str | None = None,
|
| 114 |
status: str | None = None,
|
| 115 |
+
study_type: str | None = None,
|
| 116 |
) -> list[dict]:
|
| 117 |
"""Return trials with at least one site matching the location filters.
|
| 118 |
|
| 119 |
`status`: "Recruiting" keeps only trials with an open overall status; "Not recruiting"
|
| 120 |
+
keeps only closed ones; anything else (None / "All") keeps all.
|
| 121 |
+
`study_type`: "Interventional" keeps interventional trials; "Expanded Access" keeps
|
| 122 |
+
expanded-access (investigational-use) programs; anything else (None / "All") keeps both.
|
| 123 |
+
Each returned trial is a shallow copy with `matched_sites` attached; recruiting first.
|
| 124 |
"""
|
| 125 |
if not any([facility, city, state, country]):
|
| 126 |
return []
|
| 127 |
|
| 128 |
status_filter = (status or "").strip().lower()
|
| 129 |
+
type_filter = (study_type or "").strip().lower()
|
| 130 |
results: list[dict] = []
|
| 131 |
for trial in trials:
|
| 132 |
+
is_eap = bool(trial.get("is_expanded_access"))
|
| 133 |
+
if type_filter == "interventional" and is_eap:
|
| 134 |
+
continue
|
| 135 |
+
if type_filter == "expanded access" and not is_eap:
|
| 136 |
+
continue
|
| 137 |
+
|
| 138 |
matched_sites = [
|
| 139 |
s for s in trial.get("locations", [])
|
| 140 |
if _site_matches(s, facility, city, state, country)
|
|
|
|
| 168 |
|
| 169 |
|
| 170 |
def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]:
|
| 171 |
+
"""Distinct facility + city names with their TRIAL counts, each sorted busiest first.
|
| 172 |
|
| 173 |
+
Counts distinct trials, not site rows: a single trial that lists an anonymized placeholder
|
| 174 |
+
like "GSK Investigational Site" 49 times counts once, so placeholders don't balloon to the
|
| 175 |
+
top of the ranking and the shown count matches what a search can return. Built once at
|
| 176 |
+
startup; feeds location_choices() which becomes the combobox `choices`.
|
| 177 |
"""
|
| 178 |
from collections import Counter
|
| 179 |
|
| 180 |
city_counts: Counter = Counter()
|
| 181 |
facility_counts: Counter = Counter()
|
| 182 |
for t in trials:
|
| 183 |
+
cities_here: set[str] = set()
|
| 184 |
+
facilities_here: set[str] = set()
|
| 185 |
for s in t.get("locations", []):
|
| 186 |
city = (s.get("city") or "").strip()
|
| 187 |
facility = (s.get("facility") or "").strip()
|
| 188 |
if city:
|
| 189 |
+
cities_here.add(city)
|
| 190 |
if facility:
|
| 191 |
+
facilities_here.add(facility)
|
| 192 |
+
for c in cities_here:
|
| 193 |
+
city_counts[c] += 1
|
| 194 |
+
for f in facilities_here:
|
| 195 |
+
facility_counts[f] += 1
|
| 196 |
|
| 197 |
def _ranked(counter: Counter) -> list[tuple[str, int]]:
|
| 198 |
# busiest first, then alphabetical for stable ties
|
|
|
|
| 201 |
return {"cities": _ranked(city_counts), "facilities": _ranked(facility_counts)}
|
| 202 |
|
| 203 |
|
| 204 |
+
def location_choices(index: dict) -> dict[str, list[str]]:
|
| 205 |
+
"""Gradio combobox choices — plain names, busiest first.
|
| 206 |
|
| 207 |
+
Just the names (no "· N trials" suffix): with a filterable/custom-value Dropdown the
|
| 208 |
+
displayed option text becomes the field value, so any suffix would leak into the search
|
| 209 |
+
term. Ranking is preserved by list order; the count is used only for that ordering.
|
| 210 |
"""
|
| 211 |
+
return {
|
| 212 |
+
"cities": [name for name, _ in index.get("cities", [])],
|
| 213 |
+
"facilities": [name for name, _ in index.get("facilities", [])],
|
| 214 |
+
}
|
| 215 |
|
| 216 |
|
| 217 |
def _find_supporting_papers(
|
|
|
|
| 453 |
"url": trial.get("url", ""),
|
| 454 |
"target_entities": trial.get("target_entities", []),
|
| 455 |
"mechanism": mechanism,
|
| 456 |
+
"eligibility": trial.get("eligibility", {}) or {},
|
| 457 |
"matched_sites": trial.get("matched_sites", []),
|
| 458 |
"key_papers": key_papers,
|
| 459 |
"sibling_trials": sibling_trials,
|
|
|
|
| 510 |
)
|
| 511 |
|
| 512 |
|
| 513 |
+
def _eligibility_html(elig: dict) -> str:
|
| 514 |
+
"""Collapsible enrollment-criteria block: age/sex summary + inclusion/exclusion text.
|
| 515 |
+
|
| 516 |
+
Rendered only for active/recruiting trials (the caller gates on status), since that's when
|
| 517 |
+
a physician assesses whether a patient qualifies.
|
| 518 |
+
"""
|
| 519 |
+
criteria = (elig.get("criteria") or "").strip()
|
| 520 |
+
if not criteria:
|
| 521 |
+
return ""
|
| 522 |
+
bits = []
|
| 523 |
+
age = " – ".join(x for x in (elig.get("min_age"), elig.get("max_age")) if x) or None
|
| 524 |
+
if age:
|
| 525 |
+
bits.append(f"Age {html.escape(age)}")
|
| 526 |
+
sex = elig.get("sex")
|
| 527 |
+
if sex and sex != "ALL":
|
| 528 |
+
bits.append(html.escape(sex.title()))
|
| 529 |
+
elif sex == "ALL":
|
| 530 |
+
bits.append("All sexes")
|
| 531 |
+
if elig.get("healthy_volunteers"):
|
| 532 |
+
bits.append("Accepts healthy volunteers")
|
| 533 |
+
summary = "Eligibility" + (f" · {' · '.join(bits)}" if bits else "")
|
| 534 |
+
body = html.escape(criteria).replace("\n", "<br>")
|
| 535 |
+
return (
|
| 536 |
+
'<details style="margin-top:6px;font-size:0.82rem;color:#555;">'
|
| 537 |
+
f'<summary style="cursor:pointer;color:#0984E3;">{summary}</summary>'
|
| 538 |
+
f'<div style="margin:4px 0 0 4px;line-height:1.4;">{body}</div>'
|
| 539 |
+
'</details>'
|
| 540 |
+
)
|
| 541 |
+
|
| 542 |
+
|
| 543 |
def render_trials_html(enriched: list[dict], match_count: int) -> str:
|
| 544 |
"""Render enriched location-search results as an HTML card list."""
|
| 545 |
if not enriched:
|
|
|
|
| 556 |
mech = t.get("mechanism", "")
|
| 557 |
mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
|
| 558 |
rationale_html = _tier_rationale_html(ev)
|
| 559 |
+
# Enrollment criteria only for active/recruiting trials — the enrollable ones.
|
| 560 |
+
elig_html = _eligibility_html(t.get("eligibility", {})) if t.get("is_recruiting") else ""
|
| 561 |
phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
|
| 562 |
title = html.escape(t.get("title", "")[:140])
|
| 563 |
nct = html.escape(t.get("nct_id", ""))
|
|
|
|
| 606 |
f'<div style="font-weight:600;">{nct_link} — {title}</div>'
|
| 607 |
f'{rationale_html}'
|
| 608 |
f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
|
| 609 |
+
f'{elig_html}{papers_html}{siblings_html}'
|
| 610 |
'</div>'
|
| 611 |
)
|
| 612 |
|