KevinIsInCoding Claude Opus 4.8 commited on
Commit
bc94ae4
·
unverified ·
1 Parent(s): 1c7b1c5

feat(trials): filters, polish & eligibility on main (recover #29 + #30) (#32)

Browse files

#28 (autocomplete) reached main, but stacked PRs #29 (filters + polish) and #30
(eligibility) merged into their base branches, not up to main. This lands their
content on main so the full Clinical Trials tab work ships: study-type/status
filters + defaults, layout & number-consistency polish, and enrollment/eligibility
(ingestion capture, backfill script, active-trial display).

Data change: run scripts/upload_data.py --only trials before deploying.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

Files changed (6) hide show
  1. README.md +5 -5
  2. app.py +49 -13
  3. ingestion/clinicaltrials.py +18 -0
  4. models.py +3 -0
  5. scripts/backfill_eligibility.py +104 -0
  6. trials_query.py +68 -15
README.md CHANGED
@@ -13,11 +13,11 @@ pinned: false
13
 
14
  **ALS Research Intelligence for Physicians**
15
 
16
- Candle-fire is a physician-facing tool that synthesizes evidence from ~500 curated ALS research papers and a biomedical knowledge graph. Ask a free-text question about ALS biology, drug targets, or clinical trials — get a structured, cited answer in under 30 seconds.
17
 
18
  ## What It Does
19
 
20
- - **Two-layer retrieval**: Knowledge graph expansion (BioLORD-2023-C embeddings + NetworkX) → RAG over ~500 ALS papers
21
  - **Citation-weighted ranking**: Highly-cited papers surface first
22
  - **Structured synthesis**: Claude Sonnet produces mechanism summaries, entity tables, evidence strength assessments, and trial links
23
  - **Biomedical synonyms**: BioLORD understands that "TDP-43" = "TARDBP" = "TAR DNA-binding protein 43"
@@ -47,7 +47,7 @@ cp .env.example .env
47
  Build the knowledge assets before launching the app. Each step is resumable.
48
 
49
  ```bash
50
- # 1. Ingest ~500 ALS papers from PubMed + PMC full text + citation counts (~15 min)
51
  uv run python scripts/ingest_papers.py
52
 
53
  # 2. Ingest ALS clinical trials from ClinicalTrials.gov (< 1 min, run in parallel)
@@ -131,10 +131,10 @@ Physician query
131
 
132
  | Source | Content | Volume |
133
  |---|---|---|
134
- | PubMed Entrez | ALS paper abstracts + metadata | ~500 papers (2018–2024) |
135
  | PubMed Central | Full text for Open Access papers | ~50% coverage |
136
  | Semantic Scholar | Citation counts per paper | All papers |
137
- | ClinicalTrials.gov v2 | Active ALS recruiting trials | ~112 trials |
138
 
139
  ## Disclaimer
140
 
 
13
 
14
  **ALS Research Intelligence for Physicians**
15
 
16
+ Candle-fire is a physician-facing tool that synthesizes evidence from ~10,000 curated ALS research papers and a biomedical knowledge graph. Ask a free-text question about ALS biology, drug targets, or clinical trials — get a structured, cited answer in under 30 seconds.
17
 
18
  ## What It Does
19
 
20
+ - **Two-layer retrieval**: Knowledge graph expansion (BioLORD-2023-C embeddings + NetworkX) → RAG over ~10,000 ALS papers
21
  - **Citation-weighted ranking**: Highly-cited papers surface first
22
  - **Structured synthesis**: Claude Sonnet produces mechanism summaries, entity tables, evidence strength assessments, and trial links
23
  - **Biomedical synonyms**: BioLORD understands that "TDP-43" = "TARDBP" = "TAR DNA-binding protein 43"
 
47
  Build the knowledge assets before launching the app. Each step is resumable.
48
 
49
  ```bash
50
+ # 1. Ingest ~10,000 ALS papers from PubMed + PMC full text + citation counts (~15 min)
51
  uv run python scripts/ingest_papers.py
52
 
53
  # 2. Ingest ALS clinical trials from ClinicalTrials.gov (< 1 min, run in parallel)
 
131
 
132
  | Source | Content | Volume |
133
  |---|---|---|
134
+ | PubMed Entrez | ALS paper abstracts + metadata | ~10,000 papers |
135
  | PubMed Central | Full text for Open Access papers | ~50% coverage |
136
  | Semantic Scholar | Citation counts per paper | All papers |
137
+ | ClinicalTrials.gov v2 | Interventional + expanded-access ALS trials | ~720 trials |
138
 
139
  ## Disclaimer
140
 
app.py CHANGED
@@ -1,6 +1,7 @@
1
  """Gradio web UI for candle-fire — physician-facing ALS research intelligence."""
2
  from __future__ import annotations
3
 
 
4
  import json
5
  from pathlib import Path
6
 
@@ -103,7 +104,13 @@ _trials = _load_trials()
103
  _client = anthropic.Anthropic()
104
 
105
  _n_chunks = _collection.count() if _collection else 0
106
- _n_trials = len(_trials)
 
 
 
 
 
 
107
  _kg_nodes = _graph.number_of_nodes() if _graph else 0
108
 
109
  # Experimental therapy landscape (offline-built artifact; loaded once)
@@ -171,12 +178,15 @@ footer { display: none !important; }
171
  /* Autocomplete gate: hide a combobox's attached option list until ≥3 chars (see
172
  _AUTOCOMPLETE_GATE_JS). The script toggles .ac-hide on the input's wrapper by length. */
173
  #facility_combo.ac-hide ul, #city_combo.ac-hide ul { display: none !important; }
 
 
 
174
  """
175
 
176
  _TITLE_MD = """# 🕯️ Candle-Fire
177
  ### ALS Research Intelligence for Physicians
178
  Ask a free-text question about ALS biology, drug targets, or clinical trials.
179
- Answers are synthesized from ~500 curated ALS papers and enriched by a biomedical knowledge graph.
180
  """
181
 
182
  _DISCLAIMER_MD = """<div class="disclaimer">
@@ -223,23 +233,24 @@ _LOC_CHOICES = trials_query.location_choices(_LOC_INDEX)
223
  _AUTOCOMPLETE_GATE_JS = f"""
224
  () => {{
225
  const MIN = {trials_query.MIN_AUTOCOMPLETE_CHARS};
226
- const gate = (id) => {{
227
  const root = document.getElementById(id);
228
  if (!root) return;
229
  const input = root.querySelector('input');
230
  if (!input) return;
 
231
  const apply = () => root.classList.toggle('ac-hide', input.value.trim().length < MIN);
232
  input.addEventListener('input', apply);
233
  input.addEventListener('focus', apply);
234
  apply();
235
  }};
236
- gate('facility_combo');
237
- gate('city_combo');
238
  }}
239
  """
240
 
241
 
242
- def _search_trials(facility: str, state: str, city: str, status: str) -> str:
243
  facility = (facility or "").strip() or None
244
  city = (city or "").strip() or None
245
  state = None if (not state or state == "All") else state
@@ -248,8 +259,28 @@ def _search_trials(facility: str, state: str, city: str, status: str) -> str:
248
  return '<div style="color:#888;padding:12px 0;">Enter a facility, state, or city to search.</div>'
249
 
250
  matches = trials_query.search_trials_by_location(
251
- _trials, facility=facility, city=city, state=state, status=status
 
252
  )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
253
  enriched = [
254
  trials_query.enrich_trial(t, _collection, _graph, _trials)
255
  for t in matches[:_TRIAL_ENRICH_CAP]
@@ -269,7 +300,7 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
269
  gr.HTML(
270
  f'<div class="status-bar">'
271
  f'{_n_chunks} paper chunks &nbsp;·&nbsp; '
272
- f'{_n_trials} clinical trials &nbsp;·&nbsp; '
273
  f'{_kg_nodes} knowledge graph nodes'
274
  f'</div>'
275
  )
@@ -380,12 +411,11 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
380
  "**location** (state / city). Each result is enriched with recruiting status, an "
381
  "evidence-strength tier, key supporting papers, and related trials for the same compound."
382
  )
383
- with gr.Row():
384
  facility_tb = gr.Dropdown(
385
  choices=_LOC_CHOICES["facilities"], value=None,
386
  label="Facility / institution", scale=2,
387
  filterable=True, allow_custom_value=True, elem_id="facility_combo",
388
- info="Type ≥3 letters and pick a match (e.g. Mass General).",
389
  )
390
  state_dd = gr.Dropdown(
391
  choices=_US_STATES, value="All", label="State", scale=1,
@@ -394,11 +424,17 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
394
  choices=_LOC_CHOICES["cities"], value=None,
395
  label="City", scale=1,
396
  filterable=True, allow_custom_value=True, elem_id="city_combo",
397
- info="Type ≥3 letters; busiest trial cities first.",
 
 
 
 
 
 
398
  )
399
  trial_status_dd = gr.Dropdown(
400
  choices=["All", "Recruiting", "Not recruiting"],
401
- value="All", label="Recruitment status", scale=1,
402
  )
403
  search_btn = gr.Button("Search trials", variant="primary")
404
  trial_results = gr.HTML(
@@ -408,7 +444,7 @@ with gr.Blocks(title="Candle-Fire — ALS Research Intelligence") as demo:
408
  # Facility/city are typeable comboboxes (filterable Dropdowns) — the physician
409
  # types and picks from the attached, busiest-first list. No per-keystroke server
410
  # event needed; the search reads the selected/typed value directly.
411
- _trial_search_inputs = [facility_tb, state_dd, city_tb, trial_status_dd]
412
  search_btn.click(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
413
  # Picking a facility/city from its list also runs the search immediately.
414
  facility_tb.select(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
 
1
  """Gradio web UI for candle-fire — physician-facing ALS research intelligence."""
2
  from __future__ import annotations
3
 
4
+ import html
5
  import json
6
  from pathlib import Path
7
 
 
104
  _client = anthropic.Anthropic()
105
 
106
  _n_chunks = _collection.count() if _collection else 0
107
+ # Headline trial count = recruiting, interventional trials only (the actionable set), not the
108
+ # full corpus (which includes completed/terminated studies and expanded-access programs).
109
+ _RECRUITING = {"RECRUITING", "NOT_YET_RECRUITING", "ENROLLING_BY_INVITATION", "AVAILABLE"}
110
+ _n_trials = sum(
111
+ 1 for t in _trials
112
+ if t.get("study_type") == "INTERVENTIONAL" and t.get("status") in _RECRUITING
113
+ )
114
  _kg_nodes = _graph.number_of_nodes() if _graph else 0
115
 
116
  # Experimental therapy landscape (offline-built artifact; loaded once)
 
178
  /* Autocomplete gate: hide a combobox's attached option list until ≥3 chars (see
179
  _AUTOCOMPLETE_GATE_JS). The script toggles .ac-hide on the input's wrapper by length. */
180
  #facility_combo.ac-hide ul, #city_combo.ac-hide ul { display: none !important; }
181
+ /* Smaller filter labels so long ones (e.g. "Recruitment status") stay on one line and the
182
+ dropdown chevron doesn't overlap the text. */
183
+ .trial-filters label span { font-size: 0.78rem !important; white-space: nowrap; }
184
  """
185
 
186
  _TITLE_MD = """# 🕯️ Candle-Fire
187
  ### ALS Research Intelligence for Physicians
188
  Ask a free-text question about ALS biology, drug targets, or clinical trials.
189
+ Answers are synthesized from a curated ALS research corpus and enriched by a biomedical knowledge graph.
190
  """
191
 
192
  _DISCLAIMER_MD = """<div class="disclaimer">
 
233
  _AUTOCOMPLETE_GATE_JS = f"""
234
  () => {{
235
  const MIN = {trials_query.MIN_AUTOCOMPLETE_CHARS};
236
+ const gate = (id, hint) => {{
237
  const root = document.getElementById(id);
238
  if (!root) return;
239
  const input = root.querySelector('input');
240
  if (!input) return;
241
+ if (hint) input.setAttribute('placeholder', hint); // in-box hint; hides once they type
242
  const apply = () => root.classList.toggle('ac-hide', input.value.trim().length < MIN);
243
  input.addEventListener('input', apply);
244
  input.addEventListener('focus', apply);
245
  apply();
246
  }};
247
+ gate('facility_combo', 'Type ≥3 letters, e.g. Mass General');
248
+ gate('city_combo', 'Type ≥3 letters, busiest cities first');
249
  }}
250
  """
251
 
252
 
253
+ def _search_trials(facility: str, state: str, city: str, study_type: str, status: str) -> str:
254
  facility = (facility or "").strip() or None
255
  city = (city or "").strip() or None
256
  state = None if (not state or state == "All") else state
 
259
  return '<div style="color:#888;padding:12px 0;">Enter a facility, state, or city to search.</div>'
260
 
261
  matches = trials_query.search_trials_by_location(
262
+ _trials, facility=facility, city=city, state=state,
263
+ status=status, study_type=study_type,
264
  )
265
+
266
+ # If the active filters hide everything, say whether broader filters would find trials —
267
+ # e.g. a facility with only completed studies under the default Recruiting + Interventional.
268
+ if not matches and (status != "All" or study_type != "All"):
269
+ broad = trials_query.search_trials_by_location(
270
+ _trials, facility=facility, city=city, state=state, status="All", study_type="All",
271
+ )
272
+ if broad:
273
+ where = ", ".join(p for p in (facility, city, state) if p)
274
+ return (
275
+ '<div style="background:#fff6e5;border:1px solid #ffe0a3;border-radius:8px;'
276
+ 'padding:10px 12px;margin:6px 0;color:#7a5b00;font-size:0.9rem;">'
277
+ f'No <b>{html.escape((study_type or "").lower())}</b> trials that are '
278
+ f'<b>{html.escape((status or "").lower())}</b> at {html.escape(where)}. '
279
+ f'{len(broad)} trial(s) exist there under broader filters — set '
280
+ '<b>Study type</b> and <b>Recruitment status</b> to <b>All</b> to see them.'
281
+ '</div>'
282
+ )
283
+
284
  enriched = [
285
  trials_query.enrich_trial(t, _collection, _graph, _trials)
286
  for t in matches[:_TRIAL_ENRICH_CAP]
 
300
  gr.HTML(
301
  f'<div class="status-bar">'
302
  f'{_n_chunks} paper chunks &nbsp;·&nbsp; '
303
+ f'{_n_trials} recruiting interventional trials &nbsp;·&nbsp; '
304
  f'{_kg_nodes} knowledge graph nodes'
305
  f'</div>'
306
  )
 
411
  "**location** (state / city). Each result is enriched with recruiting status, an "
412
  "evidence-strength tier, key supporting papers, and related trials for the same compound."
413
  )
414
+ with gr.Row(elem_classes="trial-filters"):
415
  facility_tb = gr.Dropdown(
416
  choices=_LOC_CHOICES["facilities"], value=None,
417
  label="Facility / institution", scale=2,
418
  filterable=True, allow_custom_value=True, elem_id="facility_combo",
 
419
  )
420
  state_dd = gr.Dropdown(
421
  choices=_US_STATES, value="All", label="State", scale=1,
 
424
  choices=_LOC_CHOICES["cities"], value=None,
425
  label="City", scale=1,
426
  filterable=True, allow_custom_value=True, elem_id="city_combo",
427
+ )
428
+ # Filters on their own row so the labels/values have full width — no wrapping,
429
+ # no value running under the chevron.
430
+ with gr.Row(elem_classes="trial-filters"):
431
+ study_type_dd = gr.Dropdown(
432
+ choices=["Interventional", "Expanded Access", "All"],
433
+ value="Interventional", label="Study type", scale=1,
434
  )
435
  trial_status_dd = gr.Dropdown(
436
  choices=["All", "Recruiting", "Not recruiting"],
437
+ value="Recruiting", label="Recruitment status", scale=1,
438
  )
439
  search_btn = gr.Button("Search trials", variant="primary")
440
  trial_results = gr.HTML(
 
444
  # Facility/city are typeable comboboxes (filterable Dropdowns) — the physician
445
  # types and picks from the attached, busiest-first list. No per-keystroke server
446
  # event needed; the search reads the selected/typed value directly.
447
+ _trial_search_inputs = [facility_tb, state_dd, city_tb, study_type_dd, trial_status_dd]
448
  search_btn.click(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
449
  # Picking a facility/city from its list also runs the search immediately.
450
  facility_tb.select(_search_trials, inputs=_trial_search_inputs, outputs=[trial_results])
ingestion/clinicaltrials.py CHANGED
@@ -70,6 +70,22 @@ def fetch_als_trials(
70
  return trials
71
 
72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
73
  def _flatten_trial(study: dict) -> dict:
74
  proto = study.get("protocolSection", {})
75
  id_mod = proto.get("identificationModule", {})
@@ -79,6 +95,7 @@ def _flatten_trial(study: dict) -> dict:
79
  arms_mod = proto.get("armsInterventionsModule", {})
80
  status_mod = proto.get("statusModule", {})
81
  contacts_mod = proto.get("contactsLocationsModule", {})
 
82
 
83
  nct_id = id_mod.get("nctId", "")
84
  interventions = [
@@ -126,6 +143,7 @@ def _flatten_trial(study: dict) -> dict:
126
  "locations": locations,
127
  "contact_phone": contact_phone,
128
  "contact_email": contact_email,
 
129
  }
130
 
131
 
 
70
  return trials
71
 
72
 
73
+ def _extract_eligibility(elig_mod: dict) -> dict:
74
+ """Enrollment/eligibility fields from a CT.gov v2 eligibilityModule.
75
+
76
+ Shared by _flatten_trial (ingestion) and scripts/backfill_eligibility.py so the stored
77
+ shape is identical whichever path populated it.
78
+ """
79
+ return {
80
+ "criteria": elig_mod.get("eligibilityCriteria", ""), # free text: Inclusion/Exclusion
81
+ "sex": elig_mod.get("sex", ""), # ALL / MALE / FEMALE
82
+ "min_age": elig_mod.get("minimumAge", ""), # e.g. "18 Years"
83
+ "max_age": elig_mod.get("maximumAge", ""),
84
+ "healthy_volunteers": elig_mod.get("healthyVolunteers"),
85
+ "std_ages": elig_mod.get("stdAges", []), # e.g. ["ADULT", "OLDER_ADULT"]
86
+ }
87
+
88
+
89
  def _flatten_trial(study: dict) -> dict:
90
  proto = study.get("protocolSection", {})
91
  id_mod = proto.get("identificationModule", {})
 
95
  arms_mod = proto.get("armsInterventionsModule", {})
96
  status_mod = proto.get("statusModule", {})
97
  contacts_mod = proto.get("contactsLocationsModule", {})
98
+ elig_mod = proto.get("eligibilityModule", {})
99
 
100
  nct_id = id_mod.get("nctId", "")
101
  interventions = [
 
143
  "locations": locations,
144
  "contact_phone": contact_phone,
145
  "contact_email": contact_email,
146
+ "eligibility": _extract_eligibility(elig_mod),
147
  }
148
 
149
 
models.py CHANGED
@@ -118,6 +118,9 @@ class TrialSummary:
118
  locations: list[dict] = field(default_factory=list)
119
  contact_phone: str = ""
120
  contact_email: str = ""
 
 
 
121
 
122
 
123
  @dataclass
 
118
  locations: list[dict] = field(default_factory=list)
119
  contact_phone: str = ""
120
  contact_email: str = ""
121
+ # Enrollment/eligibility from CT.gov: {criteria, sex, min_age, max_age,
122
+ # healthy_volunteers, std_ages}. Shown for active/recruiting trials.
123
+ eligibility: dict = field(default_factory=dict)
124
 
125
 
126
  @dataclass
scripts/backfill_eligibility.py ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Backfill enrollment/eligibility criteria into an existing trials.jsonl.
2
+
3
+ Adds the `eligibility` field to trials that don't have it by fetching each trial's
4
+ eligibilityModule from ClinicalTrials.gov v2 — WITHOUT re-running the expensive LLM target
5
+ extraction that a full `ingest_trials.py` would. Fetches in batches via `filter.ids`, merges
6
+ in place, and rewrites trials.jsonl. Idempotent: re-running only fills trials still missing it.
7
+
8
+ Usage:
9
+ uv run python scripts/backfill_eligibility.py # backfill missing eligibility
10
+ uv run python scripts/backfill_eligibility.py --force # refetch for ALL trials
11
+ uv run python scripts/backfill_eligibility.py --dry-run # report only, no write
12
+
13
+ After it writes trials.jsonl, deploy per the runbook:
14
+ uv run python scripts/upload_data.py --only trials # push to the HF dataset repo
15
+ git push hf main
16
+ """
17
+ from __future__ import annotations
18
+
19
+ import argparse
20
+ import json
21
+ import sys
22
+ import time
23
+ from pathlib import Path
24
+
25
+ import httpx
26
+
27
+ sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
28
+
29
+ from config import CTGOV_BASE, TRIALS_PATH # noqa: E402
30
+ from ingestion.clinicaltrials import _extract_eligibility # noqa: E402
31
+
32
+ _BATCH = 50 # NCT ids per request (filter.ids)
33
+ _PAUSE_S = 0.34 # ~3 req/s, polite to CT.gov
34
+
35
+
36
+ def _fetch_eligibility(nct_ids: list[str]) -> dict[str, dict]:
37
+ """{nct_id: eligibility dict} for a batch of NCT ids."""
38
+ resp = httpx.get(
39
+ CTGOV_BASE,
40
+ params={
41
+ "filter.ids": ",".join(nct_ids),
42
+ "fields": "NCTId,EligibilityModule",
43
+ "pageSize": len(nct_ids),
44
+ },
45
+ timeout=30,
46
+ )
47
+ resp.raise_for_status()
48
+ out: dict[str, dict] = {}
49
+ for study in resp.json().get("studies", []):
50
+ proto = study.get("protocolSection", {})
51
+ nct = proto.get("identificationModule", {}).get("nctId", "")
52
+ if nct:
53
+ out[nct] = _extract_eligibility(proto.get("eligibilityModule", {}))
54
+ return out
55
+
56
+
57
+ def main() -> int:
58
+ ap = argparse.ArgumentParser(description="Backfill eligibility criteria into trials.jsonl.")
59
+ ap.add_argument("--force", action="store_true", help="Refetch for all trials, not just missing.")
60
+ ap.add_argument("--dry-run", action="store_true", help="Report what would change; write nothing.")
61
+ args = ap.parse_args()
62
+
63
+ if not TRIALS_PATH.exists():
64
+ print(f"ERROR — {TRIALS_PATH} not found (run ingest_trials.py first).", file=sys.stderr)
65
+ return 1
66
+
67
+ trials = [json.loads(line) for line in TRIALS_PATH.read_text().splitlines() if line.strip()]
68
+ todo = [
69
+ t for t in trials
70
+ if t.get("nct_id") and (args.force or not (t.get("eligibility") or {}).get("criteria"))
71
+ ]
72
+ print(f"{len(trials)} trials; {len(todo)} to fetch eligibility for"
73
+ f"{' (force)' if args.force else ''}.")
74
+ if not todo:
75
+ print("Nothing to do.")
76
+ return 0
77
+ if args.dry_run:
78
+ print("--dry-run: no fetch, no write.")
79
+ return 0
80
+
81
+ by_nct = {t["nct_id"]: t for t in trials}
82
+ fetched = 0
83
+ for i in range(0, len(todo), _BATCH):
84
+ ids = [t["nct_id"] for t in todo[i : i + _BATCH]]
85
+ try:
86
+ for nct, elig in _fetch_eligibility(ids).items():
87
+ if nct in by_nct:
88
+ by_nct[nct]["eligibility"] = elig
89
+ fetched += 1
90
+ except httpx.HTTPError as exc:
91
+ print(f" batch {i // _BATCH} failed: {exc}", file=sys.stderr)
92
+ print(f" {min(i + _BATCH, len(todo))}/{len(todo)}")
93
+ if i + _BATCH < len(todo):
94
+ time.sleep(_PAUSE_S)
95
+
96
+ with_crit = sum(1 for t in trials if (t.get("eligibility") or {}).get("criteria"))
97
+ TRIALS_PATH.write_text("".join(json.dumps(t, ensure_ascii=False) + "\n" for t in trials))
98
+ print(f"Fetched {fetched}; {with_crit}/{len(trials)} trials now have eligibility criteria.")
99
+ print(f"Wrote {TRIALS_PATH}. Next: upload_data.py --only trials, then git push hf main.")
100
+ return 0
101
+
102
+
103
+ if __name__ == "__main__":
104
+ raise SystemExit(main())
trials_query.py CHANGED
@@ -112,19 +112,29 @@ def search_trials_by_location(
112
  state: str | None = None,
113
  country: str | None = None,
114
  status: str | None = None,
 
115
  ) -> list[dict]:
116
  """Return trials with at least one site matching the location filters.
117
 
118
  `status`: "Recruiting" keeps only trials with an open overall status; "Not recruiting"
119
- keeps only closed ones; anything else (None / "All") keeps all. Each returned trial is a
120
- shallow copy with `matched_sites` attached; recruiting trials are ranked first.
 
 
121
  """
122
  if not any([facility, city, state, country]):
123
  return []
124
 
125
  status_filter = (status or "").strip().lower()
 
126
  results: list[dict] = []
127
  for trial in trials:
 
 
 
 
 
 
128
  matched_sites = [
129
  s for s in trial.get("locations", [])
130
  if _site_matches(s, facility, city, state, country)
@@ -158,22 +168,31 @@ MIN_AUTOCOMPLETE_CHARS = 3 # client-side gate: the attached list stays hidden u
158
 
159
 
160
  def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]:
161
- """Distinct facility + city names with their trial-site counts, each sorted busiest first.
162
 
163
- Built once at startup; feeds location_choices() which becomes the combobox `choices`.
 
 
 
164
  """
165
  from collections import Counter
166
 
167
  city_counts: Counter = Counter()
168
  facility_counts: Counter = Counter()
169
  for t in trials:
 
 
170
  for s in t.get("locations", []):
171
  city = (s.get("city") or "").strip()
172
  facility = (s.get("facility") or "").strip()
173
  if city:
174
- city_counts[city] += 1
175
  if facility:
176
- facility_counts[facility] += 1
 
 
 
 
177
 
178
  def _ranked(counter: Counter) -> list[tuple[str, int]]:
179
  # busiest first, then alphabetical for stable ties
@@ -182,16 +201,17 @@ def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]
182
  return {"cities": _ranked(city_counts), "facilities": _ranked(facility_counts)}
183
 
184
 
185
- def location_choices(index: dict) -> dict[str, list[tuple[str, str]]]:
186
- """Gradio combobox (label, value) choices for cities and facilities, busiest first.
187
 
188
- Label carries the site count ("New York · 100 sites"); value is the clean name used for
189
- searching. Order is preserved by Gradio's client-side filter, so ranking holds as you type.
 
190
  """
191
- def _fmt(pairs: list[tuple[str, int]]) -> list[tuple[str, str]]:
192
- return [(f"{name} · {c} site{'s' if c != 1 else ''}", name) for name, c in pairs]
193
-
194
- return {"cities": _fmt(index.get("cities", [])), "facilities": _fmt(index.get("facilities", []))}
195
 
196
 
197
  def _find_supporting_papers(
@@ -433,6 +453,7 @@ def enrich_trial(
433
  "url": trial.get("url", ""),
434
  "target_entities": trial.get("target_entities", []),
435
  "mechanism": mechanism,
 
436
  "matched_sites": trial.get("matched_sites", []),
437
  "key_papers": key_papers,
438
  "sibling_trials": sibling_trials,
@@ -489,6 +510,36 @@ def _tier_rationale_html(ev: dict) -> str:
489
  )
490
 
491
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
492
  def render_trials_html(enriched: list[dict], match_count: int) -> str:
493
  """Render enriched location-search results as an HTML card list."""
494
  if not enriched:
@@ -505,6 +556,8 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
505
  mech = t.get("mechanism", "")
506
  mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
507
  rationale_html = _tier_rationale_html(ev)
 
 
508
  phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
509
  title = html.escape(t.get("title", "")[:140])
510
  nct = html.escape(t.get("nct_id", ""))
@@ -553,7 +606,7 @@ def render_trials_html(enriched: list[dict], match_count: int) -> str:
553
  f'<div style="font-weight:600;">{nct_link} — {title}</div>'
554
  f'{rationale_html}'
555
  f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
556
- f'{papers_html}{siblings_html}'
557
  '</div>'
558
  )
559
 
 
112
  state: str | None = None,
113
  country: str | None = None,
114
  status: str | None = None,
115
+ study_type: str | None = None,
116
  ) -> list[dict]:
117
  """Return trials with at least one site matching the location filters.
118
 
119
  `status`: "Recruiting" keeps only trials with an open overall status; "Not recruiting"
120
+ keeps only closed ones; anything else (None / "All") keeps all.
121
+ `study_type`: "Interventional" keeps interventional trials; "Expanded Access" keeps
122
+ expanded-access (investigational-use) programs; anything else (None / "All") keeps both.
123
+ Each returned trial is a shallow copy with `matched_sites` attached; recruiting first.
124
  """
125
  if not any([facility, city, state, country]):
126
  return []
127
 
128
  status_filter = (status or "").strip().lower()
129
+ type_filter = (study_type or "").strip().lower()
130
  results: list[dict] = []
131
  for trial in trials:
132
+ is_eap = bool(trial.get("is_expanded_access"))
133
+ if type_filter == "interventional" and is_eap:
134
+ continue
135
+ if type_filter == "expanded access" and not is_eap:
136
+ continue
137
+
138
  matched_sites = [
139
  s for s in trial.get("locations", [])
140
  if _site_matches(s, facility, city, state, country)
 
168
 
169
 
170
  def build_location_index(trials: list[dict]) -> dict[str, list[tuple[str, int]]]:
171
+ """Distinct facility + city names with their TRIAL counts, each sorted busiest first.
172
 
173
+ Counts distinct trials, not site rows: a single trial that lists an anonymized placeholder
174
+ like "GSK Investigational Site" 49 times counts once, so placeholders don't balloon to the
175
+ top of the ranking and the shown count matches what a search can return. Built once at
176
+ startup; feeds location_choices() which becomes the combobox `choices`.
177
  """
178
  from collections import Counter
179
 
180
  city_counts: Counter = Counter()
181
  facility_counts: Counter = Counter()
182
  for t in trials:
183
+ cities_here: set[str] = set()
184
+ facilities_here: set[str] = set()
185
  for s in t.get("locations", []):
186
  city = (s.get("city") or "").strip()
187
  facility = (s.get("facility") or "").strip()
188
  if city:
189
+ cities_here.add(city)
190
  if facility:
191
+ facilities_here.add(facility)
192
+ for c in cities_here:
193
+ city_counts[c] += 1
194
+ for f in facilities_here:
195
+ facility_counts[f] += 1
196
 
197
  def _ranked(counter: Counter) -> list[tuple[str, int]]:
198
  # busiest first, then alphabetical for stable ties
 
201
  return {"cities": _ranked(city_counts), "facilities": _ranked(facility_counts)}
202
 
203
 
204
+ def location_choices(index: dict) -> dict[str, list[str]]:
205
+ """Gradio combobox choices — plain names, busiest first.
206
 
207
+ Just the names (no "· N trials" suffix): with a filterable/custom-value Dropdown the
208
+ displayed option text becomes the field value, so any suffix would leak into the search
209
+ term. Ranking is preserved by list order; the count is used only for that ordering.
210
  """
211
+ return {
212
+ "cities": [name for name, _ in index.get("cities", [])],
213
+ "facilities": [name for name, _ in index.get("facilities", [])],
214
+ }
215
 
216
 
217
  def _find_supporting_papers(
 
453
  "url": trial.get("url", ""),
454
  "target_entities": trial.get("target_entities", []),
455
  "mechanism": mechanism,
456
+ "eligibility": trial.get("eligibility", {}) or {},
457
  "matched_sites": trial.get("matched_sites", []),
458
  "key_papers": key_papers,
459
  "sibling_trials": sibling_trials,
 
510
  )
511
 
512
 
513
+ def _eligibility_html(elig: dict) -> str:
514
+ """Collapsible enrollment-criteria block: age/sex summary + inclusion/exclusion text.
515
+
516
+ Rendered only for active/recruiting trials (the caller gates on status), since that's when
517
+ a physician assesses whether a patient qualifies.
518
+ """
519
+ criteria = (elig.get("criteria") or "").strip()
520
+ if not criteria:
521
+ return ""
522
+ bits = []
523
+ age = " – ".join(x for x in (elig.get("min_age"), elig.get("max_age")) if x) or None
524
+ if age:
525
+ bits.append(f"Age {html.escape(age)}")
526
+ sex = elig.get("sex")
527
+ if sex and sex != "ALL":
528
+ bits.append(html.escape(sex.title()))
529
+ elif sex == "ALL":
530
+ bits.append("All sexes")
531
+ if elig.get("healthy_volunteers"):
532
+ bits.append("Accepts healthy volunteers")
533
+ summary = "Eligibility" + (f" · {' · '.join(bits)}" if bits else "")
534
+ body = html.escape(criteria).replace("\n", "<br>")
535
+ return (
536
+ '<details style="margin-top:6px;font-size:0.82rem;color:#555;">'
537
+ f'<summary style="cursor:pointer;color:#0984E3;">{summary}</summary>'
538
+ f'<div style="margin:4px 0 0 4px;line-height:1.4;">{body}</div>'
539
+ '</details>'
540
+ )
541
+
542
+
543
  def render_trials_html(enriched: list[dict], match_count: int) -> str:
544
  """Render enriched location-search results as an HTML card list."""
545
  if not enriched:
 
556
  mech = t.get("mechanism", "")
557
  mech_pill = _pill(f'Mechanism: {mech}', "#6C5CE7") if mech else ""
558
  rationale_html = _tier_rationale_html(ev)
559
+ # Enrollment criteria only for active/recruiting trials — the enrollable ones.
560
+ elig_html = _eligibility_html(t.get("eligibility", {})) if t.get("is_recruiting") else ""
561
  phase = html.escape((t.get("phase") or "—").replace("PHASE", "Ph"))
562
  title = html.escape(t.get("title", "")[:140])
563
  nct = html.escape(t.get("nct_id", ""))
 
606
  f'<div style="font-weight:600;">{nct_link} — {title}</div>'
607
  f'{rationale_html}'
608
  f'<div style="margin-top:4px;font-size:0.82rem;color:#555;"><b>Site(s):</b><br>{sites_html}</div>'
609
+ f'{elig_html}{papers_html}{siblings_html}'
610
  '</div>'
611
  )
612