Spaces:
Sleeping
Sleeping
File size: 13,191 Bytes
0ccfe4a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 | # Candle-Fire: ALS Research Landscape Tool
## Project Initiative
Candle-fire is a physician-facing ALS research intelligence platform β a sibling to beacon. Where beacon helps patients find clinical trials, candle-fire helps physicians understand the *research evidence behind* those trials and ALS biology more broadly.
**The problem**: A physician encounters a clinical trial for an antisense oligonucleotide targeting SOD1. To evaluate it, they need to know: What is the evidence base for SOD1 as a target? What are the known mechanisms? What else has been tried? Today this requires hours of manual literature review.
**The solution**: A physician types a free-text question ("What's the evidence for tofersen targeting SOD1 in ALS?") and gets a synthesized, cited answer grounded in ~500 curated ALS papers, enriched by a knowledge graph linking genes, proteins, compounds, pathways, and clinical trials β delivered in under 30 seconds.
**End state**: A Gradio web app deployable on HuggingFace Spaces. Two layers of intelligence: (1) a vector RAG layer for semantic paper retrieval, (2) a knowledge graph layer for entity-level relationship traversal. Claude Sonnet synthesizes both into a structured research landscape report with citations, mechanism summaries, and related trial links.
---
## Critique of Original Plan
1. **RAG alone is semantically blind.** Query for "riluzole" must surface "glutamate excitotoxicity" via the KG (riluzole β INHIBITS β glutamate excitotoxicity). KG expansion *precedes* RAG retrieval β not optional.
2. **Entity normalization is a prerequisite for KG construction.** TDP-43 / TARDBP / TDP43 must resolve to one canonical node before graph build. The normalizer runs as part of extraction and writes `canonical_ids.json` as the single source of truth.
3. **All heavy compute is offline batch.** Ingestion, extraction, KG build, RAG indexing β run once. Query time: NetworkX pickle (read) + ChromaDB (read) + Anthropic API only.
4. **Citation counts weight evidence quality.** A paper cited 500 times is stronger evidence than one cited 5 times. Fetch citation counts from Semantic Scholar API (free, no auth). Use to re-rank RAG results and weight KG edge confidence.
5. **PMC XML full text, not raw PDFs.** ~50-60% of recent ALS papers are PMC Open Access. Ingest structured XML (section-labeled: Introduction, Methods, Results, Discussion) via PubMed Entrez. Always fall back to abstract. Avoid PDF parsing β too brittle and legally ambiguous for paywalled papers.
---
## Architecture
### Tech Stack
- **Language**: Python 3.11+, uv package manager
- **LLM**: Claude Sonnet 4-6 (entity extraction + research synthesis)
- **Vector store**: ChromaDB (SQLite-backed, persistent, no separate service)
- **Graph**: NetworkX DiGraph (pickled, loaded once at startup, scales to 20K+ papers)
- **Embeddings**: `all-MiniLM-L6-v2` via sentence-transformers (local, no API needed)
- **UI**: Gradio 6 (HuggingFace Spaces deployment, mirrors beacon)
- **Paper source**: PubMed Entrez API + PMC XML full text (hybrid)
- **Citation counts**: Semantic Scholar API (free, by DOI/PMID)
### Data Flow
```
OFFLINE PIPELINE (run once, in order):
Stage 1: Ingestion
scripts/ingest_papers.py
β ingestion/pubmed.py (PubMed Entrez, batch 200 PMIDs, abstract always)
β ingestion/pmc.py (PMC XML full text for OA papers, ~50-60% coverage)
β ingestion/semantic_scholar.py (citation counts per PMID/DOI)
β data/papers/papers.jsonl [ALSPaper: title, abstract, full_text?, citation_count, ...]
scripts/ingest_trials.py [parallel with above]
β ingestion/clinicaltrials.py (ClinicalTrials.gov v2, condition=ALS)
β data/trials/trials.jsonl
Stage 2: Entity Extraction
scripts/extract_entities.py
β extraction/extractor.py (Claude Sonnet, 10 papers/call, retry+backoff, resumable)
β extraction/normalizer.py (HGNC alias table + PubChem + HGNC REST fallback)
β data/extracted/entities.jsonl
β data/extracted/canonical_ids.json
Stage 3: Knowledge Graph Build
scripts/build_graph.py
β graph/builder.py (NetworkX DiGraph, upsert nodes+edges, weight by citation_count)
β data/graph/als_graph.pkl + als_graph.json
Stage 4: RAG Index Build
scripts/build_index.py
β rag/indexer.py (ChromaDB "als_papers", chunk by section if full text, else abstract)
β data/chroma/
ONLINE QUERY PIPELINE (per physician request):
Physician free-text question
β agents/research_agent.py
1. Claude: extract query entities β ["SOD1", "antisense oligonucleotide"]
2. graph/query.py: 1-hop KG expansion β ["SOD1", "TARDBP", "tofersen", "RNA splicing", ...]
3. rag/retriever.py: semantic search + entity-filtered search β top 15 papers
(re-ranked by: ChromaDB distance Γ log(citation_count + 1))
4. graph/query.py: find_trials_for_target() for each query entity
5. Claude Sonnet (streaming): synthesize ResearchLandscape
β Gradio UI (streaming response with citations)
```
---
## File Structure
```
candle-fire/
βββ pyproject.toml
βββ .env.example # ANTHROPIC_API_KEY, ENTREZ_EMAIL, NCBI_API_KEY
βββ CLAUDE.md # Architecture guide (module responsibilities)
β
βββ config.py # Model names, paths, ALS seed entities, API endpoints
βββ models.py # Dataclasses: ALSPaper, ExtractedEntity, EntityRelationship, ResearchLandscape
βββ prompts.py # System prompts: extraction + synthesis
βββ tools.py # Tool schema loader (mirrors beacon/tools.py)
βββ llm.py # LLM provider abstraction (copied from beacon/llm.py)
βββ logging_config.py # Structured JSON logging (adapted from beacon/beacon_logging.py)
βββ app.py # Gradio UI entry point (graph + ChromaDB loaded at startup)
βββ main.py # CLI entry point (Rich console)
β
βββ ingestion/
β βββ pubmed.py # PubMed Entrez: fetch abstracts + metadata by query or PMID list
β βββ pmc.py # PMC XML full text: fetch & parse structured sections for OA papers
β βββ clinicaltrials.py # ClinicalTrials.gov v2 (adapted from beacon/trials_api.py, no geo)
β βββ semantic_scholar.py # Citation counts by PMID/DOI (batch API, free tier)
β
βββ extraction/
β βββ extractor.py # Claude Sonnet NER: batch 10 papers, retry+backoff, .progress.json
β βββ normalizer.py # Canonical ID mapping (HGNC alias table + PubChem + REST fallback)
β
βββ graph/
β βββ builder.py # NetworkX DiGraph: upsert nodes+edges, seed ALS entities, citation weighting
β βββ query.py # Traversal: expand_query_entities, find_trials_for_target, get_entity_evidence
β βββ serializer.py # Save/load: pickle (fast) + JSON (human-readable)
β
βββ rag/
β βββ indexer.py # ChromaDB collection builder: section-aware chunking, citation_count metadata
β βββ retriever.py # search(), search_by_entities(), citation-weighted re-ranking
β
βββ agents/
β βββ research_agent.py # Multi-step synthesis agent (streaming, mirrors beacon/agents/research.py)
β
βββ scripts/
β βββ ingest_papers.py # CLI: PubMed + PMC XML + citation counts β papers.jsonl
β βββ ingest_trials.py # CLI: ClinicalTrials.gov β trials.jsonl
β βββ extract_entities.py # CLI: papers.jsonl β entities.jsonl (resumable)
β βββ build_graph.py # CLI: entities.jsonl + trials.jsonl β als_graph.pkl
β βββ build_index.py # CLI: papers.jsonl + entities.jsonl β data/chroma/
β
βββ data/
β βββ papers/papers.jsonl
β βββ trials/trials.jsonl
β βββ extracted/
β β βββ entities.jsonl
β β βββ canonical_ids.json
β β βββ .progress.json # Extraction resumability tracker
β βββ graph/
β β βββ als_graph.pkl
β β βββ als_graph.json
β βββ chroma/ # ChromaDB SQLite store
β βββ tools/
β βββ extract_entities.json
β βββ search_landscape.json
β
βββ tests/
βββ test_pubmed.py
βββ test_pmc.py
βββ test_semantic_scholar.py
βββ test_extractor.py
βββ test_normalizer.py
βββ test_graph_builder.py
βββ test_graph_query.py
βββ test_retriever.py
βββ test_research_agent.py
```
---
## Staged Implementation Plan
### Stage 1 β Project Foundation
**Goal**: Runnable skeleton with all dependencies wired.
- `pyproject.toml` with all dependencies (anthropic, gradio, chromadb, networkx, biopython, httpx, sentence-transformers, rich, python-dotenv)
- `config.py`, `models.py`, `prompts.py` (stubs), `tools.py`, `llm.py` (copied from beacon), `logging_config.py`
- `data/tools/extract_entities.json` and `search_landscape.json` tool schemas
- All `__init__.py` files, `.env.example`, `CLAUDE.md`
**Done when**: `uv run python -c "import anthropic, chromadb, networkx, Bio"` passes with no errors.
---
### Stage 2 β Paper Ingestion Pipeline
**Goal**: Populate `data/papers/papers.jsonl` with ~500 ALS papers including citation counts.
- `ingestion/pubmed.py`: PubMed Entrez client (fetch by MeSH query or PMID file, batch 200)
- `ingestion/pmc.py`: PMC XML full-text fetcher for OA papers (parse by section)
- `ingestion/semantic_scholar.py`: Citation count enrichment (batch by PMID)
- `scripts/ingest_papers.py`: Orchestrates all three, writes `papers.jsonl`
- `scripts/ingest_trials.py` + `ingestion/clinicaltrials.py`: ALS trials β `trials.jsonl`
PubMed seed query: `"amyotrophic lateral sclerosis"[MeSH Major Topic] AND ("2018"[PDAT]:"2024"[PDAT]) AND hasabstract[text]`
**Done when**: `papers.jsonl` has 500 lines with `citation_count` populated; `trials.jsonl` has 20+ records.
---
### Stage 3 β RAG Pipeline + Working v0
**Goal**: End-to-end working query pipeline using RAG only (no KG yet).
- `rag/indexer.py`: ChromaDB collection builder (section-aware chunking, citation_count in metadata)
- `rag/retriever.py`: `search()`, `search_by_entities()`, citation-weighted re-ranking
- `scripts/build_index.py`
- `agents/research_agent.py` (RAG-only version, no KG expansion step yet)
- `prompts.py` synthesis prompt
- `main.py`: CLI interface
**Done when**: `uv run python main.py "What compounds target glutamate excitotoxicity in ALS?"` returns a synthesized answer with cited PMIDs.
---
### Stage 4 β Entity Extraction + Knowledge Graph
**Goal**: Offline pipeline produces a populated KG; query agent upgrades to KG+RAG.
- `extraction/extractor.py`: Claude Sonnet NER, 10 papers/call, resumable via `.progress.json`
- `extraction/normalizer.py`: HGNC alias table (~50 ALS genes) + PubChem fallback + REST fallback
- `scripts/extract_entities.py`
- `graph/builder.py`: NetworkX DiGraph with upsert, citation-weighted edge confidence, seed entities
- `graph/query.py`: `expand_query_entities()`, `find_trials_for_target()`, `get_entity_evidence()`
- `graph/serializer.py`
- `scripts/build_graph.py`
- Upgrade `agents/research_agent.py` to include KG expansion step before RAG
**Done when**: `G.number_of_nodes() > 100`; query for "tofersen" surfaces SOD1 as the mechanism link (not just literal tofersen matches).
---
### Stage 5 β Gradio UI + Deployment Polish
**Goal**: Browser-accessible web app ready for HuggingFace Spaces.
- `app.py`: Gradio UI with streaming, example questions sidebar, citation display, disclaimer
- Module-level graph + ChromaDB client initialization (once at startup)
- `requirements.txt` (HF Spaces mirror of pyproject.toml)
- `README.md`: setup instructions, offline pipeline run order, env vars
- Tests: `uv run pytest` all pass
**Done when**: `uv run gradio app.py` β physician asks "What's the mechanism of tofersen in ALS?" β streaming response includes SOD1 mechanism, 3+ cited papers, at least 1 NCT trial link, and a disclaimer.
---
## Key Reuse from Beacon
| Beacon file | Candle-fire file | Change |
|---|---|---|
| `beacon/llm.py` | `llm.py` | Copy verbatim; model constants from `config.py` |
| `beacon/beacon_logging.py` | `logging_config.py` | Namespace β `candle_fire` |
| `beacon/trials_api.py` | `ingestion/clinicaltrials.py` | Remove geo/distance; add `extract_target_entities()` |
| `beacon/tools.py` | `tools.py` | Copy pattern; update tool names |
| `beacon/agents/research.py` | `agents/research_agent.py` | Adapt streaming loop; replace tool handlers |
| `beacon/app.py` | `app.py` | Single-turn instead of multi-turn state machine |
---
## Estimated Costs
| Item | Cost |
|---|---|
| Entity extraction (500 papers, 50 Sonnet calls ~10K tokens each) | ~$1.50 one-time |
| Semantic Scholar citation counts | Free |
| PMC XML full text | Free |
| Per physician query (Sonnet, ~15K tokens in+out) | ~$0.05/query |
|