Spaces:
Running
A newer version of the Gradio SDK is available: 6.22.0
Candle-Fire: ALS Research Landscape Tool
Project Initiative
Candle-fire is a physician-facing ALS research intelligence platform β a sibling to beacon. Where beacon helps patients find clinical trials, candle-fire helps physicians understand the research evidence behind those trials and ALS biology more broadly.
The problem: A physician encounters a clinical trial for an antisense oligonucleotide targeting SOD1. To evaluate it, they need to know: What is the evidence base for SOD1 as a target? What are the known mechanisms? What else has been tried? Today this requires hours of manual literature review.
The solution: A physician types a free-text question ("What's the evidence for tofersen targeting SOD1 in ALS?") and gets a synthesized, cited answer grounded in ~500 curated ALS papers, enriched by a knowledge graph linking genes, proteins, compounds, pathways, and clinical trials β delivered in under 30 seconds.
End state: A Gradio web app deployable on HuggingFace Spaces. Two layers of intelligence: (1) a vector RAG layer for semantic paper retrieval, (2) a knowledge graph layer for entity-level relationship traversal. Claude Sonnet synthesizes both into a structured research landscape report with citations, mechanism summaries, and related trial links.
Critique of Original Plan
RAG alone is semantically blind. Query for "riluzole" must surface "glutamate excitotoxicity" via the KG (riluzole β INHIBITS β glutamate excitotoxicity). KG expansion precedes RAG retrieval β not optional.
Entity normalization is a prerequisite for KG construction. TDP-43 / TARDBP / TDP43 must resolve to one canonical node before graph build. The normalizer runs as part of extraction and writes
canonical_ids.jsonas the single source of truth.All heavy compute is offline batch. Ingestion, extraction, KG build, RAG indexing β run once. Query time: NetworkX pickle (read) + ChromaDB (read) + Anthropic API only.
Citation counts weight evidence quality. A paper cited 500 times is stronger evidence than one cited 5 times. Fetch citation counts from Semantic Scholar API (free, no auth). Use to re-rank RAG results and weight KG edge confidence.
PMC XML full text, not raw PDFs. ~50-60% of recent ALS papers are PMC Open Access. Ingest structured XML (section-labeled: Introduction, Methods, Results, Discussion) via PubMed Entrez. Always fall back to abstract. Avoid PDF parsing β too brittle and legally ambiguous for paywalled papers.
Architecture
Tech Stack
- Language: Python 3.11+, uv package manager
- LLM: Claude Sonnet 4-6 (entity extraction + research synthesis)
- Vector store: ChromaDB (SQLite-backed, persistent, no separate service)
- Graph: NetworkX DiGraph (pickled, loaded once at startup, scales to 20K+ papers)
- Embeddings:
all-MiniLM-L6-v2via sentence-transformers (local, no API needed) - UI: Gradio 6 (HuggingFace Spaces deployment, mirrors beacon)
- Paper source: PubMed Entrez API + PMC XML full text (hybrid)
- Citation counts: Semantic Scholar API (free, by DOI/PMID)
Data Flow
OFFLINE PIPELINE (run once, in order):
Stage 1: Ingestion
scripts/ingest_papers.py
β ingestion/pubmed.py (PubMed Entrez, batch 200 PMIDs, abstract always)
β ingestion/pmc.py (PMC XML full text for OA papers, ~50-60% coverage)
β ingestion/semantic_scholar.py (citation counts per PMID/DOI)
β data/papers/papers.jsonl [ALSPaper: title, abstract, full_text?, citation_count, ...]
scripts/ingest_trials.py [parallel with above]
β ingestion/clinicaltrials.py (ClinicalTrials.gov v2, condition=ALS)
β data/trials/trials.jsonl
Stage 2: Entity Extraction
scripts/extract_entities.py
β extraction/extractor.py (Claude Sonnet, 10 papers/call, retry+backoff, resumable)
β extraction/normalizer.py (HGNC alias table + PubChem + HGNC REST fallback)
β data/extracted/entities.jsonl
β data/extracted/canonical_ids.json
Stage 3: Knowledge Graph Build
scripts/build_graph.py
β graph/builder.py (NetworkX DiGraph, upsert nodes+edges, weight by citation_count)
β data/graph/als_graph.pkl + als_graph.json
Stage 4: RAG Index Build
scripts/build_index.py
β rag/indexer.py (ChromaDB "als_papers", chunk by section if full text, else abstract)
β data/chroma/
ONLINE QUERY PIPELINE (per physician request):
Physician free-text question
β agents/research_agent.py
1. Claude: extract query entities β ["SOD1", "antisense oligonucleotide"]
2. graph/query.py: 1-hop KG expansion β ["SOD1", "TARDBP", "tofersen", "RNA splicing", ...]
3. rag/retriever.py: semantic search + entity-filtered search β top 15 papers
(re-ranked by: ChromaDB distance Γ log(citation_count + 1))
4. graph/query.py: find_trials_for_target() for each query entity
5. Claude Sonnet (streaming): synthesize ResearchLandscape
β Gradio UI (streaming response with citations)
File Structure
candle-fire/
βββ pyproject.toml
βββ .env.example # ANTHROPIC_API_KEY, ENTREZ_EMAIL, NCBI_API_KEY
βββ CLAUDE.md # Architecture guide (module responsibilities)
β
βββ config.py # Model names, paths, ALS seed entities, API endpoints
βββ models.py # Dataclasses: ALSPaper, ExtractedEntity, EntityRelationship, ResearchLandscape
βββ prompts.py # System prompts: extraction + synthesis
βββ tools.py # Tool schema loader (mirrors beacon/tools.py)
βββ llm.py # LLM provider abstraction (copied from beacon/llm.py)
βββ logging_config.py # Structured JSON logging (adapted from beacon/beacon_logging.py)
βββ app.py # Gradio UI entry point (graph + ChromaDB loaded at startup)
βββ main.py # CLI entry point (Rich console)
β
βββ ingestion/
β βββ pubmed.py # PubMed Entrez: fetch abstracts + metadata by query or PMID list
β βββ pmc.py # PMC XML full text: fetch & parse structured sections for OA papers
β βββ clinicaltrials.py # ClinicalTrials.gov v2 (adapted from beacon/trials_api.py, no geo)
β βββ semantic_scholar.py # Citation counts by PMID/DOI (batch API, free tier)
β
βββ extraction/
β βββ extractor.py # Claude Sonnet NER: batch 10 papers, retry+backoff, .progress.json
β βββ normalizer.py # Canonical ID mapping (HGNC alias table + PubChem + REST fallback)
β
βββ graph/
β βββ builder.py # NetworkX DiGraph: upsert nodes+edges, seed ALS entities, citation weighting
β βββ query.py # Traversal: expand_query_entities, find_trials_for_target, get_entity_evidence
β βββ serializer.py # Save/load: pickle (fast) + JSON (human-readable)
β
βββ rag/
β βββ indexer.py # ChromaDB collection builder: section-aware chunking, citation_count metadata
β βββ retriever.py # search(), search_by_entities(), citation-weighted re-ranking
β
βββ agents/
β βββ research_agent.py # Multi-step synthesis agent (streaming, mirrors beacon/agents/research.py)
β
βββ scripts/
β βββ ingest_papers.py # CLI: PubMed + PMC XML + citation counts β papers.jsonl
β βββ ingest_trials.py # CLI: ClinicalTrials.gov β trials.jsonl
β βββ extract_entities.py # CLI: papers.jsonl β entities.jsonl (resumable)
β βββ build_graph.py # CLI: entities.jsonl + trials.jsonl β als_graph.pkl
β βββ build_index.py # CLI: papers.jsonl + entities.jsonl β data/chroma/
β
βββ data/
β βββ papers/papers.jsonl
β βββ trials/trials.jsonl
β βββ extracted/
β β βββ entities.jsonl
β β βββ canonical_ids.json
β β βββ .progress.json # Extraction resumability tracker
β βββ graph/
β β βββ als_graph.pkl
β β βββ als_graph.json
β βββ chroma/ # ChromaDB SQLite store
β βββ tools/
β βββ extract_entities.json
β βββ search_landscape.json
β
βββ tests/
βββ test_pubmed.py
βββ test_pmc.py
βββ test_semantic_scholar.py
βββ test_extractor.py
βββ test_normalizer.py
βββ test_graph_builder.py
βββ test_graph_query.py
βββ test_retriever.py
βββ test_research_agent.py
Staged Implementation Plan
Stage 1 β Project Foundation
Goal: Runnable skeleton with all dependencies wired.
pyproject.tomlwith all dependencies (anthropic, gradio, chromadb, networkx, biopython, httpx, sentence-transformers, rich, python-dotenv)config.py,models.py,prompts.py(stubs),tools.py,llm.py(copied from beacon),logging_config.pydata/tools/extract_entities.jsonandsearch_landscape.jsontool schemas- All
__init__.pyfiles,.env.example,CLAUDE.md
Done when: uv run python -c "import anthropic, chromadb, networkx, Bio" passes with no errors.
Stage 2 β Paper Ingestion Pipeline
Goal: Populate data/papers/papers.jsonl with ~500 ALS papers including citation counts.
ingestion/pubmed.py: PubMed Entrez client (fetch by MeSH query or PMID file, batch 200)ingestion/pmc.py: PMC XML full-text fetcher for OA papers (parse by section)ingestion/semantic_scholar.py: Citation count enrichment (batch by PMID)scripts/ingest_papers.py: Orchestrates all three, writespapers.jsonlscripts/ingest_trials.py+ingestion/clinicaltrials.py: ALS trials βtrials.jsonl
PubMed seed query: "amyotrophic lateral sclerosis"[MeSH Major Topic] AND ("2018"[PDAT]:"2024"[PDAT]) AND hasabstract[text]
Done when: papers.jsonl has 500 lines with citation_count populated; trials.jsonl has 20+ records.
Stage 3 β RAG Pipeline + Working v0
Goal: End-to-end working query pipeline using RAG only (no KG yet).
rag/indexer.py: ChromaDB collection builder (section-aware chunking, citation_count in metadata)rag/retriever.py:search(),search_by_entities(), citation-weighted re-rankingscripts/build_index.pyagents/research_agent.py(RAG-only version, no KG expansion step yet)prompts.pysynthesis promptmain.py: CLI interface
Done when: uv run python main.py "What compounds target glutamate excitotoxicity in ALS?" returns a synthesized answer with cited PMIDs.
Stage 4 β Entity Extraction + Knowledge Graph
Goal: Offline pipeline produces a populated KG; query agent upgrades to KG+RAG.
extraction/extractor.py: Claude Sonnet NER, 10 papers/call, resumable via.progress.jsonextraction/normalizer.py: HGNC alias table (~50 ALS genes) + PubChem fallback + REST fallbackscripts/extract_entities.pygraph/builder.py: NetworkX DiGraph with upsert, citation-weighted edge confidence, seed entitiesgraph/query.py:expand_query_entities(),find_trials_for_target(),get_entity_evidence()graph/serializer.pyscripts/build_graph.py- Upgrade
agents/research_agent.pyto include KG expansion step before RAG
Done when: G.number_of_nodes() > 100; query for "tofersen" surfaces SOD1 as the mechanism link (not just literal tofersen matches).
Stage 5 β Gradio UI + Deployment Polish
Goal: Browser-accessible web app ready for HuggingFace Spaces.
app.py: Gradio UI with streaming, example questions sidebar, citation display, disclaimer- Module-level graph + ChromaDB client initialization (once at startup)
requirements.txt(HF Spaces mirror of pyproject.toml)README.md: setup instructions, offline pipeline run order, env vars- Tests:
uv run pytestall pass
Done when: uv run gradio app.py β physician asks "What's the mechanism of tofersen in ALS?" β streaming response includes SOD1 mechanism, 3+ cited papers, at least 1 NCT trial link, and a disclaimer.
Key Reuse from Beacon
| Beacon file | Candle-fire file | Change |
|---|---|---|
beacon/llm.py |
llm.py |
Copy verbatim; model constants from config.py |
beacon/beacon_logging.py |
logging_config.py |
Namespace β candle_fire |
beacon/trials_api.py |
ingestion/clinicaltrials.py |
Remove geo/distance; add extract_target_entities() |
beacon/tools.py |
tools.py |
Copy pattern; update tool names |
beacon/agents/research.py |
agents/research_agent.py |
Adapt streaming loop; replace tool handlers |
beacon/app.py |
app.py |
Single-turn instead of multi-turn state machine |
Estimated Costs
| Item | Cost |
|---|---|
| Entity extraction (500 papers, 50 Sonnet calls ~10K tokens each) | ~$1.50 one-time |
| Semantic Scholar citation counts | Free |
| PMC XML full text | Free |
| Per physician query (Sonnet, ~15K tokens in+out) | ~$0.05/query |