Spaces:
Running
A newer version of the Gradio SDK is available: 6.22.0
Design Doc: candle-fire Vector DB β Deployment Fix & Migration Path
Status: Draft for review Β· Author: Β· Reviewers: Β· Date: 2026-07 This is a design doc to review/modify before implementation. It ends with open decisions (Section 10) you should resolve; the recommendation is phased so Phase 0 can proceed independently of the Phase 1 decision.
1. Context & Problem
candle-fire's RAG layer uses an embedded ChromaDB (SQLite) vector index. The concern raised: the vector DB "keeps getting larger" β is it time to migrate (e.g., to AWS)?
Current numbers (measured):
- 30,967 vectors (chunks) from ~10k papers; 1.2 GB on disk (
data/chroma/chroma.sqlite3). - Corpus capped at 20k papers (
config.PUBMED_DEFAULT_MAX) β ~62k chunks / ~2.4 GB at the cap. - Embeddings: BioLORD-2023-C (768-dim). Reranker:
cross-encoder/ms-marco-MiniLM-L-6-v2. - Deployment: the index is stored in the HF dataset repo (
KevinIsCoding/candle-fire-data) and downloaded to the Space at cold start (app.py:_ensure_data). The 1.2 GB download amplified the recent 503 outage.
2. Key Finding β this is a deployment problem, not a scale problem
- 31kβ62k vectors is trivially small for any vector store (they scale to millions/billions). Chroma is not being outgrown in capability.
- The 1.2 GB is mostly text, not vectors. Vectors alone are ~100 MB (31k Γ 768 Γ 4 bytes).
The rest is full document text + a full-text index β because the retriever relies on Chroma's
substring search (
where_document $contains) to ground drug codes / gene IDs. That text is load-bearing for citation grounding; it can't just be dropped to shrink the index. - Therefore the pain is where the DB lives (bundled + downloaded per cold start), not its size.
3. Goals / Non-Goals
Goals
- Remove the 1.2 GB cold-start download that couples the DB to the Space lifecycle.
- Preserve the hybrid retrieval (semantic ANN + exact substring + metadata filter) β load-bearing for grounding.
- Keep a low-risk path that also scales if the corpus grows.
Non-Goals
- Shrinking the index by dropping document text (breaks grounding).
- Rewriting the ranking pipeline (RRF, cross-encoder, citation/recency boost) β it is backend-independent and stays as-is.
4. Current Architecture (what any migration must preserve)
Retriever (rag/retriever.py) splits cleanly into two layers:
- Candidate fetch β backend-specific (the only code that must change on migration):
search()/search_by_entities(): semantic ANN viacollection.query(query_texts=...).search_by_keyword(),is_grounded_in_corpus(),is_grounded_in_abstract(): exact substring viawhere_document={"$contains": ...}, expanded over_term_variants()(SPG302 / SPG 302 / SPG-302).get_paper(),paper_texts_for_pmids(): metadata fetch viacollection.get(where=...).
- Ranking β backend-independent (REUSE UNCHANGED):
_parse_raw()normalizes any backend response to the result-dict shape.rrf_merge(),cross_encoder_rerank(),apply_citation_boost()operate purely on result dicts.
Indexer (rag/indexer.py): build_collection() chunks papers (β€6 chunks/paper), embeds via a
SentenceTransformerEmbeddingFunction baked into the Chroma collection, IDs {pmid}_s{i}, scalarized metadata.
Key design lever: only the candidate-fetch functions touch the backend. If a new backend keeps the
_parse_rawoutput shape, the entire ranking/fusion/boost pipeline is reused verbatim. This bounds migration scope to ~6 functions + the indexer writer.
5. Options Considered
| Option | Fit | Verdict |
|---|---|---|
| A. Cheap deploy fix (persistent storage on the Space, and/or Chroma client/server on a small host) | Keeps all Chroma features; ~no code change | Recommended first (Phase 0) |
| B. Postgres + pgvector (AWS Aurora/RDS, or Neon/Supabase) | One store for vectors (<=>) + full-text (ILIKE/tsvector) + metadata (SQL); preserves the hybrid |
Recommended migration target (Phase 1) |
| C. AWS OpenSearch (kNN + BM25 + filters) | Fits the hybrid; AWS-native | Viable but heavier ops/tuning; overkill at 31k vectors |
| D. Vector-only managed (Pinecone / Qdrant Cloud) | No native substring/full-text | Rejected β would require a separate keyword store to keep grounding |
| E. Amazon Bedrock Knowledge Bases | Managed end-to-end RAG | Rejected β replaces the custom hybrid + grounding + citation-weighting (the product's differentiators) |
6. Recommendation (phased)
- Phase 0 β now (do regardless of the migration decision): Option A. Enable persistent storage
on the Space so
_ensure_datastops re-downloading 1.2 GB each restart; optionally run Chroma in client/server mode so the Space queries it over HTTP. Removes the cold-start pain with little/no rewrite. (Pairs with the separate COE action item to make data refresh not require a factory-reboot.) - Phase 1 β when a trigger hits: Option B (pgvector). Triggers: corpus scaled toward 100k+ chunks / continuous ingestion; OR a need to decouple the DB from the Space (multi-instance, independent updates); OR wanting to stop shipping a data blob entirely. Abstract the backend first (Section 7) so the swap is localized and reversible.
7. Phase 1 Detailed Design (pgvector)
Backend abstraction (do this even before migrating): introduce rag/backends/ with a thin interface β
semantic_candidates(query, n), substring_candidates(term_variants, n), fetch_by_pmids(pmids),
abstract_contains(variant), count() β implemented by chroma_backend (wraps today's code) and
pg_backend. retriever.py calls the interface; the ranking layer is untouched. indexer.py gains a pg writer.
Schema (single table):
CREATE TABLE chunks (
chunk_id text PRIMARY KEY,
pmid text,
section text,
chunk_index int,
title text,
year int,
doi text,
citation_count int,
entity_names text,
has_full_text boolean,
document text,
embedding vector(768),
doc_tsv tsvector
);
-- indexes: HNSW on embedding (vector_cosine_ops); GIN on doc_tsv; btree on pmid
Query mapping (preserve current semantics):
collection.query(query_texts)β embed query with BioLORD explicitly, thenSELECT ... ORDER BY embedding <=> $qvec LIMIT n.where_document $containsβWHERE document ILIKE '%'||$variant||'%'(exact substring β matches Chroma$containsbehavior for drug codes; use ILIKE, not tsvector, to keep grounding parity).where={pmid $in}/chunk_index/getβ plain SQLWHERE.- Query embeddings use the same
configmodel; the ranking pipeline consumes_parse_raw-shaped dicts unchanged.
Config: DB URL via env var; embedding model unchanged. Graph/agent layers untouched.
8. Rollout / Migration
- Stand up Postgres + pgvector (RDS/Aurora/Neon);
CREATE EXTENSION vector. - Backfill with a one-off script that reuses the existing
rag/indexer.pychunker onpapers.jsonl, embeds, and writes rows (idempotent onchunk_id). - Dual-read validation (Section 11) β compare Chroma vs pgvector on a fixed query set.
- Flip the Space to the pg backend via an env flag, keeping the Chroma path behind the flag for rollback.
- Once stable, drop the 1.2 GB blob from the dataset repo.
9. Risks & Mitigations
- Substring-grounding parity: ILIKE must reproduce
$containsover_term_variants(). Port the variant expansion; add tests comparing grounding booleans and keyword hits against Chroma. - Per-request query fan-out:
search_by_entities()issues up to 12 queries/request β batch into a single SQL round-trip or reduceRETRIEVAL_ENTITY_QUERY_CAP; HNSW keeps ANN fast. - DB cost/ops: smallest managed tier is ample at this scale.
- Embedding drift: pin the embedding model; re-embed if it changes.
10. Open Questions / Decisions for Reviewer
- Growth: staying ~10β20k papers, or scaling toward full PubMed / continuous ingestion? (Decides whether Phase 1 is needed at all.)
- Host: AWS specifically (Aurora/RDS), or is managed-elsewhere (Neon/Supabase) acceptable β often cheaper/simpler at this scale?
- Driver: is decoupling the DB from the Space a goal in itself, or is cold-start relief (Phase 0) enough for now?
- Ops/budget: appetite for an always-on DB instance vs the current zero-infra file model.
11. Verification
- Parity harness: fixed query set; assert the final top-15 cited PMIDs and the grounding booleans
(
is_grounded_in_corpus,is_grounded_in_abstract) match Chroma within tolerance. - Latency: p50/p95 per query under both backends.
- Cold-start: Space boot time before/after β Phase 0 target: no 1.2 GB re-download; Phase 1 target: no data blob shipped at all.