Spaces:
Running
Download implementationplan.md from Samad14/bio-nexus-api: direct link, hf CLI and curl.
- Browser
- Download file 8.19 kB
-
https://huggingface.co/spaces/Samad14/bio-nexus-api/resolve/main/implementationplan.md
- Command line
-
hf download hf://spaces/Samad14/bio-nexus-api/implementationplan.md
-
curl -L -o implementationplan.md https://huggingface.co/spaces/Samad14/bio-nexus-api/resolve/main/implementationplan.md
implementationplan.md β BioFlow AI Implementation Plan
Current status (v4.0, 2026-08-05): Track A prototype completed. Track B sprints 1β10 completed. Additional sprints 11β15 (structure suite, docking, MD, ADMET, function, sequencing, design system) shipped. See Track B below.
Track A β Prototype Sprint (18 Days) β Completed
One golden path only: Sequence In β BLAST β AI-interpreted report. Everything else is Track B.
Day 0 β Pre-flight (~2 hrs, today)
- Create
bioflow-frontendandbioflow-backendrepos - Create Supabase project, run schema.md migrations
- Create Cloudflare R2 bucket
- Get API keys: Gemini (Resend optional for prototype)
- Confirm NCBI BLAST URL API access β no key required, but max 1 request every 10s; build this constraint in from day one, not as a fix later
Days 1β2 β Foundations
Backend
- FastAPI scaffold +
/health - Supabase client wrapper (
core/db.py), R2 wrapper (core/storage.py) - Pydantic models for
jobs/job_steps POST /jobs(create, status=queued),GET /jobs/{id}
Frontend
- Next.js + TS + Tailwind scaffold, design tokens from design.md β
tailwind.config.ts/globals.css - Supabase client (anon key) + anonymous session on first visit
- Layout shell, static landing page
Job CRUD is the spine everything attaches to β get it working with dummy data before touching NCBI.
Days 3β4 β NCBI BLAST Integration (core engine)
integrations/ncbi/blast.py:submit_blast(),check_status(),fetch_results()against the QBLAST URL API- Background task (start with
asyncio.create_taskβ sufficient for a single-instance prototype): on job create, submit β poll every 15s β onREADY, fetch XML, push to R2, updatejob_steps - XML parser β hit list (accession, description, % identity, E-value, bit score, alignment)
Risk flag: NCBI nr searches can take 1β5 min and occasionally queue longer. Build a demo-mode fallback now β pick 1β2 well-characterized sequences (e.g., human insulin), pre-run them, cache the result. This is your insurance for Day 18.
Day 5 β Sequence Input + Validation
integrations/ncbi/efetch.pyβ fetch sequence by accession- Sequence-type detection (nucleotide vs. protein, composition-based)
POST /sequences/validate- Frontend: wizard shell, step indicator, Step 1 (operation grid), Step 2 (paste/accession tabs + validation feedback)
Days 6β7 β Wizard β Job Creation β Processing Screen
- Step 3 (confirm/run) β
POST /jobsβ/jobs/[id] - Processing screen:
useJobStatuspolling hook, animated status indicator, friendly per-state copy - Backend: wire job creation to the Day 3β4 background task, update
job_stepsat each stage
Checkpoint: by end of Day 7 you can paste a sequence, hit run, and watch real NCBI status updates. First "it's alive" moment.
Days 8β9 β Results Rendering (raw, no AI yet)
- Frontend: hits table, score-bar visualization (confidence bands), alignment view for top hit
- Backend:
GET /jobs/{id}/resultsβ parsed hits + top-hit alignment from stored XML
Days 10β11 β AI Interpretation Layer
services/interpreter.pyβ Gemini call- Prompt template at
/prompts/blast_interpretation.md: takes parsed hits, returns a 2β3 sentence summary + per-term explanations (E-value, bit score, % identity) in context of this specific result. System instruction bakes in the "faithful execution, not 100% accuracy" framing from the PRD. - New step
interpretingβ result stored injob_steps.result_json - Frontend: AI Summary component, "What does this mean?" expandables wired to the AI response
Day 12 β Guest β Account Flow
- Banner component (appflow Β§3.5/3.6)
- Sign-up modal β Supabase
linkIdentityupgrade - Post-conversion redirect handling
Day 13 β Dashboard
GET /jobs?user_id=- Dashboard page β job cards, empty state, click-through
Day 14 β Landing Page Polish
- Full design system applied
- Sequence typewriter hero animation
- CTA wiring
Day 15 β Error States & Edge Cases
- Invalid sequence input messaging
- NCBI timeout/failure β
failedstatus + retry - Zero significant hits β its own AI framing ("no strong matches β here's what that can mean")
- Rate-limit queueing if multiple jobs fire close together
Day 16 β Deploy
- Frontend β Vercel, Backend β Railway or Render
- Env vars wired per techspec.md
- Full flow test on deployed URLs
Day 17 β Demo Prep + Buffer
- Finalize 2β3 demo sequences with rich, interesting hits; pre-run and cache them
- Fix whatever broke on Day 16
Day 18 β Demo Day
- Final run-through with cached demo-mode results as backup
- Buffer only β no new features
Track B β Phase 1 Full Build (post-prototype, ~10β12 weeks)
Sprint 1β2: Pairwise Alignment + Pipeline Chaining Foundation β
- Add Clustal Omega pairwise alignment as the second live operation β
- Implement
parent_job_idchaining (schema already supports this) β - Operation grid: two live cards β
Sprint 3β4: UniProt + PDB β
- UniProt annotation lookup β
- PDB structure fetch + Mol* 3D viewer β (3Dmol.js viewer)
- Chain: "from this BLAST hit β fetch structure" β
Sprint 5β6: MSA + Phylogenetic Tree β
- Clustal Omega multi-sequence alignment β (multi-method: ClustalOmega/MUSCLE/Kalign/MAFFT/T-Coffee)
- Basic tree construction + visualization β
- Completes the BLAST β shortlist β MSA β tree workflow from your syllabus mapping β
Sprint 7: Pathway Integration β
- Gene/protein β pathway lookup via Reactome/WikiPathways β
- Pathway diagram viewer β
Sprint 8: Onboarding + /learn β
- First-run tutorial (TutorialWalkthrough component, 5 steps, localStorage flag) β
- 10+ documentation pages at
/learn(BLAST, Alignment, Domains, Phylo, Structure, Pathways, Interactions, Primers, Tools, Glossary) β - LearnPopover component for inline
(?)help tooltips β
Sprint 9β10: Hardening β
- PDF report export (
GET /api/export/job/{id}?format=pdf|json) β - Cache-hit check with
from_cacheflag,/api/admin/cache-statsendpoint β @ttl_cachecoverage: pathway enrichment, NCBI search, BLAST, UniProt, AlphaFold β- Sentry error monitoring (
@sentry/nextjsfrontend +sentry-sdkbackend) β
Sprint 11: Pipeline v2 Engine + Domain/Structure Depth β
- 8-step in-memory pipeline: BLAST β UniProt β MSA β Phylo β Domains β Pathway Enrichment β AlphaFold β AI β
- Pipeline wizard with step checkboxes, progressive reveal results β
- Global/local BLAST modes, DNA query support, poll cap 65 min β
- Pairwise: global/local via "Align pair" from BLAST hits + standalone tool + full-length view β
- Domains: PROSITE raw-sequence scan, reviewed/organism UniProt filters β
- Structure analysis suite: Ramachandran, secondary structure, Foldseek comparison β
Sprint 12: Drug Discovery Compute β
- AutoDock Vina docking: full 1.2.7 log (RMSD l.b./u.b., version, seed), RMSD table, run-config UI β
- MD simulation: verified force field Γ solvent matrix, startup probe, 25-min budget β
- ADMET: RDKit descriptor computation + traffic-light output β
- Function prediction + protein interactions β
Sprint 13: Sequencing MVP β
- FASTQ upload β QC β trimming β assembly/consensus β variant calling β annotation β
- SARS-CoV-2 reference, job persistence, progress polling β
Sprint 14: Reliability β
- AI model fallback chain (Groq β Gemini β Ollama), honest failure banner β
- Share links fixed for all job types + enriched share message β
- Wizard jobs persisted to dashboard history β
- BLAST params honored (program/db/max_hits), long MD jobs not abandoned β
Sprint 15: Design System (v4.0) β
- Dark-only OLED theme, semantic color tokens, AA text tiers, 4-band confidence bands β
- Geist fonts, Phosphor icons, HUD glass components, tiered motion β
- Landing rebuild: DNA-helix hero, bento features, route-style pipeline β
- Dashboard/tools/results converted to native dark tokens β