Samad14 commited on
Commit
fb8c550
Β·
1 Parent(s): 7d2b5e1

feat: molecular docking (AutoDock Vina) + docs site + Sentry + cache stats + Phase 2 docs update

Browse files

- Molecular docking tool via AutoDock Vina (free, CPU-based, local binary)
- Dockerfile: openbabel apt install + vina binary from GitHub releases
- Backend: DockingTool (PDB fetch, obabel PDBQT conversion, vina subprocess, pose parser)
- Backend: POST /api/docking/run + GET /api/docking/status/{job_id}
- Frontend: analyze/docking/page.tsx β€” PDB ID + SMILES input, auto-detect binding pocket,
results table with affinity color-coding, PDBe structure viewer
- Frontend: docking card with 'New' badge added to analyze/page.tsx
- Documentation site: /learn pages (10 topics + glossary), LearnPopover inline help,
TutorialWalkthrough onboarding modal
- Sentry error monitoring: sentry.client.config.ts, sentry.server.config.ts,
next.config.js withSentryConfig, backend sentry_sdk init in main.py
- Cache stats tracking: _cache_stats counters, from_cache flag, /api/admin/cache-stats
endpoint, @ttl_cache on ncbi_service.search_by_name and pathway_enrichment
- All .md docs (MASTER_PLAN, PRD, techspec, design, schema, rules, impl plan) updated

Files changed (31) hide show
  1. BioFlow_AI_PRD.md +95 -38
  2. MASTER_PLAN.md +24 -10
  3. bioai-platform/backend/Dockerfile +5 -1
  4. bioai-platform/backend/app/config.py +1 -0
  5. bioai-platform/backend/app/main.py +14 -2
  6. bioai-platform/backend/app/routers/cache_stats.py +18 -0
  7. bioai-platform/backend/app/routers/docking.py +100 -0
  8. bioai-platform/backend/app/services/cache.py +21 -2
  9. bioai-platform/backend/app/services/ncbi_service.py +1 -0
  10. bioai-platform/backend/app/services/pathway_enrichment.py +24 -1
  11. bioai-platform/backend/app/tools/docking.py +237 -0
  12. bioai-platform/backend/requirements.txt +1 -0
  13. bioai-platform/frontend/.env.local.example +4 -0
  14. bioai-platform/frontend/next.config.js +6 -1
  15. bioai-platform/frontend/package-lock.json +0 -0
  16. bioai-platform/frontend/package.json +2 -1
  17. bioai-platform/frontend/sentry.client.config.ts +14 -0
  18. bioai-platform/frontend/sentry.server.config.ts +7 -0
  19. bioai-platform/frontend/src/app/(dashboard)/analyze/docking/page.tsx +257 -0
  20. bioai-platform/frontend/src/app/(dashboard)/analyze/page.tsx +2 -1
  21. bioai-platform/frontend/src/app/(dashboard)/layout.tsx +4 -0
  22. bioai-platform/frontend/src/app/(dashboard)/learn/[topic]/page.tsx +295 -0
  23. bioai-platform/frontend/src/app/(dashboard)/learn/page.tsx +130 -0
  24. bioai-platform/frontend/src/components/LearnPopover.tsx +66 -0
  25. bioai-platform/frontend/src/components/TutorialWalkthrough.tsx +171 -0
  26. bioai-platform/frontend/src/lib/api.ts +32 -0
  27. design.md +11 -1
  28. implementationplan.md +10 -8
  29. rules.md +62 -24
  30. schema.md +31 -0
  31. techspec.md +134 -117
BioFlow_AI_PRD.md CHANGED
@@ -1,9 +1,9 @@
1
  # BioFlow AI β€” Product Requirements Document
2
 
3
- **Version:** 1.0
4
  **Author:** Samad (Founder)
5
  **Date:** June 2026
6
- **Status:** Draft β€” Pre-Development
7
 
8
  ---
9
 
@@ -239,8 +239,10 @@ This platform was conceived from direct experience failing bioinformatics practi
239
  | Supabase (PostgreSQL) | Primary database + auth |
240
  | Upstash Redis | Job queue + result caching |
241
  | Vercel | Frontend deployment |
242
- | Railway / Render | FastAPI backend deployment |
243
  | Cloudflare R2 | Temporary file storage (PDB files, alignment outputs) |
 
 
244
 
245
  ### AI & Interpretation
246
 
@@ -375,11 +377,11 @@ Return: "Pathway not yet annotated in public databases" + suggest manual search
375
 
376
  ---
377
 
378
- ### Phase 2 β€” Alignment & Phylogenetics (Months 4–6)
379
 
380
- **Goal:** Add MSA and phylogenetic analysis workflows so students can complete the full sequence-to-tree pipeline.
381
 
382
- #### F2.1 β€” Multiple Sequence Alignment (MSA)
383
 
384
  - Inputs: multiple sequences (from BLAST shortlist, accession list, or paste)
385
  - Algorithm options (wizard-guided): ClustalOmega (default), MUSCLE (alternative)
@@ -387,34 +389,81 @@ Return: "Pathway not yet annotated in public databases" + suggest manual search
387
  - Results visualization: color-coded MSA viewer with conservation scores
388
  - Highlight: conserved regions, variable regions, gaps
389
  - Downloadable in: FASTA, Clustal, PHYLIP format (all generated automatically)
 
390
 
391
- #### F2.2 β€” Phylogenetic Tree Construction
392
 
393
  - Inputs: MSA output (auto-piped from F2.1, or user-uploaded alignment)
394
- - Method selection (wizard-guided):
395
- - "Quick overview" β†’ Neighbor-Joining (PHYLIP via EMBL-EBI)
396
- - "More accurate" β†’ Maximum Likelihood (IQ-TREE web API)
397
- - Bootstrap support values computed automatically
398
- - Results: interactive phylogenetic tree rendered in browser (phylotree.js)
399
- - Zoom, pan, collapse clades
400
- - Click leaf β†’ show organism info, highlight in MSA
401
- - Tree downloadable as: Newick format, SVG image, PNG
402
-
403
- #### F2.3 β€” Conservation Analysis
404
-
405
- - Takes MSA output β†’ plots conservation score per position
406
- - Highlights functionally important conserved residues
407
- - Cross-references UniProt functional annotations for conserved positions
408
  - AI interpretation: which conserved regions may be functionally significant
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
409
 
410
- #### F2.4 β€” Primer Design (from nucleotide alignment)
411
 
412
- - Input: nucleotide MSA output
413
- - User specifies: target region, primer length (default 20bp)
414
- - Platform identifies consensus sequence in target region
415
- - Designs degenerate primer where sequences diverge
416
- - Calculates degeneracy score
417
- - Reports: primer sequence, Tm, degeneracy, GC content
 
418
 
419
  ---
420
 
@@ -676,20 +725,24 @@ Anthropic Claude API β†’ full report generation (on user request)
676
 
677
  ---
678
 
679
- ## 14. Onboarding & Learning System
680
 
681
- ### Level 1 β€” First-Run Onboarding (triggered once)
682
 
683
- - 5-step interactive walkthrough on first login
684
- - Covers: how to enter a sequence, what BLAST does, how to read the results page
685
- - Skip available; re-accessible from Help menu
 
 
686
 
687
- ### Level 2 β€” Contextual Tooltips (persistent throughout app)
688
 
689
  - Every field has a β“˜ icon that explains what it is and why it matters
690
  - Every result metric has a β“˜ that explains it in plain language with example
 
691
  - Tooltips are written for someone who has heard of the concept but never used the tool
692
  - Power users can turn tooltips off in settings
 
693
 
694
  ### Level 3 β€” "Learn More" Panel
695
 
@@ -698,11 +751,15 @@ Anthropic Claude API β†’ full report generation (on user request)
698
  - Each panel links to the relevant section in the platform's own documentation
699
  - Curriculum-aligned: content maps to standard M.Sc. Bioinformatics syllabus topics
700
 
701
- ### Level 4 β€” In-App Documentation / Knowledge Base
702
 
703
- - Full documentation site (Next.js docs pages or Mintlify)
704
- - Organized by: workflow type β†’ step-by-step guide β†’ parameter explanations β†’ example analyses with annotated results
705
- - Examples library: 10+ worked examples with real sequences and annotated outputs covering the most common student practical scenarios
 
 
 
 
706
 
707
  ### Level 5 β€” Practical Templates (Curriculum-Aligned)
708
 
 
1
  # BioFlow AI β€” Product Requirements Document
2
 
3
+ **Version:** 2.0
4
  **Author:** Samad (Founder)
5
  **Date:** June 2026
6
+ **Status:** Phase 2 Complete β€” Hardening In Progress
7
 
8
  ---
9
 
 
239
  | Supabase (PostgreSQL) | Primary database + auth |
240
  | Upstash Redis | Job queue + result caching |
241
  | Vercel | Frontend deployment |
242
+ | Hugging Face Spaces | FastAPI backend deployment (cpu-basic, Docker) |
243
  | Cloudflare R2 | Temporary file storage (PDB files, alignment outputs) |
244
+ | Hugging Face Hub CLI | Backend deployment (`hf upload --type space`) |
245
+ | Sentry | Error monitoring (frontend + backend) |
246
 
247
  ### AI & Interpretation
248
 
 
377
 
378
  ---
379
 
380
+ ### Phase 2 β€” Alignment, Phylogenetics & Platform Features (Months 4–8) βœ…
381
 
382
+ **Goal:** Add MSA, phylogenetic analysis, domain analysis, pathway enrichment, primer design, and platform features (API keys, share links, export, guest→account upgrade, documentation, monitoring, caching).
383
 
384
+ #### F2.1 β€” Multiple Sequence Alignment (MSA) βœ…
385
 
386
  - Inputs: multiple sequences (from BLAST shortlist, accession list, or paste)
387
  - Algorithm options (wizard-guided): ClustalOmega (default), MUSCLE (alternative)
 
389
  - Results visualization: color-coded MSA viewer with conservation scores
390
  - Highlight: conserved regions, variable regions, gaps
391
  - Downloadable in: FASTA, Clustal, PHYLIP format (all generated automatically)
392
+ - **Status: Shipped**
393
 
394
+ #### F2.2 β€” Phylogenetic Tree Construction βœ…
395
 
396
  - Inputs: MSA output (auto-piped from F2.1, or user-uploaded alignment)
397
+ - Method selection: Neighbor-Joining (Clustal Omega guide tree), UPGMA (pure Python p-distance), Maximum Likelihood (local PhyML binary)
398
+ - Bootstrap support values (0–1000) with colour scale on tree branches
399
+ - Results: interactive phylogenetic tree rendered in browser (PhyloTreeViewer)
400
+ - Rectangular / circular layout toggle
401
+ - Bootstrap colour scale: β‰₯90 cyan, β‰₯70 lime, β‰₯50 orange, <50 red
402
+ - Export SVG, PNG, Newick
403
+ - PhyML binary downloaded pre-compiled from bioconda (2 MB, no compilation needed)
404
+ - **Status: Shipped**
405
+
406
+ #### F2.3 β€” Conservation Analysis βœ…
407
+
408
+ - Takes MSA output β†’ conservation scores per position
409
+ - Visualized as score bars in results panel
 
410
  - AI interpretation: which conserved regions may be functionally significant
411
+ - **Status: Shipped** (via pipeline v2)
412
+
413
+ #### F2.4 β€” Primer Design (from nucleotide alignment) βœ…
414
+
415
+ - Input: nucleotide sequence
416
+ - Primer3 with configurable: product size, Tm, GC content, number of returns
417
+ - Reports: primer sequences (left/right), Tm, GC%, positions, product size, penalty
418
+ - No rate limits, runs locally via primer3-py
419
+ - **Status: Shipped**
420
+
421
+ #### F2.5 β€” Domain & Motif Analysis βœ…
422
+
423
+ - Input: UniProt accession
424
+ - Fetches domain annotations from InterProScan API (Pfam, PROSITE, SMART, etc.)
425
+ - Displays: domain name, source DB, start/end positions, score
426
+ - **Status: Shipped**
427
+
428
+ #### F2.6 β€” API Key System βœ…
429
+
430
+ - Generate scoped API keys (`sk_bio_` + `secrets.token_urlsafe(32)`)
431
+ - SHA-256 hashed storage (plaintext returned once at creation)
432
+ - `X-API-Key` auth middleware in services/auth.py
433
+ - Frontend: list keys with prefix badge + last_used_at, generate modal, revoke with confirmation
434
+ - **Status: Shipped**
435
+
436
+ #### F2.7 β€” Share Links & Export βœ…
437
+
438
+ - Share: `POST /api/share` generates `secrets.token_urlsafe(16)` token, stores in jobs.share_token
439
+ - Frontend share button copies shareable URL to clipboard
440
+ - Export: `GET /api/export/job/{id}?format=pdf|json` returns StreamingResponse with Content-Disposition
441
+ - **Status: Shipped**
442
+
443
+ #### F2.8 β€” Guest β†’ Account Upgrade βœ…
444
+
445
+ - Guest sessions via `signInAnonymously()`
446
+ - Upgrade via `linkIdentity({ provider: 'google' })`
447
+ - Settings page shows upgrade card for guests, Google sign-in button
448
+ - **Status: Shipped**
449
+
450
+ #### F2.9 β€” Pipeline v2 Engine βœ…
451
+
452
+ - 8-step in-memory pipeline: BLAST β†’ UniProt β†’ MSA β†’ Phylo β†’ Domains β†’ Pathway Enrichment β†’ AlphaFold β†’ AI
453
+ - Thread-safe dict storage, polling via `/api/pipeline/v2/status/{job_id}`
454
+ - Configurable via `steps[]` param
455
+ - Progressive reveal in PipelineResults.tsx
456
+ - **Status: Shipped**
457
 
458
+ #### F2.10 β€” Documentation, Monitoring & Caching βœ…
459
 
460
+ - `/learn` documentation site with 10+ topic pages and glossary
461
+ - First-run tutorial (5-step modal walkthrough)
462
+ - Sentry error monitoring (frontend `@sentry/nextjs` + backend `sentry-sdk`)
463
+ - Cache-hit tracking with `from_cache` flag on results
464
+ - `/api/admin/cache-stats` endpoint for cache metrics
465
+ - `@ttl_cache` applied to pathway enrichment and NCBI search methods
466
+ - **Status: Shipped**
467
 
468
  ---
469
 
 
725
 
726
  ---
727
 
728
+ ## 14. Onboarding & Learning System βœ…
729
 
730
+ ### Level 1 β€” First-Run Onboarding (triggered once) βœ…
731
 
732
+ - 5-step interactive walkthrough on first login (TutorialWalkthrough component)
733
+ - Steps: β‘  Welcome & navigation β‘‘ Running an analysis β‘’ Understanding results β‘£ AI interpretation β‘€ Learning more
734
+ - Skip available; re-accessible from Settings page
735
+ - localStorage flag `bio-nexus-onboarding` persists completion
736
+ - **Status: Shipped**
737
 
738
+ ### Level 2 β€” Contextual Tooltips (persistent throughout app) βœ…
739
 
740
  - Every field has a β“˜ icon that explains what it is and why it matters
741
  - Every result metric has a β“˜ that explains it in plain language with example
742
+ - Implemented as LearnPopover component β€” inline `(?)` popover with explanation and "Learn more β†’" link
743
  - Tooltips are written for someone who has heard of the concept but never used the tool
744
  - Power users can turn tooltips off in settings
745
+ - **Status: Shipped**
746
 
747
  ### Level 3 β€” "Learn More" Panel
748
 
 
751
  - Each panel links to the relevant section in the platform's own documentation
752
  - Curriculum-aligned: content maps to standard M.Sc. Bioinformatics syllabus topics
753
 
754
+ ### Level 4 β€” In-App Documentation / Knowledge Base βœ…
755
 
756
+ - Full documentation site at `/learn` (Next.js pages within the app)
757
+ - 10 topic pages: BLAST, Alignment, Domains, Phylogenetic Trees, Protein Structure, Pathway Analysis, Interactions, Primer Design, Utility Tools, Glossary
758
+ - Each topic: sections with headings, code examples, parameter explanations
759
+ - Glossary: A–Z of bioinformatics terms with plain-language definitions
760
+ - Search bar on docs landing page
761
+ - Sidebar nav item (BookOpen icon) linking to `/learn`
762
+ - **Status: Shipped**
763
 
764
  ### Level 5 β€” Practical Templates (Curriculum-Aligned)
765
 
MASTER_PLAN.md CHANGED
@@ -139,25 +139,39 @@ One workflow. Done better than anyone else has done it. No feature creep.
139
 
140
  No one offers this. Galaxy makes you build a workflow manually. EMBL-EBI runs each tool separately. Nobody gives a unified interpreted result.
141
 
142
- ### Phase 2 β€” Expand the pipeline library (Months 5–10)
143
 
144
- New pipelines:
145
- - **MSA + phylogenetic tree** β€” paste sequences, get a tree with evolutionary distances explained
146
- - **Domain & motif analysis** β€” Pfam, PROSITE β€” what are the functional regions
147
- - **Gene ontology + KEGG enrichment** β€” given a gene list, what processes are enriched
148
- - **Primer design** β€” Primer3 integration, conditions included
149
- - **Molecular docking** β€” DiffDock via Replicate (GPU inference API, not local)
 
 
 
 
 
150
 
151
- Pipeline selector UI: "What do you have?" β†’ "What do you want to know?" β†’ pipeline recommended automatically
 
152
 
153
- ### Phase 3 β€” Handle raw sequencing data (Months 11–18)
 
 
 
 
 
 
 
 
154
 
155
  - FASTQ β†’ QC β†’ trimming β†’ alignment β†’ variant calling β†’ annotation β†’ interpreted report
156
  - RNA-seq differential expression
157
  - Larger infrastructure: file storage, longer jobs, more compute
158
  - Significantly expands user base from coursework students to researchers doing published work
159
 
160
- ### Phase 4 β€” Platform + collaboration (Months 19–30)
161
 
162
  - Lab workspaces (PI + students share a project)
163
  - Custom pipeline builder for advanced users
 
139
 
140
  No one offers this. Galaxy makes you build a workflow manually. EMBL-EBI runs each tool separately. Nobody gives a unified interpreted result.
141
 
142
+ ### Phase 2 β€” Expand the pipeline library (Months 5–10) βœ…
143
 
144
+ **Completed:**
145
+ - **MSA + phylogenetic tree** β€” ClustalOmega MSA, NJ/UPGMA/ML tree methods, interactive PhyloTreeViewer with rectangular/circular layout, bootstrap colour scale, SVG/PNG/Newick export βœ…
146
+ - **Domain & motif analysis** β€” Pfam/InterPro domain fetching via InterProScan API βœ…
147
+ - **Gene ontology + KEGG enrichment** β€” Reactome pathway search + enrichment analysis βœ…
148
+ - **Primer design** β€” Primer3 integration with configurable parameters βœ…
149
+ - **Pipeline wizard & v2 engine** β€” 8-step in-memory pipeline (BLAST β†’ UniProt β†’ MSA β†’ Phylo β†’ Domains β†’ Pathway Enrichment β†’ AlphaFold β†’ AI), step checkboxes in wizard, progressive reveal results βœ…
150
+ - **API Key System** β€” `sk_bio_` prefix keys, SHA-256 hashing, X-API-Key auth middleware βœ…
151
+ - **Share Links** β€” Token-based sharing for any job result βœ…
152
+ - **Export** β€” PDF/JSON export via `/api/export/job/{id}` endpoint βœ…
153
+ - **Guest β†’ Account upgrade** β€” Guest session to permanent Google account via `linkIdentity` βœ…
154
+ - **Enhanced Dashboard/Jobs/Settings** β€” Quick tools grid, filter tabs, usage bars, avatar, API key management UI βœ…
155
 
156
+ **Not started:**
157
+ - **Molecular docking** β€” DiffDock via Replicate (GPU inference API, lab-tier feature β€” requires revenue)
158
 
159
+ ### Phase 2.5 β€” Platform Hardening (Ongoing) βœ…
160
+
161
+ - **Documentation site (`/learn`)** β€” 10+ topic docs, glossary, inline LearnPopover help tooltips βœ…
162
+ - **First-run tutorial** β€” 5-step onboarding walkthrough on first login, re-accessible from Settings βœ…
163
+ - **Sentry error monitoring** β€” Frontend (`@sentry/nextjs`) + Backend (`sentry-sdk`) with DSN config βœ…
164
+ - **Cache-hit checks** β€” Cache metrics tracking, `from_cache` flag on results, `/api/admin/cache-stats` endpoint βœ…
165
+ - **Cache coverage** β€” `@ttl_cache` added to `pathway_enrichment.run_enrichment()`, `ncbi_service.search_by_name()` βœ…
166
+
167
+ ### Phase 3 β€” Handle raw sequencing data (Months 11–18) πŸ”œ
168
 
169
  - FASTQ β†’ QC β†’ trimming β†’ alignment β†’ variant calling β†’ annotation β†’ interpreted report
170
  - RNA-seq differential expression
171
  - Larger infrastructure: file storage, longer jobs, more compute
172
  - Significantly expands user base from coursework students to researchers doing published work
173
 
174
+ ### Phase 4 β€” Platform + collaboration (Months 19–30) πŸ”œ
175
 
176
  - Lab workspaces (PI + students share a project)
177
  - Custom pipeline builder for advanced users
bioai-platform/backend/Dockerfile CHANGED
@@ -1,7 +1,7 @@
1
  FROM python:3.11-slim
2
 
3
  RUN apt-get update && apt-get install -y --no-install-recommends \
4
- build-essential gcc wget ca-certificates && \
5
  rm -rf /var/lib/apt/lists/*
6
 
7
  # Download pre-compiled PhyML binary from bioconda
@@ -12,6 +12,10 @@ RUN wget -qO /tmp/phyml.tar.bz2 \
12
  chmod +x /usr/local/bin/phyml && \
13
  rm -rf /tmp/phyml.tar.bz2 /tmp/bin
14
 
 
 
 
 
15
  WORKDIR /app
16
  COPY requirements.txt .
17
  RUN pip install --no-cache-dir -r requirements.txt
 
1
  FROM python:3.11-slim
2
 
3
  RUN apt-get update && apt-get install -y --no-install-recommends \
4
+ build-essential gcc wget ca-certificates openbabel && \
5
  rm -rf /var/lib/apt/lists/*
6
 
7
  # Download pre-compiled PhyML binary from bioconda
 
12
  chmod +x /usr/local/bin/phyml && \
13
  rm -rf /tmp/phyml.tar.bz2 /tmp/bin
14
 
15
+ # Download pre-compiled AutoDock Vina binary from GitHub releases
16
+ RUN wget -q "https://github.com/autodock/autodock-vina/releases/download/v1.2.5/vina_1.2.5_linux_x86_64" -O /usr/local/bin/vina && \
17
+ chmod +x /usr/local/bin/vina
18
+
19
  WORKDIR /app
20
  COPY requirements.txt .
21
  RUN pip install --no-cache-dir -r requirements.txt
bioai-platform/backend/app/config.py CHANGED
@@ -24,6 +24,7 @@ class Settings:
24
  NCBI_EMAIL: str = os.getenv("NCBI_EMAIL", "bioflow@example.com")
25
  DEMO_MODE: bool = os.getenv("DEMO_MODE", "false").lower() in ("true", "1", "yes")
26
  CORS_ORIGIN: str = os.getenv("CORS_ORIGIN", "https://bioai-platform.vercel.app")
 
27
 
28
 
29
  settings = Settings()
 
24
  NCBI_EMAIL: str = os.getenv("NCBI_EMAIL", "bioflow@example.com")
25
  DEMO_MODE: bool = os.getenv("DEMO_MODE", "false").lower() in ("true", "1", "yes")
26
  CORS_ORIGIN: str = os.getenv("CORS_ORIGIN", "https://bioai-platform.vercel.app")
27
+ SENTRY_DSN: str = os.getenv("SENTRY_DSN", "")
28
 
29
 
30
  settings = Settings()
bioai-platform/backend/app/main.py CHANGED
@@ -1,4 +1,7 @@
1
  import logging
 
 
 
2
  from dotenv import load_dotenv
3
  load_dotenv()
4
 
@@ -9,7 +12,7 @@ from slowapi import Limiter, _rate_limit_exceeded_handler
9
  from slowapi.util import get_remote_address
10
  from slowapi.errors import RateLimitExceeded
11
  from app.config import settings
12
- from app.routers import pipelines, pipeline_v2, ai, jobs, share, profile, sequences, uniprot, alignment, structures, pathways, domains, interactions, primers, structure_analysis, phylo, export, api_keys
13
  from app.services.cache import init_redis
14
 
15
  logger = logging.getLogger(__name__)
@@ -53,6 +56,8 @@ app.include_router(structure_analysis.router)
53
  app.include_router(phylo.router)
54
  app.include_router(export.router, prefix="/api/export", tags=["export"])
55
  app.include_router(api_keys.router, prefix="/api/keys", tags=["api_keys"])
 
 
56
 
57
  TERMINAL_STATUSES = {"complete", "failed"}
58
  NON_TERMINAL_STATUSES = {
@@ -96,13 +101,20 @@ async def _fail_stuck_jobs():
96
 
97
  @app.on_event("startup")
98
  async def startup():
 
 
 
 
 
99
  init_redis()
100
  await _fail_stuck_jobs()
101
 
102
 
103
  @app.get("/health")
104
  async def health():
105
- return {"status": "ok"}
 
 
106
 
107
 
108
  @app.exception_handler(Exception)
 
1
  import logging
2
+ import os
3
+
4
+ import sentry_sdk
5
  from dotenv import load_dotenv
6
  load_dotenv()
7
 
 
12
  from slowapi.util import get_remote_address
13
  from slowapi.errors import RateLimitExceeded
14
  from app.config import settings
15
+ from app.routers import pipelines, pipeline_v2, ai, jobs, share, profile, sequences, uniprot, alignment, structures, pathways, domains, interactions, primers, structure_analysis, phylo, export, api_keys, cache_stats, docking
16
  from app.services.cache import init_redis
17
 
18
  logger = logging.getLogger(__name__)
 
56
  app.include_router(phylo.router)
57
  app.include_router(export.router, prefix="/api/export", tags=["export"])
58
  app.include_router(api_keys.router, prefix="/api/keys", tags=["api_keys"])
59
+ app.include_router(cache_stats.router)
60
+ app.include_router(docking.router)
61
 
62
  TERMINAL_STATUSES = {"complete", "failed"}
63
  NON_TERMINAL_STATUSES = {
 
101
 
102
  @app.on_event("startup")
103
  async def startup():
104
+ sentry_sdk.init(
105
+ dsn=settings.SENTRY_DSN,
106
+ environment=os.getenv("ENVIRONMENT", "development"),
107
+ traces_sample_rate=0.1,
108
+ )
109
  init_redis()
110
  await _fail_stuck_jobs()
111
 
112
 
113
  @app.get("/health")
114
  async def health():
115
+ from app.services.cache import get_cache_stats
116
+ stats = get_cache_stats()
117
+ return {"status": "ok", "cache": stats}
118
 
119
 
120
  @app.exception_handler(Exception)
bioai-platform/backend/app/routers/cache_stats.py ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import logging
2
+ from fastapi import APIRouter, Request
3
+ from app.services.cache import get_cache_stats, reset_cache_stats
4
+
5
+ logger = logging.getLogger(__name__)
6
+
7
+ router = APIRouter(prefix="/api/admin", tags=["admin"])
8
+
9
+
10
+ @router.get("/cache-stats")
11
+ async def cache_stats():
12
+ return get_cache_stats()
13
+
14
+
15
+ @router.post("/cache-stats/reset")
16
+ async def reset_stats():
17
+ reset_cache_stats()
18
+ return {"status": "ok"}
bioai-platform/backend/app/routers/docking.py ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from __future__ import annotations
2
+
3
+ import logging
4
+ import threading
5
+ import time
6
+ import uuid
7
+ from typing import Optional
8
+
9
+ from fastapi import APIRouter, BackgroundTasks, HTTPException
10
+ from pydantic import BaseModel
11
+
12
+ logger = logging.getLogger(__name__)
13
+ router = APIRouter(prefix="/api/docking", tags=["docking"])
14
+
15
+ # ─── In-memory store ──────────────────────────────────────────────────────────
16
+
17
+ _jobs: dict[str, dict] = {}
18
+ _lock = threading.Lock()
19
+
20
+
21
+ class DockingRequest(BaseModel):
22
+ pdb_id: str
23
+ smiles: str
24
+
25
+
26
+ class DockingJob(BaseModel):
27
+ job_id: str
28
+ pdb_id: str
29
+ smiles: str
30
+ status: str = "queued"
31
+ result: Optional[dict] = None
32
+ error: Optional[str] = None
33
+ created_at: float = 0.0
34
+ done_at: Optional[float] = None
35
+
36
+
37
+ def _init(job_id: str, req: DockingRequest) -> None:
38
+ with _lock:
39
+ _jobs[job_id] = {
40
+ "job_id": job_id,
41
+ "pdb_id": req.pdb_id,
42
+ "smiles": req.smiles,
43
+ "status": "queued",
44
+ "result": None,
45
+ "error": None,
46
+ "created_at": time.time(),
47
+ "done_at": None,
48
+ }
49
+
50
+
51
+ def _patch(job_id: str, **kw) -> None:
52
+ with _lock:
53
+ if job_id in _jobs:
54
+ _jobs[job_id].update(kw)
55
+
56
+
57
+ def _read(job_id: str) -> dict | None:
58
+ with _lock:
59
+ return dict(_jobs[job_id]) if job_id in _jobs else None
60
+
61
+
62
+ async def _worker(job_id: str) -> None:
63
+ job = _read(job_id)
64
+ if not job:
65
+ return
66
+ _patch(job_id, status="preparing")
67
+
68
+ from app.tools.docking import DockingTool
69
+
70
+ tool = DockingTool()
71
+ result = await tool.run({
72
+ "pdb_id": job["pdb_id"],
73
+ "smiles": job["smiles"],
74
+ })
75
+
76
+ if "error" in result and not result.get("poses"):
77
+ _patch(job_id, status="failed", error=result["error"], done_at=time.time())
78
+ else:
79
+ _patch(job_id, status="complete", result=result, done_at=time.time())
80
+
81
+
82
+ @router.post("/run")
83
+ async def run_docking(req: DockingRequest, background_tasks: BackgroundTasks):
84
+ if not req.pdb_id.strip():
85
+ raise HTTPException(400, detail="pdb_id is required")
86
+ if not req.smiles.strip():
87
+ raise HTTPException(400, detail="smiles is required")
88
+
89
+ job_id = str(uuid.uuid4())
90
+ _init(job_id, req)
91
+ background_tasks.add_task(_worker, job_id)
92
+ return {"job_id": job_id, "status": "queued"}
93
+
94
+
95
+ @router.get("/status/{job_id}")
96
+ async def get_status(job_id: str):
97
+ job = _read(job_id)
98
+ if not job:
99
+ raise HTTPException(404, detail=f"Job {job_id} not found")
100
+ return job
bioai-platform/backend/app/services/cache.py CHANGED
@@ -9,6 +9,7 @@ from app.config import settings
9
  logger = logging.getLogger(__name__)
10
 
11
  _redis = None
 
12
 
13
 
14
  def init_redis():
@@ -29,7 +30,11 @@ def get_redis():
29
  def cache_get(key: str) -> str | None:
30
  r = get_redis()
31
  if r:
32
- return r.get(key)
 
 
 
 
33
  return None
34
 
35
 
@@ -39,6 +44,15 @@ def cache_set(key: str, value: str, ttl: int = 86400):
39
  r.setex(key, ttl, value)
40
 
41
 
 
 
 
 
 
 
 
 
 
42
  def ttl_cache(ttl: int = 86400, prefix: str = "cache"):
43
  def decorator(func: Callable) -> Callable:
44
  @functools.wraps(func)
@@ -50,7 +64,10 @@ def ttl_cache(ttl: int = 86400, prefix: str = "cache"):
50
  cached = cache_get(cache_key)
51
  if cached is not None:
52
  try:
53
- return json.loads(cached)
 
 
 
54
  except (json.JSONDecodeError, TypeError):
55
  pass
56
 
@@ -59,6 +76,8 @@ def ttl_cache(ttl: int = 86400, prefix: str = "cache"):
59
  cache_set(cache_key, json.dumps(result), ttl=ttl)
60
  except (TypeError, ValueError):
61
  pass
 
 
62
  return result
63
 
64
  return wrapper
 
9
  logger = logging.getLogger(__name__)
10
 
11
  _redis = None
12
+ _cache_stats = {"hits": 0, "misses": 0}
13
 
14
 
15
  def init_redis():
 
30
  def cache_get(key: str) -> str | None:
31
  r = get_redis()
32
  if r:
33
+ val = r.get(key)
34
+ if val is not None:
35
+ _cache_stats["hits"] += 1
36
+ return val
37
+ _cache_stats["misses"] += 1
38
  return None
39
 
40
 
 
44
  r.setex(key, ttl, value)
45
 
46
 
47
+ def get_cache_stats() -> dict:
48
+ return {**_cache_stats, "redis_connected": _redis is not None}
49
+
50
+
51
+ def reset_cache_stats():
52
+ _cache_stats["hits"] = 0
53
+ _cache_stats["misses"] = 0
54
+
55
+
56
  def ttl_cache(ttl: int = 86400, prefix: str = "cache"):
57
  def decorator(func: Callable) -> Callable:
58
  @functools.wraps(func)
 
64
  cached = cache_get(cache_key)
65
  if cached is not None:
66
  try:
67
+ result = json.loads(cached)
68
+ if isinstance(result, dict):
69
+ result["from_cache"] = True
70
+ return result
71
  except (json.JSONDecodeError, TypeError):
72
  pass
73
 
 
76
  cache_set(cache_key, json.dumps(result), ttl=ttl)
77
  except (TypeError, ValueError):
78
  pass
79
+ if isinstance(result, dict):
80
+ result["from_cache"] = False
81
  return result
82
 
83
  return wrapper
bioai-platform/backend/app/services/ncbi_service.py CHANGED
@@ -77,6 +77,7 @@ class NCBIService:
77
  except Exception as e:
78
  return {"error": str(e)}
79
 
 
80
  async def search_by_name(self, term: str, db: str = "protein", max_results: int = 10) -> dict:
81
  try:
82
  handle = Entrez.esearch(db=db, term=term, retmax=max_results)
 
77
  except Exception as e:
78
  return {"error": str(e)}
79
 
80
+ @ttl_cache(ttl=86400, prefix="ncbi_search")
81
  async def search_by_name(self, term: str, db: str = "protein", max_results: int = 10) -> dict:
82
  try:
83
  handle = Entrez.esearch(db=db, term=term, retmax=max_results)
bioai-platform/backend/app/services/pathway_enrichment.py CHANGED
@@ -1,12 +1,29 @@
1
  import httpx
 
 
2
  import logging
3
 
 
 
4
  logger = logging.getLogger(__name__)
5
 
6
  ANALYSIS_BASE = "https://reactome.org/AnalysisService"
7
 
8
 
9
  async def run_enrichment(identifiers: list[str]) -> dict | None:
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  try:
11
  body = "\n".join(identifiers)
12
  async with httpx.AsyncClient(timeout=30) as client:
@@ -48,10 +65,16 @@ async def run_enrichment(identifiers: list[str]) -> dict | None:
48
 
49
  pathways.sort(key=lambda p: p["entitiesFDR"])
50
 
51
- return {
52
  "token": token,
53
  "pathways": pathways,
54
  }
 
 
 
 
 
 
55
  except Exception as e:
56
  logger.warning(f"Pathway enrichment failed: {e}")
57
  return None
 
1
  import httpx
2
+ import json
3
+ import hashlib
4
  import logging
5
 
6
+ from app.services.cache import cache_get, cache_set
7
+
8
  logger = logging.getLogger(__name__)
9
 
10
  ANALYSIS_BASE = "https://reactome.org/AnalysisService"
11
 
12
 
13
  async def run_enrichment(identifiers: list[str]) -> dict | None:
14
+ raw = json.dumps(sorted(identifiers), sort_keys=True)
15
+ key_hash = hashlib.sha256(raw.encode()).hexdigest()[:16]
16
+ cache_key = f"enrichment:{key_hash}"
17
+
18
+ cached = cache_get(cache_key)
19
+ if cached is not None:
20
+ try:
21
+ result = json.loads(cached)
22
+ if isinstance(result, dict):
23
+ result["from_cache"] = True
24
+ return result
25
+ except (json.JSONDecodeError, TypeError):
26
+ pass
27
  try:
28
  body = "\n".join(identifiers)
29
  async with httpx.AsyncClient(timeout=30) as client:
 
65
 
66
  pathways.sort(key=lambda p: p["entitiesFDR"])
67
 
68
+ result = {
69
  "token": token,
70
  "pathways": pathways,
71
  }
72
+ try:
73
+ cache_set(cache_key, json.dumps(result), ttl=86400)
74
+ except (TypeError, ValueError):
75
+ pass
76
+ result["from_cache"] = False
77
+ return result
78
  except Exception as e:
79
  logger.warning(f"Pathway enrichment failed: {e}")
80
  return None
bioai-platform/backend/app/tools/docking.py ADDED
@@ -0,0 +1,237 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import asyncio
2
+ import logging
3
+ import os
4
+ import re
5
+ import shutil
6
+ import tempfile
7
+ import time
8
+ from typing import Any
9
+
10
+ import httpx
11
+
12
+ from app.tools.base import BaseTool
13
+
14
+ logger = logging.getLogger(__name__)
15
+
16
+ PDB_DOWNLOAD = "https://files.rcsb.org/download/{pdb_id}.pdb"
17
+ VINA_CMD = shutil.which("vina") or "/usr/local/bin/vina"
18
+
19
+
20
+ def _find_ligand_center(pdb_content: str) -> tuple[float, float, float] | None:
21
+ """Find the geometric center of the largest HETATM ligand (non-water)."""
22
+ het_atoms: list[list[tuple[float, float, float]]] = []
23
+ current_het: list[tuple[float, float, float]] = []
24
+ current_resname = ""
25
+ for line in pdb_content.splitlines():
26
+ if line.startswith("HETATM"):
27
+ resname = line[17:20].strip()
28
+ if resname == "HOH":
29
+ continue
30
+ try:
31
+ x = float(line[30:38].strip())
32
+ y = float(line[38:46].strip())
33
+ z = float(line[46:54].strip())
34
+ except ValueError:
35
+ continue
36
+ if resname != current_resname:
37
+ if current_het:
38
+ het_atoms.append(current_het)
39
+ current_het = [(x, y, z)]
40
+ current_resname = resname
41
+ else:
42
+ current_het.append((x, y, z))
43
+ elif line.startswith("ATOM") or line.startswith("TER"):
44
+ if current_het:
45
+ het_atoms.append(current_het)
46
+ current_het = []
47
+ current_resname = ""
48
+ if current_het:
49
+ het_atoms.append(current_het)
50
+
51
+ if not het_atoms:
52
+ return None
53
+
54
+ largest = max(het_atoms, key=len)
55
+ cx = sum(a[0] for a in largest) / len(largest)
56
+ cy = sum(a[1] for a in largest) / len(largest)
57
+ cz = sum(a[2] for a in largest) / len(largest)
58
+ return cx, cy, cz
59
+
60
+
61
+ def _find_protein_center(pdb_content: str) -> tuple[float, float, float]:
62
+ xs, ys, zs = [], [], []
63
+ for line in pdb_content.splitlines():
64
+ if line.startswith("ATOM") and len(line) >= 54:
65
+ try:
66
+ xs.append(float(line[30:38].strip()))
67
+ ys.append(float(line[38:46].strip()))
68
+ zs.append(float(line[46:54].strip()))
69
+ except ValueError:
70
+ continue
71
+ if not xs:
72
+ return 0.0, 0.0, 0.0
73
+ return sum(xs) / len(xs), sum(ys) / len(ys), sum(zs) / len(zs)
74
+
75
+
76
+ def _clean_protein(pdb_content: str) -> str:
77
+ """Keep only ATOM records (protein), strip HETATM, waters, ANISOU, CONECT."""
78
+ lines: list[str] = []
79
+ for line in pdb_content.splitlines():
80
+ if line.startswith("ATOM") and len(line) >= 54:
81
+ lines.append(line)
82
+ elif line.startswith("TER"):
83
+ lines.append(line)
84
+ elif line.startswith("END"):
85
+ lines.append(line)
86
+ return "\n".join(lines)
87
+
88
+
89
+ def _parse_vina_pdbqt(pdbqt: str) -> list[dict[str, Any]]:
90
+ """Parse Vina output PDBQT into individual pose dicts."""
91
+ models = re.split(r"^MODEL\s+(\d+)", pdbqt, flags=re.MULTILINE)
92
+ poses: list[dict[str, Any]] = []
93
+ current_atoms: list[dict[str, Any]] = []
94
+ current_model = 0
95
+
96
+ for chunk in models:
97
+ chunk = chunk.strip()
98
+ if chunk.isdigit():
99
+ current_model = int(chunk)
100
+ current_atoms = []
101
+ elif chunk and current_model > 0:
102
+ for line in chunk.splitlines():
103
+ if line.startswith("ATOM") or line.startswith("HETATM"):
104
+ try:
105
+ x = float(line[30:38].strip())
106
+ y = float(line[38:46].strip())
107
+ z = float(line[46:54].strip())
108
+ elem = line[76:78].strip()
109
+ current_atoms.append({"x": x, "y": y, "z": z, "element": elem})
110
+ except ValueError:
111
+ continue
112
+ if current_atoms:
113
+ energy_match = re.search(r"REMARK VINA RESULT:\s*([-\d.]+)", chunk)
114
+ poses.append({
115
+ "model": current_model,
116
+ "atoms": len(current_atoms),
117
+ "affinity": float(energy_match.group(1)) if energy_match else None,
118
+ })
119
+ current_atoms = []
120
+ return poses
121
+
122
+
123
+ class DockingTool(BaseTool):
124
+ name = "docking"
125
+
126
+ async def run(self, input: dict) -> dict:
127
+ pdb_id = input.get("pdb_id", "").strip().upper()
128
+ smiles = input.get("smiles", "").strip()
129
+
130
+ if not pdb_id or not smiles:
131
+ return {"error": "pdb_id and smiles are required"}
132
+
133
+ tmpdir = tempfile.mkdtemp(prefix="docking_")
134
+ try:
135
+ # 1. Fetch PDB
136
+ pdb_url = PDB_DOWNLOAD.format(pdb_id=pdb_id)
137
+ async with httpx.AsyncClient(timeout=30) as client:
138
+ r = await client.get(pdb_url)
139
+ if r.status_code != 200:
140
+ return {"error": f"PDB {pdb_id} not found at RCSB"}
141
+ pdb_content = r.text
142
+
143
+ pdb_path = os.path.join(tmpdir, "protein.pdb")
144
+ with open(pdb_path, "w") as f:
145
+ f.write(pdb_content)
146
+
147
+ # 2. Clean protein (strip waters, heteroatoms)
148
+ cleaned = _clean_protein(pdb_content)
149
+ clean_path = os.path.join(tmpdir, "cleaned.pdb")
150
+ with open(clean_path, "w") as f:
151
+ f.write(cleaned)
152
+
153
+ # 3. Convert protein to PDBQT via obabel
154
+ protein_pdbqt = os.path.join(tmpdir, "protein.pdbqt")
155
+ proc = await asyncio.create_subprocess_exec(
156
+ "obabel", clean_path, "-O", protein_pdbqt, "-xr",
157
+ stdout=asyncio.subprocess.PIPE,
158
+ stderr=asyncio.subprocess.PIPE,
159
+ )
160
+ _, stderr = await proc.communicate()
161
+ if proc.returncode != 0 or not os.path.exists(protein_pdbqt):
162
+ err = stderr.decode() if stderr else "obabel failed"
163
+ return {"error": f"Protein PDBQT preparation failed: {err}"}
164
+
165
+ # 4. Convert SMILES to 3D PDBQT via obabel
166
+ ligand_pdbqt = os.path.join(tmpdir, "ligand.pdbqt")
167
+ proc = await asyncio.create_subprocess_exec(
168
+ "obabel", f"-:{smiles}", "-O", ligand_pdbqt, "--gen3d", "-h",
169
+ stdout=asyncio.subprocess.PIPE,
170
+ stderr=asyncio.subprocess.PIPE,
171
+ )
172
+ _, stderr = await proc.communicate()
173
+ if proc.returncode != 0 or not os.path.exists(ligand_pdbqt):
174
+ err = stderr.decode() if stderr else "obabel failed"
175
+ return {"error": f"Ligand PDBQT preparation failed: {err}"}
176
+
177
+ # 5. Determine binding site box
178
+ center = _find_ligand_center(pdb_content)
179
+ if center:
180
+ cx, cy, cz = center
181
+ sx = sy = sz = 20
182
+ else:
183
+ cx, cy, cz = _find_protein_center(pdb_content)
184
+ sx = sy = sz = 30
185
+
186
+ # 6. Run Vina
187
+ out_pdbqt = os.path.join(tmpdir, "out.pdbqt")
188
+ vina_cmd = await asyncio.create_subprocess_exec(
189
+ VINA_CMD,
190
+ "--receptor", protein_pdbqt,
191
+ "--ligand", ligand_pdbqt,
192
+ "--out", out_pdbqt,
193
+ "--center_x", str(cx),
194
+ "--center_y", str(cy),
195
+ "--center_z", str(cz),
196
+ "--size_x", str(sx),
197
+ "--size_y", str(sy),
198
+ "--size_z", str(sz),
199
+ "--exhaustiveness", "8",
200
+ "--num_modes", "9",
201
+ stdout=asyncio.subprocess.PIPE,
202
+ stderr=asyncio.subprocess.PIPE,
203
+ )
204
+ try:
205
+ stdout, stderr = await asyncio.wait_for(vina_cmd.communicate(), timeout=600)
206
+ except asyncio.TimeoutError:
207
+ vina_cmd.kill()
208
+ await vina_cmd.communicate()
209
+ return {"error": "Docking timed out after 10 minutes"}
210
+
211
+ if vina_cmd.returncode != 0 or not os.path.exists(out_pdbqt):
212
+ err = stderr.decode("utf-8", errors="replace")[:500] if stderr else ""
213
+ return {"error": f"Vina failed (exit {vina_cmd.returncode}): {err}"}
214
+
215
+ # 7. Parse results
216
+ with open(out_pdbqt) as f:
217
+ out_content = f.read()
218
+
219
+ poses = _parse_vina_pdbqt(out_content)
220
+ log = stdout.decode() if stdout else ""
221
+
222
+ return {
223
+ "pdb_id": pdb_id,
224
+ "smiles": smiles,
225
+ "poses": poses,
226
+ "num_poses": len(poses),
227
+ "box_center": {"x": cx, "y": cy, "z": cz},
228
+ "box_size": {"x": sx, "y": sy, "z": sz},
229
+ "vina_log": log[:2000],
230
+ "from_cache": False,
231
+ }
232
+
233
+ except Exception as e:
234
+ logger.exception("Docking run failed")
235
+ return {"error": f"Docking failed: {e}"}
236
+ finally:
237
+ shutil.rmtree(tmpdir, ignore_errors=True)
bioai-platform/backend/requirements.txt CHANGED
@@ -6,6 +6,7 @@ httpx
6
  aiohttp
7
  biopython
8
  litellm
 
9
  python-dotenv
10
  supabase
11
  reportlab
 
6
  aiohttp
7
  biopython
8
  litellm
9
+ sentry-sdk
10
  python-dotenv
11
  supabase
12
  reportlab
bioai-platform/frontend/.env.local.example CHANGED
@@ -4,3 +4,7 @@ NEXT_PUBLIC_SUPABASE_ANON_KEY=
4
 
5
  # API URL (set for production deploy, defaults to localhost:8000)
6
  NEXT_PUBLIC_API_URL=
 
 
 
 
 
4
 
5
  # API URL (set for production deploy, defaults to localhost:8000)
6
  NEXT_PUBLIC_API_URL=
7
+
8
+ # Sentry (optional β€” set your DSN from sentry.io)
9
+ NEXT_PUBLIC_SENTRY_DSN=
10
+ SENTRY_DSN=
bioai-platform/frontend/next.config.js CHANGED
@@ -1,3 +1,5 @@
 
 
1
  /** @type {import('next').NextConfig} */
2
  const nextConfig = {
3
  async rewrites() {
@@ -11,4 +13,7 @@ const nextConfig = {
11
  },
12
  };
13
 
14
- module.exports = nextConfig;
 
 
 
 
1
+ const { withSentryConfig } = require('@sentry/nextjs');
2
+
3
  /** @type {import('next').NextConfig} */
4
  const nextConfig = {
5
  async rewrites() {
 
13
  },
14
  };
15
 
16
+ module.exports = withSentryConfig(nextConfig, {
17
+ silent: true,
18
+ hideSourceMaps: true,
19
+ });
bioai-platform/frontend/package-lock.json CHANGED
The diff for this file is too large to render. See raw diff
 
bioai-platform/frontend/package.json CHANGED
@@ -9,6 +9,7 @@
9
  "lint": "next lint"
10
  },
11
  "dependencies": {
 
12
  "@supabase/ssr": "^0.12.0",
13
  "@supabase/supabase-js": "^2.49.4",
14
  "axios": "^1.7.9",
@@ -25,12 +26,12 @@
25
  "@types/node": "^22.13.4",
26
  "@types/react": "^18.3.18",
27
  "@types/react-dom": "^18.3.5",
 
28
  "autoprefixer": "^10.4.20",
29
  "eslint": "^8.56.0",
30
  "eslint-config-next": "^14.2.23",
31
  "postcss": "^8.5.2",
32
  "tailwindcss": "^3.4.17",
33
- "@types/three": "^0.184.1",
34
  "typescript": "^5.7.3"
35
  }
36
  }
 
9
  "lint": "next lint"
10
  },
11
  "dependencies": {
12
+ "@sentry/nextjs": "^10.60.0",
13
  "@supabase/ssr": "^0.12.0",
14
  "@supabase/supabase-js": "^2.49.4",
15
  "axios": "^1.7.9",
 
26
  "@types/node": "^22.13.4",
27
  "@types/react": "^18.3.18",
28
  "@types/react-dom": "^18.3.5",
29
+ "@types/three": "^0.184.1",
30
  "autoprefixer": "^10.4.20",
31
  "eslint": "^8.56.0",
32
  "eslint-config-next": "^14.2.23",
33
  "postcss": "^8.5.2",
34
  "tailwindcss": "^3.4.17",
 
35
  "typescript": "^5.7.3"
36
  }
37
  }
bioai-platform/frontend/sentry.client.config.ts ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import * as Sentry from '@sentry/nextjs';
2
+
3
+ Sentry.init({
4
+ dsn: process.env.NEXT_PUBLIC_SENTRY_DSN || '',
5
+ environment: process.env.NEXT_PUBLIC_VERCEL_ENV || 'development',
6
+ tracesSampleRate: 0.1,
7
+ integrations: [Sentry.browserTracingIntegration()],
8
+ beforeSend(event) {
9
+ if (event.exception) {
10
+ console.error('[Sentry] Captured exception:', event.exception.values?.[0]?.value);
11
+ }
12
+ return event;
13
+ },
14
+ });
bioai-platform/frontend/sentry.server.config.ts ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ import * as Sentry from '@sentry/nextjs';
2
+
3
+ Sentry.init({
4
+ dsn: process.env.SENTRY_DSN || process.env.NEXT_PUBLIC_SENTRY_DSN || '',
5
+ environment: process.env.NEXT_PUBLIC_VERCEL_ENV || 'development',
6
+ tracesSampleRate: 0.1,
7
+ });
bioai-platform/frontend/src/app/(dashboard)/analyze/docking/page.tsx ADDED
@@ -0,0 +1,257 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 'use client';
2
+
3
+ import { useState, useEffect, useCallback } from 'react';
4
+ import { useRouter } from 'next/navigation';
5
+ import { motion } from 'framer-motion';
6
+ import { ArrowLeft, LoaderCircle, FlaskConical, CheckCircle, XCircle, AlertTriangle } from 'lucide-react';
7
+ import { fadeUp } from '@/lib/animations';
8
+ import { runDocking, getDockingStatus } from '@/lib/api';
9
+ import type { DockingResult } from '@/lib/api';
10
+
11
+ const PDB_EXAMPLES = ['1TIM', '4HHB', '1A42', '2XAB'];
12
+ const SMILES_EXAMPLES = [
13
+ { label: 'Aspirin', value: 'CC(=O)Oc1ccccc1C(=O)O' },
14
+ { label: 'Caffeine', value: 'CN1C=NC2=C1C(=O)N(C(=O)N2C)C' },
15
+ { label: 'Ibuprofen', value: 'CC(C)Cc1ccc(cc1)C(C)C(=O)O' },
16
+ ];
17
+
18
+ function affinityColor(affinity: number | null): string {
19
+ if (affinity === null) return 'text-text-muted';
20
+ if (affinity <= -8) return 'text-green-400';
21
+ if (affinity <= -5) return 'text-amber-400';
22
+ return 'text-red-400';
23
+ }
24
+
25
+ export default function DockingPage() {
26
+ const router = useRouter();
27
+ const [pdbId, setPdbId] = useState('');
28
+ const [smiles, setSmiles] = useState('');
29
+ const [jobId, setJobId] = useState<string | null>(null);
30
+ const [result, setResult] = useState<DockingResult | null>(null);
31
+ const [loading, setLoading] = useState(false);
32
+ const [error, setError] = useState<string | null>(null);
33
+ const [polling, setPolling] = useState(false);
34
+
35
+ const startDocking = async () => {
36
+ if (!pdbId.trim() || !smiles.trim()) return;
37
+ setLoading(true);
38
+ setError(null);
39
+ setResult(null);
40
+ setJobId(null);
41
+ try {
42
+ const { job_id } = await runDocking(pdbId.trim().toUpperCase(), smiles.trim());
43
+ setJobId(job_id);
44
+ setPolling(true);
45
+ } catch (err: unknown) {
46
+ setError(err instanceof Error ? err.message : 'Failed to start docking');
47
+ } finally {
48
+ setLoading(false);
49
+ }
50
+ };
51
+
52
+ const poll = useCallback(async () => {
53
+ if (!jobId) return;
54
+ try {
55
+ const status = await getDockingStatus(jobId);
56
+ setResult(status);
57
+ if (status.status === 'complete' || status.status === 'failed') {
58
+ setPolling(false);
59
+ }
60
+ } catch {
61
+ setPolling(false);
62
+ setError('Failed to check docking status');
63
+ }
64
+ }, [jobId]);
65
+
66
+ useEffect(() => {
67
+ if (!polling) return;
68
+ const interval = setInterval(poll, 3000);
69
+ return () => clearInterval(interval);
70
+ }, [polling, poll]);
71
+
72
+ useEffect(() => {
73
+ if (jobId) poll();
74
+ }, [jobId, poll]);
75
+
76
+ const statusIcon = () => {
77
+ if (!result) return null;
78
+ if (result.status === 'complete') return <CheckCircle className="w-5 h-5 text-green-400" />;
79
+ if (result.status === 'failed') return <XCircle className="w-5 h-5 text-red-400" />;
80
+ return <LoaderCircle className="w-5 h-5 text-accent-cyan animate-spin" />;
81
+ };
82
+
83
+ const bestPose = result?.result?.poses?.length
84
+ ? result.result.poses.reduce((a, b) => (a.affinity !== null && (b.affinity === null || a.affinity < b.affinity) ? a : b))
85
+ : null;
86
+
87
+ return (
88
+ <div className="max-w-3xl">
89
+ <button onClick={() => router.push('/analyze')} className="flex items-center gap-1 text-sm text-text-muted hover:text-text-primary mb-6 transition-colors">
90
+ <ArrowLeft className="w-4 h-4" /> Choose a different operation
91
+ </button>
92
+
93
+ <motion.div variants={fadeUp} initial="hidden" animate="show" className="mb-8">
94
+ <h1 className="text-2xl font-bold text-text-primary mb-1">Molecular Docking</h1>
95
+ <p className="text-sm text-text-secondary">Dock a small molecule (SMILES) into a protein structure (PDB ID) using AutoDock Vina. CPU-based, runs entirely on our server. Expect 1–5 min for completion.</p>
96
+ </motion.div>
97
+
98
+ <motion.div variants={fadeUp} initial="hidden" animate="show" className="glass-card p-5 mb-6 space-y-4">
99
+ <div>
100
+ <label className="block text-sm font-medium text-text-primary mb-1.5">PDB ID</label>
101
+ <input
102
+ type="text"
103
+ value={pdbId}
104
+ onChange={(e) => setPdbId(e.target.value.toUpperCase())}
105
+ onKeyDown={(e) => e.key === 'Enter' && startDocking()}
106
+ placeholder="e.g. 1TIM"
107
+ className="w-full px-4 py-3 rounded-xl border border-glass-border focus:border-accent-cyan/40 focus:ring-2 focus:ring-accent-cyan/10 outline-none transition text-sm font-mono bg-surface-1 text-text-primary"
108
+ />
109
+ <div className="flex gap-2 mt-2 flex-wrap">
110
+ <span className="text-xs text-text-muted">Examples:</span>
111
+ {PDB_EXAMPLES.map((pdb) => (
112
+ <button
113
+ key={pdb}
114
+ onClick={() => setPdbId(pdb)}
115
+ className="px-2 py-1 text-xs rounded bg-accent-cyan/10 text-accent-cyan hover:bg-accent-cyan/20 transition font-mono"
116
+ >
117
+ {pdb}
118
+ </button>
119
+ ))}
120
+ </div>
121
+ </div>
122
+
123
+ <div>
124
+ <label className="block text-sm font-medium text-text-primary mb-1.5">Ligand SMILES</label>
125
+ <input
126
+ type="text"
127
+ value={smiles}
128
+ onChange={(e) => setSmiles(e.target.value)}
129
+ onKeyDown={(e) => e.key === 'Enter' && startDocking()}
130
+ placeholder="e.g. CC(=O)Oc1ccccc1C(=O)O"
131
+ className="w-full px-4 py-3 rounded-xl border border-glass-border focus:border-accent-cyan/40 focus:ring-2 focus:ring-accent-cyan/10 outline-none transition text-sm font-mono bg-surface-1 text-text-primary"
132
+ />
133
+ <div className="flex gap-2 mt-2 flex-wrap">
134
+ <span className="text-xs text-text-muted">Examples:</span>
135
+ {SMILES_EXAMPLES.map((ex) => (
136
+ <button
137
+ key={ex.label}
138
+ onClick={() => setSmiles(ex.value)}
139
+ className="px-2 py-1 text-xs rounded bg-accent-cyan/10 text-accent-cyan hover:bg-accent-cyan/20 transition"
140
+ >
141
+ {ex.label}
142
+ </button>
143
+ ))}
144
+ </div>
145
+ </div>
146
+
147
+ <button onClick={startDocking} disabled={loading || !pdbId.trim() || !smiles.trim() || polling}
148
+ className="btn-primary w-full py-3 flex items-center justify-center gap-2 disabled:opacity-50">
149
+ {loading ? <LoaderCircle className="w-4 h-4 animate-spin" /> : <FlaskConical className="w-4 h-4" />}
150
+ {loading ? 'Starting...' : polling ? 'Running...' : 'Run Docking'}
151
+ </button>
152
+ </motion.div>
153
+
154
+ {error && (
155
+ <motion.div variants={fadeUp} initial="hidden" animate="show" className="glass-card p-4 mb-6 border border-red-400/20">
156
+ <div className="flex items-start gap-3">
157
+ <AlertTriangle className="w-5 h-5 text-red-400 flex-shrink-0 mt-0.5" />
158
+ <p className="text-sm text-red-400">{error}</p>
159
+ </div>
160
+ </motion.div>
161
+ )}
162
+
163
+ {result && (
164
+ <motion.div variants={fadeUp} initial="hidden" animate="show" className="space-y-4">
165
+ <div className="glass-card p-5">
166
+ <div className="flex items-center justify-between mb-3">
167
+ <div className="flex items-center gap-2">
168
+ {statusIcon()}
169
+ <span className="text-sm font-medium text-text-primary capitalize">{result.status}</span>
170
+ </div>
171
+ <span className="text-xs text-text-muted font-mono">{result.result?.pdb_id} + {result.result?.smiles?.slice(0, 20)}...</span>
172
+ </div>
173
+
174
+ {result.status === 'failed' && result.error && (
175
+ <div className="p-3 rounded-lg bg-red-400/5 border border-red-400/20 mt-3">
176
+ <pre className="text-xs text-red-400 whitespace-pre-wrap font-mono">{result.error}</pre>
177
+ </div>
178
+ )}
179
+ </div>
180
+
181
+ {result.result?.poses && result.result.poses.length > 0 && (
182
+ <div className="glass-card p-5">
183
+ <h3 className="text-sm font-semibold text-text-primary mb-3">Docking Results</h3>
184
+
185
+ <div className="grid grid-cols-2 gap-4 mb-4">
186
+ <div className="p-3 rounded-xl bg-surface-1">
187
+ <p className="text-xs text-text-muted">Best Affinity</p>
188
+ <p className={`text-lg font-bold font-mono ${affinityColor(bestPose?.affinity ?? null)}`}>
189
+ {bestPose?.affinity?.toFixed(2) ?? 'β€”'} <span className="text-xs font-normal">kcal/mol</span>
190
+ </p>
191
+ </div>
192
+ <div className="p-3 rounded-xl bg-surface-1">
193
+ <p className="text-xs text-text-muted">Poses Generated</p>
194
+ <p className="text-lg font-bold text-text-primary font-mono">{result.result.num_poses}</p>
195
+ </div>
196
+ </div>
197
+
198
+ <div className="overflow-x-auto">
199
+ <table className="w-full text-sm">
200
+ <thead>
201
+ <tr className="text-xs text-text-muted uppercase border-b border-glass-border">
202
+ <th className="text-left py-2 pr-4">Pose</th>
203
+ <th className="text-left py-2 pr-4">Atoms</th>
204
+ <th className="text-left py-2">Affinity (kcal/mol)</th>
205
+ </tr>
206
+ </thead>
207
+ <tbody className="divide-y divide-glass-border">
208
+ {result.result.poses.map((pose) => (
209
+ <tr key={pose.model} className="text-text-primary">
210
+ <td className="py-2 pr-4 font-mono">{pose.model}</td>
211
+ <td className="py-2 pr-4 font-mono">{pose.atoms}</td>
212
+ <td className={`py-2 font-mono ${affinityColor(pose.affinity)}`}>
213
+ {pose.affinity !== null ? pose.affinity.toFixed(2) : 'β€”'}
214
+ </td>
215
+ </tr>
216
+ ))}
217
+ </tbody>
218
+ </table>
219
+ </div>
220
+ </div>
221
+ )}
222
+
223
+ {result.result?.box_center && (
224
+ <div className="glass-card p-5">
225
+ <h3 className="text-sm font-semibold text-text-primary mb-3">Binding Site Search Box</h3>
226
+ <div className="grid grid-cols-2 gap-4 text-sm">
227
+ <div>
228
+ <p className="text-xs text-text-muted mb-1">Center (x, y, z)</p>
229
+ <p className="text-text-primary font-mono">
230
+ {result.result.box_center.x.toFixed(2)}, {result.result.box_center.y.toFixed(2)}, {result.result.box_center.z.toFixed(2)}
231
+ </p>
232
+ </div>
233
+ <div>
234
+ <p className="text-xs text-text-muted mb-1">Size (x, y, z)</p>
235
+ <p className="text-text-primary font-mono">
236
+ {result.result.box_size.x}Γ…, {result.result.box_size.y}Γ…, {result.result.box_size.z}Γ…
237
+ </p>
238
+ </div>
239
+ </div>
240
+ </div>
241
+ )}
242
+
243
+ {result.result?.pdb_id && (
244
+ <div className="glass-card p-5">
245
+ <h3 className="text-sm font-semibold text-text-primary mb-3">Structure</h3>
246
+ <iframe
247
+ src={`https://www.ebi.ac.uk/pdbe/entry/pdb/${result.result.pdb_id}/embedded/`}
248
+ className="w-full h-80 rounded-xl border-0"
249
+ title="Protein structure"
250
+ />
251
+ </div>
252
+ )}
253
+ </motion.div>
254
+ )}
255
+ </div>
256
+ );
257
+ }
bioai-platform/frontend/src/app/(dashboard)/analyze/page.tsx CHANGED
@@ -1,7 +1,7 @@
1
  'use client';
2
 
3
  import { useRouter } from 'next/navigation';
4
- import { Dna, Layout, Search, Globe, GitBranch, Beaker, Layers, Share2, FlaskConical, Shuffle, GitFork } from 'lucide-react';
5
  import { motion } from 'framer-motion';
6
  import { fadeUp, stagger, cardHover } from '@/lib/animations';
7
  import { ReactNode } from 'react';
@@ -33,6 +33,7 @@ const groups: { title: string; items: Operation[] }[] = [
33
  { id: 'pathway', name: 'Pathway Analysis', description: 'Map genes to biological pathways from Reactome/KEGG.', icon: GitBranch, active: true },
34
  { id: 'interactions', name: 'Protein Interactions', description: 'Explore interaction partners from the STRING database.', icon: Share2, active: true },
35
  { id: 'compare', name: 'Structure Compare', description: 'Find structurally similar proteins via PDBeFold (TM-align).', icon: Shuffle, active: true },
 
36
  ],
37
  },
38
  {
 
1
  'use client';
2
 
3
  import { useRouter } from 'next/navigation';
4
+ import { Dna, Layout, Search, Globe, GitBranch, Beaker, Layers, Share2, FlaskConical, Shuffle, GitFork, Atom } from 'lucide-react';
5
  import { motion } from 'framer-motion';
6
  import { fadeUp, stagger, cardHover } from '@/lib/animations';
7
  import { ReactNode } from 'react';
 
33
  { id: 'pathway', name: 'Pathway Analysis', description: 'Map genes to biological pathways from Reactome/KEGG.', icon: GitBranch, active: true },
34
  { id: 'interactions', name: 'Protein Interactions', description: 'Explore interaction partners from the STRING database.', icon: Share2, active: true },
35
  { id: 'compare', name: 'Structure Compare', description: 'Find structurally similar proteins via PDBeFold (TM-align).', icon: Shuffle, active: true },
36
+ { id: 'docking', name: 'Molecular Docking', description: 'Dock a small molecule into a protein using AutoDock Vina (free, CPU-based).', icon: Atom, active: true, badge: 'New' },
37
  ],
38
  },
39
  {
bioai-platform/frontend/src/app/(dashboard)/layout.tsx CHANGED
@@ -11,6 +11,7 @@ import {
11
  History,
12
  Search,
13
  Settings,
 
14
  ChevronRight,
15
  LogOut,
16
  Dna,
@@ -19,6 +20,7 @@ import {
19
  import { useAuth } from '@/contexts/auth';
20
  import { ThemeToggle } from '@/components/ThemeToggle';
21
  import { ErrorBoundary } from '@/components/ErrorBoundary';
 
22
 
23
  const NAV_ITEMS = [
24
  { href: '/dashboard', icon: LayoutDashboard, label: 'Dashboard' },
@@ -26,6 +28,7 @@ const NAV_ITEMS = [
26
  { href: '/retrieve', icon: Search, label: 'Retrieve' },
27
  { href: '/jobs', icon: Clock, label: 'Jobs' },
28
  { href: '/history', icon: History, label: 'History' },
 
29
  { href: '/settings', icon: Settings, label: 'Settings' },
30
  ] as const;
31
 
@@ -303,6 +306,7 @@ export default function DashboardLayout({ children }: { children: React.ReactNod
303
  </div>
304
  </main>
305
  </div>
 
306
  </div>
307
  );
308
  }
 
11
  History,
12
  Search,
13
  Settings,
14
+ BookOpen,
15
  ChevronRight,
16
  LogOut,
17
  Dna,
 
20
  import { useAuth } from '@/contexts/auth';
21
  import { ThemeToggle } from '@/components/ThemeToggle';
22
  import { ErrorBoundary } from '@/components/ErrorBoundary';
23
+ import { TutorialWalkthrough } from '@/components/TutorialWalkthrough';
24
 
25
  const NAV_ITEMS = [
26
  { href: '/dashboard', icon: LayoutDashboard, label: 'Dashboard' },
 
28
  { href: '/retrieve', icon: Search, label: 'Retrieve' },
29
  { href: '/jobs', icon: Clock, label: 'Jobs' },
30
  { href: '/history', icon: History, label: 'History' },
31
+ { href: '/learn', icon: BookOpen, label: 'Learn' },
32
  { href: '/settings', icon: Settings, label: 'Settings' },
33
  ] as const;
34
 
 
306
  </div>
307
  </main>
308
  </div>
309
+ <TutorialWalkthrough />
310
  </div>
311
  );
312
  }
bioai-platform/frontend/src/app/(dashboard)/learn/[topic]/page.tsx ADDED
@@ -0,0 +1,295 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 'use client';
2
+
3
+ import { useParams, useRouter } from 'next/navigation';
4
+ import Link from 'next/link';
5
+ import { motion } from 'framer-motion';
6
+ import { fadeUp } from '@/lib/animations';
7
+ import { ArrowLeft } from 'lucide-react';
8
+
9
+ type Section = {
10
+ heading: string;
11
+ content: string;
12
+ code?: string;
13
+ };
14
+
15
+ type TopicData = {
16
+ title: string;
17
+ description: string;
18
+ sections: Section[];
19
+ };
20
+
21
+ const topics: Record<string, TopicData> = {
22
+ blast: {
23
+ title: 'BLAST Search',
24
+ description: 'Basic Local Alignment Search Tool β€” the most widely used method for finding sequence similarity.',
25
+ sections: [
26
+ {
27
+ heading: 'What is BLAST?',
28
+ content: 'BLAST compares a query sequence against a database of sequences and finds regions of local similarity. It uses a heuristic approach that is much faster than full dynamic programming (Smith-Waterman) while remaining sensitive enough for most searches. BLAST comes in several variants: BLASTP (protein-protein), BLASTN (nucleotide-nucleotide), BLASTX (translated nucleotide query against protein database), TBLASTN (protein query against translated nucleotide database), and TBLASTX (translated nucleotide against translated nucleotide).',
29
+ },
30
+ {
31
+ heading: 'How to read E-values',
32
+ content: 'The E-value (Expect value) describes how many matches you would expect to see by chance when searching a database of a given size. A lower E-value means a more significant match. An E-value of 0.05 means there is a 5% chance of seeing that match by chance alone. A good rule of thumb: E-values below 1e-5 (0.00001) are typically considered significant for homology searches. Values between 0.001 and 0.1 may indicate distant homology and should be investigated further.',
33
+ code: 'E-value = K Γ— m Γ— n Γ— e^(βˆ’Ξ»S)\n\n K = search-space constant\n m = query length\n n = database length\n S = raw alignment score\n Ξ» = scoring-system lambda',
34
+ },
35
+ {
36
+ heading: 'Understanding bit scores',
37
+ content: 'The bit score is a normalized, log-scaled version of the raw alignment score. It is independent of database size and scoring matrix, making it comparable across different searches. A bit score of 50 or higher typically indicates a biologically relevant match. Bit scores are calculated as S\' = (Ξ»S βˆ’ ln K) / ln 2, where S is the raw score, Ξ» and K are statistical parameters of the scoring system.',
38
+ },
39
+ {
40
+ heading: 'Interpreting identity percentage',
41
+ content: 'Percent identity is simply the fraction of aligned positions where the residues match exactly, expressed as a percentage. It is the most intuitive metric but can be misleading for divergent sequences. For proteins, 30% identity over a full-length alignment is often considered the "twilight zone" below which inferring homology becomes unreliable. However, short regions of high identity can be functionally significant even when overall identity is low.',
42
+ },
43
+ ],
44
+ },
45
+ alignment: {
46
+ title: 'Sequence Alignment',
47
+ description: 'Comparing sequences to identify regions of similarity β€” the foundation of bioinformatics.',
48
+ sections: [
49
+ {
50
+ heading: 'Pairwise vs Multiple Alignment',
51
+ content: 'Pairwise alignment compares two sequences at a time. It can be global (Needleman-Wunsch) which aligns the entire length of both sequences, or local (Smith-Waterman) which finds the best matching subsequence. Multiple sequence alignment (MSA) extends this to three or more sequences, revealing conserved regions across a family. Common MSA tools include Clustal Omega, MAFFT, and MUSCLE. MSA is the basis for building phylogenetic trees, identifying conserved motifs, and improving structure prediction.',
52
+ },
53
+ {
54
+ heading: 'Scoring matrices',
55
+ content: 'Scoring matrices define the score for aligning any two residues. BLOSUM (BLOcks SUbstitution Matrix) matrices are the most common for proteins. BLOSUM62 is the default for most searches β€” it assumes sequences with ~62% identity. Higher numbers (BLOSUM80) are better for closely related sequences; lower numbers (BLOSUM45) are better for distantly related ones. For nucleotides, simple match/mismatch scores are typically used (e.g., +1/-1 or +2/-3).',
56
+ code: 'BLOSUM62 example (positive scores = conserved substitutions):\n\n A R N D C Q E G ...\n A 4 -1 -2 -2 0 -1 -1 0\n R -1 5 0 -2 -3 1 0 -2\n N -2 0 6 1 -3 0 0 0\n D -2 -2 1 6 -3 0 2 -1',
57
+ },
58
+ {
59
+ heading: 'Gap penalties',
60
+ content: 'Gap penalties control the cost of inserting gaps in an alignment. They consist of two components: a gap-open penalty (cost for starting a gap) and a gap-extension penalty (cost for extending an existing gap). Typical values for protein alignments are 10-12 for opening and 1-2 for extension. High gap penalties produce shorter, more compact alignments; low penalties allow longer gaps but risk over-fitting.',
61
+ },
62
+ {
63
+ heading: 'Reading alignment output',
64
+ content: 'Standard alignment output uses a three-line format for each block: the query sequence, a match line showing identical (|), conserved (:), and gap ( ) symbols, and the subject sequence. Identical residues indicate perfect conservation; conserved substitutions (similar biochemical properties) are shown with colons; non-conserved substitutions have no symbol. Gaps introduced in either sequence are shown as dashes.',
65
+ code: 'Query: MKLLVLFLLGLVALSECDIYNYNA...KLCGVL\n ||:||| || | ::|.||: ...||:..|\nSubject: MKLLILFLLGLVALLLCEPSLYNYNA...NYCTAL',
66
+ },
67
+ ],
68
+ },
69
+ domains: {
70
+ title: 'Domain Analysis',
71
+ description: 'Identifying conserved functional and structural units within proteins.',
72
+ sections: [
73
+ {
74
+ heading: 'What are protein domains?',
75
+ content: 'A protein domain is a conserved, independently folding region of a protein that carries a specific function. Domains are the evolutionary building blocks of proteins β€” they can be shuffled, duplicated, and combined in different arrangements to create proteins with new functions. Most eukaryotic proteins contain multiple domains. Identifying domains helps predict protein function, even when the overall sequence has no known homologs.',
76
+ },
77
+ {
78
+ heading: 'Pfam and InterPro',
79
+ content: 'Pfam is a comprehensive database of protein domain families, each represented by a multiple sequence alignment and a hidden Markov model (HMM) profile. InterPro combines multiple domain databases (Pfam, SMART, PROSITE, CDD, etc.) into a single resource. When you run a domain analysis, your query is scanned against these HMM profiles to identify known domains. Each hit includes an E-value, bitscore, and the region of the query that matches the domain model.',
80
+ },
81
+ {
82
+ heading: 'Domain architecture',
83
+ content: 'Domain architecture refers to the linear arrangement of domains along a protein sequence. Many proteins have a modular architecture where different domains work together. For example, a signaling protein might have a receptor domain, a kinase domain, and a protein-protein interaction domain. Analyzing domain architecture helps predict function, evolutionary relationships, and potential interactions with other proteins.',
84
+ code: 'Example domain architecture:\n\n Protein: EGFR (Epidermal Growth Factor Receptor)\n \n [Receptor L]──[Furin-like]──[GF_recep]──[TM]──[PKinase_Tyr]\n | | | | |\n Ligand-binding | Growth Trans- Tyrosine\n (extracellular) | factor membrane kinase\n Cysteine-rich rec. (cytoplasmic)\n domain domain',
85
+ },
86
+ ],
87
+ },
88
+ phylo: {
89
+ title: 'Phylogenetic Trees',
90
+ description: 'Reconstructing evolutionary relationships from molecular sequences.',
91
+ sections: [
92
+ {
93
+ heading: 'Phylogenetic trees',
94
+ content: 'A phylogenetic tree is a branching diagram showing the evolutionary relationships among species, genes, or sequences. Trees consist of branches (edges) and nodes (branch points). Terminal nodes (leaves) represent extant sequences; internal nodes represent hypothetical ancestors. Trees can be rooted (with a known common ancestor) or unrooted. The topology describes the branching order, while branch lengths typically represent evolutionary distance.',
95
+ },
96
+ {
97
+ heading: 'NJ vs UPGMA vs Maximum Likelihood',
98
+ content: 'Neighbor-Joining (NJ) is a fast distance-based method that builds a tree by iteratively joining the closest pair of sequences. UPGMA is another distance method that assumes a constant molecular clock (same rate across all lineages). Maximum Likelihood (ML) is a more sophisticated method that evaluates different tree topologies and selects the one that makes the sequence data most likely under a given substitution model. ML is slower but more accurate. Modern ML tools include RAxML-NG, IQ-TREE, and PhyML.',
99
+ },
100
+ {
101
+ heading: 'Reading bootstrap values',
102
+ content: 'Bootstrap values indicate how strongly the data supports a given branch. The original sequences are resampled (with replacement) hundreds or thousands of times, a tree is built from each replicate, and the fraction of replicates that recover the same branch is the bootstrap value. Values above 70% are considered moderately supported; above 95% is strongly supported. Bootstrap values below 50% suggest the branching order at that node is unreliable.',
103
+ code: 'Example tree with bootstrap values:\n\n β”Œβ”€β”€β”€ Human\n β”Œβ”€β”€β”€ 98 ───\n β”‚ └─── Chimp\n ── 100 ──\n β”‚ β”Œβ”€β”€β”€ Mouse\n └─── 72 ───\n └─── Rat\n\n 100 = very strong support for human/chimp clade\n 72 = moderate support for mouse/rat clade',
104
+ },
105
+ {
106
+ heading: 'Branch lengths',
107
+ content: 'Branch lengths represent the amount of evolutionary change along a branch. The units are typically substitutions per site β€” the expected number of residue changes per position along that lineage. Longer branches mean more divergence. In distance-based trees, branch lengths are additive: the distance between two sequences is the sum of branch lengths along the path connecting them. In ML trees, branch lengths are optimized to maximize the likelihood of the data.',
108
+ },
109
+ ],
110
+ },
111
+ structure: {
112
+ title: 'Protein Structure',
113
+ description: 'Understanding the three-dimensional shapes of proteins and how to analyze them.',
114
+ sections: [
115
+ {
116
+ heading: 'PDB format',
117
+ content: 'The Protein Data Bank (PDB) format is the standard file format for macromolecular structures. Each line in a PDB file contains specific information identified by a record type (ATOM, HETATM, HELIX, SHEET, etc.). The ATOM records contain the coordinates (x, y, z) of each atom, along with the atom name, residue name, chain identifier, residue number, and occupancy/temperature factors. Modern alternatives include mmCIF and PDBx/mmCIF, but PDB remains widely supported.',
118
+ code: 'Example PDB ATOM record:\n\nATOM 1 N ALA A 1 21.894 16.287 5.352 1.00 9.58 N\nATOM 2 CA ALA A 1 22.482 15.026 5.846 1.00 9.74 C\nATOM 3 C ALA A 1 23.176 14.242 4.744 1.00 9.46 C\nATOM 4 O ALA A 1 23.121 14.636 3.579 1.00 9.22 O\n\nCol 1-6: Record name\nCol 7-11: Serial number\nCol 13-16: Atom name\nCol 18-20: Residue name\nCol 22: Chain ID\nCol 23-26: Residue number\nCol 31-38: X coordinate\nCol 39-46: Y coordinate\nCol 47-54: Z coordinate',
119
+ },
120
+ {
121
+ heading: 'AlphaFold',
122
+ content: 'AlphaFold is a deep learning system developed by DeepMind that predicts protein structures from amino acid sequences with accuracy comparable to experimental methods. AlphaFold2 won the CASP14 competition in 2020. Its successor, AlphaFold3, extends predictions to protein-ligand, protein-nucleic acid, and protein-small molecule complexes. The AlphaFold Database contains over 200 million predicted protein structures covering nearly all known proteins.',
123
+ },
124
+ {
125
+ heading: 'Reading pLDDT scores',
126
+ content: 'pLDDT (predicted Local Distance Difference Test) is AlphaFold\'s per-residue confidence score, ranging from 0 to 100. A pLDDT above 90 indicates very high confidence (comparable to experimental structures). Values between 70 and 90 indicate good backbone prediction. Values between 50 and 70 indicate low confidence, and below 50 indicates very low confidence β€” likely unstructured or disordered regions. The pLDDT score is stored in the B-factor column of the PDB file in AlphaFold predictions.',
127
+ code: 'pLDDT confidence interpretation:\n\n > 90 β€” Very high (comparable to experiment)\n 70–90 β€” Good backbone prediction\n 50–70 β€” Low confidence\n < 50 β€” Very low (likely disordered)',
128
+ },
129
+ {
130
+ heading: 'Structure visualization',
131
+ content: 'Protein structures can be visualized in several representations: cartoon/ribbon (shows secondary structure), surface (shows solvent-accessible surface), sticks (shows atomic bonds), and spheres (space-filling). Web-based viewers like Mol* (MolStar), NGL Viewer, and 3Dmol.js enable interactive visualization directly in the browser. Bio Nexus uses Mol* for structure rendering, supporting PDB and mmCIF files with customizable color schemes and selection highlighting.',
132
+ },
133
+ ],
134
+ },
135
+ pathways: {
136
+ title: 'Pathway Analysis',
137
+ description: 'Mapping genes and proteins to the biological pathways they participate in.',
138
+ sections: [
139
+ {
140
+ heading: 'What are pathways?',
141
+ content: 'A biological pathway is a series of molecular interactions and reactions that produce a specific cellular outcome. Metabolic pathways involve chemical transformations (e.g., glycolysis, citric acid cycle). Signaling pathways transmit signals from the cell surface to the nucleus (e.g., MAPK/ERK, Wnt). Gene regulatory pathways control gene expression. Pathway analysis helps interpret high-throughput data (RNA-seq, proteomics) by identifying which pathways are enriched in a set of differentially expressed genes.',
142
+ },
143
+ {
144
+ heading: 'Reactome vs KEGG',
145
+ content: 'Reactome is a free, open-source, manually curated pathway database with detailed molecular-level annotations. It provides excellent cross-references to other databases and supports pathway overrepresentation analysis (ORA). KEGG (Kyoto Encyclopedia of Genes and Genomes) is a comprehensive resource containing pathway maps, ortholog information, and chemical reactions. While KEGG remains popular, its licensing has become more restrictive. Reactome is generally preferred for open academic use.',
146
+ },
147
+ {
148
+ heading: 'Enrichment analysis',
149
+ content: 'Enrichment analysis determines whether a set of genes (e.g., upregulated in an RNA-seq experiment) contains more genes from a particular pathway than expected by chance. The standard method is Fisher\'s exact test or a hypergeometric test, corrected for multiple testing (Benjamini-Hochberg FDR). The result is a list of pathways ranked by significance, with enrichment ratios and adjusted p-values. Bio Nexus performs pathway enrichment against both Reactome and KEGG databases.',
150
+ code: 'Enrichment analysis results example:\n\nPathway Genes Expected Ratio p-value FDR\n──────────────────────────────────────────────────────────────────────\nDNA Replication 12 2.1 5.7 8e-12 2e-9\nCell Cycle 18 4.3 4.2 2e-10 3e-8\np53 Signaling 8 1.2 6.7 5e-8 4e-6\n\nRatio = observed / expected count\nFDR = false discovery rate (corrected p-value)',
151
+ },
152
+ ],
153
+ },
154
+ interactions: {
155
+ title: 'Protein Interactions',
156
+ description: 'Exploring the network of physical and functional associations between proteins.',
157
+ sections: [
158
+ {
159
+ heading: 'STRING database',
160
+ content: 'STRING (Search Tool for the Retrieval of Interacting Genes/Proteins) is a comprehensive database of known and predicted protein-protein interactions. It covers over 67 million proteins from more than 14,000 organisms. Interactions are derived from four sources: experimental evidence, curated databases, text mining of scientific literature, and computational predictions (gene neighborhood, gene fusions, gene co-occurrence). Each interaction is scored by how well the evidence supports it.',
161
+ },
162
+ {
163
+ heading: 'Interaction networks',
164
+ content: 'An interaction network consists of nodes (proteins) and edges (interactions). Networks can be visualized with different layout algorithms: force-directed (Fruchterman-Reingold), circular, or hierarchical. The network topology reveals hub proteins (highly connected), bottlenecks, and clusters corresponding to functional modules. Bio Nexus uses the STRING API to fetch interaction data and renders interactive networks using a force-directed layout.',
165
+ },
166
+ {
167
+ heading: 'Confidence scores',
168
+ content: 'STRING assigns each interaction a confidence score from 0 to 1,000, with higher values indicating stronger evidence. Scores are divided into three tiers: low confidence (< 150), medium confidence (150–700), and high confidence (> 700). The combined score integrates evidence from all sources using a naive Bayes approach. For most analyses, filtering at medium confidence (β‰₯ 400) provides a good balance of sensitivity and specificity.',
169
+ code: 'STRING confidence tiers:\n\n > 700 β€” High confidence (strong experimental + database evidence)\n 400–700 β€” Medium confidence (good for most analyses)\n 150–400 β€” Low confidence (primarily text-mining)\n < 150 β€” Very low (likely noise)',
170
+ },
171
+ ],
172
+ },
173
+ primers: {
174
+ title: 'Primer Design',
175
+ description: 'Designing oligonucleotide primers for PCR amplification.',
176
+ sections: [
177
+ {
178
+ heading: 'PCR basics',
179
+ content: 'The Polymerase Chain Reaction (PCR) amplifies a specific DNA region between two primer binding sites. Each cycle consists of three steps: denaturation (95Β°C β€” separate DNA strands), annealing (50–65Β°C β€” primers bind), and extension (72Β°C β€” DNA polymerase extends). After 30–35 cycles, the target region is amplified by over a billion-fold. Successful PCR depends on well-designed primers that are specific, have appropriate melting temperatures, and do not form secondary structures.',
180
+ },
181
+ {
182
+ heading: 'Primer3',
183
+ content: 'Primer3 is the most widely used primer design software. It picks PCR primers from a template sequence, optimizing for melting temperature, GC content, primer length, and avoiding problematic features like hairpins, self-dimers, and cross-dimers. Bio Nexus uses Primer3 via its backend API to design primers for any input sequence. The tool evaluates hundreds of candidate primer pairs and returns the best ones ranked by a quality score.',
184
+ },
185
+ {
186
+ heading: 'Melting temperature',
187
+ content: 'The melting temperature (Tm) of a primer is the temperature at which half of the primer molecules are annealed to the template. It depends on primer length, GC content, and salt concentration. A common rule of thumb: Tm = 2Β°C Γ— (A+T) + 4Β°C Γ— (G+C). For PCR, primers should have Tm values between 55Β°C and 65Β°C, and the forward and reverse primers should have Tm values within 2–5Β°C of each other.',
188
+ code: 'Tm estimation (nearest-neighbor, simplified):\n\n Tm = Ξ”H / (Ξ”S + R Γ— ln(C/4)) βˆ’ 273.15 + 16.6 Γ— log([Na+])\n\n Ξ”H = enthalpy change\n Ξ”S = entropy change\n R = gas constant (1.987 cal/molΒ·K)\n C = primer concentration\n\nRule of thumb:\n Tm β‰ˆ 2(A+T) + 4(G+C)',
189
+ },
190
+ {
191
+ heading: 'GC content',
192
+ content: 'GC content β€” the percentage of guanine and cytosine bases in a primer β€” affects both melting temperature and secondary structure formation. Ideal primers have 40–60% GC content. Too high GC content (> 65%) increases the risk of non-specific binding and stable secondary structures. Too low GC content (< 35%) results in weak binding and low Tm. Primers with balanced GC content across the 3\' end provide the most reliable amplification.',
193
+ },
194
+ ],
195
+ },
196
+ tools: {
197
+ title: 'Format Converter',
198
+ description: 'Converting between common bioinformatics sequence formats.',
199
+ sections: [
200
+ {
201
+ heading: 'Format conversion',
202
+ content: 'Bio Nexus supports conversion between FASTA, GenBank, EMBL, and plain text formats. FASTA is the simplest format β€” a header line starting with ">" followed by the sequence. GenBank and EMBL are richer formats that include annotations, features, and references. When converting between formats, only the sequence and basic header information are preserved. Annotations and features are kept when converting between GenBank and EMBL.',
203
+ code: 'FASTA format:\n\n >seq_id description\n ATGCGATCGTAGCTAGCTAGCTAGCATCGATCG\n GCTAGCTAGCATCGATCGATCGATCGATCGTAG\n\nGenBank format:\n\n LOCUS NM_001 1234 bp DNA linear\n DEFINITION Sample sequence.\n ORIGIN\n 1 atgcgatcgt agctagctag ctagcatcga tcg\n 61 gctagctagc atcgatcgat cgatcgtagg tagcta\n //',
204
+ },
205
+ {
206
+ heading: 'Sequence validation',
207
+ content: 'Sequence validation checks that your input contains only valid residues for the specified molecule type. For DNA, valid characters are A, C, G, T, and U (uracil is converted to thymine). For RNA, valid characters are A, C, G, and U. For protein, valid characters are the 20 standard amino acids (plus B, Z, X, and * for selenocysteine/pyrrolysine/stop). The validator also detects common issues like whitespace, line breaks, and numeric characters embedded in the sequence.',
208
+ },
209
+ ],
210
+ },
211
+ glossary: {
212
+ title: 'Glossary',
213
+ description: 'A–Z reference of bioinformatics terms with plain-English definitions.',
214
+ sections: [
215
+ {
216
+ heading: 'A–C',
217
+ content: 'Alignment β€” The arrangement of sequences to identify regions of similarity.\nAmino acid β€” One of 20 organic compounds that form proteins.\nBLAST β€” Basic Local Alignment Search Tool for finding sequence similarity.\nBit score β€” Normalized, database-size-independent score from a sequence search.\nBootstrap β€” Resampling method to assess confidence in phylogenetic tree branches.\nCDS β€” Coding Sequence, the region of a gene that is translated into protein.\nConserved β€” A residue or region that remains unchanged across evolution.\nContig β€” A contiguous sequence assembled from overlapping sequencing reads.',
218
+ },
219
+ {
220
+ heading: 'D–H',
221
+ content: 'Domain β€” A conserved, independently folding functional unit of a protein.\nE-value β€” Expect value: number of chance matches expected in a database search.\nEnrichment β€” Statistical overrepresentation of a pathway in a gene set.\nFASTA β€” Text-based sequence format using a single-line header starting with ">".\nFDR β€” False Discovery Rate, a correction for multiple hypothesis testing.\nGap β€” A space inserted in an alignment to compensate for insertions/deletions.\nGC content β€” Percentage of guanine and cytosine bases in a sequence.\nHMM β€” Hidden Markov Model, a statistical model used for profile searches.',
222
+ },
223
+ {
224
+ heading: 'I–M',
225
+ content: 'Identity β€” The percentage of exactly matching residues in an alignment.\nInterPro β€” Integrated database of protein domains, families, and functional sites.\nKEGG β€” Kyoto Encyclopedia of Genes and Genomes, a pathway database.\nLocal alignment β€” Alignment of only the most similar subsequences (Smith-Waterman).\nMelting temperature (Tm) β€” Temperature at which half of DNA duplex dissociates.\nML β€” Maximum Likelihood, a phylogenetic method that optimizes tree topology.\nMSA β€” Multiple Sequence Alignment, alignment of three or more sequences.\nMutation β€” A change in the nucleotide sequence of a genome.',
226
+ },
227
+ {
228
+ heading: 'N–R',
229
+ content: 'NJ β€” Neighbor-Joining, a fast distance-based phylogenetic tree-building method.\nORF β€” Open Reading Frame, a region of DNA potentially coding for a protein.\nOrtholog β€” Genes in different species that evolved from a common ancestral gene.\nPCR β€” Polymerase Chain Reaction, a method to amplify specific DNA sequences.\nPDB β€” Protein Data Bank, the global repository of 3D macromolecular structures.\nPfam β€” A database of protein domain families with associated HMM profiles.\nPhylogeny β€” The evolutionary history and relationships among organisms/sequences.\npLDDT β€” Predicted Local Distance Difference Test, AlphaFold\'s per-residue confidence.',
230
+ },
231
+ {
232
+ heading: 'S–Z',
233
+ content: 'Scoring matrix β€” A table of scores for aligning each pair of residues.\nSmith-Waterman β€” An algorithm for local sequence alignment.\nSTRING β€” Database of known and predicted protein-protein interactions.\nSubstitution β€” A residue replaced by another during evolution.\nTopology β€” The branching pattern of a phylogenetic tree (not including branch lengths).\nTwilight zone β€” Region of sequence similarity (~20–35% identity) where homology is uncertain.\nUPGMA β€” Unweighted Pair Group Method with Arithmetic Mean, a distance-based clustering method.\nVariant β€” A specific form of a genetic sequence that differs from the reference.',
234
+ },
235
+ ],
236
+ },
237
+ };
238
+
239
+ export default function TopicPage() {
240
+ const params = useParams();
241
+ const router = useRouter();
242
+ const topic = params.topic as string;
243
+ const data = topics[topic];
244
+
245
+ if (!data) {
246
+ return (
247
+ <div className="text-center py-20">
248
+ <h1 className="text-2xl font-bold text-text-primary mb-2">Topic not found</h1>
249
+ <p className="text-text-muted mb-6">No documentation available for &ldquo;{topic}&rdquo;.</p>
250
+ <button
251
+ onClick={() => router.push('/learn')}
252
+ className="btn-primary px-5 py-2.5 text-sm"
253
+ >
254
+ Back to Documentation
255
+ </button>
256
+ </div>
257
+ );
258
+ }
259
+
260
+ return (
261
+ <div className="max-w-3xl">
262
+ <motion.div variants={fadeUp} initial="hidden" animate="show">
263
+ <Link
264
+ href="/learn"
265
+ className="inline-flex items-center gap-1.5 text-sm text-text-muted hover:text-text-primary transition mb-6"
266
+ >
267
+ <ArrowLeft className="w-4 h-4" />
268
+ Back to Documentation
269
+ </Link>
270
+
271
+ <h1 className="text-2xl font-bold text-text-primary mb-2">{data.title}</h1>
272
+ <p className="text-text-muted mb-10 text-sm">{data.description}</p>
273
+ </motion.div>
274
+
275
+ <div className="space-y-10">
276
+ {data.sections.map((section, i) => (
277
+ <motion.section
278
+ key={i}
279
+ variants={fadeUp}
280
+ initial="hidden"
281
+ animate="show"
282
+ >
283
+ <h2 className="text-lg font-semibold text-text-primary mb-3">{section.heading}</h2>
284
+ <p className="text-sm text-text-secondary leading-relaxed whitespace-pre-line">{section.content}</p>
285
+ {section.code && (
286
+ <pre className="mt-4 p-4 rounded-xl bg-surface-1 border border-glass-border overflow-x-auto text-xs font-mono text-text-secondary leading-relaxed">
287
+ <code>{section.code}</code>
288
+ </pre>
289
+ )}
290
+ </motion.section>
291
+ ))}
292
+ </div>
293
+ </div>
294
+ );
295
+ }
bioai-platform/frontend/src/app/(dashboard)/learn/page.tsx ADDED
@@ -0,0 +1,130 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 'use client';
2
+
3
+ import Link from 'next/link';
4
+ import {
5
+ Search, Dna, Layout, Layers, GitFork, Globe,
6
+ GitBranch, Share2, FlaskConical, Beaker, BookOpen, ArrowRight,
7
+ } from 'lucide-react';
8
+ import { motion } from 'framer-motion';
9
+ import { fadeUp, stagger, cardHover } from '@/lib/animations';
10
+ import { useState } from 'react';
11
+ import type { LucideIcon } from 'lucide-react';
12
+
13
+ type Topic = {
14
+ id: string;
15
+ title: string;
16
+ description: string;
17
+ icon: LucideIcon;
18
+ };
19
+
20
+ const groups: { title: string; items: Topic[] }[] = [
21
+ {
22
+ title: 'Sequence Analysis',
23
+ items: [
24
+ { id: 'blast', title: 'BLAST Search', description: 'Find similar sequences and understand E-values, bit scores, and identity.', icon: Search },
25
+ { id: 'alignment', title: 'Sequence Alignment', description: 'Pairwise and multiple alignment β€” scoring matrices, gap penalties, and output.', icon: Layout },
26
+ { id: 'domains', title: 'Domain Analysis', description: 'Protein domains, Pfam, InterPro, and domain architecture.', icon: Layers },
27
+ { id: 'phylo', title: 'Phylogenetic Trees', description: 'Tree-building methods, bootstrap values, and branch lengths.', icon: GitFork },
28
+ ],
29
+ },
30
+ {
31
+ title: 'Structure & Networks',
32
+ items: [
33
+ { id: 'structure', title: 'Protein Structure', description: 'PDB format, AlphaFold, pLDDT scores, and structure visualization.', icon: Dna },
34
+ { id: 'pathways', title: 'Pathway Analysis', description: 'Reactome vs KEGG, pathway mapping, and enrichment analysis.', icon: GitBranch },
35
+ { id: 'interactions', title: 'Protein Interactions', description: 'STRING database, interaction networks, and confidence scores.', icon: Globe },
36
+ ],
37
+ },
38
+ {
39
+ title: 'Utilities',
40
+ items: [
41
+ { id: 'primers', title: 'Primer Design', description: 'PCR basics, Primer3, melting temperature, and GC content.', icon: FlaskConical },
42
+ { id: 'tools', title: 'Format Converter', description: 'Sequence format conversion and validation utilities.', icon: Beaker },
43
+ { id: 'glossary', title: 'Glossary', description: 'A–Z of bioinformatics terms with plain-English definitions.', icon: BookOpen },
44
+ ],
45
+ },
46
+ ];
47
+
48
+ function TopicCard({ topic }: { topic: Topic }) {
49
+ const Icon = topic.icon;
50
+ return (
51
+ <motion.div variants={fadeUp} whileHover={cardHover}>
52
+ <Link
53
+ href={`/learn/${topic.id}`}
54
+ className="relative block p-5 rounded-2xl border border-glass-border bg-glass-card hover:bg-surface-1 transition h-full"
55
+ >
56
+ <div className="flex items-start gap-4">
57
+ <div className="p-3 rounded-xl bg-accent-cyan/10 flex-shrink-0">
58
+ <Icon className="w-5 h-5 text-accent-cyan" />
59
+ </div>
60
+ <div className="flex-1 min-w-0">
61
+ <h3 className="font-semibold text-text-primary">{topic.title}</h3>
62
+ <p className="text-sm text-text-muted mt-1 leading-relaxed">{topic.description}</p>
63
+ <div className="flex items-center gap-1 mt-3 text-xs text-accent-cyan font-medium">
64
+ Learn more <ArrowRight className="w-3 h-3" />
65
+ </div>
66
+ </div>
67
+ </div>
68
+ </Link>
69
+ </motion.div>
70
+ );
71
+ }
72
+
73
+ export default function LearnPage() {
74
+ const [query, setQuery] = useState('');
75
+
76
+ const filtered = query.trim()
77
+ ? groups.map(g => ({
78
+ ...g,
79
+ items: g.items.filter(t =>
80
+ t.title.toLowerCase().includes(query.toLowerCase()) ||
81
+ t.description.toLowerCase().includes(query.toLowerCase())
82
+ ),
83
+ })).filter(g => g.items.length > 0)
84
+ : groups;
85
+
86
+ return (
87
+ <div>
88
+ <motion.div variants={fadeUp} initial="hidden" animate="show">
89
+ <h1 className="text-2xl font-bold text-text-primary mb-1">Documentation & Learning</h1>
90
+ <p className="text-text-muted mb-6">Learn the concepts behind every tool in Bio Nexus.</p>
91
+ </motion.div>
92
+
93
+ <motion.div variants={fadeUp} className="relative mb-10 max-w-xl">
94
+ <Search className="absolute left-4 top-1/2 -translate-y-1/2 w-4 h-4 text-text-muted" />
95
+ <input
96
+ type="text"
97
+ value={query}
98
+ onChange={e => setQuery(e.target.value)}
99
+ placeholder="Search topics..."
100
+ className="w-full pl-11 pr-4 py-3 rounded-2xl border border-glass-border bg-glass-card text-text-primary text-sm placeholder:text-text-muted/50 outline-none focus:border-accent-cyan/30 transition"
101
+ />
102
+ </motion.div>
103
+
104
+ {filtered.map(group => (
105
+ <div key={group.title} className="mb-10">
106
+ <motion.h2 variants={fadeUp} className="text-sm font-semibold text-text-muted uppercase tracking-wider mb-4">
107
+ {group.title}
108
+ </motion.h2>
109
+ <motion.div
110
+ variants={stagger}
111
+ initial="hidden"
112
+ whileInView="show"
113
+ viewport={{ once: true, margin: '-40px' }}
114
+ className="grid md:grid-cols-2 gap-4"
115
+ >
116
+ {group.items.map(topic => (
117
+ <TopicCard key={topic.id} topic={topic} />
118
+ ))}
119
+ </motion.div>
120
+ </div>
121
+ ))}
122
+
123
+ {filtered.length === 0 && (
124
+ <motion.p variants={fadeUp} className="text-text-muted text-sm text-center py-12">
125
+ No topics found for &ldquo;{query}&rdquo;.
126
+ </motion.p>
127
+ )}
128
+ </div>
129
+ );
130
+ }
bioai-platform/frontend/src/components/LearnPopover.tsx ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 'use client';
2
+
3
+ import { useState, useRef, useEffect } from 'react';
4
+ import Link from 'next/link';
5
+ import { HelpCircle } from 'lucide-react';
6
+
7
+ interface LearnPopoverProps {
8
+ term: string;
9
+ explanation: string;
10
+ topic?: string;
11
+ children: React.ReactNode;
12
+ }
13
+
14
+ export function LearnPopover({ term, explanation, topic, children }: LearnPopoverProps) {
15
+ const [open, setOpen] = useState(false);
16
+ const wrapperRef = useRef<HTMLSpanElement>(null);
17
+
18
+ useEffect(() => {
19
+ function handleClickOutside(e: MouseEvent) {
20
+ if (wrapperRef.current && !wrapperRef.current.contains(e.target as Node)) {
21
+ setOpen(false);
22
+ }
23
+ }
24
+ if (open) {
25
+ document.addEventListener('mousedown', handleClickOutside);
26
+ }
27
+ return () => document.removeEventListener('mousedown', handleClickOutside);
28
+ }, [open]);
29
+
30
+ return (
31
+ <span ref={wrapperRef} className="inline-flex items-center gap-0.5 relative">
32
+ <span
33
+ className="cursor-pointer border-b border-dotted border-accent-cyan/40 hover:border-accent-cyan transition"
34
+ onClick={() => setOpen(!open)}
35
+ role="button"
36
+ tabIndex={0}
37
+ onKeyDown={e => { if (e.key === 'Enter' || e.key === ' ') { e.preventDefault(); setOpen(!open); } }}
38
+ aria-label={`Learn about ${term}`}
39
+ >
40
+ {children}
41
+ </span>
42
+ <HelpCircle
43
+ size={12}
44
+ className="inline-block text-text-muted cursor-pointer hover:text-accent-cyan transition flex-shrink-0"
45
+ onClick={() => setOpen(!open)}
46
+ />
47
+ {open && (
48
+ <div
49
+ className="absolute z-50 top-full left-0 mt-2 p-4 rounded-xl border border-glass-border bg-surface-2 shadow-glass-md max-w-xs text-sm"
50
+ style={{ backdropFilter: 'blur(16px)' }}
51
+ >
52
+ <p className="font-semibold text-text-primary text-xs mb-1">{term}</p>
53
+ <p className="text-text-secondary text-xs leading-relaxed">{explanation}</p>
54
+ {topic && (
55
+ <Link
56
+ href={`/learn/${topic}`}
57
+ className="inline-block mt-2 text-xs text-accent-cyan hover:underline font-medium"
58
+ >
59
+ Learn more &rarr;
60
+ </Link>
61
+ )}
62
+ </div>
63
+ )}
64
+ </span>
65
+ );
66
+ }
bioai-platform/frontend/src/components/TutorialWalkthrough.tsx ADDED
@@ -0,0 +1,171 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 'use client';
2
+
3
+ import { useState, useEffect, useCallback } from 'react';
4
+ import { motion, AnimatePresence } from 'framer-motion';
5
+ import { X, ChevronLeft, ChevronRight, BookOpen } from 'lucide-react';
6
+
7
+ const STEPS = [
8
+ {
9
+ title: 'Welcome to Bio Nexus',
10
+ description: 'Bio Nexus is your all-in-one bioinformatics platform. From BLAST searches to protein structure visualization, everything is designed to be fast and intuitive. This short tour will show you the essentials.',
11
+ highlight: 'sidebar',
12
+ },
13
+ {
14
+ title: 'Running an analysis',
15
+ description: 'Click any tool from the sidebar β€” BLAST, Alignment, Domain Analysis, and more. Paste a sequence, adjust parameters, and hit run. Results appear in seconds.',
16
+ highlight: 'nav-analyze',
17
+ },
18
+ {
19
+ title: 'Understanding results',
20
+ description: 'Results are displayed with clean, interactive visualizations. Tables, graphs, and 3D viewers help you interpret your data at a glance. Each result can be downloaded or shared.',
21
+ highlight: 'results',
22
+ },
23
+ {
24
+ title: 'AI interpretation',
25
+ description: 'Click "Interpret with AI" on any result page to get a plain-English explanation of what your results mean. The AI understands BLAST hits, pathway enrichment, domain architectures, and more.',
26
+ highlight: 'ai',
27
+ },
28
+ {
29
+ title: 'Learning more',
30
+ description: 'Visit the Learn section for in-depth documentation on every concept β€” E-values, bootstrap values, pLDDT scores, and dozens more. Terms marked with a (?) icon have instant explanations.',
31
+ highlight: 'nav-learn',
32
+ },
33
+ ];
34
+
35
+ const STORAGE_KEY = 'bio-nexus-onboarding';
36
+
37
+ export function TutorialWalkthrough() {
38
+ const [step, setStep] = useState(0);
39
+ const [visible, setVisible] = useState(false);
40
+
41
+ useEffect(() => {
42
+ const seen = localStorage.getItem(STORAGE_KEY);
43
+ if (!seen) {
44
+ setVisible(true);
45
+ }
46
+ }, []);
47
+
48
+ const dismiss = useCallback(() => {
49
+ setVisible(false);
50
+ localStorage.setItem(STORAGE_KEY, 'true');
51
+ }, []);
52
+
53
+ const next = useCallback(() => {
54
+ if (step < STEPS.length - 1) {
55
+ setStep(s => s + 1);
56
+ } else {
57
+ dismiss();
58
+ }
59
+ }, [step, dismiss]);
60
+
61
+ const prev = useCallback(() => {
62
+ if (step > 0) {
63
+ setStep(s => s - 1);
64
+ }
65
+ }, []);
66
+
67
+ const current = STEPS[step];
68
+
69
+ return (
70
+ <AnimatePresence>
71
+ {visible && (
72
+ <motion.div
73
+ key="tutorial-backdrop"
74
+ initial={{ opacity: 0 }}
75
+ animate={{ opacity: 1 }}
76
+ exit={{ opacity: 0 }}
77
+ transition={{ duration: 0.2 }}
78
+ className="fixed inset-0 z-[200] flex items-center justify-center"
79
+ style={{ background: 'rgba(4,4,10,0.75)', backdropFilter: 'blur(8px)' }}
80
+ >
81
+ <motion.div
82
+ key={`step-${step}`}
83
+ initial={{ opacity: 0, y: 24, scale: 0.97 }}
84
+ animate={{ opacity: 1, y: 0, scale: 1 }}
85
+ exit={{ opacity: 0, y: -16, scale: 0.97 }}
86
+ transition={{ duration: 0.35, ease: [0.25, 1, 0.5, 1] }}
87
+ className="relative w-full max-w-lg mx-4 p-8 rounded-2xl border border-glass-border shadow-glass-lg"
88
+ style={{
89
+ background: 'rgba(13,13,26,0.92)',
90
+ backdropFilter: 'blur(32px) saturate(180%)',
91
+ }}
92
+ >
93
+ <button
94
+ onClick={dismiss}
95
+ className="absolute top-4 right-4 p-1.5 rounded-lg text-text-muted hover:text-text-primary hover:bg-surface-1 transition"
96
+ aria-label="Skip tutorial"
97
+ >
98
+ <X className="w-4 h-4" />
99
+ </button>
100
+
101
+ <div className="flex items-center gap-3 mb-6">
102
+ <div className="p-2.5 rounded-xl bg-accent-cyan/10">
103
+ <BookOpen className="w-5 h-5 text-accent-cyan" />
104
+ </div>
105
+ <div>
106
+ <h2 className="text-lg font-semibold text-text-primary">{current.title}</h2>
107
+ <p className="text-xs text-text-muted">Step {step + 1} of {STEPS.length}</p>
108
+ </div>
109
+ </div>
110
+
111
+ <p className="text-sm text-text-secondary leading-relaxed mb-8">
112
+ {current.description}
113
+ </p>
114
+
115
+ {/* Step dots */}
116
+ <div className="flex items-center justify-center gap-1.5 mb-6">
117
+ {STEPS.map((_, i) => (
118
+ <div
119
+ key={i}
120
+ className={`h-1.5 rounded-full transition-all duration-300 ${
121
+ i === step
122
+ ? 'w-6 bg-accent-cyan'
123
+ : i < step
124
+ ? 'w-1.5 bg-accent-cyan/40'
125
+ : 'w-1.5 bg-glass-border'
126
+ }`}
127
+ />
128
+ ))}
129
+ </div>
130
+
131
+ <div className="flex items-center justify-between">
132
+ <button
133
+ onClick={prev}
134
+ disabled={step === 0}
135
+ className="btn-ghost px-4 py-2 text-sm disabled:opacity-30"
136
+ >
137
+ <ChevronLeft className="w-4 h-4" />
138
+ Back
139
+ </button>
140
+
141
+ <div className="flex items-center gap-3">
142
+ <button
143
+ onClick={dismiss}
144
+ className="text-xs text-text-muted hover:text-text-primary transition"
145
+ >
146
+ Skip
147
+ </button>
148
+
149
+ <button
150
+ onClick={next}
151
+ className="btn-primary px-5 py-2 text-sm"
152
+ >
153
+ {step < STEPS.length - 1 ? (
154
+ <>Next <ChevronRight className="w-4 h-4" /></>
155
+ ) : (
156
+ 'Get Started'
157
+ )}
158
+ </button>
159
+ </div>
160
+ </div>
161
+ </motion.div>
162
+ </motion.div>
163
+ )}
164
+ </AnimatePresence>
165
+ );
166
+ }
167
+
168
+ export function startTutorial() {
169
+ localStorage.removeItem(STORAGE_KEY);
170
+ window.location.reload();
171
+ }
bioai-platform/frontend/src/lib/api.ts CHANGED
@@ -245,3 +245,35 @@ export async function runEnrichment(identifiers: string[]): Promise<EnrichmentRe
245
  const res = await api.post('/api/pathways/enrichment', { identifiers });
246
  return res.data;
247
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
245
  const res = await api.post('/api/pathways/enrichment', { identifiers });
246
  return res.data;
247
  }
248
+
249
+ export type DockingPose = {
250
+ model: number;
251
+ atoms: number;
252
+ affinity: number | null;
253
+ };
254
+
255
+ export type DockingResult = {
256
+ job_id: string;
257
+ status: string;
258
+ result?: {
259
+ pdb_id: string;
260
+ smiles: string;
261
+ poses: DockingPose[];
262
+ num_poses: number;
263
+ box_center: { x: number; y: number; z: number };
264
+ box_size: { x: number; y: number; z: number };
265
+ vina_log?: string;
266
+ from_cache?: boolean;
267
+ };
268
+ error?: string;
269
+ };
270
+
271
+ export async function runDocking(pdbId: string, smiles: string): Promise<{ job_id: string; status: string }> {
272
+ const res = await api.post('/api/docking/run', { pdb_id: pdbId, smiles });
273
+ return res.data;
274
+ }
275
+
276
+ export async function getDockingStatus(jobId: string): Promise<DockingResult> {
277
+ const res = await api.get(`/api/docking/status/${jobId}`);
278
+ return res.data;
279
+ }
design.md CHANGED
@@ -82,6 +82,16 @@ Type scale: `12 / 14 / 16 / 20 / 28 / 40 / 56px`, line-height 1.5 for body, 1.1
82
  ## 7. Iconography
83
  - Lucide icons (already in your stack), 1.5px stroke, sizes 16/20/24.
84
 
85
- ## 8. Accessibility
 
 
 
 
 
 
 
 
86
  - All `--accent` on `--bg-base` text combinations pass WCAG AA at β‰₯14px.
87
  - Confidence bands are never color-only β€” always paired with the text label (colorblind-safe, and this is scientific data where ambiguity matters).
 
 
 
82
  ## 7. Iconography
83
  - Lucide icons (already in your stack), 1.5px stroke, sizes 16/20/24.
84
 
85
+ ## 8. Documentation Site Design (`/learn`)
86
+
87
+ - **Layout**: Centered content (max-width 880px), topic pages with back navigation, sticky right-side section nav for long pages
88
+ - **Topic cards**: Same glass-card pattern as /analyze β€” icon + title + description in a 2-column grid
89
+ - **Code examples**: `font-mono` on `bg-elevated` background, with horizontal scroll for long lines
90
+ - **Glossary**: A–Z listing with letter dividers, term in `font-mono` semibold, definition in `text-text-secondary`
91
+ - **LearnPopover**: Small `(?)` icon in `text-accent-cyan/60` next to term β†’ click opens floating glass-card popover (max-width 320px) with 1-2 sentence explanation + optional "Learn more β†’" link in accent cyan
92
+
93
+ ## 9. Accessibility
94
  - All `--accent` on `--bg-base` text combinations pass WCAG AA at β‰₯14px.
95
  - Confidence bands are never color-only β€” always paired with the text label (colorblind-safe, and this is scientific data where ambiguity matters).
96
+ - Tutorial walkthrough has keyboard navigation (Tab/Enter + Escape to close)
97
+ - Sequence data rendered in monospace (`font-mono`) β€” never sans-serif
implementationplan.md CHANGED
@@ -111,11 +111,13 @@ One golden path only: **Sequence In β†’ BLAST β†’ AI-interpreted report.** Every
111
  - Gene/protein β†’ pathway lookup via Reactome/WikiPathways
112
  - Pathway diagram viewer
113
 
114
- ### Sprint 8: Onboarding + `/learn`
115
- - First-run tutorial
116
- - Documentation pages generated from existing "what does this mean" content β€” write once, reuse
117
-
118
- ### Sprint 9–10: Hardening
119
- - PDF report export
120
- - Cache-hit check before re-calling external APIs (raw responses already stored from Day 1)
121
- - Error monitoring (Sentry free tier)
 
 
 
111
  - Gene/protein β†’ pathway lookup via Reactome/WikiPathways
112
  - Pathway diagram viewer
113
 
114
+ ### Sprint 8: Onboarding + `/learn` βœ…
115
+ - First-run tutorial (TutorialWalkthrough component, 5 steps, localStorage flag) βœ…
116
+ - 10+ documentation pages at `/learn` (BLAST, Alignment, Domains, Phylo, Structure, Pathways, Interactions, Primers, Tools, Glossary) βœ…
117
+ - LearnPopover component for inline `(?)` help tooltips βœ…
118
+
119
+ ### Sprint 9–10: Hardening βœ…
120
+ - PDF report export (`GET /api/export/job/{id}?format=pdf|json`) βœ…
121
+ - Cache-hit check with `from_cache` flag, `/api/admin/cache-stats` endpoint βœ…
122
+ - `@ttl_cache` coverage: pathway enrichment, NCBI search, BLAST, UniProt, AlphaFold βœ…
123
+ - Sentry error monitoring (`@sentry/nextjs` frontend + `sentry-sdk` backend) βœ…
rules.md CHANGED
@@ -1,15 +1,15 @@
1
  # BioFlow AI β€” Rules
2
 
3
- **Version:** 1.0
4
- **Scope:** Both repos β€” `bioflow-frontend` and `bioflow-backend`
5
  **Last Updated:** June 2026
6
 
7
  ---
8
 
9
  ## 0 β€” The Prime Rule
10
 
11
- **The prototype ships June 30. Every decision is evaluated against this.**
12
- If a rule conflicts with shipping the demo, ship the demo. Document the deviation. Fix it after.
13
 
14
  ---
15
 
@@ -473,23 +473,61 @@ Data:
473
 
474
  ---
475
 
476
- ## 8 β€” What NOT To Build (Pre-Demo)
477
-
478
- This list exists because scope creep during the 18-day sprint will kill the demo.
479
-
480
- ❌ MSA wizard (Phase 2)
481
- ❌ Phylogenetic tree viewer (Phase 2)
482
- ❌ Protein 3D structure viewer (Phase 3)
483
- ❌ AlphaFold integration (Phase 3)
484
- ❌ Drug docking features (Phase 4)
485
- ❌ PDF export (post-demo Phase 1)
486
- ❌ Onboarding tutorial (post-demo Phase 1)
487
- ❌ Email notifications
488
- ❌ Rate limiting UI
489
- ❌ Admin panel
490
- ❌ Analytics dashboard
491
- ❌ Dark/light mode toggle (dark by default β€” toggle later)
492
- ❌ Mobile app
493
- ❌ Public API
494
-
495
- If you find yourself working on any of the above before June 30: stop, commit your current work, and return to the demo track.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # BioFlow AI β€” Rules
2
 
3
+ **Version:** 2.0
4
+ **Scope:** Both repos β€” `bioflow-frontend` and `bioflow-backend` (monorepo at `bio-nexus/bioai-platform/`)
5
  **Last Updated:** June 2026
6
 
7
  ---
8
 
9
  ## 0 β€” The Prime Rule
10
 
11
+ **The prototype shipped June 30 β€” Phase 2 is complete.**
12
+ Hardening (docs, Sentry, cache checks) is done. Phase 3+ items should be planned before building.
13
 
14
  ---
15
 
 
473
 
474
  ---
475
 
476
+ ## 8 β€” Monitoring & Observability Rules (Sentry)
477
+
478
+ 1. **Errors must be captured.** Every unhandled exception in production should reach Sentry. Frontend: `@sentry/nextjs` with `beforeSend` filtering. Backend: `sentry_sdk.init()` in startup.
479
+
480
+ 2. **No secrets in Sentry.** Ensure `beforeSend` strips auth tokens, API keys, and sequence data from Sentry events.
481
+
482
+ 3. **Traces** at `tracesSampleRate: 0.1` (10%) β€” enough for debugging, cheap enough for free tier.
483
+
484
+ 4. **Environment tagging.** All Sentry events must be tagged with `environment: development | production`.
485
+
486
+ ---
487
+
488
+ ## 9 β€” Caching Rules
489
+
490
+ 1. **Cache-first architecture.** Every external API call must check the corresponding cache before executing. Use `@ttl_cache` decorator from `services/cache.py`.
491
+
492
+ 2. **Cache key format.** `{prefix}:{sha256_first_16_chars_of_json_input}` β€” consistent, deterministic.
493
+
494
+ 3. **TTL guidelines:**
495
+ - BLAST results: 24h
496
+ - UniProt records: 24h
497
+ - AlphaFold predictions: 30 days
498
+ - Pathway enrichment: 12h
499
+ - NCBI sequence/search: 24h
500
+
501
+ 4. **Cache misses are tracked.** `get_cache_stats()` exposes hit/miss counts. Monitor via `/api/admin/cache-stats`.
502
+
503
+ 5. **`from_cache` flag.** All cached results include `from_cache: true/false` in the response dict for observability.
504
+
505
+ 6. **Graceful fallback.** If Redis is unavailable (`_redis = None`), caching is silently disabled β€” the app still works.
506
+
507
+ ---
508
+
509
+ ## 10 β€” Documentation & Learning Rules
510
+
511
+ 1. **`/learn` is the canonical docs source.** All inline "Learn more β†’" links must point to a valid `/learn/{topic}` route.
512
+
513
+ 2. **LearnPopover consistency.** Every scientific term shown to users (E-value, bit score, pLDDT, bootstrap, etc.) must have a LearnPopover component available.
514
+
515
+ 3. **First-run tutorial.** New users see the TutorialWalkthrough once. It must be re-accessible from the Settings page.
516
+
517
+ 4. **Plain language.** All docs and help text must be understandable by a first-year M.Sc. student. No jargon without explanation.
518
+
519
+ ---
520
+
521
+ ## 11 β€” What NOT To Build (Current)
522
+
523
+ This list keeps scope in check for Phase 3+.
524
+
525
+ ❌ **Phase 2 items** β€” already built (MSA, Phylo, Domains, Pathways, Primers, API keys, Share, Export, Guest upgrade, Docs, Sentry, Cache checks)
526
+ ❌ **Molecular docking / DiffDock** β€” requires revenue for paid Replicate API
527
+ ❌ **RNA-seq pipeline** β€” Phase 3, requires file storage infrastructure
528
+ ❌ **FASTQ / variant calling** β€” Phase 3, requires compute
529
+ ❌ **Lab workspaces** β€” Phase 4, requires institution licensing
530
+ ❌ **Custom pipeline builder** β€” Phase 4
531
+ ❌ **Mobile app** β€” not planned
532
+ ❌ **Email notifications** β€” not planned until Phase 4
533
+ ❌ **Admin panel** β€” not needed until 100+ users
schema.md CHANGED
@@ -586,6 +586,37 @@ These are stored in `jobs.input_params` as JSONB:
586
 
587
  ## Migration Notes
588
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
589
  - Run migrations in order via Supabase SQL editor or `supabase db push`
590
  - Never modify enum types after data exists β€” add new values only
591
  - Cache tables do not need migrations for TTL changes β€” update application config
 
586
 
587
  ## Migration Notes
588
 
589
+ ## Additional Tables (Phase 2)
590
+
591
+ ### API Keys
592
+
593
+ Applied via migration `006_api_keys.sql`:
594
+
595
+ ```sql
596
+ create table if not exists api_keys (
597
+ id uuid primary key default gen_random_uuid(),
598
+ user_id uuid not null,
599
+ name text not null,
600
+ key_hash text not null,
601
+ key_prefix text not null,
602
+ created_at timestamptz default now(),
603
+ last_used_at timestamptz
604
+ );
605
+
606
+ create index if not exists idx_api_keys_user on api_keys(user_id);
607
+ create index if not exists idx_api_keys_hash on api_keys(key_hash);
608
+
609
+ alter table api_keys enable row level security;
610
+
611
+ create policy "Users can manage own API keys" on api_keys for all
612
+ using (auth.uid() = user_id)
613
+ with check (auth.uid() = user_id);
614
+ ```
615
+
616
+ ---
617
+
618
+ ## Migration Notes
619
+
620
  - Run migrations in order via Supabase SQL editor or `supabase db push`
621
  - Never modify enum types after data exists β€” add new values only
622
  - Cache tables do not need migrations for TTL changes β€” update application config
techspec.md CHANGED
@@ -1,125 +1,141 @@
1
  # BioFlow AI β€” Technical Specification
2
 
3
- **Version:** 1.0
4
- **Repos:** `bioflow-frontend` (Next.js 14) Β· `bioflow-backend` (FastAPI)
5
  **Last Updated:** June 2026
6
 
7
  ---
8
 
9
  ## Repository Structure
10
 
11
- ### `bioflow-frontend`
12
 
13
  ```
14
- bioflow-frontend/
15
- β”œβ”€β”€ app/
16
- β”‚ β”œβ”€β”€ (auth)/
17
- β”‚ β”‚ β”œβ”€β”€ login/
18
- β”‚ β”‚ β”‚ └── page.tsx
19
- β”‚ β”‚ └── signup/
20
- β”‚ β”‚ └── page.tsx
21
- β”‚ β”œβ”€β”€ (app)/
22
- β”‚ β”‚ β”œβ”€β”€ dashboard/
23
- β”‚ β”‚ β”‚ └── page.tsx # Job history, saved analyses
24
- β”‚ β”‚ β”œβ”€β”€ job/
25
- β”‚ β”‚ β”‚ └── [jobId]/
26
- β”‚ β”‚ β”‚ └── page.tsx # Live job status + results
27
- β”‚ β”‚ └── layout.tsx # App shell with nav
28
- β”‚ β”œβ”€β”€ wizard/
29
- β”‚ β”‚ β”œβ”€β”€ page.tsx # Workflow selection entry
30
- β”‚ β”‚ └── [workflowType]/
31
- β”‚ β”‚ └── page.tsx # Specific wizard flow
32
- β”‚ β”œβ”€β”€ api/
33
- β”‚ β”‚ └── auth/
34
- β”‚ β”‚ └── [...nextauth]/
35
- β”‚ β”‚ └── route.ts # NextAuth handlers
36
- β”‚ β”œβ”€β”€ layout.tsx # Root layout
37
- β”‚ β”œβ”€β”€ page.tsx # Landing page
38
- β”‚ └── globals.css
39
- β”œβ”€β”€ components/
40
- β”‚ β”œβ”€β”€ ui/ # Primitive components (button, input, card, etc.)
41
- β”‚ β”œβ”€β”€ visualizations/
42
- β”‚ β”‚ β”œβ”€β”€ SequenceViewer.tsx # Colored sequence display with annotations
43
- β”‚ β”‚ β”œβ”€β”€ BlastResultsTable.tsx # Interactive BLAST hits table
44
- β”‚ β”‚ β”œβ”€β”€ AlignmentViewer.tsx # Color-coded pairwise alignment
45
- β”‚ β”‚ β”œβ”€β”€ MSAViewer.tsx # Multiple sequence alignment block view
46
- β”‚ β”‚ β”œβ”€β”€ StructureViewer.tsx # NGL.js / Mol* 3D structure wrapper
47
- β”‚ β”‚ β”œβ”€β”€ PhyloTreeViewer.tsx # phylotree.js tree renderer
48
- β”‚ β”‚ β”œβ”€β”€ RamachandranPlot.tsx # D3.js phi/psi scatter plot
49
- β”‚ β”‚ β”œβ”€β”€ ConservationPlot.tsx # Per-position conservation bar chart
50
- β”‚ β”‚ └── PathwayViewer.tsx # Reactome pathway embed
51
- β”‚ β”œβ”€β”€ wizard/
52
- β”‚ β”‚ β”œβ”€β”€ WizardShell.tsx # Step container, progress bar, navigation
53
- β”‚ β”‚ β”œβ”€β”€ WizardStep.tsx # Individual step wrapper
54
- β”‚ β”‚ └── steps/
55
- β”‚ β”‚ β”œβ”€β”€ SequenceInputStep.tsx
56
- β”‚ β”‚ β”œβ”€β”€ DatabaseSelectStep.tsx
57
- β”‚ β”‚ β”œβ”€β”€ AlgorithmSelectStep.tsx
58
- β”‚ β”‚ └── ConfirmRunStep.tsx
59
- β”‚ β”œβ”€β”€ job/
60
- β”‚ β”‚ β”œβ”€β”€ JobStatusCard.tsx # Live step-by-step progress display
61
- β”‚ β”‚ β”œβ”€β”€ StepTimeline.tsx # Visual pipeline step timeline
62
- β”‚ β”‚ └── ResultsPage.tsx # Full results layout wrapper
63
- β”‚ └── layout/
64
- β”‚ β”œβ”€β”€ Navbar.tsx
65
- β”‚ β”œβ”€β”€ Sidebar.tsx
66
- β”‚ └── GuestBanner.tsx # Shown to guest users prompting signup
67
- β”œβ”€β”€ lib/
68
- β”‚ β”œβ”€β”€ api.ts # Type-safe backend API client
69
- β”‚ β”œβ”€β”€ supabase.ts # Supabase client (browser)
70
- β”‚ β”œβ”€β”€ supabase-server.ts # Supabase client (server components)
71
- β”‚ β”œβ”€β”€ types.ts # All shared TypeScript types
72
- β”‚ └── constants.ts # Workflow types, step labels, etc.
73
- β”œβ”€β”€ hooks/
74
- β”‚ β”œβ”€β”€ useJobPolling.ts # React Query polling for job status
75
- β”‚ β”œβ”€β”€ useGuestSession.ts # Guest session cookie management
76
- β”‚ └── useWorkflowWizard.ts # Wizard state machine
77
- └── public/
78
- └── fonts/ # Self-hosted Space Grotesk + JetBrains Mono
79
  ```
80
 
81
- ### `bioflow-backend`
82
 
83
  ```
84
- bioflow-backend/
85
  β”œβ”€β”€ app/
86
- β”‚ β”œβ”€β”€ main.py # FastAPI app, CORS, lifespan events
87
- β”‚ β”œβ”€β”€ api/
88
- β”‚ β”‚ β”œβ”€β”€ deps.py # Auth dependencies, DB session
89
- β”‚ β”‚ └── routes/
90
- β”‚ β”‚ β”œβ”€β”€ jobs.py # POST /jobs, GET /jobs/{id}, GET /jobs/{id}/steps
91
- β”‚ β”‚ β”œβ”€β”€ sequences.py # POST /sequences/fetch, POST /sequences/validate
92
- β”‚ β”‚ β”œβ”€β”€ blast.py # POST /blast/submit, GET /blast/{job_id}/status
93
- β”‚ β”‚ β”œβ”€β”€ alignment.py # POST /align/pairwise, POST /align/msa
94
- β”‚ β”‚ β”œβ”€β”€ structures.py # POST /structures/fetch, POST /structures/predict
95
- β”‚ β”‚ └── dashboard.py # GET /dashboard/jobs (paginated job history)
 
 
 
 
 
 
 
 
 
 
 
 
96
  β”‚ β”œβ”€β”€ services/
97
- β”‚ β”‚ β”œβ”€β”€ ncbi_service.py # Entrez API wrapper (fetch, BLAST)
98
- β”‚ β”‚ β”œβ”€β”€ embl_service.py # EMBL-EBI tools (BLAST, ClustalOmega, MUSCLE)
99
- β”‚ β”‚ β”œβ”€β”€ pdb_service.py # RCSB PDB REST API
100
- β”‚ β”‚ β”œβ”€β”€ uniprot_service.py # UniProt REST API
101
- β”‚ β”‚ β”œβ”€β”€ alphafold_service.py # AlphaFold EBI API
102
- β”‚ β”‚ β”œβ”€β”€ reactome_service.py # Reactome REST + WikiPathways fallback
103
- β”‚ β”‚ β”œβ”€β”€ groq_service.py # Groq API for AI interpretation
104
- β”‚ β”‚ β”œβ”€β”€ r2_service.py # Cloudflare R2 upload/download
105
- β”‚ β”‚ └── cache_service.py # Cache read/write logic
 
 
 
 
 
 
 
 
106
  β”‚ β”œβ”€β”€ workers/
107
- β”‚ β”‚ └── pipeline_worker.py # Celery task: executes pipeline steps
108
- β”‚ β”œβ”€β”€ parsers/
109
- β”‚ β”‚ β”œβ”€β”€ blast_parser.py # NCBI BLAST XML β†’ structured dict
110
- β”‚ β”‚ β”œβ”€β”€ alignment_parser.py # ClustalW/MSA output β†’ structured dict
111
- β”‚ β”‚ β”œβ”€β”€ pdb_parser.py # PDB/mmCIF parsing via BioPython
112
- β”‚ β”‚ └── newick_parser.py # Newick tree validation/normalization
113
- β”‚ β”œβ”€β”€ models/
114
- β”‚ β”‚ └── schemas.py # Pydantic v2 request/response models
115
- β”‚ └── core/
116
- β”‚ β”œβ”€β”€ config.py # Settings (env vars via pydantic-settings)
117
- β”‚ β”œβ”€β”€ database.py # Supabase + async DB client
118
- β”‚ └── exceptions.py # Custom exception classes
119
- β”œβ”€β”€ requirements.txt
120
- β”œβ”€β”€ .env.example
121
- β”œβ”€β”€ Dockerfile
122
- └── celery_worker.py # Celery worker entrypoint
123
  ```
124
 
125
  ---
@@ -666,19 +682,20 @@ RATE_LIMIT_EXCEEDED β€” user has exceeded daily job quota
666
  ## Deployment
667
 
668
  ### Frontend β€” Vercel
669
- - Repo: `bioflow-frontend`
670
- - Framework preset: Next.js
671
  - Environment variables: set in Vercel dashboard
672
- - Domain: bioflow.ai (or subdomain for staging: staging.bioflow.ai)
673
 
674
- ### Backend β€” Railway
675
- - Repo: `bioflow-backend`
676
- - Service 1: FastAPI web (`uvicorn app.main:app --host 0.0.0.0 --port $PORT`)
677
- - Service 2: Celery worker (`celery -A celery_worker worker --loglevel=info`)
678
- - Both services share same repo, different start commands
679
- - Redis: Upstash (external, connects via REDIS_URL)
 
680
 
681
  ### Staging vs Production
682
- - Two Vercel deployments: `main` β†’ production, `dev` β†’ staging
683
- - Two Railway environments: production and staging
684
- - Two Supabase projects: `bioflow-prod` and `bioflow-dev`
 
1
  # BioFlow AI β€” Technical Specification
2
 
3
+ **Version:** 2.0
4
+ **Repos:** Monorepo at `bio-nexus/bioai-platform/` β€” `frontend/` (Next.js 14) Β· `backend/` (FastAPI)
5
  **Last Updated:** June 2026
6
 
7
  ---
8
 
9
  ## Repository Structure
10
 
11
+ ### `bioai-platform/frontend` (Current Structure)
12
 
13
  ```
14
+ bioai-platform/frontend/
15
+ β”œβ”€β”€ src/
16
+ β”‚ β”œβ”€β”€ app/
17
+ β”‚ β”‚ β”œβ”€β”€ (auth)/
18
+ β”‚ β”‚ β”‚ β”œβ”€β”€ auth/
19
+ β”‚ β”‚ β”‚ β”‚ β”œβ”€β”€ callback/page.tsx
20
+ β”‚ β”‚ β”‚ β”‚ └── page.tsx
21
+ β”‚ β”‚ β”‚ └── layout.tsx
22
+ β”‚ β”‚ β”œβ”€β”€ (dashboard)/
23
+ β”‚ β”‚ β”‚ β”œβ”€β”€ layout.tsx # App shell with collapsible sidebar + header
24
+ β”‚ β”‚ β”‚ β”œβ”€β”€ dashboard/page.tsx # Stats, quick tools grid, recent jobs
25
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/page.tsx # Operation hub (all tools listed)
26
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/blast/page.tsx
27
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/uniprot/page.tsx
28
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/structure/page.tsx
29
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/alignment/page.tsx
30
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/domains/page.tsx
31
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/phylo/page.tsx
32
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/pathway/page.tsx
33
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/interactions/page.tsx
34
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/compare/page.tsx
35
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/tools/page.tsx
36
+ β”‚ β”‚ β”‚ β”œβ”€β”€ analyze/primers/page.tsx
37
+ β”‚ β”‚ β”‚ β”œβ”€β”€ wizard/page.tsx # 4-step guided pipeline wizard
38
+ β”‚ β”‚ β”‚ β”œβ”€β”€ jobs/page.tsx # Job list with filter tabs
39
+ β”‚ β”‚ β”‚ β”œβ”€β”€ jobs/[jobId]/page.tsx # Job detail + share
40
+ β”‚ β”‚ β”‚ β”œβ”€β”€ results/[jobId]/page.tsx
41
+ β”‚ β”‚ β”‚ β”œβ”€β”€ report/[jobId]/page.tsx # Print-to-PDF report
42
+ β”‚ β”‚ β”‚ β”œβ”€β”€ history/page.tsx
43
+ β”‚ β”‚ β”‚ β”œβ”€β”€ retrieve/page.tsx
44
+ β”‚ β”‚ β”‚ β”œβ”€β”€ settings/page.tsx # API keys, profile, guest upgrade, usage
45
+ β”‚ β”‚ β”‚ β”œβ”€β”€ shared/[token]/page.tsx
46
+ β”‚ β”‚ β”‚ └── learn/ # Documentation site
47
+ β”‚ β”‚ β”‚ β”œβ”€β”€ page.tsx # Docs landing with topic grid
48
+ β”‚ β”‚ β”‚ └── [topic]/page.tsx # Dynamic topic pages
49
+ β”‚ β”‚ β”œβ”€β”€ layout.tsx # Root layout (fonts, providers)
50
+ β”‚ β”‚ β”œβ”€β”€ providers.tsx # Theme + Auth providers
51
+ β”‚ β”‚ └── globals.css # Tailwind + glassmorphism overrides
52
+ β”‚ β”œβ”€β”€ components/
53
+ β”‚ β”‚ β”œβ”€β”€ phylo/PhyloTreeViewer.tsx
54
+ β”‚ β”‚ β”œβ”€β”€ results/PipelineResults.tsx
55
+ β”‚ β”‚ β”œβ”€β”€ learn/LearnPopover.tsx # Inline help popover
56
+ β”‚ β”‚ β”œβ”€β”€ TutorialWalkthrough.tsx # First-run onboarding modal
57
+ β”‚ β”‚ β”œβ”€β”€ ErrorBoundary.tsx
58
+ β”‚ β”‚ β”œβ”€β”€ GuestBanner.tsx
59
+ β”‚ β”‚ β”œβ”€β”€ ThemeToggle.tsx
60
+ β”‚ β”‚ └── ... (BlastPanel, ScoreBars, UniprotPanel, DomainSummary, etc.)
61
+ β”‚ β”œβ”€β”€ contexts/
62
+ β”‚ β”‚ β”œβ”€β”€ auth.tsx # Auth context (Supabase session)
63
+ β”‚ β”‚ └── theme.tsx # Theme context + localStorage
64
+ β”‚ β”œβ”€β”€ lib/
65
+ β”‚ β”‚ β”œβ”€β”€ api.ts # Type-safe backend API client
66
+ β”‚ β”‚ β”œβ”€β”€ supabase.ts # Supabase client (browser)
67
+ β”‚ β”‚ β”œβ”€β”€ types.ts
68
+ β”‚ β”‚ └── animations.ts # Framer motion variants
69
+ β”‚ └── hooks/
70
+ β”‚ └── useJobPolling.ts
71
+ β”œβ”€β”€ sentry.client.config.ts # Sentry client config
72
+ β”œβ”€β”€ sentry.server.config.ts # Sentry server config
73
+ β”œβ”€β”€ next.config.js # Sentry-wrapped Next config
74
+ └── package.json
 
 
 
 
75
  ```
76
 
77
+ ### `bioai-platform/backend` (Current Structure)
78
 
79
  ```
80
+ bioai-platform/backend/
81
  β”œβ”€β”€ app/
82
+ β”‚ β”œβ”€β”€ main.py # FastAPI app, CORS, lifespan (Sentry init, Redis init)
83
+ β”‚ β”œβ”€β”€ config.py # Settings via pydantic-settings + dotenv
84
+ β”‚ β”œβ”€β”€ routers/ # 19 route modules
85
+ β”‚ β”‚ β”œβ”€β”€ pipelines.py # POST /api/pipelines/run
86
+ β”‚ β”‚ β”œβ”€β”€ pipeline_v2.py # POST /api/pipeline/v2/run, GET /status/{job_id}
87
+ β”‚ β”‚ β”œβ”€β”€ ai.py # POST /api/ai/interpret, /interpret/stream
88
+ β”‚ β”‚ β”œβ”€β”€ jobs.py # GET/POST/DELETE /api/jobs
89
+ β”‚ β”‚ β”œβ”€β”€ share.py # POST /api/share, GET /api/share/{token}
90
+ β”‚ β”‚ β”œβ”€β”€ profile.py # GET/PUT /api/profile
91
+ β”‚ β”‚ β”œβ”€β”€ sequences.py # POST /api/sequences/fetch, /validate, /search
92
+ β”‚ β”‚ β”œβ”€β”€ uniprot.py # POST /api/uniprot/search, /detail
93
+ β”‚ β”‚ β”œβ”€β”€ alignment.py # POST /api/alignment/run
94
+ β”‚ β”‚ β”œβ”€β”€ structures.py # POST /api/structures/fetch, /search
95
+ β”‚ β”‚ β”œβ”€β”€ pathways.py # POST /api/pathways/search, /detail, /kegg/search, /enrichment
96
+ β”‚ β”‚ β”œβ”€β”€ domains.py # GET /api/domains/{accession}
97
+ β”‚ β”‚ β”œβ”€β”€ interactions.py # GET /api/interactions/{gene_name}
98
+ β”‚ β”‚ β”œβ”€β”€ primers.py # POST /api/primers/design
99
+ β”‚ β”‚ β”œβ”€β”€ structure_analysis.py # GET /api/structure_analysis/ramachandran, /secondary, /compare
100
+ β”‚ β”‚ β”œβ”€β”€ phylo.py # POST /phylo/run, GET /status/{job_id}, /models
101
+ β”‚ β”‚ β”œβ”€β”€ export.py # GET /api/export/job/{id}?format=pdf|json
102
+ β”‚ β”‚ β”œβ”€β”€ api_keys.py # GET/POST /api/keys, DELETE /api/keys/{id}
103
+ β”‚ β”‚ └── cache_stats.py # GET /api/admin/cache-stats, POST /reset
104
  β”‚ β”œβ”€β”€ services/
105
+ β”‚ β”‚ β”œβ”€β”€ cache.py # Redis cache wrapper, @ttl_cache decorator, stats tracking
106
+ β”‚ β”‚ β”œβ”€β”€ auth.py # JWT auth, X-API-Key middleware
107
+ β”‚ β”‚ β”œβ”€β”€ export.py # PDF/JSON report generation (reportlab)
108
+ β”‚ β”‚ β”œβ”€β”€ ncbi_service.py # NCBI Entrez (fetch, search) β€” @ttl_cache on both
109
+ β”‚ β”‚ β”œβ”€β”€ pathway_enrichment.py # Reactome enrichment β€” cached via cache_get/set
110
+ β”‚ β”‚ β”œβ”€β”€ supabase.py # Supabase REST client
111
+ β”‚ β”‚ β”œβ”€β”€ rate_limit.py # Per-user rate limiting
112
+ β”‚ β”‚ β”œβ”€β”€ redis.py # Redis connection
113
+ β”‚ β”‚ β”œβ”€β”€ sequence_utils.py # Sequence validation, type detection
114
+ β”‚ β”‚ └── validators.py # Input validation
115
+ β”‚ β”œβ”€β”€ tools/ # Tool classes with @ttl_cache on run()
116
+ β”‚ β”‚ β”œβ”€β”€ blast.py # EBI BLAST submit/poll/parse
117
+ β”‚ β”‚ β”œβ”€β”€ uniprot.py # UniProt REST lookup
118
+ β”‚ β”‚ β”œβ”€β”€ alphafold.py # AlphaFold DB query
119
+ β”‚ β”‚ β”œβ”€β”€ base.py # Abstract BaseTool
120
+ β”‚ β”‚ └── registration.py # Tool registry
121
+ β”‚ β”œβ”€β”€ pipeline/ # Pipeline v1 engine (deprecated in favor of v2)
122
  β”‚ β”œβ”€β”€ workers/
123
+ β”‚ β”‚ β”œβ”€β”€ pipeline_worker.py # Thread-based pipeline execution
124
+ β”‚ β”‚ └── celery_app.py # Celery app config (unused, kept for reference)
125
+ β”‚ β”œβ”€β”€ ai/ # AI interpretation layer
126
+ β”‚ β”‚ β”œβ”€β”€ interpreter.py
127
+ β”‚ β”‚ β”œβ”€β”€ llm_client.py # LiteLLM wrapper (Groq)
128
+ β”‚ β”‚ └── prompts.py # Prompt templates
129
+ β”‚ β”œβ”€β”€ models/responses.py # Pydantic response models
130
+ β”‚ β”œβ”€β”€ integrations/ncbi/ # NCBI-specific modules
131
+ β”‚ β”‚ β”œβ”€β”€ blast.py # BLAST submission & polling
132
+ β”‚ β”‚ └── parser.py # XML parsing
133
+ β”‚ β”œβ”€β”€ data/demo_results.py # Demo mode fallback sequences
134
+ β”‚ └── core/storage.py # R2 storage wrapper
135
+ β”œβ”€β”€ requirements.txt # + sentry-sdk
136
+ β”œβ”€β”€ .env.deploy # Deployment env template (+ SENTRY_DSN)
137
+ β”œβ”€β”€ Dockerfile # Pre-compiled PhyML binary download
138
+ └── railway.json / render.yaml # Deploy configs
139
  ```
140
 
141
  ---
 
682
  ## Deployment
683
 
684
  ### Frontend β€” Vercel
685
+ - Deployed from `bioai-platform/` (Root Directory: auto-detect)
686
+ - Production URL: https://bioai-platform.vercel.app
687
  - Environment variables: set in Vercel dashboard
688
+ - Sentry DSN set as `NEXT_PUBLIC_SENTRY_DSN` + `SENTRY_DSN`
689
 
690
+ ### Backend β€” Hugging Face Spaces
691
+ - Space: `Samad14/bio-nexus-api`
692
+ - Public URL: https://samad14-bio-nexus-api.hf.space
693
+ - SDK: Docker (cpu-basic, sleeps after 48h)
694
+ - Deployed via `hf upload --type space ...` from local
695
+ - Env vars set in HF Space dashboard (secrets)
696
+ - PhyML binary: downloaded pre-compiled from bioconda in Dockerfile
697
 
698
  ### Staging vs Production
699
+ - Single Vercel deployment: `main` β†’ production
700
+ - Single HF Space: `samad14-bio-nexus-api`
701
+ - Supabase project: `bjbktegnmkljhuzlsvrf` (single project, RLS on tables)