Feature Extraction
Transformers
Safetensors
English
eva-rna
biology
transcriptomics
rna-seq
gene-expression
foundation-model
single-cell
bulk-rna
immunology
custom_code
Instructions to use ScientaLab/eva-rna with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ScientaLab/eva-rna with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="ScientaLab/eva-rna", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: other | |
| license_name: scienta-lab-eva-model-license | |
| license_link: LICENSE | |
| language: | |
| - en | |
| tags: | |
| - biology | |
| - transcriptomics | |
| - rna-seq | |
| - gene-expression | |
| - foundation-model | |
| - single-cell | |
| - bulk-rna | |
| - immunology | |
| library_name: transformers | |
| pipeline_tag: feature-extraction | |
| # EVA-RNA: Foundation Model for Transcriptomics | |
| Transformer-based foundation model that produces sample-level and gene-level | |
| embeddings from RNA-seq profiles (bulk, microarray, pseudobulked single-cell) in | |
| human and mouse. | |
| **You can find a complete end-to-end tutorial notebook [here](https://colab.research.google.com/#fileId=https%3A//huggingface.co/ScientaLab/eva-rna.ipynb).** | |
| ## Installation | |
| We recommend proceeding with the [uv package manager](https://docs.astral.sh/uv/getting-started/installation/). | |
| ```bash | |
| uv venv --python 3.10 | |
| source .venv/bin/activate | |
| uv pip install transformers torch==2.6.0 scanpy anndata tqdm scipy scikit-misc pysrab scipy | |
| ``` | |
| ### Optional: Flash Attention | |
| To handle larger gene contexts, EVA-RNA automatically runs on Flash Attention if | |
| available -- only available for post-Ampere GPUs (A100 and beyond). We recommend | |
| using the following wheel. | |
| ```bash | |
| uv pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl | |
| ``` | |
| ## Quick Start | |
| A full example is provided in a notebook (see "Use this model" -> "Google colab"). | |
| ```python | |
| import scanpy as sc | |
| from transformers import AutoModel, AutoTokenizer | |
| model = AutoModel.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True) | |
| tokenizer = AutoTokenizer.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True) | |
| # Load example dataset (bulk dataset, GSE) | |
| import anndata as ad | |
| import pandas as pd | |
| counts_url = "https://www.ncbi.nlm.nih.gov/geo/download/?type=rnaseq_counts&acc=GSE193677&format=file&file=GSE193677_raw_counts_GRCh38.p13_NCBI.tsv.gz" | |
| counts_df = pd.read_csv(counts_url, sep="\t", index_col=0, compression="gzip").transpose() | |
| adata = ad.AnnData(X=counts_df) | |
| # Select 2,000 genes for efficiency (kbest recommended, otherwise highly variable genes) | |
| sc.pp.highly_variable_genes(adata, n_top_genes=2000, flavor="seurat_v3") | |
| # Preprocess counts (EVA expects a target_sum of 1e6 before gene selection and log1p normalization) | |
| sc.pp.normalize_total(adata, target_sum=1e6) | |
| sc.pp.log1p(adata, base=2) | |
| adata = adata[:, adata.var.highly_variable].copy() | |
| # Encode (gene symbols auto-converted, GPU used if available) | |
| embeddings = model.encode_anndata(tokenizer, adata) | |
| adata.obsm["X_eva"] = embeddings | |
| ``` | |
| ### Options | |
| `model.encode_anndata()` accepts the following parameters: | |
| - `gene_column` β column in `adata.var` with gene identifiers (default: uses `adata.var_names`) | |
| - `species` β `"human"` or `"mouse"` for gene ID conversion (default: auto-detected) | |
| - `batch_size` β samples per inference batch (default: 32) | |
| - `device` β `"cpu"`, `"cuda"`, etc. (default: CUDA if available) | |
| - `show_progress` β show a progress bar (default: True) | |
| ## Advanced: Raw Tensor API | |
| For users who need direct control over inputs (mixed precision is applied automatically): | |
| ```python | |
| import torch | |
| from transformers import AutoModel, AutoTokenizer | |
| model = AutoModel.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True) | |
| tokenizer = AutoTokenizer.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True) | |
| device = "cuda" if torch.cuda.is_available() else "cpu" | |
| model = model.to(device).eval() | |
| # Gene IDs must be NCBI GeneIDs as strings | |
| gene_ids = ["7157", "675", "672"] # TP53, BRCA2, BRCA1 | |
| # Expression values should sum to 1e6 (before gene selection) and be log1p-normalized | |
| expression_values = [5.5, 3.2, 4.1] | |
| inputs = tokenizer(gene_ids, expression_values, padding=True, return_tensors="pt") | |
| inputs = {k: v.to(device) for k, v in inputs.items()} | |
| with torch.inference_mode(): | |
| outputs = model(**inputs) | |
| sample_embedding = outputs.cls_embedding # (1, 256) | |
| gene_embeddings = outputs.gene_embeddings # (1, 3, 256) | |
| ``` | |
| ### Batch Processing | |
| ```python | |
| batch_gene_ids = [ | |
| ["7157", "675", "672"], | |
| ["7157", "1956", "5290"], | |
| ] | |
| batch_expression = [ | |
| [5.5, 3.2, 4.1], | |
| [2.1, 6.3, 1.8], | |
| ] | |
| inputs = tokenizer(batch_gene_ids, batch_expression, padding=True, return_tensors="pt") | |
| inputs = {k: v.to(device) for k, v in inputs.items()} | |
| with torch.inference_mode(): | |
| outputs = model(**inputs) | |
| sample_embeddings = outputs.cls_embedding # (2, 256) | |
| ``` | |
| ## Expression Decoder | |
| EVA-RNA includes a pre-trained deterministic expression decoder that maps | |
| gene embeddings back to predicted expression values. | |
| ```python | |
| with torch.inference_mode(): | |
| # Encode | |
| output = model.encode(**inputs) | |
| # output.cls_embedding β sample-level embedding (batch, hidden_size) | |
| # output.gene_embeddings β per-gene embeddings (batch, n_genes, hidden_size) | |
| # Decode expression values | |
| predicted_expression = model.decode(output.gene_embeddings) | |
| # predicted_expression β (batch, n_genes) | |
| ``` | |
| ## GPU and Precision | |
| EVA-RNA automatically applies mixed precision for optimal performance: | |
| - **Ampere+ GPUs** (A100, H100, RTX 30/40 series): bfloat16 | |
| - **Older CUDA GPUs** (V100, RTX 20 series): float16 | |
| - **CPU**: full precision (float32) | |
| No manual `torch.autocast()` is needed. | |
| > **Note β Flash Attention constraints:** When flash attention is installed and an | |
| > Ampere+ GPU is detected, the model uses flash attention layers. These layers | |
| > **require CUDA and half-precision inputs**. If you move the model to CPU you will | |
| > get a clear error asking you to move it back to GPU. If you pass `autocast=False`, | |
| > autocast is re-enabled automatically with a warning since flash attention cannot | |
| > run in full precision. | |
| ### Disabling Automatic Mixed Precision | |
| For advanced use cases requiring manual precision control, pass `autocast=False`. | |
| This only takes effect when flash attention is **not** active (i.e., on older GPUs or | |
| when flash attention is not installed): | |
| ```python | |
| model = model.to("cuda").eval() | |
| with torch.inference_mode(): | |
| # Disable automatic mixed precision (ignored when flash attention is active) | |
| outputs = model(**inputs, autocast=False) | |
| # Or via sample_embedding | |
| embedding = model.sample_embedding( | |
| gene_ids=gene_ids, | |
| expression_values=values, | |
| autocast=False, | |
| ) | |
| ``` | |
| ## Converting Gene Symbols to NCBI Gene IDs | |
| The tokenizer vocabulary uses NCBI GeneIDs. A built-in gene mapper is included to | |
| convert gene symbols or Ensembl IDs: | |
| ```python | |
| tokenizer = AutoTokenizer.from_pretrained("ScientaLab/eva-rna", trust_remote_code=True) | |
| # Available mappings: | |
| # "symbol_to_ncbi" β human gene symbols β NCBI GeneIDs | |
| # "ensembl_to_ncbi" β human Ensembl IDs β NCBI GeneIDs | |
| # "symbol_to_ncbi_mouse" β mouse gene symbols β NCBI GeneIDs | |
| mapper = tokenizer.gene_mapper["symbol_to_ncbi"] | |
| gene_symbols = ["TP53", "BRCA2", "BRCA1"] | |
| gene_ids = [mapper[s] for s in gene_symbols] | |
| # gene_ids = ["7157", "675", "672"] | |
| expression_values = [5.5, 3.2, 4.1] | |
| inputs = tokenizer(gene_ids, expression_values, padding=True, return_tensors="pt") | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @article{eva-rna, | |
| title={EVA: Towards a universal model of the immune system}, | |
| author={Scienta Team}, | |
| journal={arXiv}, | |
| year={2026}, | |
| } | |
| ``` | |
| ## License | |
| [Scienta Lab EVA Model License](LICENSE) | |