dotan1111 commited on
Commit
6562d10
·
verified ·
1 Parent(s): abb39d6

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +7 -3
README.md CHANGED
@@ -9,15 +9,19 @@ tags:
9
  - text-generation-inference
10
  ---
11
 
12
- # Protein2Text: Providing Rich Descriptions from Protein Sequences
 
 
 
13
  ## Abstract:
14
- Understanding the functionality of proteins has been a focal point of biological research due to their critical roles in various biological processes. Unraveling protein functions is essential for advancements in medicine, agriculture, and biotechnology, enabling the development of targeted therapies, engineered crops, and novel biomaterials. However, this endeavor is challenging due to the complex nature of proteins, requiring sophisticated experimental designs and extended timelines to uncover their specific functions. Public large language models (LLMs), though proficient in natural language processing, struggle with biological sequences due to the unique and intricate nature of biochemical data. These models often fail to accurately interpret and predict the functional and structural properties of proteins, limiting their utility in bioinformatics. To address this gap, we introduce BetaDescribe, a collection of models designed to generate detailed and rich textual descriptions of proteins, encompassing properties such as function, catalytic activity, involvement in specific metabolic pathways, subcellular localizations, and the presence of particular domains. The trained BetaDescribe model receives protein sequences as input and outputs a textual description of these properties. BetaDescribe’s starting point was the LLAMA2 model, which was trained on trillions of tokens. Next, we trained our model on datasets containing both biological and English text, allowing biological knowledge to be incorporated. We demonstrate the utility of BetaDescribe by providing descriptions for proteins that share little to no sequence similarity to proteins with functional descriptions in public datasets. We also show that BetaDescribe can be harnessed to conduct *in-silico* mutagenesis procedures to identify regions important for protein functionality without needing homologous sequences for the inference. Altogether, BetaDescribe offers a powerful tool to explore protein functionality, augmenting existing approaches such as annotation transfer based on sequence or structure similarity.
15
 
16
  ![image/png](https://cdn-uploads.huggingface.co/production/uploads/63047e2d412a1b9d381b045d/S51dW59cdeY2lbgXbXSx9.png)
17
 
18
  BetaDescribe workflow. The generator processes the protein sequences and creates multiple candidate descriptions. Independently, the validators provide simple textual properties of the protein. The judge receives the candidate descriptions (from the generator) and the predicted properties (from the validators) and rejects or accepts each description. Finally, BetaDescribe provides up to three alternative descriptions for each protein.
19
 
20
- ## Preprint: https://www.biorxiv.org/content/10.1101/2024.12.04.626777v1.full.pdf+html
 
21
 
22
  ## Examples of descriptions of unknown proteins:
23
 
 
9
  - text-generation-inference
10
  ---
11
 
12
+ # BetaDescribe: Providing rich descriptions from protein sequences
13
+ ## Significance:
14
+ Determining protein function remains a major bottleneck in biology, particularly for proteins lacking close homologs with known annotations. This work introduces BetaDescribe, a protein-to-text modeling framework that generates rich, human-readable functional descriptions directly from protein sequences. The proposed architecture relies on a specialized multicomponent framework featuring protein-to-text generation, multiple validators, and a judge model. We show that this integrated system significantly boosts performance and ensures description reliability. By integrating biological and natural language knowledge, BetaDescribe captures functional signals beyond sequence similarity and highlights regions critical for protein activity. This approach provides a complementary strategy to traditional annotation methods, enabling deeper exploration of protein function across uncharacterized regions of the proteome universe.
15
+
16
  ## Abstract:
17
+ Deciphering protein function is fundamental to advancements in medicine and biotechnology. However, conventional experimental characterization remains resource-intensive. Public large language models (LLMs), though proficient in natural language processing, often fail to accurately interpret and predict the functional and structural properties of proteins, limiting their utility in bioinformatics. To address this gap, we introduce BetaDescribe, designed to generate detailed and rich textual descriptions of proteins, including their function, catalytic activity, involvement in specific metabolic pathways, subcellular localizations, and the presence of specific domains. The trained BetaDescribe model receives protein sequences as input and outputs a textual description of these properties. BetaDescribe starting point was the LLAMA2 model, which was trained on trillions of tokens. Our model was next trained on datasets containing both biological and English text, which allowed the incorporation of biological knowledge. In addition to the description generator, BetaDescribe comprises multiple validator models and a judge, which together enable accurate ranking of alternative generated descriptions. We demonstrate the utility of BetaDescribe by providing descriptions for proteins that share little to no sequence similarity to proteins with functional descriptions in public datasets. Using in silico mutagenesis, we further show that BetaDescribe relies on functionally important regions, as part of its prediction, suggesting that the model identifies regions of importance for the protein functionality without needing homologous sequence. BetaDescribe offers a powerful tool to explore protein functionality, augmenting existing approaches such as annotation transfer based on sequence or structure similarity.
18
 
19
  ![image/png](https://cdn-uploads.huggingface.co/production/uploads/63047e2d412a1b9d381b045d/S51dW59cdeY2lbgXbXSx9.png)
20
 
21
  BetaDescribe workflow. The generator processes the protein sequences and creates multiple candidate descriptions. Independently, the validators provide simple textual properties of the protein. The judge receives the candidate descriptions (from the generator) and the predicted properties (from the validators) and rejects or accepts each description. Finally, BetaDescribe provides up to three alternative descriptions for each protein.
22
 
23
+ ### *Dotan, E., Lyubman, I., Ehrlich, M., Bacharach, E., Pupko, T., and Belinkov, Y. 2026.* BetaDescribe: providing rich descriptions from protein sequences. *Proceedings of the National Academy of Sciences (PNAS)*
24
+ ### Link: https://www.pnas.org/doi/10.1073/pnas.2537345123
25
 
26
  ## Examples of descriptions of unknown proteins:
27