Update README.md
Browse files
README.md
CHANGED
|
@@ -9,15 +9,19 @@ tags:
|
|
| 9 |
- text-generation-inference
|
| 10 |
---
|
| 11 |
|
| 12 |
-
#
|
|
|
|
|
|
|
|
|
|
| 13 |
## Abstract:
|
| 14 |
-
|
| 15 |
|
| 16 |

|
| 17 |
|
| 18 |
BetaDescribe workflow. The generator processes the protein sequences and creates multiple candidate descriptions. Independently, the validators provide simple textual properties of the protein. The judge receives the candidate descriptions (from the generator) and the predicted properties (from the validators) and rejects or accepts each description. Finally, BetaDescribe provides up to three alternative descriptions for each protein.
|
| 19 |
|
| 20 |
-
##
|
|
|
|
| 21 |
|
| 22 |
## Examples of descriptions of unknown proteins:
|
| 23 |
|
|
|
|
| 9 |
- text-generation-inference
|
| 10 |
---
|
| 11 |
|
| 12 |
+
# BetaDescribe: Providing rich descriptions from protein sequences
|
| 13 |
+
## Significance:
|
| 14 |
+
Determining protein function remains a major bottleneck in biology, particularly for proteins lacking close homologs with known annotations. This work introduces BetaDescribe, a protein-to-text modeling framework that generates rich, human-readable functional descriptions directly from protein sequences. The proposed architecture relies on a specialized multicomponent framework featuring protein-to-text generation, multiple validators, and a judge model. We show that this integrated system significantly boosts performance and ensures description reliability. By integrating biological and natural language knowledge, BetaDescribe captures functional signals beyond sequence similarity and highlights regions critical for protein activity. This approach provides a complementary strategy to traditional annotation methods, enabling deeper exploration of protein function across uncharacterized regions of the proteome universe.
|
| 15 |
+
|
| 16 |
## Abstract:
|
| 17 |
+
Deciphering protein function is fundamental to advancements in medicine and biotechnology. However, conventional experimental characterization remains resource-intensive. Public large language models (LLMs), though proficient in natural language processing, often fail to accurately interpret and predict the functional and structural properties of proteins, limiting their utility in bioinformatics. To address this gap, we introduce BetaDescribe, designed to generate detailed and rich textual descriptions of proteins, including their function, catalytic activity, involvement in specific metabolic pathways, subcellular localizations, and the presence of specific domains. The trained BetaDescribe model receives protein sequences as input and outputs a textual description of these properties. BetaDescribe starting point was the LLAMA2 model, which was trained on trillions of tokens. Our model was next trained on datasets containing both biological and English text, which allowed the incorporation of biological knowledge. In addition to the description generator, BetaDescribe comprises multiple validator models and a judge, which together enable accurate ranking of alternative generated descriptions. We demonstrate the utility of BetaDescribe by providing descriptions for proteins that share little to no sequence similarity to proteins with functional descriptions in public datasets. Using in silico mutagenesis, we further show that BetaDescribe relies on functionally important regions, as part of its prediction, suggesting that the model identifies regions of importance for the protein functionality without needing homologous sequence. BetaDescribe offers a powerful tool to explore protein functionality, augmenting existing approaches such as annotation transfer based on sequence or structure similarity.
|
| 18 |
|
| 19 |

|
| 20 |
|
| 21 |
BetaDescribe workflow. The generator processes the protein sequences and creates multiple candidate descriptions. Independently, the validators provide simple textual properties of the protein. The judge receives the candidate descriptions (from the generator) and the predicted properties (from the validators) and rejects or accepts each description. Finally, BetaDescribe provides up to three alternative descriptions for each protein.
|
| 22 |
|
| 23 |
+
### *Dotan, E., Lyubman, I., Ehrlich, M., Bacharach, E., Pupko, T., and Belinkov, Y. 2026.* BetaDescribe: providing rich descriptions from protein sequences. *Proceedings of the National Academy of Sciences (PNAS)*
|
| 24 |
+
### Link: https://www.pnas.org/doi/10.1073/pnas.2537345123
|
| 25 |
|
| 26 |
## Examples of descriptions of unknown proteins:
|
| 27 |
|