Learning Latent Protein Languages for Autoregressive Generation
Abstract
Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates. We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample. These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.
Community
Hi everyone, I'm one of the authors of this paper.
We study how learned target representations affect autoregressive protein generation. We introduce two discrete latent languages that remain decodable to protein data: PLL represents contextual sequence information from a frozen ESM-2 encoder, and SLL represents backbone geometry through an adapted GCP-VQVAE Lite tokenizer. We train separate autoregressive transformers on both languages using next-token prediction.
Some key results:
- Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid model.
- PLLM reduces low-complexity generations by 54% relative to the amino acid model across sampling temperatures, using a heuristic residue-composition entropy threshold.
- In a matched sequence-to-structure comparison, SLL reduces best validation perplexity by 34% relative to the original GCP-VQVAE Lite tokenizer.
- SLLM produces diverse and novel backbones. We also observe early evidence that selecting sampled structure predictions using internal token confidence improves prediction quality beyond a single decoded sample.
Get this paper in your agent:
hf papers read 2610.03978 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper