Does Learning Protein Folding Generalize to Broader Reasoning?
Abstract
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
Community
Hi everyone — sharing our recent work, Fold2Reason 👋
Can protein structure supervision improve reasoning beyond proteins?
We introduce FoldingCorpus and Fold2Reason, using both structural reasoning tasks and continuous 3D geometry for post-training.
Fold2Reason achieves 2.7–3.5× better structure prediction on FoldBench, while also improving all 10 downstream benchmarks across spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33%.
Control experiments suggest the gains come from the real structural information in protein geometry, rather than simply adding more training data.
Our broader takeaway: science may not only benefit from AI — scientific problems themselves may teach AI better ways to reason.
Happy to discuss!
Get this paper in your agent:
hf papers read 2609.38879 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper

