BioForm-LM: an audited generative formulation-design pipeline
Read this first
This checkpoint's headline architectural claim (protein-conditional generative formulation design) is not supported by the evidence collected so far. That finding, and how it was reached, is the actual contribution β see the companion paper ("What Determines a Biologic's Formulation?") for the full analysis. This card describes what the model demonstrably does and does not do; it does not repeat an earlier draft's fabricated performance table.
What changed from the original release
An earlier version of this card reported recall@10 = 0.167, perfect calibration
(r = 1.0) on a held-out protein, and 4.7x formulation-space coverage versus
random sampling. Those numbers came from an evaluation path that never actually
called the model (verified: the code path fell through to np.random before
reaching a model forward pass) and from a benchmark that, at the time, silently
dropped 49 of 67 rows via a CSV parser bug and kept 18 unsourced placeholder
rows. Both defects are now fixed, and the numbers below are freshly measured,
model-produced, and reproducible from scripts/evaluate_real.py.
Architecture (unchanged)
Three-stage pipeline:
- Mechanistic simulator. DLVO colloidal theory + Lumry-Eyring aggregation
kinetics generate synthetic (protein, recipe) to stability triples for
pretraining. An audit found the original version (v1) has two structurally
dead recipe slots (buffer species/concentration change predicted stability
by exactly 0.0) and corner-collapsing optima in 3 of 4 continuous variables
(100% monotone for ionic strength, osmolarity, temperature).
v2(simulator/mechanistic_sim_v2.py) adds nine literature-anchored competing degradation pathways that partly fix this: pH optima become interior in 95.5% of sampled proteins and both dead slots are revived, but ionic strength (38.5% interior) and osmolarity (13%) are only partly repaired and temperature stays monotone. Reproduce withscripts/diagnose_v2.py. - In-context generative transformer. Decoder-only, 6 layers, 8 heads, 256 hidden dim, 4.9M parameters, 199-token vocabulary. Pretrained on the synthetic corpus; adapts to a novel protein from 2-10 real measurements supplied in context, no gradient update.
- Stability-classification head. Scores a full (protein, recipe) sequence for a discretized stability bin; used for best-of-N reranking.
Real, verified results (BioFormBench-Real)
The protein-specificity number below was revised after this card's first release. The original evaluation grouped "one protein" by a (molecular weight, pI)-string match, which silently merged up to 4 different real antibodies (Trastuzumab, Omalizumab, and two others) into one fake pooled protein wherever their placeholder pI values collided β so the original 0.509 was measured on a benchmark that could not, by construction, cleanly separate several of its own proteins. Fixing the grouping and correcting pI for 5 of 13 proteins gives:
| Metric | Value | 95% CI / p | Chance | Folds |
|---|---|---|---|---|
| Likelihood percentile | 0.797 | CI [0.729, 0.865] | 0.5 | 9 (original grouping) |
| ...unconditional baseline | 0.872 | - | 0.5 | 9 |
| Protein specificity (original, buggy grouping) | 0.509 | CI [0.358, 0.669] | 0.5 | 9 |
| Protein specificity (identity-fixed) | 0.737β0.781 | p = 0.008β0.012 (one-sided) | 0.5 | 8 (7 excl. one unresolved protein) |
| ...untrained control, seed 0 | 0.323 | p = 0.92 (n.s.) | 0.5 | 8 |
| ...untrained control, 10 seeds | 0.32β0.71 (mean 0.57) | 1 of 10 seeds p < 0.05 | 0.5 | 8 |
| Recall@10 (exact match) | 0.011β0.025 | CI incl. 0 | - | 8β9 |
Read this table correctly, in both directions. The likelihood-percentile
caveat still holds: it is winnable from the marginal recipe distribution alone
(the unconditional baseline scores higher), so it does not by itself evidence
conditional knowledge. The corrected protein-specificity result is real
progress, but suggestive rather than established: nominally significant and
unanimous (7/7) once the one still-unresolved protein, mAb2, is excluded, and
compared against two controls: an untrained same-architecture model, and a
second model
(checkpoint_unconditional_control.pt) trained on the identical corpus with
protein descriptors independently permuted against recipes and outcomes, so
that P(recipe|protein) is destroyed by construction. The decorrelated control
stays at chance (p β₯ 0.27). A single untrained initialisation is a noisy
control: across ten random seeds (scripts/random_control_seeds.py) untrained
specificity ranges 0.32β0.71 and one seed reaches nominal significance, so the
primary model's 0.737 exceeds every untrained seed, but narrowly. A second, independently-trained
checkpoint (checkpoints_icl) replicates the direction of the positive
result (Spearman rank agreement 0.73, p=0.039 across the two checkpoints'
per-protein values) and reaches significance at n=7 (p=0.039) but not at n=8
(p=0.105). An earlier release combined the two models with Stouffer's method
(p=0.0016β0.0064); that is invalid because the models share a corpus, proteins
and descriptors (score correlation 0.78). Brown's dependence-corrected method
gives p=0.045 (see the paper). All p-values here are exact one-sided sign-flip
tests (scripts/consolidate_specificity.py). Treat this as genuine,
mechanistically-explained evidence that the earlier
chance-level verdict was partly a benchmark bug, not as proof the model has
learned formulation physics.
ICL requires ICL-structured training
Holding the pretraining corpus fixed and scoring the same held-out rows with and
without two real in-context examples (identity-corrected benchmark), a model
trained on in-context-structured sequences improves in likelihood percentile on
6/6 proteins (mean +0.029), while a model trained only on flat triples gets
worse on 6/6 (mean -0.094); exact p = 0.031 each. With one in-context example
(all 8 proteins) the direction holds but is weaker (7/8 vs 1/8; sign test
p = 0.07). An earlier release reported 9/9 (p = 0.0039); that analysis predates
the identity fix and compared different held-out rows, and is superseded.
Reproduce with scripts/icl_necessity.py. This is an architecture-level result,
not a domain-specific claim about formulation design.
Intended use
What this model is not, currently: a validated tool for proposing formulations for a novel protein. The evidence above does not support that use.
What it is useful for: a reference implementation of the sim-to-real + in-context architecture, a demonstration of the ICL-training-necessity finding above (independent of the formulation domain), and a component to build on if the underlying data problem (Discussion, companion paper) is addressed β specifically, real measured stability outcomes across many proteins, not just the 13 currently available in BioFormBench-Real, or a richer protein representation than the current 3 scalar descriptors.
Limitations
- BioFormBench-Real supports only 9 LOPO folds; too few for a definitive protein-conditionality verdict on stability outcomes.
- Protein descriptors used by this checkpoint are 3 quantized scalars (MW, pI, baseline Tm), not full sequence. A follow-up analysis using real VH/VL sequence-derived descriptors on 80 approved antibodies (BioFormBench-Marketed) also found no detectable protein effect on marketed formulation choice, which argues the null result is not merely an artifact of the impoverished descriptor, but this has not been tested with sequence descriptors fed directly into this model architecture.
- No wet-lab validation at any stage.
- Simulator v2's pathway weights are literature-informed, round numbers, not fit by optimization against any dataset β treat outputs as documented, falsifiable priors, not calibrated probabilities.
Training details
- Framework: PyTorch, bf16 autocast, TF32 matmul
- Optimizer: AdamW, OneCycleLR
- Loss: joint next-recipe-token CE (masked to simulator-preferred rows) + stability-bin CE
- Hardware: RTX 3050 (8GB) + 28-core CPU for parallel data generation/encoding
- Sequence length: 13 tokens (5 protein-descriptor prefix + 8 recipe slots)
Ethical considerations
Generated formulation recipes are computational hypotheses only, and the evidence in this card is a reason for additional scrutiny before wet-lab screening, not a substitute for it. Do not use outputs for therapeutic decisions.
Cite this model
@article{kumar2026bioformlm,
title={What Determines a Biologic's Formulation? Two Open Benchmarks, an
Audited Mechanistic Simulator, and a Well-Powered
Platform-Convergence Result},
author={Kumar, Bonthada Sravan},
year={2026}
}
Related resources
- GitHub: https://github.com/Maheshbonthada/bioform-lm
- Dataset: Sravankumarbonthada/BioFormBench (BioFormBench-Real + BioFormBench-Marketed)
Author
Bonthada Sravan Kumar, Independent Researcher β sravansaijohn@gmail.com
License
MIT
Last updated: September 8, 2026
- Downloads last month
- 82