PepPA / docs /SOPHIE_RUNBOOK.md
pranamanam's picture
Upload 97 files
98bde72 verified
|
Raw History Blame Contribute Delete
15.7 kB

Sophie’s PepPA execution runbook

The scientific objective is to balance therapeutic properties during generation, verify robustness computationally across the intended development contexts, and commit to final sequences before experimental characterization. The two computational proposal cycles are complete before any prospective assay data are collected.

E0 — Scientific worker preflight

Install the CPU package and run:

python -m pip install -e .
python scripts/run_compiled_example.py
python -m unittest discover -s tests -v

Use docs/WORKER_INTEGRATION.md and configs/model_registry.json to install each native worker separately. Freeze code, checkpoint, tokenizer, preprocessing, input chemistry, software environment, and hardware. Run one known input and one deliberately invalid input through every worker. Save native output, normalized output, timing, model lineage, and validation diagnostics under runs/preflight/TOOL/. Verify moPPIt motif-index conventions, AlloGen state labels, PeptiVerse units/uncertainty, pepADMET domains, and explicit PTM/bond handling for AF3. A sequence-changing EvoBind run must create a new candidate.

Complete chemical standardization and monomer mapping before release. For each allowed residue/closure, verify stereochemistry, graph identity, formula, mass, protecting groups, monomer source, and route annotation. Have the synthesis lead review the permitted chemistry registry once before the comparative runs.

Artifact: registry.lock.json with known-input reports and scientific validators. Resolve every required worker before the corresponding task is admitted.

E1 — Data, task, and development profile freeze

Curate eight affinity/selectivity, eight motif, and eight conformation tasks. A ranking pool must have the measurements needed for its declared evaluation; retain endpoint masks. Record any task-category shortfall before running comparisons. Add human and preclinical target contexts, countertargets, isoforms/states, route, and formulation requirements. For xenografts, separate the human tumor target from host-species exposure and off-target contexts.

Cluster proteins at 30% identity with 80% bidirectional coverage and peptides at 80% global identity. Join records sharing target cluster, peptide cluster, or primary publication. Assign connected components to 70% model fit, 10% ensemble fit, 10% calibration, and 10% test. Record actual sizes. Freeze labels and test-outcome passages away from the agent's inputs. Use a time cutoff for retrieved literature and record pretrained-model exposure separately.

Two authors annotate constructs, residue maps, mechanisms, required tools, endpoint definitions, source support/opposition, and counterexamples. Adjudicate disagreements. Up to three hypotheses per task are allowed. Native structural thresholds are selected on mechanistically annotated validation positives/decoys at a prespecified 5% false-positive operating point, with sensitivity and sample size retained. Endpoint property thresholds and species applicability are written explicitly in each task. A small or missing validation set receives an unresolved status and does not support a calibrated claim.

Use six query templates: target + peptide binding; target + binding motif; target + selectivity/countertargets; target + state/modification; target + mechanism/partner; target + species/developability. Retain at most ten records per template, DOI/accession deduplication, exact passages, and source dates.

Ensemble fitting: deduplicate identical checkpoint/input/preprocessing evaluations. Fit compatible endpoints separately, choose ridge strength from {0.001, 0.01, 0.1, 1} by grouped cross-validation, then calibrate on the untouched calibration partition. Use total marginal error allocation 0.1 across required calibrated endpoints. Save prediction matrices and model membership. Evaluate selected-distribution coverage separately.

Artifacts: tasks.jsonl, sources.jsonl, partitions.json, context_profiles.json, ensemble_fit/, calibration/, thresholds.json, registry.lock.json.

E2 — Computational comparisons

Seven core conditions: fixed workflow; ReAct/Muse; compiled/reduced checks; full PepPA/Muse; full PepPA/GPT-6 Astra Max; full PepPA/Claude Opus Max; full PepPA/Muse without retrieval. With 24 tasks and seeds {2027,2028,2029}, run 504 episodes. The same task/seed pair shares tools, raw proposal allowance, structural cap, and GPU budget.

Per episode: two proposal cycles of 192 raw outputs, 384 total; at most 18 structurally evaluated candidates per cycle, 36 total; 24 GPU-hours on the declared hardware. Canonical primary comparisons use 18 residues. Run 12- and 24-residue transfer strata separately. Each candidate/context structure uses three model seeds and five diffusion samples per seed. Native model/search hyperparameters must be pinned before the comparison.

Divide proposal slots equally among compatible generators, then accepted hypotheses, using stable round-robin allocation. Four objective vectors are balanced and affinity-, specificity-, and developability-emphasized. For K objective groups, balanced weights are 1/K; each emphasized vector assigns 0.6 to its selected group and 0.4/(K-1) to each other group. Normalize subobjective weights within each group. For K=1, use weight one. Preserve the weights and ordering in the plan.

The outer checker evaluates structures/docking, external ADMET, chemistry, synthesis, and CMC. Classify first-cycle failures as binding, specificity, exposure/developability, structural, or chemical-manufacturing. Assign second-cycle quotas proportional to (failure count + 1). Round by largest remainder, using category IDs for ties. Visit categories in decreasing failure-count order and emit slots until their quotas are exhausted. Each category maps to a frozen objective vector and permitted chemistry transformations. Every modified identity is rescored. Store both cycle reports and the final release.

Additional experiments: six component ablations × 24 tasks × three seeds = 432 episodes. Five fresh compilations × 24 tasks × three backends = 360 planning runs. Twelve fault classes × ten fixtures × seven conditions = 840 fault episodes. Reuse cached scientific outputs for the fault suite. Compare single-context with multi-context guidance using paired seeds and all other settings fixed.

Metrics: exact requirement recovery; executable graph rate; complete-report rate; precision@12 and recovery@12 on labeled pools; nominal and selected-set coverage; worst-context feasibility; canonical/noncanonical validity; duplicate-lineage rank change; diversity; false release; success on all five repeated compilations; wall time/tokens/GPU use. For fresh runs report rank Jaccard, score deviation, and release agreement; cached replay requires identical hashes.

Artifacts: one directory per task/condition/seed with frozen plan, source snapshot, proposals, scores, structures, failures, resource ledger, trace, and release. Figures 2 and the aggregation figure use these artifacts. Keep every empirical cell blank until its source analysis is available.

E3 — Final prospective commitment

Six directions: β-catenin over p53; p53 over β-catenin; phospho-ERK2 over unmodified ERK2; VHL recruitment to PAX3::FOXO1, SS18::SSX1, and EWS::FLI1. Three conditions: fixed workflow, ReAct/Muse, full PepPA/Muse. Each direction/condition has 18 slots, giving 324 nominations. Twelve shared control slots per direction give 72 controls. Include at least 12 canonical slots and up to six noncanonical analogues where the corresponding workers support the chemistry. Each analogue occupies its own slot.

Freeze the complete peptide identities, ranking, predicted scores, hypothesis/context matrix, and assay allocations. Hash the bundle. Keep empty slots and failed candidates visible; do not silently replace them after screening. Phage characterization precedes free-peptide assays for ERK2 and ternary tasks, but the full frozen synthesis panel is already committed.

Each synthesis record contains monomer order, stereochemistry, termini, closure atoms, formula/mass, proposed route, requested amount, analytical method, and buffer/handling requirements. Use Liberty Blue synthesis, RAZR/Razor cleavage, and Prodigy purification. Require ≥95% analytical HPLC purity and expected LC–MS identity for assay-ready material. Record failed synthesis and recovery against the original nomination denominator.

E4 — Experimental characterization

Reciprocal binding and regulation

Freeze full-length β-catenin, its N-terminal IDR boundaries, isolated IDR, IDR-deletion construct, full-length p53, tags, and sequence maps. BLI measures each β-catenin-directed peptide against all four proteins and the reciprocal p53 guides against p53 and β-catenin. Use 12 concentrations over 0.1 nM–10 µM, reference sensors, blanks, matched loading, and a second orientation or solution competition. Fit kinetic or registered equilibrium models with residual inspection. Criteria: target KD ≤100 nM, ≥10-fold countertarget discrimination; IDR binding also requires ≥10-fold affinity loss after deletion and retained isolated-IDR binding.

Insert canonical β-catenin guides into peptide–GSGSG–CHIPΔTPR in the published pcDNA3 uAb backbone. Sequence-verify the guide, linker, and effector. Insert reciprocal p53 guides into peptide–GSGSG–OTUB1 catalytic-domain duAbs. Use DLD1 for β-catenin and HeLa for p53. Confirm target and countertarget alleles and include the expressed sequences in the computational context panel before generation. Freeze construct boundaries, promoter, DNA dose, cell density, and delivery conditions after a reference-construct pilot using the supplied SaLT&PepPr, PepPrCLIP, and duAb protocols. Include non-targeting, effector-only, catalytically inactive CHIP, OTUB1 C91S, OTUB1 D88A/C91S/H265A, and PR-619 controls. Collect both protein bands at 24, 48, and 72 h, effector expression, viability, turnover/ubiquitination, and transcript readouts. Quantify equal-protein whole-cell lysates with BCA, SDS–PAGE, membrane transfer, and unsaturated target/countertarget/effector images. Normalize background-subtracted bands to total-lane protein and the matched non-targeting control. Keep compartment-specific measurements separate. Fix the primary molecular time point in the reference-construct pilot and record later TOPFlash/FOPFlash, p53 transcription, and PARP-cleavage responses as functional outcomes. Primary uAb criterion: ≥50% β-catenin loss with p53 within 20% of control. Primary duAb criterion: ≥2-fold p53 increase with β-catenin within 20% of control; viability ≥80%.

Peptide properties

PAMPA: 10 µM donor, pH 7.4, 25°C, 4 h; freeze membrane lipid, area, volumes, and cosolvent. Quantify both compartments by peptide-calibrated LC–MS. Retain membrane retention and mass recovery. Use the manuscript equation, require 80–120% recovery for quantitative interpretation, and preserve detection bounds. Proposed success threshold: Pe ≥10^-6 cm/s.

Solubility: supernatant concentration after 24 h in the registered buffer; threshold ≥10 µM. Stability: intact peptide at 0, 15, 30, 60, 120, and 240 min at 37°C in 50% human serum; threshold half-life ≥60 min. Add matched preclinical-species serum and formulation conditions for the declared transfer analysis. Fit exponential loss with diagnostics. Keep PAMPA, cellular uptake, intracellular exposure, and systemic exposure as separate endpoints.

Phospho-ERK2

Use occupancy-verified pThr185/pTyr187 and matched unmodified WT-sequence ERK2. Common low-valency pIII display and three independent phage selections, each with three panning rounds; retain input/display/amplification controls, tag-only controls, and phosphatase-treated target. Sequence all pools. Analyze differential enrichment with pseudocount one and library-size normalization as in Methods. Test the complete frozen peptide panel by paired BLI. Criteria: phospho-KD ≤100 nM and ≥10-fold preference. Verify occupancy before and after binding assays. Cellular pull-downs quantify phospho- and total ERK2 normalized to input with pathway activation/inhibition and scrambled controls.

VHL ternary assembly and degradation

Freeze P3F1, SS18::SSX1, and EWS::FLI1 breakpoint sequences and parental countertargets. Use VHL–Elongin B–Elongin C. Phage co-capture uses both partners, each alone, each omitted, tag-only, non-targeting, and interface-competition controls in reciprocal capture orientations. Record valency, protein concentrations, input and output counts.

Measure free-peptide binary KD by BLI and ternary concentration matrices by a calibrated proximity assay, such as TR-FRET between labeled target and VHL complex. Include donor-only, acceptor-only, tag-matched, and partner-omission controls and confirm assembly by reciprocal co-capture. Fit the mass-balanced equilibrium model only with identifiable binary constants and calibrated observation model. Retain hook-effect data and interval/identifiability diagnostics. Use RH4 for P3F1, HS-SY-II for SS18::SSX1, and A673 for EWS::FLI1, with confirmed breakpoints and VHL expression. Dose eight concentrations spanning 1 nM–10 µM at 2, 6, and 24 h. Quantify fusion, parents, loading, intracellular peptide exposure, target engagement, and viability. Criterion: ≥50% degradation at ≤1 µM with parental proteins within 20% of control and viability ≥80%. Test proteasome/neddylation inhibition, VHL competition or depletion, target ubiquitination, and target transcripts.

Use three independent assay days, phage selections, or cellular preparations. Average technical triplicates within the independent unit. Randomize by assay-day block and blind condition identities. Retain full unsaturated western blots and raw curves.

E5 — Outcome analysis

Link experimental records to the frozen release hash in a separate evaluation directory. Report nomination-slot success as primary, interpretable-measurement success as secondary. Separate binding, PTM discrimination, ternary assembly, permeability, and cellular regulation. Use equal task weights, average seeds within task, and 2,000 task-bootstrap resamples for 95% intervals. Report all losses, unfilled slots, synthesis failures, censored values, and missing endpoints.

The six prospective directions are exploratory at this allocation. Later animal pharmacology, toxicology, formulation/scale-up, and clinical studies require their own designs and evidence. Preserve the same committed molecular identity when assessing transfer; document any subsequently necessary redesign as a new development candidate.

Minimum candidate bundle

candidates.jsonl, scores.csv, contexts.json, synthesis_manifest.csv, structures/, evidence.jsonl, plan.lock.json, model_registry.lock.json, calibration_manifest.json, trace.jsonl, failures.jsonl, resource_use.json, and checksums.json. Every score includes endpoint/unit/context, native value, transformed value, interval or structural range, support type, applicability, model lineage, and source artifact. Every required metric must be populated before release.

Controller settings

Compare Muse Glimmer-30B with GPT-6 Astra Max and Claude Opus Max. The latter two use gpt-6-astra and claude-opus-5 with maximum reasoning effort. Freeze the exact serving snapshot, output-token cap, input context, and total token budget before evaluation. Save returned model IDs and provider request settings. The adapter tests verify request settings with local mocks; live model comparisons remain part of E2.