Repository navigation
feat: confidence-gap safeguards before wet-lab synthesis - #15
Merged
Merged
Conversation
Three gaps closed to prevent wasting the $10k synthesis budget: Gap 1 — Safety scorer v0.4: add hydrophobic moment (μH) as a hemolysis signal. Pilot-panel SEED-005 variants had μH 0.69–0.87 but were trivially scored safety=1.0 because their hydrophobic fraction sat below the old 0.65 threshold. Now μH > 0.55 adds a risk penalty proportional to the excess, reducing those candidates to safety 0.52–0.77. Gap 2 — Retrospective AUROC benchmark: compare 44 known AMPs against 44 composition-matched shuffled decoys (RNG seed=42). Only order-dependent feature (hydrophobic_moment) can separate them. Gate: >0.70 proceed, 0.55–0.70 caution, <0.55 STOP. Current result: AUROC=0.5305 (POOR — near-random). Reported honestly via `make validate-scoring`. Gap 3 — External predictor checklist: generates pilot FASTA and a structured markdown table for manual submission to CAMPR4, AMPScanner v2, and dbAMP 2.0. Decision gate: ≥12/20 agree → synthesise; 6–11 → wave-1 only; <6 → STOP. Run via `make external-predict`. Tests: 23 new tests in test_safety_moment.py and test_retrospective.py. Full suite: 364 passed.
cschanhniem
force-pushed
the
feat/confidence-gaps
branch
from
June 27, 2026 15:56
47df7cb to
417cec1
Compare
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three safeguards added to close critical confidence gaps before committing to the $10k wet-lab synthesis budget.
Gap 1 — Safety scorer v0.4: hydrophobic moment as hemolysis signal
Pilot-panel SEED-005 variants had μH 0.69–0.87 but were trivially scored
safety=1.0because their hydrophobic fraction sat below the old 0.65 threshold. Amphipathicity (μH) is the primary predictor of non-selective membrane disruption per Dathe & Wieprecht (1999).if mu_h > 0.55: risk += (mu_h - 0.55) * 1.5Gap 2 — Retrospective AUROC benchmark (critical gate)
Compares 44 known AMPs against 44 composition-identical shuffled decoys. Since all composition features (charge, hydrophobic fraction, Boman, GRAVY, length) are identical per pair, only order-dependent features (hydrophobic_moment) can separate them. This tests whether the scoring signal is real.
Result: AUROC = 0.5305 (gate: >0.70 proceed, 0.55–0.70 caution, <0.55 STOP)
This is a hard finding. The model sits in the POOR zone — it is near-random at discriminating known AMPs from their composition-matched shuffles. Root cause: hydrophobic_moment contributes only ~15% of the activity score; the remaining ~85% is composition-based (identical for each AMP/decoy pair).
Implication: the current scoring model's order-sensitive discriminative power is weak. Run
make validate-scoringfor the full report.Gap 3 — External predictor checklist
Generates a pilot FASTA + structured markdown table for manual submission to three independent published tools:
Decision gate: ≥12/20 agree → synthesise; 6–11 → wave-1 only; <6 → STOP. Run
make external-predict.New CLI commands
What to do next (honest)
The AUROC=0.5305 result means the scoring model is not sufficiently discriminative in order-sensitive features. Before spending $10k:
make external-predict). If ≥12/20 agree, the external consensus overrides the weak internal signal.Test plan
test_safety_moment.py(7 tests),test_retrospective.py(16 tests)make validate-scoringruns cleanly and outputs AUROC=0.5305 with interpretationoutputs/pilot_panel.fastato CAMPR4, AMPScanner v2, dbAMP 2.0 and record results in checklist🤖 Generated with Claude Code