Repository navigation
feat: Phase 3+4 — 89 AMP nominees, dual-scorer consensus, full nomination report - #12
Closed
cschanhniem wants to merge 25 commits into
Closed
cschanhniem wants to merge 25 commits into
cschanhniem wants to merge 25 commits into
Conversation
…tion - Add hydrophobic_moment() to physchem.py using Eisenberg (1984) consensus scale at 100°/residue helical projection; literature-cited correlate of AMP activity - Expand activity_likeness_score() to incorporate amphipathicity (15% weight) with reduced charge/hydrophobicity weights to keep total at 1.0 - Add recall_at_k(), random_recall_at_k(), enrichment_factor(), benchmark_summary() to benchmark/evaluate.py with honest disclaimer in every output - Add 'openamp-foundry bench baseline' CLI subcommand for pipeline vs random recall - Add 'make bench-baseline' Makefile target - 20 new tests: amphipathicity feature, hydrophobic moment edge cases, recall@k boundary conditions, enrichment factor, benchmark summary structure
- Add examples/benchmark/mixed_candidates.csv (20 sequences: 5 known-active AMPs + 15 non-AMP control sequences) for proper enrichment benchmarking - Add examples/benchmark/active_labels.csv (5 known-active IDs matching above) - Add make bench-hidden-active target using bench baseline CLI - 12 new tests in test_hidden_active_recovery.py: - all positives rank in top half - recall@5 = 1.0 (perfect recovery) - enrichment factor >= 2.0 at k=5 (actual EF=4.0) - pipeline verdict correctly says 'outperforms random' - negatives score lower than positives on average - CLI integration test for bench baseline command - benchmark data integrity checks - Pipeline achieves EF=4.0 at k=5: all 5 known AMPs recovered in top 5 of 20 vs 25% expected from random — meets Phase 2 criterion from AGENTS.md
- Expand CI to validate evidence certificates, run leakage check, and gate on hidden-active EF >= 1.5 at k=5 (currently achieves 4.0) - Add test_negative_penalization.py: 20 tests verifying that problematic sequences (extreme hydrophobicity, high-cysteine, purely negative charge, long repeat runs) score lower than known AMP-like sequences on activity, safety, and synthesis dimensions - Fix all 10 ruff lint warnings (unused imports) across 7 files - 89 tests passing, ruff clean
- Add build_batch_report() to pipeline.py — generates a machine-readable batch_report.json alongside the markdown report; validates against schema - Expand batch_report.schema.json to require disclaimer, score_averages, and selected_ids fields - Add test_ablation.py: 9 tests verifying that removing safety/novelty filters degrades selection quality (per AGENTS.md Phase 2 ablation requirement): - Novelty filter correctly excludes near-duplicates of references - Ablation of novelty filter causes near-duplicates to be selected (worse) - Safety filter excludes high-risk sequences; ablation includes them - Batch report JSON generated automatically alongside markdown - Batch report validates against batch_report.schema.json - Counts in batch report match actual output - Disclaimer field present and non-empty - 98 tests passing, ruff clean
- Add cluster_by_similarity() and cluster_split() to splits.py: greedy
single-linkage clustering groups near-duplicate sequences so benchmark
reference and test sets are never contaminated by each other
- Add find_contaminated_references() to evaluate.py: identifies reference
sequences that are near-duplicates of test positives (Phase 2 leakage check)
- Add recall_at_k(), random_recall_at_k(), enrichment_factor(),
benchmark_summary() to evaluate.py (full benchmark evaluation suite)
- Add test_cluster_split.py: 21 tests covering clustering properties,
split partitioning, contamination detection, and end-to-end enrichment
verification — all 3 positives rank above negatives without references,
confirming scoring is feature-based not reference-proximity-based
- Add examples/benchmark/cluster_split_{pool,refs}.csv test data
- Fix ruff F401 unused imports in pipeline.py, test_cli.py, test_pipeline_filters.py
- 58 tests passing, ruff clean
…ves, 53 tests Phase 2: Negative-set robustness — pipeline must enrich AMP-like sequences above negatives regardless of negative type, not just against easy all-repeat controls. - Add examples/negative/poly_cationic.csv: 5 poly-K/R sequences with high charge density but no hydrophobic face (a harder negative set than degenerate repeats) - Add examples/negative/poly_hydrophobic.csv: 5 poly-L/I/V sequences with high hydrophobic fraction but zero charge (another harder negative class) - Add examples/benchmark/robustness_positives.csv: 3 canonical AMP-like positives used across all robustness tests - Add test_negative_robustness.py: 16 tests verifying: - Poly-cationic sequences have correct physicochemical properties (high charge, zero hydro) - Poly-hydrophobic sequences have correct properties (high hydro, zero charge) - Safety scorer penalizes both classes (excess charge / excess hydrophobicity) - EF > 1.0 vs degenerate, poly-cationic, AND poly-hydrophobic negatives - recall@3 = 1.0 (all 3 positives in top-3) vs both hard negative sets - Mean positive ensemble score > mean negative ensemble across all negative types - Add recall_at_k, random_recall_at_k, enrichment_factor, benchmark_summary to evaluate.py - Fix ruff F401 unused imports in pipeline.py, test_cli.py, test_pipeline_filters.py - 53 tests passing, ruff clean
Phase 2: Novelty pressure — top candidates must not be mere copies of known AMP motifs. - Add examples/benchmark/novelty_pressure_pool.csv: 11 sequences including 3 near-dups of known references (NOV-DUP-*), 3 genuinely novel AMP-like sequences (NOV-NEW-*), and 5 non-AMP negatives (NOV-NEG-*) - Add test_novelty_pressure.py: 13 tests verifying: - Exact reference copies receive novelty = 0.0 - 1-substitution near-dups receive novelty < 0.20 (below min_novelty threshold) - Genuinely novel sequences receive novelty >= 0.20 - Novelty decreases monotonically with increasing reference similarity - nearest_reference field is populated for near-duplicate candidates - Pipeline min_novelty filter excludes near-dups from selection - Exact reference copies never appear in selected batch - Novel AMP-like candidates preferentially selected over near-dups - Near-dups do not dominate the top-5 ranked (novelty weighting has effect) - Add recall_at_k, random_recall_at_k, enrichment_factor, benchmark_summary to evaluate.py - Fix ruff F401 unused imports across pipeline.py, test_cli.py, test_pipeline_filters.py - 50 tests passing, ruff clean
Phase 2: Reproducibility — rankings from same inputs must be reproducible.
- Update schemas/run_manifest.schema.json: add generated_at and input_hashes
as required fields (previously absent from schema but present in output)
- Add test_reproducibility.py: 18 tests verifying:
- Two independent runs produce identical JSONL ranking order
- Two runs produce identical scores for every candidate
- Two runs produce identical selected candidate set
- run_manifest.json is generated alongside ranked.jsonl
- Explicit manifest_path argument is respected
- Manifest validates against updated JSON Schema
- All required fields present (run_id, pipeline_version, config_hash,
generated_at, inputs, input_hashes, outputs)
- pipeline_version in manifest matches installed package __version__
- run_id is valid UUID format
- Two runs produce different run_ids (unique per run)
- SHA-256 hashes in manifest match actual files
- Candidate and reference paths recorded in manifest
- stable_json_hash() is deterministic for same config
- stable_json_hash() changes when config changes
- build_run_manifest() produces correct structure
- SHA-256 output is 64-char lowercase hex
- Different file content → different SHA-256
- file_sha256() matches stdlib hashlib computation
- Add recall_at_k, enrichment_factor, benchmark_summary to evaluate.py
- Fix ruff F401/E741 issues across pipeline.py, test_cli.py, test_pipeline_filters.py
- 55 tests passing, ruff clean
Phase 2: Toxicity penalty — predicted hemolytic/toxic candidates are down-ranked. - Add test_toxicity_penalty.py: 13 tests covering the full toxicity penalty mechanism: Safety scorer penalty signals: - Hydrophobic fraction > 0.65 → hemolysis proxy penalty - Charge density > 0.55 → toxicity proxy penalty - Length > 35 aa → stability and synthesis penalty - Cysteine fraction > 0.25 → disulfide complexity penalty - Longest repeat run ≥ 6 → degenerate composition penalty Tests verify: - Extreme hydrophobicity reduces safety score (<0.6) - Extreme charge density reduces safety score (<0.6) - High cysteine fraction reduces safety (<0.9) - Very long sequences penalized - Long repeat runs penalized - All balanced AMP candidates score higher safety than poly-hydrophobic sequences - All balanced AMP candidates score higher safety than poly-cationic sequences - Monotonic safety reduction with increasing excess hydrophobicity - Double penalty (charge + repeat run) on poly-K - Mean AMP ensemble > mean high-risk ensemble in full pipeline - All AMPs outrank all high-risk candidates - Toxicity penalty propagates correctly through pipeline safety field - High-risk candidates excluded with strict max_safety_risk filter - Fix ruff F401 unused imports (pipeline.py, test_cli.py, test_pipeline_filters.py) - 50 tests passing, ruff clean
… pre-registered selection rule - Add template_mutator.py: conservative substitution generator (single, double, charge-enhanced variants) from AMP-like seed sequences. Deterministic, no ML model. - Add amp_seeds.csv: 5 AMP-like template seeds for Phase 3 generation. - Add phase3_pool.csv: 383 candidates generated from 5 seeds (rng_seed=2024). - Add generate-batch CLI command: takes seeds CSV → candidate pool CSV. - Add Makefile targets: `make generate` (pool generation) and `make phase3` (full pipeline). - Add configs/phase3.yaml: Phase 3-specific config (min_novelty=0.05, max_safety_risk=0.40). - Add docs/SELECTION_RULE.md: pre-registered pass/fail criteria locked before generation. - 89 candidates selected from 383, all validated against candidate.schema.json. - Run manifest captures SHA-256 of all inputs for reproducibility. - 213 tests pass, lint clean. Disclaimer: Generated candidates have no demonstrated biological activity. All scores are computational heuristics. The lab is the judge.
…ports + risk review Completes all 10 Phase 3 definition-of-done requirements from AGENTS.md: 1. 89 selected candidates ✓ (prior commit) 2. Evidence certificates ✓ (prior commit) 3. Diversity clustering report ✓ (32 clusters, 40.6% singletons) 4. Novelty report ✓ (mean novelty 0.139, per-candidate nearest-reference) 5. Toxicity/hemolysis risk report ✓ (safety flags, risk thresholds documented) 6. Synthesis feasibility report ✓ (length, cys, pro, repeat analysis) 7. Pre-registered selection rule ✓ (prior commit) 8. Pre-registered pass/fail criteria ✓ (prior commit) 9. Risk review ✓ (docs/RISK_REVIEW.md, 5 human review gates defined) 10. Independent expert review: PENDING (human gate, not automatable) New files: - src/openamp_foundry/reports/batch_pack.py — four sub-report generators - tests/test_batch_pack.py — 38 tests for all sub-reports - docs/RISK_REVIEW.md — computational and dual-use risk assessment CLI: added `batch-pack` subcommand Makefile: `make phase3` now includes batch-pack step 251 tests pass, lint clean.
…lab results schema, methods appendix
Builds the bridge between computational nomination (Phase 3) and wet-lab validation (Phase 4).
Changes:
- examples/known_reference/amp_curated_references.csv: 45 diverse known AMPs from published
literature (magainin, buforin, temporin, aurein, cecropin, indolicidin, cathelicidin families)
for meaningful novelty scoring. Phase 3 re-run: mean novelty 0.139 → 0.172 vs real AMP space.
- schemas/lab_result.schema.json: JSON schema for ingesting wet-lab assay results
(MIC, MBC, hemolysis, cytotoxicity) — enables active-learning loop when data arrives.
- src/openamp_foundry/data/lab_results.py: loader, validator, summariser, and candidate mapper
for lab results ingestion.
- docs/EXPERT_REVIEW_PACK.md: complete expert review pack ready to send to a qualified
microbiologist. Includes batch stats, top-20 table, reviewer questions, next-step checklist.
- docs/METHODS.md: publication-quality methods appendix covering generation, scoring,
selection, reproducibility, benchmark validation, and known failure modes.
- Makefile: phase3 now references amp_curated_references.csv instead of seeds.
15 new lab_results tests. 266 total tests pass. Lint clean.
make test && make demo both pass.
Remaining Phase 4 human gates (not automatable):
- Expert review sign-off (docs/EXPERT_REVIEW_PACK.md)
- CRO/lab partner selection
- Synthesis and assay execution
- Results ingestion via schemas/lab_result.schema.json
Implements the Boman (2003) interaction-potential-based activity scorer as a second independent predictor alongside the existing physicochemical heuristic. Adds model disagreement signal to flag uncertain nominations. - scoring/boman.py: boman_index(), boman_activity_score(), gravy_score(), model_disagreement() - Published Boman 2003 Table 1 potentials (transparent, no training) - tanh normalization to [0,1]; disagreement = |activity − boman_activity| - features/physchem.py: boman_index and gravy added to compute_features() output - pipeline.py: boman_activity and disagreement stored in raw_scores - scoring/ensemble.py: disagreement-aware selection reasons and failure modes - schemas/candidate.schema.json: boman_activity and disagreement as optional score fields - tests/test_boman_scorer.py: 37 tests covering all four functions + pipeline integration - docs/METHODS.md, EXPERT_REVIEW_PACK.md: document second scorer and uncertainty signal Candidates with disagreement < 0.20 have dual-scorer consensus (more robust). Candidates with disagreement >= 0.30 are flagged for extra scrutiny.
Surfaces the Boman index vs. activity-likeness disagreement signal in the human-readable batch pack and expert review documents. - batch_pack.py: scorer_consensus_report() — new 5th sub-report - Labels each candidate: high_consensus (<0.20), moderate, uncertain (≥0.30) - Sorted by disagreement ascending (strongest consensus first) - Gracefully handles candidates without boman_activity in scores - generate_batch_pack() → batch_pack_version 1.1, includes scorer_consensus - Summary adds n_high_consensus, n_uncertain_disagreement, mean_scorer_disagreement - write_batch_pack_markdown() → Section 5 Scorer Consensus table - tests/test_batch_pack.py: 11 new tests (TestScorerConsensusReport) - Updated _make_candidate helper to include boman_activity/disagreement - Updated TestWriteBatchPackMarkdown to assert "Scorer Consensus" in markdown - Updated TestGenerateBatchPack to assert scorer_consensus key present 314 tests pass.
Collaborator
Author
v0.2 Update — Boman index second scorer + scorer consensus reportCommits pushed to this branch since last update:
|
…results Reflects the actual Phase 3 run output: 89 candidates selected, all evidence certificates schema-validated, dual-scorer consensus data from v0.2 pipeline. Key changes: - Top-20 table now includes Boman activity and disagreement columns (live data) - Batch statistics include mean Boman (0.503), mean disagreement (0.311) - Explains scientifically why high disagreement is expected for helical AMPs: activity scorer rewards amphipathic character; Boman index penalizes hydrophobic residues — these are different mechanistic models, not a data quality issue - Adds "Suggested pilot candidates" table ranked by Boman × low-disagreement: SEED-003 tryptophan-rich 11-mers prioritized for first synthesis round - Reviewer questions updated to cover dual-scorer methodology - Limitations table updated to reflect current two-scorer state - 6.2 asks expert: do SEED-003 11-mers look more promising than SEED-005 14-mers?
Complete scientific documentation of how the 89 Phase 3 candidates were found, scored, filtered, and selected. Suitable for sharing with expert reviewers and as a pre-publication methods record. Sections: 1. Abstract (89 nominees, computational only, no bio claims) 2. Motivation (cationic AMPs, why short variants, why this approach) 3. Seed templates (5 seeds, family rationale, published sources) 4. Candidate generation (3 strategies, conservative groups, rng_seed=2024, 383 total) 5. Scoring pipeline (6 dimensions with formulas and references) - Activity-likeness (Scorer 1: heuristic) - Boman activity (Scorer 2: Boman 2003 potentials, independent) - Model disagreement (|act - boman| uncertainty proxy) - Safety proxy (hemolysis risk flags) - Synthesis feasibility - Novelty (Levenshtein vs 45 curated references) - Ensemble (pre-registered weights) 6. Selection criteria (hard filters + ranking + greedy diversity) 7. Results - 89 selected from 383 (pass-rate 23%) - Per-seed breakdown: SEED-003 best Boman (0.538), SEED-005 best ensemble (0.863) - Novelty distribution: 9 high, 58 mid, 22 low - Top-10 table with dual-scorer columns - Suggested 5-candidate pilot (Boman × low-disagreement criterion) 8. Evidence trail (all artifacts with locations) 9. Reproducibility (make phase3 from clean checkout) 10. What we do not know (mandatory integrity section) 11. Safety and dual-use assessment 12. Next steps (human gates only) 13. References (Boman 2003, Eisenberg 1984, Kyte-Doolittle 1982, Zasloff 1987...)
Collaborator
Author
|
Closing in favor of a clean cherry-picked branch that resolves squash-merge conflicts. New PR incoming from feat/phase4-rebase. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR delivers the complete computational nomination pipeline for Phase 3 and the lab-bridge documentation for Phase 4. 89 antimicrobial peptide candidates have been nominated, scored, filtered, and certified — ready for expert review and wet-lab validation.
No biological activity has been demonstrated. All scores are computational heuristics only.
What Was Built
Phase 3 — Candidate Generation and Nomination
Generator (
src/openamp_foundry/generators/template_mutator.py)Scoring pipeline (6 dimensions)
Selection (pre-registered rule in
docs/SELECTION_RULE.md)Evidence artifacts (all committed to
outputs/but not tracked in git)schemas/candidate.schema.jsonPhase 4 — Lab Bridge
examples/known_reference/amp_curated_references.csv)schemas/lab_result.schema.json) — ready for MIC/hemolysis data ingestiondocs/EXPERT_REVIEW_PACK.md) — dual-scorer top-20 table, pilot candidate suggestions, reviewer questionsdocs/NOMINATION_REPORT.md) — complete scientific methodology (new this PR)Key Results
Per seed family:
SEED-003 (tryptophan-rich 11-mers) is recommended for pilot synthesis: highest Boman activity mean (0.538) and strongest dual-scorer consensus across the batch.
Suggested 5-candidate pilot:
How to Reproduce
Files Changed
New source files:
src/openamp_foundry/generators/template_mutator.py— conservative substitution generatorsrc/openamp_foundry/scoring/boman.py— Boman (2003) index + GRAVY + model disagreementsrc/openamp_foundry/reports/batch_pack.py— 5-section batch pack report generatorsrc/openamp_foundry/data/lab_results.py— lab results ingestion for active-learning loopNew tests (314 total):
tests/test_template_mutator.py(34 tests)tests/test_boman_scorer.py(37 tests)tests/test_batch_pack.py(49 tests)tests/test_lab_results.py(15 tests)New documentation:
docs/NOMINATION_REPORT.md— full scientific methodology report (new)docs/EXPERT_REVIEW_PACK.md— lab handoff document with dual-scorer top-20 tabledocs/METHODS.md— technical methods appendixdocs/SELECTION_RULE.md— pre-registered selection rule (locked before generation)docs/RISK_REVIEW.md— risk and dual-use reviewNew schemas:
schemas/candidate.schema.json— updated with optional boman_activity, disagreement fieldsschemas/lab_result.schema.json— for MIC/hemolysis results ingestionNew data:
examples/known_reference/amp_curated_references.csv— 45 curated AMP referencesexamples/sequences/amp_seeds.csv— 5 generator seed sequencesconfigs/phase3.yaml— Phase 3 exploration config (min_novelty=0.05)Remaining Human Gates
These steps require human action and cannot be completed computationally:
docs/EXPERT_REVIEW_PACK.mdschemas/lab_result.schema.jsonto close the active-learning loopTest Plan
make test— 314 tests passmake demo— end-to-end pipeline runs, certificates generate and validatemake phase3— 89 candidates nominated, all certificates schema-valid🤖 Generated with Claude Code