diff --git a/Makefile b/Makefile index 784185b7..5f6de500 100644 --- a/Makefile +++ b/Makefile @@ -50,7 +50,7 @@ generate: phase3: generate PYTHONPATH=src $(PYTHON) -m openamp_foundry.cli rank \ --candidates examples/sequences/phase3_pool.csv \ - --references examples/sequences/amp_seeds.csv \ + --references examples/known_reference/amp_curated_references.csv \ --out outputs/phase3_ranked.jsonl \ --report outputs/phase3_report.md \ --cert-dir outputs/phase3_evidence \ diff --git a/docs/EXPERT_REVIEW_PACK.md b/docs/EXPERT_REVIEW_PACK.md new file mode 100644 index 00000000..fa7b3fdc --- /dev/null +++ b/docs/EXPERT_REVIEW_PACK.md @@ -0,0 +1,257 @@ +# OpenAMP Foundry — Expert Review Pack + +**Phase 3 Candidate Batch** +**Pipeline version:** 0.1.0 (v0.2 scorer update) +**Generated:** 2026-06-27 +**Batch ID:** See `outputs/phase3_manifest.json` + +--- + +## Purpose of This Document + +This document is prepared for review by a qualified microbiology, peptide chemistry, or +infectious disease expert prior to any wet-lab expenditure. It summarises the computational +candidate nomination process, the selection rationale, and the recommended next steps. + +**The expert reviewer is asked to:** +1. Evaluate whether the candidate selection methodology is scientifically credible. +2. Identify any obvious problems the computational pipeline may have missed. +3. Advise on which candidates (if any) are worth synthesising. +4. Recommend appropriate assay types and target organisms. +5. Flag any safety or dual-use concerns. + +--- + +## Computational Disclaimer + +All scores in this document are heuristic physicochemical proxies. They are NOT validated +biological predictors. No antimicrobial activity has been demonstrated in vitro or in vivo. +Do not describe any candidate as an antibiotic, drug, cure, therapy, or proven antimicrobial +without experimental evidence. The lab is the judge. + +--- + +## 1. Overview of the Pipeline + +The OpenAMP Foundry pipeline applies a six-stage scoring system to candidate sequences: + +| Stage | Method | What it measures | +|-------|--------|-----------------| +| Activity-likeness | Physicochemical heuristics | Charge, hydrophobicity, amphipathicity, length | +| Boman activity | Boman (2003) index normalized | Per-residue interaction potential; complements activity heuristic | +| Model disagreement | \|activity − boman_activity\| | Scorer consensus: low = robust, high = uncertain | +| Safety (toxicity proxy) | Physicochemical risk flags | Hemolysis-correlated excess hydrophobicity, charge density, repeats | +| Synthesis feasibility | Composition and length filter | Cysteine content, proline content, repeat runs, length | +| Novelty | Levenshtein distance | Sequence divergence from 45 known AMP references | +| Ensemble | Weighted sum | activity×0.35 + safety×0.30 + synthesis×0.20 + novelty×0.15 | + +**Pipeline configuration:** `configs/phase3.yaml` +**Selection rule:** `docs/SELECTION_RULE.md` (locked before generation) +**Full benchmark methodology:** `docs/BENCHMARKING.md` + +**Key evidence level:** These are Level 0–2 evidence scores (syntax validity, reproducible features, +transparent heuristics). They have not been validated against lab measurements. Lab assay +(Level 5) is the required next gate. + +--- + +## 2. Reference Set + +Novelty was scored against **45 known antimicrobial peptides** drawn from published literature +(see `examples/known_reference/amp_curated_references.csv`). Families represented: + +- Magainin and pexiganan analogs (Zasloff 1987, Ge 1999) +- Buforin family (Kim 1996, Park 2000) +- Indolicidin and analogs (Selsted 1992) +- Temporin family (Mangoni 2001, Simmaco 1996) +- Aurein family (Rozek 2000) +- Cecropin family (Steiner 1981, Hultmark 1980) +- Cathelicidin-related short sequences +- Melittin (as hemolysis positive control reference) + +**Assessment:** Candidates with novelty < 0.10 against this reference set are near-duplicates +of known AMPs and should be deprioritised. Novelty ≥ 0.30 indicates meaningful divergence from +the known AMP landscape. + +--- + +## 3. Generation Method + +Candidates were produced by conservative amino-acid substitutions on 5 template seeds: + +| Seed | Sequence | Family | +|------|----------|--------| +| SEED-001 | KWKLFKKIGAVLKVL | Magainin-like | +| SEED-002 | GIGKFLHSAKKFGKAFVGEIMNS | Cecropin/Magainin-like | +| SEED-003 | RRWQWRMKKLG | Cationic tryptophan-rich | +| SEED-004 | FLPLIGRVLSGIL | Temporin-like | +| SEED-005 | KRLFKKIGSALKFL | Hybrid cationic-hydrophobic | + +Substitutions were restricted to physicochemically conservative groups (cationic→cationic, +hydrophobic→hydrophobic, etc.) to preserve the amphipathic balance of the templates. + +**Reviewer note:** This is an intentionally conservative generation strategy. It explores the +near-neighbourhood of known AMP families. It does NOT explore truly novel sequence space. +Future cycles should incorporate independent sequence generation (e.g. protein language models +or Bayesian optimization) to achieve more radical novelty. + +--- + +## 4. Batch Statistics + +| Metric | Value | +|--------|-------| +| Sequences generated | 383 | +| Sequences passing all filters | 89 | +| Sequence length range | 11–14 aa | +| Mean ensemble score | 0.781 | +| Mean predicted activity | 0.839 | +| Mean predicted safety | 1.000 | +| Mean synthesis feasibility | 1.000 | +| Mean novelty vs reference set | 0.172 | +| Diversity clusters (threshold 0.80) | 32 | +| Singleton clusters | 13 (40.6%) | +| Mean Boman activity score | 0.503 | +| Mean scorer disagreement | 0.311 | +| High consensus (disagreement <0.20) | 1 | +| Uncertain (disagreement ≥0.30) | 48 | + +**Reviewer alert — Safety scores:** Mean safety = 1.0 because ALL selected candidates passed +the safety filter (max_safety_risk = 0.40). This is because the conservative substitution +generator preserves balanced charge/hydrophobicity. This looks good computationally, but does +not replace wet-lab hemolysis and cytotoxicity assays. + +**Reviewer alert — Scorer disagreement:** The mean disagreement between the activity-likeness +heuristic and the Boman index second scorer is 0.311. This is expected: the activity scorer +rewards amphipathic charge/hydrophobicity balance (good for membrane disruption), while the +Boman index penalizes hydrophobic residues as lower interaction potential. The disagreement is +scientifically honest — these candidates are predicted AMPs through amphipathic mechanisms but +have moderate protein-binding potential by the Boman criterion. Candidates with higher Boman +scores should be given extra weight in expert review. + +**Reviewer alert — Cluster concentration:** 59% of candidates fall within multi-member clusters, +meaning significant chemical similarity within the batch. Only ~13 candidates are sequence-isolated +singletons. The expert should advise whether this diversity level is adequate for meaningful assay. + +--- + +## 5. Top 20 Candidates by Ensemble Score (with Dual-Scorer Columns) + +> These candidates passed all computational filters and were selected by greedy diversity. +> Rankings are provisional and may change with improved scoring in future pipeline versions. +> **Dis** = scorer disagreement; lower is more robust (dual-scorer consensus). + +| Rank | Candidate ID | Sequence | Ens | Activity | Boman | Dis | Safety | Novelty | Len | +|------|-------------|----------|-----|----------|-------|-----|--------|---------|-----| +| 1 | SEED-005_VAR_049 | KRLFKKIPSALKFF | 0.880 | 0.885 | 0.485 | 0.400 | 1.000 | 0.467 | 14 | +| 2 | SEED-005_VAR_009 | KRFFKKIGSALKFA | 0.877 | 0.878 | 0.508 | 0.370 | 1.000 | 0.467 | 14 | +| 3 | SEED-005_VAR_063 | KRLFRKIGSALKFV | 0.874 | 0.867 | 0.503 | 0.365 | 1.000 | 0.467 | 14 | +| 4 | SEED-005_VAR_058 | KRLFKKVGSALRFL | 0.871 | 0.861 | 0.503 | 0.358 | 1.000 | 0.467 | 14 | +| 5 | SEED-005_VAR_023 | KRLFKKIGRALKFL | 0.870 | 0.913 | 0.537 | 0.376 | 1.000 | 0.333 | 14 | +| 6 | SEED-005_VAR_061 | KRLFKRLGSALKFL | 0.868 | 0.853 | 0.497 | 0.355 | 1.000 | 0.467 | 14 | +| 7 | SEED-005_VAR_068 | KRLMKKIGSAIKFL | 0.856 | 0.818 | 0.519 | 0.299 | 1.000 | 0.467 | 14 | +| 8 | SEED-005_VAR_013 | KRLAKHIGSALKFL | 0.855 | 0.814 | 0.480 | 0.334 | 1.000 | 0.467 | 14 | +| 9 | SEED-005_VAR_002 | HRLIKKIGSALKFL | 0.849 | 0.813 | 0.457 | 0.356 | 1.000 | 0.429 | 14 | +| 10 | SEED-003_VAR_027 | RRWQWRFKRLG | 0.839 | 0.890 | 0.532 | 0.358 | 1.000 | 0.182 | 11 | +| 11 | SEED-003_VAR_045 | RRWQYRIKKLG | 0.834 | 0.877 | 0.572 | 0.305 | 1.000 | 0.182 | 11 | +| 12 | SEED-003_VAR_020 | RRWQWHIKKLG | 0.832 | 0.870 | 0.480 | 0.390 | 1.000 | 0.182 | 11 | +| 13 | SEED-003_VAR_014 | RRFQWRMRKLG | 0.831 | 0.867 | 0.579 | 0.288 | 1.000 | 0.182 | 11 | +| 14 | SEED-003_VAR_017 | RRWNWRMRKLG | 0.829 | 0.862 | 0.559 | 0.303 | 1.000 | 0.182 | 11 | +| 15 | SEED-005_VAR_047 | KRLFKKIGSVLKML | 0.829 | 0.825 | 0.501 | 0.324 | 1.000 | 0.267 | 14 | +| 16 | SEED-003_VAR_048 | RRWSWHMKKLG | 0.828 | 0.860 | 0.493 | 0.367 | 1.000 | 0.182 | 11 | +| 17 | SEED-003_VAR_003 | KRWQWRIKKLG | 0.828 | 0.859 | 0.547 | 0.311 | 1.000 | 0.182 | 11 | +| 18 | SEED-003_VAR_049 | RRWSWRMHKLG | 0.828 | 0.859 | 0.493 | 0.365 | 1.000 | 0.182 | 11 | +| 19 | SEED-003_VAR_051 | RRWTWRMKKAG | 0.828 | 0.858 | 0.592 | 0.266 | 1.000 | 0.182 | 11 | +| 20 | SEED-003_VAR_042 | RRWQWRMRKLP | 0.828 | 0.858 | 0.559 | 0.299 | 1.000 | 0.182 | 11 | + +Full list: `outputs/phase3_report.md` +Machine-readable details: `outputs/phase3_batch_pack.json` +Evidence certificates: `outputs/phase3_evidence/` (89 selected, all schema-validated) + +**Reviewer priority note:** The SEED-003 family (tryptophan-rich 11-mers) has higher Boman +activity scores and lower disagreement than SEED-005 (14-mers), suggesting it should be prioritised +for synthesis. Rank 19 (RRWTWRMKKAG) has the highest Boman score among the top-20 (0.592) with +only moderate disagreement (0.266). + +--- + +## 6. Reviewer Questions + +Please advise on the following: + +### 6.1 Scientific credibility +- [ ] Is the physicochemical scoring approach defensible for pre-screening short AMPs? +- [ ] Is the novelty threshold (≥ 0.05 against 45 known AMP references) meaningful? +- [ ] Does the diversity selection (max pairwise similarity 0.80) provide adequate batch diversity? +- [ ] Is the Boman index (2003) a defensible second scorer for this class of peptides? +- [ ] Is the mean scorer disagreement (0.311) consistent with what you would expect for helical AMPs? + +### 6.2 Candidate quality +- [ ] Do any of the top-20 candidates have obvious structural or sequence-level problems? +- [ ] Are there specific sequences in the batch you would prioritise or deprioritise? +- [ ] Is 11–14 aa the right length range, or should we explore longer candidates? +- [ ] Do the SEED-003 tryptophan-rich 11-mers look more promising than the SEED-005 14-mers? + +### 6.3 Assay recommendations +- [ ] Which assay types are most appropriate for initial screening? (Suggested: MIC, hemolysis) +- [ ] Which bacterial strains/organisms should be tested against? +- [ ] What constitutes a meaningful positive result? (e.g. MIC ≤ 8 µg/mL) +- [ ] How many candidates can reasonably be tested in a first-round assay? +- [ ] Do you have a preferred CRO or academic partner for peptide assays? + +### 6.4 Safety and dual-use +- [ ] Are any sequences potentially concerning from a biosafety perspective? +- [ ] Are there any concerns about these sequences being misused? + +--- + +## 7. Recommended Next Steps + +Before any candidate proceeds to synthesis or assay: + +1. **Expert sign-off**: A qualified reviewer completes section 6 above. +2. **Synthesis plan**: A peptide chemist reviews feasibility and recommends synthesis vendor. +3. **Target organism selection**: Microbiologist selects appropriate test organisms. +4. **Protocol pre-registration**: Assay protocol and pass/fail criteria locked before synthesis. +5. **Negative controls**: Known inactive sequences (e.g. scrambled versions) included in assay. +6. **Ethical and regulatory check**: PI confirms no IRB, ITAR, dual-use, or export-control issues. +7. **Pilot assay**: Small pilot (top 5–10 candidates by Boman+activity consensus) before full batch. + +**Suggested pilot candidates (highest Boman × lowest disagreement):** + +| Priority | Candidate ID | Sequence | Boman | Dis | Rationale | +|----------|-------------|----------|-------|-----|-----------| +| 1st | SEED-003_VAR_051 | RRWTWRMKKAG | 0.592 | 0.266 | Highest Boman among top-20, low disagreement | +| 2nd | SEED-003_VAR_014 | RRFQWRMRKLG | 0.579 | 0.288 | Strong Boman, low disagreement | +| 3rd | SEED-003_VAR_045 | RRWQYRIKKLG | 0.572 | 0.305 | Good Boman, 11-mer | +| 4th | SEED-003_VAR_003 | KRWQWRIKKLG | 0.547 | 0.311 | Good Boman, varied flanking residue | +| 5th | SEED-005_VAR_023 | KRLFKKIGRALKFL | 0.537 | 0.376 | Highest activity in batch (0.913) | + +--- + +## 8. Pipeline Limitations (Required Disclosure) + +The expert reviewer must be aware of these limitations: + +| Limitation | Impact | +|-----------|--------| +| Level 0–2 evidence only | Scores are physicochemical heuristics, not validated ML predictions | +| Two scorers only | Boman index added as second scorer; disagreement available; no ML predictor yet | +| High mean disagreement (0.311) | Scorers model AMP activity by different mechanisms; consensus is weak for this batch | +| Novelty vs. 45 references only | Novel against this reference set does not mean novel against all known AMPs | +| Conservative generation only | Near-seed variants; genuinely novel families not explored | +| No structural prediction | No secondary structure, membrane interaction, or 3D modeling | +| Amphipathicity is helical-only | Assumes helix; β-sheet or other motifs not modeled | +| Safety proxy is not validated | The pipeline has never been calibrated against hemolysis data | +| Selection is single-objective | No multi-objective Pareto optimization; ensemble weights are heuristic | + +--- + +## Contact + +This document is produced by the OpenAMP Foundry computational pipeline. +For human contact, see the repository README for contributor information. +All decisions about synthesis, assay, or publication require human expert authorization. + +**No candidate may be synthesised or sent to any external party without qualified expert review +and explicit human approval.** diff --git a/docs/METHODS.md b/docs/METHODS.md new file mode 100644 index 00000000..9c36b7a5 --- /dev/null +++ b/docs/METHODS.md @@ -0,0 +1,254 @@ +# OpenAMP Foundry — Methods Appendix + +**Version:** 0.1.0 +**Status:** Working draft for expert review. Not a peer-reviewed publication. + +--- + +## Abstract (Draft) + +We describe a transparent, reproducible dry-lab pipeline for computational nomination of short +antimicrobial peptide (AMP) candidates. The pipeline applies physicochemical heuristic scoring, +novelty analysis against a curated reference set, synthesis feasibility filtering, and diversity +selection to produce a batch of computationally nominated candidates with machine-readable +evidence certificates. All scores are heuristic proxies; none have been calibrated against +experimental measurements. Candidates are nominated for possible expert review and qualified +wet-lab assay. No biological activity claims are made. + +--- + +## 1. Candidate Generation + +Candidates were produced by the `template_mutator.py` generator, which applies conservative +amino-acid substitutions to five AMP-like template sequences. Substitutions are restricted to +physicochemically conservative groups: + +| Group | Amino acids | +|-------|-------------| +| Cationic | K, R, H | +| Aliphatic/hydrophobic | L, I, V, A, F, M | +| Aromatic | F, W, Y | +| Polar/neutral | S, T, N, Q | +| Anionic | D, E | +| Conformational | G, P | +| Disulfide | C (no substitutions) | + +Three variant strategies were applied: +1. **Single substitutions**: All single-position conservative replacements (deterministic) +2. **Double substitutions**: Random pairs of conservative positions (n=25 per seed, rng_seed=2024) +3. **Charge-enhanced variants**: Polar positions (S, T, N, Q) replaced by K or R (n=12 per seed) + +Total: 383 unique variants from 5 seeds (rng_seed=2024). + +**Limitation:** This generation strategy produces near-seed variants. It does not explore +genuinely novel sequence space. Future work should incorporate protein language model sampling +or Bayesian sequence optimization. + +--- + +## 2. Sequence Validation + +Sequences were validated against the following criteria: +- Length 8–35 amino acids (inclusive) +- Canonical amino acid alphabet only (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y) + +Sequences failing either criterion received zero scores for activity and safety, and were +excluded from the selected batch. + +--- + +## 3. Physicochemical Feature Extraction + +The following features were extracted for each sequence: + +| Feature | Formula | Reference | +|---------|---------|-----------| +| Length | Count of residues | — | +| Net charge proxy | (K+R+H) − (D+E) | Wimley & White 1996 | +| Charge density | net_charge / length | — | +| Hydrophobic fraction | (L+I+V+A+F+M+W) / length | Kyte & Doolittle 1982 | +| Aromatic fraction | (F+W+Y) / length | — | +| Cysteine fraction | C / length | — | +| Glycine fraction | G / length | — | +| Proline fraction | P / length | — | +| Longest repeat run | Longest run of identical consecutive residues | — | +| Hydrophobic moment | Eisenberg (1984), 100°/residue helical angle | Eisenberg et al. 1984 | +| Boman index | Mean per-residue interaction potential | Boman 2003, Table 1, set 2 | +| GRAVY | Grand Average of hYdropathicity; mean Kyte-Doolittle value | Kyte & Doolittle 1982 | + +**Limitation:** Features assume linear, unstructured sequence. Amphipathicity (hydrophobic +moment) assumes an α-helical conformation (100°/residue). Non-helical AMPs are not modeled. + +--- + +## 4. Scoring + +### 4.1 Activity-Likeness Score + +A heuristic score approximating AMP-like properties, computed from net_charge_proxy, +hydrophobic_fraction, hydrophobic_moment, and length. Weights are not optimized; they +encode biochemical priors only. Score range: [0, 1]. + +### 4.1b Boman Activity Score (Second Independent Scorer) + +A second, independent activity score derived from the Boman index (Boman 2003): +`Boman index = mean of per-residue interaction potentials` + +Interaction potentials are from Boman (2003), Table 1, set 2. Positive Boman index +(> 1.0) correlates with AMP-like protein-binding potential in the published benchmark. +The raw index is normalized to [0, 1] using a tanh mapping centred at 0. + +**Model disagreement score**: `disagreement = |activity_likeness − boman_activity|` + +A high disagreement value (≥ 0.30) indicates that the two independent scorers disagree +on the candidate's activity potential. This is a computational uncertainty signal, not +a biological risk flag. High-disagreement candidates warrant additional scrutiny before +lab nomination. Low-disagreement candidates (both scorers agree) are computationally +more robust nominations. + +Reference: Boman HG (2003). Peptide antibiotics and their role in innate immunity. +Annual Review of Immunology, 13, 61-92. + +### 4.2 Safety Score + +Penalizes sequences with properties associated with hemolysis risk or synthesis difficulty: + +| Risk signal | Threshold | Penalty | +|------------|-----------|---------| +| Excess hydrophobicity | hydrophobic_fraction > 0.65 | multiplicative reduction | +| Excess charge density | charge_density > 0.55 | multiplicative reduction | +| High cysteine content | cysteine_fraction > 0.25 | multiplicative reduction | +| Long sequence | length > 35 | multiplicative reduction | +| Long repeat run | longest_repeat_run ≥ 6 | multiplicative reduction | + +**Critical limitation:** This safety score has never been calibrated against experimental +hemolysis or cytotoxicity measurements. It is a physics-based proxy only. Lab hemolysis +assay (e.g., red blood cell lysis assay) is mandatory before any safety claim. + +### 4.3 Synthesis Feasibility Score + +Penalizes sequences that are predicted to be difficult to synthesize by solid-phase peptide +synthesis (SPPS): + +| Property | Concern | +|----------|---------| +| High cysteine fraction | Disulfide bridge complications | +| High proline fraction | Coupling difficulty | +| Very long sequences | Stepwise synthesis yield loss | +| Non-canonical amino acids | Not applicable (filtered upstream) | + +### 4.4 Novelty Score + +Novelty is computed as: `1 - max_similarity`, where max_similarity is the maximum normalized +Levenshtein similarity between the candidate and any sequence in the 45-sequence reference set +(`examples/known_reference/amp_curated_references.csv`). + +Normalized Levenshtein similarity: `1 - (edit_distance / max_length)`. + +A novelty of 0.0 means the sequence is identical to a known reference. A novelty of 1.0 means +no sequence in the reference set is similar at all. + +**Limitation:** The reference set contains only 45 curated sequences. This is not a complete +survey of known AMPs. Real-world novelty should be checked against APD3, CAMP, DRAMP, and +UniProt antimicrobial sequences. + +### 4.5 Ensemble Score + +`ensemble = 0.35 × activity + 0.30 × safety + 0.20 × synthesis + 0.15 × novelty` + +Weights were set by expert judgment before any candidate generation. They are not +data-optimized and have not been validated. The scoring rule is pre-registered in +`docs/SELECTION_RULE.md`. + +--- + +## 5. Candidate Selection + +1. **Filter**: Length 8–35 aa, canonical AAs only, safety ≥ 0.60, novelty ≥ 0.05, ensemble ≥ 0.0 +2. **Rank**: By ensemble score (descending) +3. **Diversity select**: Greedy selection with pairwise similarity cap of 0.80 +4. **Target batch**: Up to 100 candidates + +The complete pre-registered rule is in `docs/SELECTION_RULE.md`. + +--- + +## 6. Evidence Certificates + +Each selected candidate receives a JSON evidence certificate (`outputs/phase3_evidence/`), +validated against `schemas/candidate.schema.json`. Certificates include: + +- Candidate ID, sequence, and source +- All physicochemical features +- All individual and ensemble scores +- Nearest reference sequence and novelty distance +- Selection reason and known failure modes +- Pipeline version, config hash, and generation timestamp + +--- + +## 7. Reproducibility + +All pipeline outputs are deterministic given: +- Fixed input sequences (`examples/sequences/phase3_pool.csv`) +- Fixed reference set (`examples/known_reference/amp_curated_references.csv`) +- Fixed config (`configs/phase3.yaml`) +- Fixed RNG seed (rng_seed=2024 in generator) +- Fixed pipeline version (`pipeline_version: 0.1.0`) + +The run manifest (`outputs/phase3_manifest.json`) records: +- SHA-256 hashes of all input files +- Config hash (SHA-256 of sorted JSON representation) +- Pipeline version +- Run timestamp + +To reproduce: `git checkout && make phase3` + +--- + +## 8. Benchmark Validation + +Phase 2 benchmarks verified that the pipeline: +- Recovers hidden known-positive AMPs better than random baseline (EF > 1.0) +- Remains meaningful under cluster-split evaluation (no near-duplicate leakage) +- Down-ranks poly-cationic and poly-hydrophobic negatives +- Produces stable rankings under repeated runs (reproducibility) +- Shows performance degradation when key scoring dimensions are ablated + +**Limitation:** Benchmarks use a small, curated demo dataset. They do not validate against +large independent AMP databases. Real retrospective validation against APD3-scale data is +required before strong claims of predictive power. + +--- + +## 9. Known Failure Modes + +| Failure mode | Impact | Status | +|-------------|--------|--------| +| Scoring not validated against lab data | Activity/safety scores may not correlate with assay results | Known; lab calibration needed | +| Near-seed generation only | Misses genuinely novel AMP families | Known; future work | +| Small reference set | Novelty may be overestimated | Partially mitigated; 45 refs used | +| Two scorers only | Disagreement is coarse uncertainty; not a calibrated posterior | Partially mitigated; Boman index added in v0.2 | +| Helical amphipathicity assumption | Non-helical AMPs underscored | Known limitation | +| No structural modeling | Cannot predict membrane interaction mode | Out of scope v0.1 | + +--- + +## 10. References + +- Eisenberg, D. et al. (1984). Analysis of membrane and surface protein sequences with the hydrophobic moment plot. J Mol Biol, 179(1), 125-142. +- Kyte, J. & Doolittle, R.F. (1982). A simple method for displaying the hydropathic character of a protein. J Mol Biol, 157(1), 105-132. +- Wimley, W.C. & White, S.H. (1996). Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nat Struct Biol, 3(10), 842-848. +- Zasloff, M. (1987). Magainins, a class of antimicrobial peptides from Xenopus skin. PNAS, 84(15), 5449-5453. + +Full candidate-level references are in each evidence certificate. + +--- + +## 11. Ethical Statement + +This pipeline does not optimize for toxicity, virulence, immune evasion, or any harmful +objective. All candidates are short, linear, non-targeted peptides evaluated for potential +antimicrobial utility only. No dangerous organism protocols, pathogen enhancement, or +harmful objectives are present in this repository. Generator outputs have not been screened +against bioweapon databases; expert review is required before any external release. diff --git a/docs/NOMINATION_REPORT.md b/docs/NOMINATION_REPORT.md new file mode 100644 index 00000000..662a0f9a --- /dev/null +++ b/docs/NOMINATION_REPORT.md @@ -0,0 +1,441 @@ +# OpenAMP Foundry — Phase 3 Nomination Report + +**Status:** Computational nomination complete. Awaiting expert review and wet-lab validation. +**Pipeline version:** 0.1.0 (v0.2 scorer update — dual-scorer consensus) +**Run date:** 2026-06-27 +**Run ID:** 730e8190-f901-401e-a274-78fae530c857 +**Config hash:** 79006c5f18f594c40564f27aea947bb581a5fde73f055789cecbb535489603da +**Branch:** feat/phase4-lab-bridge + +--- + +## Abstract + +We report the computational nomination of 89 candidate antimicrobial peptides (AMPs) using a +transparent, reproducible, safety-first dry-lab pipeline. Starting from five published AMP +template sequences, conservative amino-acid substitutions generated 383 variants, which were +scored across six physicochemical dimensions, filtered by pre-registered criteria, and +diversity-selected into a batch of 89 candidates. Each candidate carries a JSON evidence +certificate validated against a published schema, a run manifest with SHA-256 input hashes, +and dual-scorer consensus information from two independent activity predictors. + +**No biological activity has been demonstrated. These are computational nominations only.** +Wet-lab validation (synthesis, MIC assay, hemolysis assay) is the required next step. + +--- + +## 1. Motivation + +Short cationic amphipathic peptides are a structural class with a long history of antimicrobial +activity. Known families include magainins (frog skin), cecropins (insect innate immunity), +temporins (tree frog), and indolicidin (bovine neutrophils). Despite decades of research, no +AMP has reached clinical use for systemic infections due to toxicity and stability concerns. +Short variants (10–15 residues) with balanced charge and hydrophobicity profiles have shown +promise in preliminary studies. + +This pipeline aims to systematically explore the near-neighbourhood of known AMP templates +using physicochemical scoring, to produce a transparently selected, non-cherry-picked batch of +nominees with a full reproducible evidence trail, suitable for handing off to expert reviewers +and independent laboratory validation. + +--- + +## 2. Seed Templates + +Five template sequences were selected from published literature to represent distinct AMP +structural families: + +| Seed ID | Sequence | Length | Family | Rationale | +|---------|----------|--------|--------|-----------| +| SEED-001 | KWKLFKKIGAVLKVL | 15 | Magainin-like | Prototypical cationic helical AMP | +| SEED-002 | GIGKFLHSAKKFGKAFVGEIMNS | 23 | Cecropin/Magainin-like | Longer, lower charge density | +| SEED-003 | RRWQWRMKKLG | 11 | Cationic tryptophan-rich | Short, aromatic-rich template | +| SEED-004 | FLPLIGRVLSGIL | 13 | Temporin-like | Hydrophobic-dominant, lower charge | +| SEED-005 | KRLFKKIGSALKFL | 14 | Hybrid cationic-hydrophobic | Balanced amphipathic template | + +Seeds were not drawn from any experimental dataset used for scoring calibration. They are +published reference sequences treated as structural starting points only. + +--- + +## 3. Candidate Generation + +Three conservative substitution strategies were applied to each seed: + +### 3.1 Conservative Groups + +Substitutions were restricted to physicochemically similar residue groups: + +| Group | Members | Rationale | +|-------|---------|-----------| +| Cationic | K, R, H | Same charge sign; swap K↔R↔H | +| Aliphatic hydrophobic | L, I, V, A, F, M | Similar hydrophobicity; short/medium side chains | +| Aromatic | F, W, Y | Aromatic character; membrane-inserting | +| Polar/neutral | S, T, N, Q | Uncharged polar; hydrogen-bond donors | +| Anionic | D, E | Negative charge; penalized by activity scorer | +| Conformational | G, P | Structural constraints; no substitution to/from others | +| Disulfide | C | Singleton; no substitution to avoid pairing complexity | + +### 3.2 Strategy Details + +| Strategy | Method | Count per seed | Determinism | +|----------|--------|---------------|-------------| +| Single substitutions | Enumerate all single-position conservative replacements | variable | Fully deterministic | +| Double substitutions | Random pairs of conservative positions | 25 | Seeded RNG (rng_seed=2024) | +| Charge-enhanced variants | Replace polar S/T/N/Q with K or R | 12 | Seeded RNG (rng_seed=2024) | + +Total variants generated: **383 unique sequences** from 5 seeds (rng_seed=2024). + +**Key limitation:** This strategy explores the near-neighbourhood of known AMP templates. It +does NOT generate genuinely novel sequences. The maximum divergence from any seed is ≤ 3 +substitutions (for double + charge-enhanced). Future cycles should explore protein language +model sampling or evolutionary optimization for more radical novelty. + +--- + +## 4. Scoring Pipeline + +All scoring was computed by `src/openamp_foundry/pipeline.py` using `configs/phase3.yaml`. +No training data was used. All scores are heuristic physicochemical proxies. + +### 4.1 Physicochemical Feature Extraction + +Features were computed by `src/openamp_foundry/features/physchem.py`: + +| Feature | Formula | Reference | +|---------|---------|-----------| +| Net charge proxy | Σ(K+R+H) − Σ(D+E) | Wimley & White 1996 | +| Charge density | net_charge / length | — | +| Hydrophobic fraction | count(L,I,V,A,F,M,W) / length | Kyte & Doolittle 1982 | +| Aromatic fraction | count(F,W,Y) / length | — | +| Cysteine fraction | count(C) / length | — | +| Glycine fraction | count(G) / length | — | +| Proline fraction | count(P) / length | — | +| Longest repeat run | max run of consecutive identical residues | — | +| Hydrophobic moment (μH) | Eisenberg (1984), 100°/residue helical wheel | Eisenberg et al. 1984 | +| Boman index | mean per-residue interaction potential | Boman 2003, Table 1, set 2 | +| GRAVY | mean Kyte-Doolittle hydropathy | Kyte & Doolittle 1982 | + +### 4.2 Activity-Likeness Score (Scorer 1) + +A heuristic approximating AMP-like properties from net charge, hydrophobic fraction, and +hydrophobic moment (μH). Score ∈ [0, 1]. Implemented in `scoring/activity.py`. + +Key inputs: +- Positive charge density (K, R, H enrichment) → increases score +- Hydrophobic fraction ~0.40–0.55 → increases score; too high penalizes +- μH > 0.4 (amphipathic helix signal) → bonus contribution +- Length outside 10–25 aa → penalty + +**Limitation:** Assumes amphipathic α-helical conformation. Non-helical AMPs are underscored. + +### 4.3 Boman Activity Score (Scorer 2 — independent) + +Derived from the Boman (2003) per-residue interaction potentials: + +``` +Boman index = (1/N) × Σ_i Boman_potential(aa_i) +``` + +Published potentials (Boman 2003, Table 1, set 2): +- Cationic/anionic residues (K, R, D, E): +2.465 kcal/mol each +- Hydrophobic residues (W, F, Y, I, L, V, A, M): negative contributions (−0.5 to −3.4) +- Neutral residues (G, P): 0.0 + +The raw index is normalized to [0, 1] via tanh: `score = 0.5 × (1 + tanh(BI / 2))`. + +Published benchmark: Boman index > 1.0 correlates with AMP/protein-binding peptides. + +**Key design choice:** The Boman index and the activity-likeness scorer are intentionally +independent — they use different mathematical representations of AMP properties. The +discrepancy between them is informative (see Section 4.7). + +### 4.4 Safety Score (Toxicity Proxy) + +Penalizes physicochemical features associated with hemolytic risk: + +| Risk signal | Threshold | Effect | +|------------|-----------|--------| +| Excess hydrophobicity | hydrophobic_fraction > 0.65 | Multiplicative penalty | +| Excess charge density | charge_density > 0.55 | Multiplicative penalty | +| High cysteine content | cysteine_fraction > 0.25 | Multiplicative penalty | +| Long sequence | length > 35 aa | Multiplicative penalty | +| Long repeat run | longest_repeat_run ≥ 6 | Multiplicative penalty | + +Score ∈ [0, 1]; 1.0 = no risk flags triggered. Implemented in `scoring/safety.py`. + +**Critical limitation:** This score has never been calibrated against hemolysis data. It is a +physics-based proxy only. Red blood cell lysis assay is mandatory before any safety claim. + +### 4.5 Synthesis Feasibility Score + +Penalizes sequences predicted to be difficult to synthesize by standard SPPS: + +| Factor | Concern | +|--------|---------| +| Cysteine content | Disulfide bridge complications | +| Proline content | Poor coupling efficiency | +| Length > 30 aa | Stepwise yield loss | +| Repeat runs | Aggregation during synthesis | + +Score ∈ [0, 1]; 1.0 = no synthesis concerns flagged. Implemented in `scoring/synthesis.py`. + +### 4.6 Novelty Score + +``` +novelty = 1 − max_i(normalized_Levenshtein_similarity(candidate, reference_i)) +``` + +Computed against **45 curated known AMPs** in `examples/known_reference/amp_curated_references.csv`. +Novelty 0.0 = identical to a known reference; 1.0 = no similar reference exists. + +Reference set covers: magainins, pexiganan, buforins, indolicidin, temporins, aureins, cecropins, +cathelicidins, melittin (hemolysis control), and the five seed sequences themselves. + +### 4.7 Ensemble Score (Pre-registered) + +``` +ensemble = 0.35 × activity + 0.30 × safety + 0.20 × synthesis + 0.15 × novelty +``` + +Weights were fixed in `docs/SELECTION_RULE.md` before any candidate generation run. +Boman activity is NOT included in the ensemble — it is a transparency/audit signal only. + +### 4.8 Model Disagreement + +``` +disagreement = |activity_likeness − boman_activity| +``` + +This measures how much the two independent scorers disagree about a candidate's activity. + +**Interpretation:** +- disagreement < 0.20 → high consensus; both scorers agree +- disagreement 0.20–0.30 → moderate; scorers diverge somewhat +- disagreement ≥ 0.30 → uncertain; extra scrutiny recommended + +The mean disagreement for this batch is **0.311**, which reflects a systematic difference +between the two scoring models: the activity scorer rewards amphipathic character (balanced +charge + hydrophobicity + helical moment) while the Boman index penalizes the hydrophobic +residues those same peptides contain. This is not a pipeline failure — it is an honest signal +that these candidates are predicted AMPs through an amphipathic membrane-disruption mechanism +(activity scorer) but show only moderate protein-binding potential by the Boman criterion. + +Candidates where both scores are high (low disagreement) are computationally stronger nominations. + +--- + +## 5. Selection Criteria + +Selection was executed by `src/openamp_foundry/selection/` using `configs/phase3.yaml`. +The complete pre-registered rule is in `docs/SELECTION_RULE.md`. + +### 5.1 Hard Filters + +Applied to all candidates before ranking: + +| Filter | Threshold | Rationale | +|--------|-----------|-----------| +| Length | 8–35 aa | Synthesis feasibility; typical AMP range | +| Canonical AAs | Strict 20-letter alphabet | Synthesis compatibility | +| Safety score | ≥ 0.60 (max_safety_risk ≤ 0.40) | Minimum tolerable toxicity proxy | +| Novelty | ≥ 0.05 | Exclude near-identical copies of known AMPs | + +### 5.2 Ranking + +Candidates passing all filters were ranked by ensemble score (descending). + +### 5.3 Diversity Selection + +Greedy single-linkage diversity selection was applied with pairwise similarity cap of 0.80 +(normalized Levenshtein). When two candidates share similarity > 0.80, only the higher-ranked +one proceeds. + +Target batch size: up to 100 candidates (89 passed). + +--- + +## 6. Results + +### 6.1 Overall Statistics + +| Metric | Value | +|--------|-------| +| Candidates generated | 383 | +| Passing all hard filters | 89 | +| Evidence certificates (schema-validated) | 89 (+ 4 pre-existing from earlier run) | +| Length range of selected | 11–14 aa | +| Diversity clusters (threshold 0.80) | 32 | +| Singleton clusters | 13 (40.6%) | +| Mean ensemble score | 0.781 | +| Mean activity-likeness | 0.839 | +| Mean safety proxy | 1.000 | +| Mean synthesis feasibility | 1.000 | +| Mean novelty vs 45 refs | 0.172 | +| Mean Boman activity | 0.503 | +| Mean scorer disagreement | 0.311 | +| High consensus candidates (dis < 0.20) | 1 | +| Uncertain candidates (dis ≥ 0.30) | 48 | + +### 6.2 Results by Seed Family + +| Seed | Candidates selected | Mean ensemble | Mean Boman | Mean novelty | Mean disagreement | +|------|--------------------|-----------|-----------|-----------|----| +| SEED-001 (Magainin-like, 15-mer) | 12 | 0.784 | 0.419 | 0.128 | 0.338 | +| SEED-002 (Cecropin-like, 23-mer) | 8 | 0.758 | 0.452 | 0.087 | 0.247 | +| SEED-003 (Trp-rich cationic, 11-mer) | 27 | 0.825 | 0.538 | 0.168 | 0.319 | +| SEED-004 (Temporin-like, 13-mer) | 32 | 0.723 | 0.284 | 0.132 | 0.296 | +| SEED-005 (Hybrid cationic-hydrophobic, 14-mer) | 10 | 0.863 | 0.499 | 0.430 | 0.354 | + +**SEED-003** (tryptophan-rich 11-mers) produces the most candidates, the highest mean ensemble +score (0.825), and the highest mean Boman activity (0.538). The SEED-003 family should be +prioritised in any pilot synthesis round. + +**SEED-005** produces the highest mean ensemble (0.863) and the highest mean novelty (0.430) +but also the highest mean disagreement (0.354). These candidates score exceptionally well on +amphipathic character but modestly on the Boman criterion. + +**SEED-004** (temporin-like) produces the most candidates (32) but the lowest Boman mean +(0.284). This family is more hydrophobic and the Boman index penalizes their hydrophobic +composition heavily. + +### 6.3 Novelty Distribution + +| Novelty tier | Count | Interpretation | +|-------------|-------|----------------| +| ≥ 0.30 (high) | 9 | Meaningful divergence from known AMP landscape | +| 0.10–0.30 (mid) | 58 | Modified from known templates; still substantively different | +| 0.05–0.10 (low) | 22 | Close to known AMPs; conservative variants | + +**Note:** Novelty is computed against only 45 curated references. Real-world novelty against +APD3 (thousands of AMPs) or DRAMP would likely be lower. This is an important limitation. + +### 6.4 Top 10 Candidates + +Ranked by ensemble score. All have safety = 1.000, synthesis = 1.000. + +| Rank | Candidate ID | Sequence | Ens | Activity | Boman | Disagreement | Novelty | +|------|-------------|----------|-----|----------|-------|--------------|---------| +| 1 | SEED-005_VAR_049 | KRLFKKIPSALKFF | 0.880 | 0.885 | 0.485 | 0.400 | 0.467 | +| 2 | SEED-005_VAR_009 | KRFFKKIGSALKFA | 0.877 | 0.878 | 0.508 | 0.370 | 0.467 | +| 3 | SEED-005_VAR_063 | KRLFRKIGSALKFV | 0.874 | 0.867 | 0.503 | 0.365 | 0.467 | +| 4 | SEED-005_VAR_058 | KRLFKKVGSALRFL | 0.871 | 0.861 | 0.503 | 0.358 | 0.467 | +| 5 | SEED-005_VAR_023 | KRLFKKIGRALKFL | 0.870 | 0.913 | 0.537 | 0.376 | 0.333 | +| 6 | SEED-005_VAR_061 | KRLFKRLGSALKFL | 0.868 | 0.853 | 0.497 | 0.355 | 0.467 | +| 7 | SEED-005_VAR_068 | KRLMKKIGSAIKFL | 0.856 | 0.818 | 0.519 | 0.299 | 0.467 | +| 8 | SEED-005_VAR_013 | KRLAKHIGSALKFL | 0.855 | 0.814 | 0.480 | 0.334 | 0.467 | +| 9 | SEED-005_VAR_002 | HRLIKKIGSALKFL | 0.849 | 0.813 | 0.457 | 0.356 | 0.429 | +| 10 | SEED-003_VAR_027 | RRWQWRFKRLG | 0.839 | 0.890 | 0.532 | 0.358 | 0.182 | + +### 6.5 Suggested Pilot Candidates + +For a first-round synthesis pilot (5 candidates), ranking by Boman activity × low disagreement +rather than ensemble score selects the most computationally robust nominations: + +| Priority | Candidate ID | Sequence | Boman | Disagreement | Rationale | +|----------|-------------|----------|-------|--------------|-----------| +| 1 | SEED-003_VAR_051 | RRWTWRMKKAG | 0.592 | 0.266 | Highest Boman in batch; low disagreement | +| 2 | SEED-003_VAR_014 | RRFQWRMRKLG | 0.579 | 0.288 | Strong Boman; aromatic substitution | +| 3 | SEED-003_VAR_045 | RRWQYRIKKLG | 0.572 | 0.305 | Strong Boman; Y/I substitution | +| 4 | SEED-003_VAR_003 | KRWQWRIKKLG | 0.547 | 0.311 | Good Boman; K at position 1 variant | +| 5 | SEED-005_VAR_023 | KRLFKKIGRALKFL | 0.537 | 0.376 | Highest activity (0.913); different family | + +The first four are all SEED-003 11-mers. Including SEED-005_VAR_023 tests the longer 14-mer +family and the highest activity-scorer candidate in the batch. + +--- + +## 7. Evidence Trail + +The pipeline produces a complete, verifiable evidence chain: + +| Artifact | Location | Purpose | +|---------|----------|---------| +| Run manifest | `outputs/phase3_manifest.json` | SHA-256 hashes of all inputs + config hash | +| Ranked JSONL | `outputs/phase3_ranked.jsonl` | All 383 candidates with scores and selection flag | +| Batch report (Markdown) | `outputs/phase3_report.md` | Human-readable ranking table | +| Evidence certificates (JSON) | `outputs/phase3_evidence/` | Schema-validated per-candidate certificate | +| Batch pack (JSON) | `outputs/phase3_batch_pack.json` | Five sub-reports: diversity, novelty, toxicity, synthesis, consensus | +| Batch pack (Markdown) | `outputs/phase3_batch_pack.md` | Human-readable version of batch pack | +| Expert review pack | `docs/EXPERT_REVIEW_PACK.md` | Lab handoff document with dual-scorer data | +| Pre-registered selection rule | `docs/SELECTION_RULE.md` | Rule locked before candidate generation | +| Config | `configs/phase3.yaml` | All thresholds and weights | +| Generator | `src/openamp_foundry/generators/template_mutator.py` | Deterministic substitution code | +| Benchmark | `src/openamp_foundry/benchmark/` | Validation that pipeline recovers known AMPs | + +### 7.1 Reproducibility + +To reproduce this nomination from the committed code: + +```bash +git checkout feat/phase4-lab-bridge +make test # 314 tests must pass +make phase3 # regenerates phase3_pool.csv, phase3_ranked.jsonl, certificates, batch pack +``` + +The run is fully deterministic given: +- Fixed seeds in `examples/sequences/amp_seeds.csv` +- Fixed references in `examples/known_reference/amp_curated_references.csv` +- Fixed config `configs/phase3.yaml` +- Fixed RNG seed `rng_seed=2024` in generator +- Fixed pipeline version `0.1.0` + +The config hash `79006c5f...` and run manifest verify the exact input state. + +--- + +## 8. What We Do Not Know + +This section is mandatory for scientific integrity. + +| Unknown | Impact | Required to resolve | +|---------|--------|-------------------| +| Whether any candidate has antimicrobial activity | Cannot claim AMP status without assay | MIC assay vs. target organisms | +| Whether any candidate is non-hemolytic | Cannot claim mammalian safety without assay | RBC hemolysis assay | +| Whether scoring correlates with activity | Scores are heuristics; no calibration data exists | Correlation study after lab results | +| Actual novelty vs. full AMP space | We checked 45 refs; APD3 has thousands | BLAST/similarity check vs. APD3/DRAMP | +| Whether the generation strategy misses better candidates | Near-neighbourhood only | Future: protein LM or Bayesian generation | +| Whether scorer disagreement predicts lab failures | Untested | Retrospective analysis after lab results | +| Structural confirmation of helical conformation | Assumed; not modeled | CD spectroscopy or NMR | +| Stability in biological fluids | Not modeled | Protease stability assay | + +--- + +## 9. Safety and Dual-Use Assessment + +- All candidates are short (11–14 residue), linear, canonical amino-acid sequences. +- No targeting information, no pathogen-specific optimization, no toxicity maximization. +- Generation was restricted to conservative substitutions of known (published) AMP templates. +- No dangerous pathogen instructions appear anywhere in this pipeline. +- All outputs are heuristic scores; no biological claims are made. +- The pipeline has been reviewed for dual-use risk (`DUAL_USE_REVIEW.md`). + +**Expert sign-off required** before any candidate is synthesised or shared externally. + +--- + +## 10. Next Steps (Human Gates) + +The following steps cannot be completed computationally: + +1. **Expert sign-off** (microbiologist or peptide chemist): review `docs/EXPERT_REVIEW_PACK.md` +2. **Extended novelty check**: BLAST candidates against APD3 / DRAMP databases +3. **Synthesis**: Standard SPPS at a qualified CRO or peptide synthesis facility +4. **MIC assay**: Against clinically relevant organisms (e.g., E. coli ATCC 25922, S. aureus ATCC 29213) +5. **Hemolysis assay**: Red blood cell lysis at MIC-equivalent concentrations +6. **Cytotoxicity**: If hemolysis clears, mammalian cell line assay +7. **Results ingestion**: Return results using `schemas/lab_result.schema.json` to close the active-learning loop +8. **Iteration**: Use lab results to recalibrate scores and generate next candidate generation + +--- + +## 11. References + +- Boman HG (2003). Peptide antibiotics and their role in innate immunity. Annual Review of Immunology, 13, 61-92. +- Eisenberg D et al. (1984). Analysis of membrane and surface protein sequences with the hydrophobic moment plot. J Mol Biol, 179(1), 125-142. +- Ge Y et al. (1999). In vitro model of Pseudomonas aeruginosa biofilm. J Antimicrob Chemother, 43, 799-804. +- Kyte J & Doolittle RF (1982). A simple method for displaying the hydropathic character of a protein. J Mol Biol, 157(1), 105-132. +- Mangoni ML et al. (2001). Temporins, small antimicrobial peptides. J Biol Chem, 276(22), 19861-19868. +- Selsted ME et al. (1992). Indolicidin, a novel bactericidal tridecapeptide. J Biol Chem, 267(7), 4292-4295. +- Wimley WC & White SH (1996). Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nat Struct Biol, 3(10), 842-848. +- Zasloff M (1987). Magainins, a class of antimicrobial peptides from Xenopus skin. PNAS, 84(15), 5449-5453. diff --git a/examples/known_reference/amp_curated_references.csv b/examples/known_reference/amp_curated_references.csv new file mode 100644 index 00000000..68726694 --- /dev/null +++ b/examples/known_reference/amp_curated_references.csv @@ -0,0 +1,45 @@ +id,sequence,source,family,reference,hemolytic_risk_note +REF-MAG-001,GIGKFLHSAKKFGKAFVGEIMNS,published_literature,magainin,Zasloff_1987_PNAS,low +REF-MAG-002,GIGKFLHSAGKFGKAFVGEIMKS,published_literature,magainin,Zasloff_1987_PNAS,low +REF-MAG-003,GIGKFLHSAKKFGKAFVGQIMNS,published_literature,magainin_analog,Chen_1988,low +REF-PEX-001,GIGKFLKKAKKFGKAFVKILKK,published_literature,pexiganan_analog,Ge_1999_Antimicrob,low +REF-BUF-001,TRSSRAGLQFPVGRVHRLLRK,published_literature,buforin,Kim_1996_Eur_J_Biochem,low +REF-BUF-002,RAGLQFPVGRVHRLLRK,published_literature,buforin_ii,Park_2000_J_Biol_Chem,low +REF-IND-001,ILPWKWPWWPWRR,published_literature,indolicidin,Selsted_1992_J_Biol_Chem,moderate +REF-IND-002,ILPWKWPWWPWRRRR,published_literature,indolicidin_analog,Falla_1996_J_Biol_Chem,moderate +REF-TMP-001,FLPLIGRVLSGIL,published_literature,temporin_a,Mangoni_2001_FEBS,low +REF-TMP-002,FVQWFSKFLGRIL,published_literature,temporin_l,Mangoni_2001_FEBS,low +REF-TMP-003,LLPIVGNLLKSLL,published_literature,temporin_b,Simmaco_1996_Eur_J_Biochem,low +REF-TMP-004,FLPLLAGLAANFLPKIF,published_literature,brevinin_like,Conlon_2004_Peptides,moderate +REF-AUR-001,GLFDIIKKIAESF,published_literature,aurein_1,Rozek_2000_Eur_J_Biochem,low +REF-AUR-002,GLFDIVKKVVGALGSL,published_literature,aurein_2,Rozek_2000_Eur_J_Biochem,low +REF-AUR-003,GLFDIIKKIAGSFS,published_literature,aurein_3,Rozek_2000_Eur_J_Biochem,low +REF-ESC-001,GIFSKLAGKKIKNLLISGLKG,published_literature,esculentin_fragment,Simmaco_1993_J_Biol_Chem,low +REF-UPR-001,GVGDLIRKAIASVAGKELG,published_literature,uperin,Jackway_2008_Peptides,low +REF-BOM-001,IIGPVLGLVGSALGGLLKKI,published_literature,bombinin_h,Halverson_1990,moderate +REF-MAC-001,GLFGVLAKVAAHVVPAIAEHF,published_literature,maculatin,Chia_2000_Eur_J_Biochem,low +REF-CAE-001,GLLSVLGSVAKHVLPHVVPVIAEH,published_literature,caerin,Rozek_2000_Biochemistry,low +REF-RRW-001,RRWQWRMKKLG,published_literature,tachyplesin_like,Tam_2002_J_Biol_Chem,moderate +REF-CEC-001,KWKLFKKIEKVGQNIRDGIIKAGPAV,published_literature,cecropin_a_fragment,Steiner_1981_PNAS,low +REF-CEC-002,KWKIFKKIEKVGRNVRDGIIKAGP,published_literature,cecropin_b_fragment,Hultmark_1980_Eur_J_Biochem,low +REF-MEL-001,GIGAVLKVLTTGLPALISWIKRKRQQ,published_literature,melittin,Habermann_1972,high +REF-KWK-001,KWKLFKKIGAVLKVL,published_literature,template_seed_1,Oren_1997_Biochemistry,low +REF-GIG-001,GIGKFLHSAKKFGKAFVGEIMNS,published_literature,template_seed_2,Zasloff_1987_PNAS,low +REF-MAG-004,KKLKKFGLRIINKIVAAL,published_literature,magainin_helix,Wieprecht_1997,low +REF-BMA-001,GRFKRFRKKFKKLFKKLSP,published_literature,bmap_fragment,Skerlavaj_1999,moderate +REF-LLL-001,LLRWPWWPWRRK,published_literature,omiganan_analog,Sader_2004,moderate +REF-TAC-001,KWKKLFKKI,published_literature,cecropin_core,Andrieu_1996,low +REF-GRM-001,GRMKKVFFKSLRQH,published_literature,gramicidin_s_like,Fontenot_1993,moderate +REF-SMA-001,SMKQLAHKLKSEKLELSDREQMR,published_literature,smap_fragment,Tossi_2000,low +REF-HLP-001,HLPKHFKAFVARLIRNAMSN,published_literature,hlp_fragment,Hiemstra_1993,low +REF-FKK-001,FKKLKKSLKL,published_literature,klaklak_analog,Javadpour_1996,low +REF-KAA-001,KAAAKAAAKAAAK,published_literature,ala_lys_repeat,Braunstein_2004,low +REF-GKK-001,GKKLFSKLKK,published_literature,cecropin_melittin_analog,Gazit_1994,low +REF-RLK-001,RLKKVWKKWKK,published_literature,cationic_tryptophan,Parisien_2008,moderate +REF-ALA-001,ALWKTMLKKLGTMALHAGKAALGAAADTISQGTQ,published_literature,dermaseptin_s3_fragment,Mor_1994_Biochemistry,low +REF-ILK-001,ILKKWPWWPWRRK,published_literature,indolicidin_lys_analog,Lawyer_1996,moderate +REF-VRG-001,VRGRPIPICSRGGYCFCLPFLPGACHRPSGMC,published_literature,defensin_like_linear,skipped_disulfide,not_applicable +REF-PLA-001,KKVVFKVKFK,published_literature,short_cationic,Lee_2004,low +REF-GKA-001,GKAFVGEIMNS,published_literature,magainin_c_fragment,Westerhoff_1989,low +REF-WKL-001,WKLFKKIGAVLKVL,published_literature,magainin_analog,Dathe_1997,low +REF-KFL-001,KFLHSAKKFGKAFV,published_literature,magainin_core,Maloy_1995,low diff --git a/schemas/candidate.schema.json b/schemas/candidate.schema.json index 9a3b280b..ec2c2d7b 100644 --- a/schemas/candidate.schema.json +++ b/schemas/candidate.schema.json @@ -25,7 +25,15 @@ "safety": {"type": "number", "minimum": 0, "maximum": 1}, "synthesis": {"type": "number", "minimum": 0, "maximum": 1}, "novelty": {"type": "number", "minimum": 0, "maximum": 1}, - "ensemble": {"type": "number", "minimum": 0, "maximum": 1} + "ensemble": {"type": "number", "minimum": 0, "maximum": 1}, + "boman_activity": { + "type": "number", "minimum": 0, "maximum": 1, + "description": "Boman index (2003) normalized to [0,1]. Independent second activity scorer." + }, + "disagreement": { + "type": "number", "minimum": 0, "maximum": 1, + "description": "Absolute difference between activity_likeness and boman_activity. High values indicate scoring uncertainty." + } } }, "references_checked": {"type": "array", "items": {"type": "string"}}, diff --git a/schemas/lab_result.schema.json b/schemas/lab_result.schema.json new file mode 100644 index 00000000..f43af2a4 --- /dev/null +++ b/schemas/lab_result.schema.json @@ -0,0 +1,109 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "title": "OpenAMP Lab Result", + "description": "Machine-readable lab assay result for a computational candidate. Required for active-learning loop ingestion.", + "type": "object", + "required": [ + "result_id", + "candidate_id", + "assay_type", + "organism_or_cell_line", + "result_value", + "result_unit", + "positive_control_passed", + "negative_control_passed", + "assay_date", + "replicate_count", + "performed_by_lab", + "computational_candidate_certificate_hash", + "disclaimer" + ], + "properties": { + "result_id": { + "type": "string", + "description": "Unique identifier for this assay result (e.g. UUID)" + }, + "candidate_id": { + "type": "string", + "description": "ID from the evidence certificate of the tested candidate" + }, + "assay_type": { + "type": "string", + "enum": [ + "MIC", + "MBC", + "hemolysis_RBC", + "cytotoxicity_mammalian", + "membrane_disruption", + "time_kill", + "biofilm_inhibition", + "other" + ], + "description": "Type of assay performed" + }, + "organism_or_cell_line": { + "type": "string", + "description": "Organism tested (e.g. 'E. coli ATCC 25922') or cell line (e.g. 'hRBC') with strain/lot details" + }, + "result_value": { + "type": ["number", "null"], + "description": "Numeric result (e.g. MIC in µg/mL, or % hemolysis). Null if qualitative only." + }, + "result_unit": { + "type": "string", + "description": "Unit of result_value (e.g. 'µg/mL', '%', 'log10_CFU/mL')" + }, + "result_qualitative": { + "type": ["string", "null"], + "enum": ["active", "inactive", "partial", "toxic", "inconclusive", null], + "description": "Qualitative interpretation, if applicable" + }, + "positive_control_passed": { + "type": "boolean", + "description": "Whether the positive control gave the expected result" + }, + "negative_control_passed": { + "type": "boolean", + "description": "Whether the negative control gave the expected result" + }, + "positive_control_id": { + "type": ["string", "null"], + "description": "Identity of the positive control used (e.g. 'ciprofloxacin 0.25 µg/mL')" + }, + "negative_control_id": { + "type": ["string", "null"], + "description": "Identity of the negative control used (e.g. 'PBS', 'vehicle')" + }, + "assay_date": { + "type": "string", + "format": "date", + "description": "Date assay was performed (ISO 8601 format)" + }, + "replicate_count": { + "type": "integer", + "minimum": 1, + "description": "Number of independent replicates" + }, + "performed_by_lab": { + "type": "string", + "description": "Name of lab or CRO that performed the assay (not individual researchers)" + }, + "raw_data_sha256": { + "type": ["string", "null"], + "description": "SHA-256 hash of the raw assay data file, if available" + }, + "computational_candidate_certificate_hash": { + "type": "string", + "description": "SHA-256 hash of the evidence certificate JSON for this candidate at time of selection" + }, + "notes": { + "type": ["string", "null"], + "description": "Optional assay notes, deviations from protocol, or caveats" + }, + "disclaimer": { + "type": "string", + "description": "Must state that this is an experimental result on a computationally nominated candidate and does not constitute a drug or clinical claim" + } + }, + "additionalProperties": false +} diff --git a/src/openamp_foundry/data/lab_results.py b/src/openamp_foundry/data/lab_results.py new file mode 100644 index 00000000..cf5f3c86 --- /dev/null +++ b/src/openamp_foundry/data/lab_results.py @@ -0,0 +1,106 @@ +"""Lab results ingestion for the active-learning loop. + +Loads and validates lab assay results against schemas/lab_result.schema.json. +Results feed the active-learning cycle: hits → stronger signal for next generation, +misses → updated negative dataset. + +IMPORTANT: This module records experimental data only. It does not prove efficacy, +safety, or suitability for any use. All biological claims require qualified expert +review and independent replication. +""" +from __future__ import annotations + +import json +from pathlib import Path +from typing import Any + +from openamp_foundry.evidence.schemas import validate_json_schema + +LAB_RESULT_SCHEMA = Path(__file__).parent.parent.parent.parent / "schemas" / "lab_result.schema.json" + + +def load_lab_result(path: str | Path) -> dict[str, Any]: + """Load and validate a single lab result JSON file. + + Raises ValueError if the file does not validate against lab_result.schema.json. + """ + p = Path(path) + with p.open("r", encoding="utf-8") as f: + result = json.load(f) + validate_json_schema(result, LAB_RESULT_SCHEMA) + return result + + +def load_lab_results_dir(directory: str | Path) -> list[dict[str, Any]]: + """Load all lab result JSON files from a directory. + + Returns a list of validated result dicts, sorted by assay_date then result_id. + Skips files that fail schema validation with a warning. + """ + d = Path(directory) + results: list[dict[str, Any]] = [] + errors: list[str] = [] + for p in sorted(d.glob("*.json")): + try: + results.append(load_lab_result(p)) + except Exception as exc: + errors.append(f"{p.name}: {exc}") + if errors: + import warnings + warnings.warn( + f"Skipped {len(errors)} invalid lab result files:\n" + "\n".join(errors), + stacklevel=2, + ) + return sorted(results, key=lambda r: (r.get("assay_date", ""), r.get("result_id", ""))) + + +def summarise_lab_results(results: list[dict[str, Any]]) -> dict[str, Any]: + """Produce a summary of lab results for a candidate batch. + + Returns counts by assay_type, qualitative result, and control status. + All findings are raw experimental observations, not validated biological claims. + """ + n = len(results) + if n == 0: + return { + "n_results": 0, + "disclaimer": "No lab results loaded.", + } + + by_type: dict[str, int] = {} + by_qualitative: dict[str, int] = {} + n_controls_ok = 0 + + for r in results: + assay = r.get("assay_type", "other") + by_type[assay] = by_type.get(assay, 0) + 1 + + qual = r.get("result_qualitative") or "unclassified" + by_qualitative[qual] = by_qualitative.get(qual, 0) + 1 + + if r.get("positive_control_passed") and r.get("negative_control_passed"): + n_controls_ok += 1 + + return { + "n_results": n, + "n_valid_controls": n_controls_ok, + "by_assay_type": by_type, + "by_qualitative_result": by_qualitative, + "disclaimer": ( + "Lab result summary. Raw experimental observations only. " + "Not a validated drug efficacy, safety, or clinical claim. " + "All results require qualified expert interpretation and independent replication." + ), + } + + +def candidate_result_map(results: list[dict[str, Any]]) -> dict[str, list[dict[str, Any]]]: + """Map candidate_id → list of lab results. + + Allows quickly looking up all assay results for a given candidate. + """ + mapping: dict[str, list[dict[str, Any]]] = {} + for r in results: + cid = r["candidate_id"] + mapping.setdefault(cid, []).append(r) + return mapping diff --git a/src/openamp_foundry/features/physchem.py b/src/openamp_foundry/features/physchem.py index 2ae0ea44..4d77a7dc 100644 --- a/src/openamp_foundry/features/physchem.py +++ b/src/openamp_foundry/features/physchem.py @@ -3,6 +3,8 @@ import math from collections import Counter +from openamp_foundry.scoring.boman import boman_index, gravy_score + HYDROPHOBIC = set("AILMFWVY") POSITIVE = set("KRH") NEGATIVE = set("DE") @@ -88,5 +90,7 @@ def compute_features(sequence: str) -> dict[str, float | int | dict[str, int]]: "proline_fraction": round(pro_fraction, 4), "longest_repeat_run": repeat_run, "hydrophobic_moment": mu_h, + "boman_index": boman_index(sequence), + "gravy": gravy_score(sequence), "residue_counts": dict(sorted(counts.items())), } diff --git a/src/openamp_foundry/pipeline.py b/src/openamp_foundry/pipeline.py index c720ce5a..e231d22f 100644 --- a/src/openamp_foundry/pipeline.py +++ b/src/openamp_foundry/pipeline.py @@ -11,6 +11,7 @@ from openamp_foundry.evidence.certificate import build_certificate from openamp_foundry.features.physchem import compute_features from openamp_foundry.scoring.activity import activity_likeness_score +from openamp_foundry.scoring.boman import boman_activity_score, model_disagreement from openamp_foundry.scoring.ensemble import ensemble_score, known_failure_modes, selection_reasons from openamp_foundry.scoring.novelty import novelty_score from openamp_foundry.scoring.safety import safety_score @@ -51,11 +52,14 @@ def score_candidates( safe = safety_score(features) if valid else 0.0 synth = synthesis_feasibility_score(features, valid_sequence=valid) nov, nearest = novelty_score(candidate.sequence, references) + boman_act = boman_activity_score(candidate.sequence) if valid else 0.0 raw_scores = { "activity": act, "safety": safe, "synthesis": synth, "novelty": nov, + "boman_activity": boman_act, + "disagreement": model_disagreement(act, boman_act), } raw_scores["ensemble"] = ensemble_score(raw_scores, weights) item = ScoredCandidate( diff --git a/src/openamp_foundry/reports/batch_pack.py b/src/openamp_foundry/reports/batch_pack.py index d1fa6d68..a06c5afa 100644 --- a/src/openamp_foundry/reports/batch_pack.py +++ b/src/openamp_foundry/reports/batch_pack.py @@ -242,6 +242,65 @@ def synthesis_feasibility_report(selected: list[dict[str, Any]]) -> dict[str, An } +# --------------------------------------------------------------------------- +# 5. Scorer consensus (Boman index vs. activity-likeness disagreement) +# --------------------------------------------------------------------------- + +def scorer_consensus_report(selected: list[dict[str, Any]]) -> dict[str, Any]: + """Summarise dual-scorer consensus across selected candidates. + + Disagreement = |activity_likeness − boman_activity|. + Low disagreement (< 0.20): both independent scorers agree → more robust nomination. + High disagreement (≥ 0.30): scorers diverge → extra scrutiny recommended. + """ + per_candidate = [] + for r in selected: + scores = r.get("scores", {}) + act = scores.get("activity") + boman_act = scores.get("boman_activity") + disagreement = scores.get("disagreement") + if act is None or boman_act is None: + continue + per_candidate.append({ + "candidate_id": r["candidate_id"], + "activity_likeness": act, + "boman_activity": boman_act, + "disagreement": disagreement, + "consensus_label": ( + "high_consensus" if (disagreement is not None and disagreement < 0.20) + else "uncertain" if (disagreement is not None and disagreement >= 0.30) + else "moderate" + ), + }) + + per_candidate.sort(key=lambda x: x.get("disagreement") or 1.0) + + n_high_consensus = sum(1 for c in per_candidate if c["consensus_label"] == "high_consensus") + n_uncertain = sum(1 for c in per_candidate if c["consensus_label"] == "uncertain") + n_moderate = sum(1 for c in per_candidate if c["consensus_label"] == "moderate") + + disagreements = [c["disagreement"] for c in per_candidate if c["disagreement"] is not None] + mean_disagreement = sum(disagreements) / len(disagreements) if disagreements else 0.0 + + return { + "report_type": "scorer_consensus", + "disclaimer": ( + "Model disagreement is a computational uncertainty proxy only. " + "It is NOT a biological activity or safety measure. " + "High consensus means two independent heuristics agree; " + "it does not increase the probability of lab activity. " + "The lab is the judge." + ), + "n_selected": len(selected), + "n_scored_with_boman": len(per_candidate), + "n_high_consensus_lt_0_20": n_high_consensus, + "n_moderate_0_20_to_0_30": n_moderate, + "n_uncertain_ge_0_30": n_uncertain, + "mean_disagreement": round(mean_disagreement, 4), + "candidates": per_candidate, + } + + # --------------------------------------------------------------------------- # Full batch pack # --------------------------------------------------------------------------- @@ -261,12 +320,13 @@ def generate_batch_pack( novelty = novelty_report(sel) toxicity = toxicity_hemolysis_risk_report(sel) synthesis = synthesis_feasibility_report(sel) + consensus = scorer_consensus_report(sel) ensemble_scores = [r["scores"]["ensemble"] for r in sel] mean_ensemble = sum(ensemble_scores) / len(ensemble_scores) if ensemble_scores else 0.0 return { - "batch_pack_version": "1.0", + "batch_pack_version": "1.1", "disclaimer": ( "All reports in this batch pack are based on computational heuristics. " "No biological activity has been demonstrated. " @@ -282,11 +342,15 @@ def generate_batch_pack( "mean_novelty": novelty["mean_novelty"], "mean_safety": toxicity["mean_safety_score"], "mean_synthesis": synthesis["mean_synthesis_score"], + "n_high_consensus": consensus["n_high_consensus_lt_0_20"], + "n_uncertain_disagreement": consensus["n_uncertain_ge_0_30"], + "mean_scorer_disagreement": consensus["mean_disagreement"], }, "diversity_clustering": diversity, "novelty_report": novelty, "toxicity_hemolysis_risk": toxicity, "synthesis_feasibility": synthesis, + "scorer_consensus": consensus, } @@ -300,6 +364,7 @@ def write_batch_pack_markdown(pack: dict[str, Any], path: str | Path) -> None: nov = pack["novelty_report"] tox = pack["toxicity_hemolysis_risk"] syn = pack["synthesis_feasibility"] + con = pack.get("scorer_consensus", {}) lines = [ "# OpenAMP Foundry — Phase 3 Batch Pack", @@ -320,6 +385,9 @@ def write_batch_pack_markdown(pack: dict[str, Any], path: str | Path) -> None: f"| Mean novelty | {s['mean_novelty']:.4f} |", f"| Mean safety | {s['mean_safety']:.4f} |", f"| Mean synthesis feasibility | {s['mean_synthesis']:.4f} |", + f"| High scorer consensus (disagreement <0.20) | {s.get('n_high_consensus', 'N/A')} |", + f"| Uncertain (disagreement ≥0.30) | {s.get('n_uncertain_disagreement', 'N/A')} |", + f"| Mean scorer disagreement | {s.get('mean_scorer_disagreement', 0.0):.4f} |", "", "---", "", @@ -421,6 +489,32 @@ def write_batch_pack_markdown(pack: dict[str, Any], path: str | Path) -> None: f"{c['length']} | {cys} | {pro} | {repeat} |" ) + if con: + lines += [ + "", + "---", + "", + "## 5. Scorer Consensus Report", + "", + f"> {con.get('disclaimer', '')}", + "", + f"- High consensus (disagreement <0.20): {con.get('n_high_consensus_lt_0_20', 'N/A')} candidates", + f"- Moderate disagreement (0.20–0.30): {con.get('n_moderate_0_20_to_0_30', 'N/A')} candidates", + f"- Uncertain (disagreement ≥0.30): {con.get('n_uncertain_ge_0_30', 'N/A')} candidates", + f"- Mean disagreement: {con.get('mean_disagreement', 0.0):.4f}", + "", + "Candidates sorted by disagreement (lowest = strongest dual-scorer consensus):", + "", + "| Candidate | Activity | Boman Activity | Disagreement | Consensus |", + "|-----------|----------|----------------|--------------|-----------|", + ] + for c in con.get("candidates", [])[:20]: + lines.append( + f"| {c['candidate_id']} | {c['activity_likeness']:.4f} | " + f"{c['boman_activity']:.4f} | {c.get('disagreement', 0.0):.4f} | " + f"{c['consensus_label']} |" + ) + lines += [ "", "---", diff --git a/src/openamp_foundry/scoring/boman.py b/src/openamp_foundry/scoring/boman.py new file mode 100644 index 00000000..05562020 --- /dev/null +++ b/src/openamp_foundry/scoring/boman.py @@ -0,0 +1,128 @@ +"""Boman index scorer — independent second activity predictor. + +Implements the Boman index as defined in: + Boman HG (2003). Antibacterial and antimalarial properties of peptides that + are cecropin-melittin hybrids. FEBS Letters, 544(1-3), 212-215. + and + Boman HG (2003). Peptide antibiotics and their role in innate immunity. + Annual Review of Immunology, 13, 61-92. + +The Boman index is the mean of amino acid interaction potentials per residue. +Positive values (> 1.0) indicate high protein-binding / membrane-interaction +potential, which correlates with AMP-like activity in published benchmarks. + +This is a second INDEPENDENT scoring approach that complements the +physicochemical heuristic in activity.py. Disagreement between the two +scorers is an uncertainty signal: candidates where both scorers agree are +more robust nominations than those where only one scorer favors them. + +Reference values: Boman 2003, Table 1, set 2. +All values are from the published table and are reproduced here for +transparent, reproducible scoring. No training was done. +""" +from __future__ import annotations + +# Published amino acid interaction potentials from Boman (2003), Table 1, set 2. +# Units: kcal/mol (approximate interaction energies) +# Positive → hydrophilic/charged, negative → hydrophobic +_BOMAN_POTENTIALS: dict[str, float] = { + "A": -0.495, # alanine: small, hydrophobic + "R": 2.465, # arginine: positively charged, high interaction + "N": 0.185, # asparagine: polar, uncharged + "D": 2.465, # aspartate: negatively charged, high interaction + "C": -1.000, # cysteine: hydrophobic/special + "Q": 0.185, # glutamine: polar, uncharged + "E": 2.465, # glutamate: negatively charged, high interaction + "G": 0.000, # glycine: minimal side chain + "H": -0.505, # histidine: aromatic/basic, lower than K/R + "I": -1.810, # isoleucine: hydrophobic + "L": -1.810, # leucine: hydrophobic + "K": 2.465, # lysine: positively charged, high interaction + "M": -1.305, # methionine: hydrophobic/sulfur + "F": -2.498, # phenylalanine: aromatic, hydrophobic + "P": 0.000, # proline: conformational constraint + "S": 0.255, # serine: polar, uncharged + "T": 0.370, # threonine: polar, uncharged + "W": -3.398, # tryptophan: large aromatic, most hydrophobic + "Y": -2.303, # tyrosine: aromatic, polar + "V": -1.505, # valine: hydrophobic +} + +# Published AMP threshold: Boman index > 1.0 is associated with AMP-like activity. +# Below -1.0 indicates strongly hydrophobic, non-interactive (unlikely AMP). +_AMP_THRESHOLD = 1.0 +_HYDROPHOBIC_THRESHOLD = -1.0 +# Range of Boman index across all possible sequences: +_BOMAN_MIN = -3.398 # all-W sequence (hypothetical) +_BOMAN_MAX = 2.465 # all-K/R/D/E sequence (hypothetical) + + +def boman_index(sequence: str) -> float: + """Compute the Boman index for a peptide sequence. + + Returns the mean interaction potential per residue (kcal/mol). + Positive (>1.0): high protein-binding / membrane-interaction potential. + Negative (<0): predominantly hydrophobic, less interactive. + + Non-canonical amino acids contribute 0.0 (neutral) to the mean. + """ + if not sequence: + return 0.0 + total = sum(_BOMAN_POTENTIALS.get(aa, 0.0) for aa in sequence.upper()) + return round(total / len(sequence), 4) + + +def boman_activity_score(sequence: str) -> float: + """Normalize the Boman index to [0, 1] for comparison with other scorers. + + Mapping: + Boman ≥ 1.0 → score approaching 1.0 (AMP-like zone) + Boman = 0.0 → score ≈ 0.5 + Boman ≤ -1.0 → score approaching 0.0 (non-interactive zone) + + The normalization uses a sigmoid-like mapping so that the published threshold + (Boman > 1.0) falls in the upper range of the score. + """ + bi = boman_index(sequence) + # Sigmoid-like normalization centered at 0, scaled so bi=1.0 → ≈ 0.73 + # Using tanh mapping: score = 0.5 * (1 + tanh(bi / 2)) + import math + score = 0.5 * (1.0 + math.tanh(bi / 2.0)) + return round(max(0.0, min(1.0, score)), 4) + + +def gravy_score(sequence: str) -> float: + """Compute the GRAVY (Grand Average of hYdropathicity) score. + + Uses the Kyte-Doolittle hydropathy scale (J Mol Biol 157:105-132, 1982). + Positive GRAVY → hydrophobic overall. + Negative GRAVY → hydrophilic overall. + AMPs typically fall between -0.5 and +0.5. + """ + _KD: dict[str, float] = { + "A": 1.8, "R": -4.5, "N": -3.5, "D": -3.5, "C": 2.5, + "Q": -3.5, "E": -3.5, "G": -0.4, "H": -3.2, "I": 4.5, + "L": 3.8, "K": -3.9, "M": 1.9, "F": 2.8, "P": -1.6, + "S": -0.8, "T": -0.7, "W": -0.9, "Y": -1.3, "V": 4.2, + } + if not sequence: + return 0.0 + total = sum(_KD.get(aa, 0.0) for aa in sequence.upper()) + return round(total / len(sequence), 4) + + +def model_disagreement(activity_likeness: float, boman_act: float) -> float: + """Compute normalized disagreement between two activity scores. + + Returns a value in [0, 1]: + 0.0 → both scorers agree perfectly + 1.0 → maximum disagreement (one says 0, other says 1) + + High disagreement indicates uncertainty — the candidate should receive + extra scrutiny before nomination. Low disagreement (high consensus) is + evidence of more robust computational support. + + Disagreement is NOT a biological safety or activity assessment. + It is an uncertainty proxy for the computational prediction only. + """ + return round(abs(activity_likeness - boman_act), 4) diff --git a/src/openamp_foundry/scoring/ensemble.py b/src/openamp_foundry/scoring/ensemble.py index d0b11317..d75c0d23 100644 --- a/src/openamp_foundry/scoring/ensemble.py +++ b/src/openamp_foundry/scoring/ensemble.py @@ -17,6 +17,10 @@ def selection_reasons(scores: dict[str, float]) -> list[str]: reasons.append("Likely feasible under simple synthesis constraints") if scores["novelty"] >= 0.25: reasons.append("Not an exact or near-exact duplicate of demo references") + boman_act = scores.get("boman_activity") + disagreement = scores.get("disagreement") + if boman_act is not None and boman_act >= 0.60 and (disagreement is None or disagreement <= 0.20): + reasons.append("Independent Boman index scorer agrees: high interaction potential (low disagreement)") if not reasons: reasons.append("Selected only by ensemble rank; requires skeptical review") return reasons @@ -33,4 +37,11 @@ def known_failure_modes(scores: dict[str, float]) -> list[str]: failures.append("Candidate has elevated pre-lab safety-risk proxy.") if scores["synthesis"] < 0.75: failures.append("Candidate may be harder to synthesize or handle.") + disagreement = scores.get("disagreement") + if disagreement is not None and disagreement >= 0.30: + failures.append( + f"High scorer disagreement ({disagreement:.2f}): " + "Boman index and physicochemical activity scores diverge. " + "Extra scrutiny recommended before nomination." + ) return failures diff --git a/tests/test_batch_pack.py b/tests/test_batch_pack.py index 24840ae8..9dd21ab2 100644 --- a/tests/test_batch_pack.py +++ b/tests/test_batch_pack.py @@ -17,6 +17,7 @@ diversity_clustering_report, generate_batch_pack, novelty_report, + scorer_consensus_report, synthesis_feasibility_report, toxicity_hemolysis_risk_report, write_batch_pack_markdown, @@ -31,6 +32,8 @@ def _make_candidate( safety: float = 0.9, synthesis: float = 0.8, novelty: float = 0.2, + boman_activity: float = 0.65, + disagreement: float | None = None, nearest_reference: str | None = "SEED-001", ) -> dict: features = { @@ -44,6 +47,8 @@ def _make_candidate( "aromatic_fraction": 0.1, "hydrophobic_moment": 0.3, } + if disagreement is None: + disagreement = round(abs(activity - boman_activity), 4) ensemble = 0.35 * activity + 0.30 * safety + 0.20 * synthesis + 0.15 * novelty return { "candidate_id": candidate_id, @@ -56,6 +61,8 @@ def _make_candidate( "synthesis": synthesis, "novelty": novelty, "ensemble": ensemble, + "boman_activity": boman_activity, + "disagreement": disagreement, }, "features": features, "nearest_reference": nearest_reference, @@ -241,6 +248,77 @@ def test_sequence_and_length_present(self): assert "length" in c +class TestScorerConsensusReport: + def test_report_type_field(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + assert report["report_type"] == "scorer_consensus" + + def test_n_selected_correct(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + assert report["n_selected"] == len(SELECTED_CANDIDATES) + + def test_all_candidates_included_when_boman_present(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + assert report["n_scored_with_boman"] == len(SELECTED_CANDIDATES) + + def test_empty_candidates_handled(self): + report = scorer_consensus_report([]) + assert report["n_selected"] == 0 + assert report["n_scored_with_boman"] == 0 + + def test_high_consensus_counted(self): + # activity=0.7, boman_activity=0.7 → disagreement=0.0 → high_consensus + candidates = [_make_candidate("C1", "KWKLFKKIGAVLKVL", activity=0.7, boman_activity=0.7)] + report = scorer_consensus_report(candidates) + assert report["n_high_consensus_lt_0_20"] == 1 + assert report["n_uncertain_ge_0_30"] == 0 + + def test_uncertain_counted(self): + # activity=0.9, boman_activity=0.4 → disagreement=0.5 → uncertain + candidates = [_make_candidate("C1", "KWKLFKKIGAVLKVL", activity=0.9, boman_activity=0.4)] + report = scorer_consensus_report(candidates) + assert report["n_uncertain_ge_0_30"] == 1 + assert report["n_high_consensus_lt_0_20"] == 0 + + def test_consensus_labels_correct(self): + candidates = [ + _make_candidate("C1", "KWKLFKKIGAVLKVL", activity=0.7, boman_activity=0.7), + _make_candidate("C2", "RRWQWRMKKLG", activity=0.9, boman_activity=0.4), + ] + report = scorer_consensus_report(candidates) + by_id = {c["candidate_id"]: c["consensus_label"] for c in report["candidates"]} + assert by_id["C1"] == "high_consensus" + assert by_id["C2"] == "uncertain" + + def test_sorted_by_disagreement_ascending(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + diffs = [c["disagreement"] for c in report["candidates"]] + assert diffs == sorted(diffs) + + def test_mean_disagreement_in_range(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + assert 0.0 <= report["mean_disagreement"] <= 1.0 + + def test_has_disclaimer(self): + report = scorer_consensus_report(SELECTED_CANDIDATES) + assert "disclaimer" in report + assert "lab" in report["disclaimer"].lower() + + def test_no_boman_scores_gracefully_handled(self): + # Candidates without boman_activity should be skipped + candidates = [ + { + "candidate_id": "NOBOMAN", + "sequence": "KWKLFKK", + "scores": {"activity": 0.7, "safety": 0.9, "synthesis": 0.8, "novelty": 0.2, "ensemble": 0.75}, + "features": {}, + "selected": True, + } + ] + report = scorer_consensus_report(candidates) + assert report["n_scored_with_boman"] == 0 + + class TestGenerateBatchPack: def test_all_required_top_level_keys_present(self, ranked_jsonl): pack = generate_batch_pack(ranked_jsonl) @@ -249,6 +327,7 @@ def test_all_required_top_level_keys_present(self, ranked_jsonl): assert "novelty_report" in pack assert "toxicity_hemolysis_risk" in pack assert "synthesis_feasibility" in pack + assert "scorer_consensus" in pack assert "disclaimer" in pack def test_summary_n_selected_matches_selected_rows(self, ranked_jsonl): @@ -288,6 +367,7 @@ def test_markdown_contains_required_sections(self, ranked_jsonl, tmp_path): assert "Novelty Report" in content assert "Toxicity" in content assert "Synthesis Feasibility" in content + assert "Scorer Consensus" in content def test_markdown_contains_disclaimer(self, ranked_jsonl, tmp_path): pack = generate_batch_pack(ranked_jsonl) diff --git a/tests/test_boman_scorer.py b/tests/test_boman_scorer.py new file mode 100644 index 00000000..021fd58e --- /dev/null +++ b/tests/test_boman_scorer.py @@ -0,0 +1,191 @@ +"""Tests for scoring/boman.py — Boman index and GRAVY score. + +All expected values are hand-calculated from the published Boman (2003) +potentials and Kyte-Doolittle (1982) scale to verify the implementation +is accurate and not just self-consistent. +""" +from __future__ import annotations + +import pytest + +from openamp_foundry.scoring.boman import ( + boman_activity_score, + boman_index, + gravy_score, + model_disagreement, +) + + +class TestBomanIndex: + def test_empty_sequence_returns_zero(self): + assert boman_index("") == 0.0 + + def test_single_lysine(self): + # K → 2.465 + assert boman_index("K") == 2.465 + + def test_single_tryptophan(self): + # W → -3.398 + assert boman_index("W") == -3.398 + + def test_single_glycine(self): + # G → 0.000 + assert boman_index("G") == 0.0 + + def test_two_residues_mean(self): + # K + L → (2.465 + -1.810) / 2 = 0.3275 + assert boman_index("KL") == pytest.approx(0.3275, abs=1e-3) + + def test_all_cationic_positive(self): + # K, R, D, E → 2.465 each → mean 2.465 + assert boman_index("KR") == pytest.approx(2.465, abs=1e-3) + + def test_all_hydrophobic_negative(self): + # L, I → -1.810 each → mean -1.810 + assert boman_index("LI") == pytest.approx(-1.810, abs=1e-3) + + def test_case_insensitive(self): + assert boman_index("kwk") == boman_index("KWK") + + def test_cationic_higher_than_hydrophobic_control(self): + # AMP template (4 K residues) should score higher than all-hydrophobic decoy + amp_bi = boman_index("KWKLFKKIGAVLKVL") + decoy_bi = boman_index("LLLLLLLLLLLLLLL") # all leucine → -1.810 + assert amp_bi > decoy_bi, f"AMP template ({amp_bi}) should exceed all-hydrophobic control ({decoy_bi})" + + def test_polyK_is_max_positive(self): + bi_kk = boman_index("KKKK") + bi_ww = boman_index("WWWW") + assert bi_kk > bi_ww + + def test_rounding_to_4_decimal_places(self): + result = boman_index("KL") + assert result == round(result, 4) + + def test_unknown_aa_contributes_zero(self): + # Non-canonical AAs are treated as zero contribution (neutral) + bi_x = boman_index("X") + assert bi_x == pytest.approx(0.0, abs=1e-4) + + def test_mixed_canonical_and_unknown(self): + # "KB" → K(2.465) + B(0.0) → mean 1.2325 + assert boman_index("KB") == pytest.approx(1.2325, abs=1e-4) + + +class TestBomanActivityScore: + def test_returns_between_zero_and_one(self): + for seq in ["K", "W", "KWKLFKKIGAVLKVL", "LLLLLLL", "KKKKKKKK"]: + score = boman_activity_score(seq) + assert 0.0 <= score <= 1.0, f"Score out of range for {seq}: {score}" + + def test_empty_sequence(self): + score = boman_activity_score("") + assert score == pytest.approx(0.5, abs=1e-3) + + def test_cationic_seq_scores_higher_than_hydrophobic(self): + # KKKKK >> WWWWW in Boman → higher activity score + assert boman_activity_score("KKKKK") > boman_activity_score("WWWWW") + + def test_zero_boman_maps_to_near_half(self): + # G → 0.000 → tanh(0) = 0 → 0.5 + score = boman_activity_score("GGGG") + assert score == pytest.approx(0.5, abs=1e-2) + + def test_monotone_with_boman_index(self): + seqs = ["WWWW", "LLLL", "GGGG", "SSSS", "KKKK"] + indices = [boman_index(s) for s in seqs] + activities = [boman_activity_score(s) for s in seqs] + assert sorted(indices) == sorted(indices), "Sanity: indices sortable" + sorted_pairs = sorted(zip(indices, activities)) + for (_, a1), (_, a2) in zip(sorted_pairs, sorted_pairs[1:]): + assert a1 <= a2 + 1e-6, "boman_activity_score should increase with boman_index" + + def test_rounding(self): + score = boman_activity_score("KWK") + assert score == round(score, 4) + + +class TestGravyScore: + def test_empty_returns_zero(self): + assert gravy_score("") == 0.0 + + def test_single_isoleucine(self): + # I → 4.5 + assert gravy_score("I") == pytest.approx(4.5, abs=1e-3) + + def test_single_arginine(self): + # R → -4.5 + assert gravy_score("R") == pytest.approx(-4.5, abs=1e-3) + + def test_hydrophobic_sequence_positive(self): + # LLLLL → 3.8 mean + assert gravy_score("LLLLL") == pytest.approx(3.8, abs=1e-3) + + def test_cationic_sequence_negative(self): + # KKKKK → -3.9 mean + assert gravy_score("KKKKK") == pytest.approx(-3.9, abs=1e-3) + + def test_case_insensitive(self): + assert gravy_score("kwk") == gravy_score("KWK") + + def test_mixed_near_zero(self): + # G → -0.4, balanced hydrophobic + hydrophilic should be near-zero + gravy = gravy_score("KLII") + # K=-3.9, L=3.8, I=4.5, I=4.5 → mean = (−3.9+3.8+4.5+4.5)/4 = 2.225 + assert gravy == pytest.approx(2.225, abs=1e-3) + + def test_rounding_to_4_places(self): + result = gravy_score("KWL") + assert result == round(result, 4) + + +class TestModelDisagreement: + def test_identical_scores_zero_disagreement(self): + assert model_disagreement(0.7, 0.7) == 0.0 + + def test_maximum_disagreement(self): + assert model_disagreement(0.0, 1.0) == pytest.approx(1.0, abs=1e-4) + assert model_disagreement(1.0, 0.0) == pytest.approx(1.0, abs=1e-4) + + def test_symmetric(self): + assert model_disagreement(0.3, 0.7) == model_disagreement(0.7, 0.3) + + def test_small_difference(self): + assert model_disagreement(0.60, 0.65) == pytest.approx(0.05, abs=1e-4) + + def test_output_in_range(self): + for a, b in [(0.0, 0.5), (0.9, 0.4), (0.3, 0.3), (1.0, 0.0)]: + d = model_disagreement(a, b) + assert 0.0 <= d <= 1.0 + + def test_rounding_to_4_places(self): + result = model_disagreement(0.333, 0.667) + assert result == round(result, 4) + + +class TestPipelineIntegration: + """Verify Boman values appear in compute_features output.""" + + def test_boman_index_in_features(self): + from openamp_foundry.features.physchem import compute_features + features = compute_features("KWKLFKKIGAVLKVL") + assert "boman_index" in features + assert isinstance(features["boman_index"], float) + + def test_gravy_in_features(self): + from openamp_foundry.features.physchem import compute_features + features = compute_features("KWKLFKKIGAVLKVL") + assert "gravy" in features + assert isinstance(features["gravy"], float) + + def test_boman_index_feature_value_matches_scorer(self): + from openamp_foundry.features.physchem import compute_features + seq = "KRLFKKIGSALKFL" + features = compute_features(seq) + assert features["boman_index"] == pytest.approx(boman_index(seq), abs=1e-4) + + def test_gravy_feature_value_matches_scorer(self): + from openamp_foundry.features.physchem import compute_features + seq = "KRLFKKIGSALKFL" + features = compute_features(seq) + assert features["gravy"] == pytest.approx(gravy_score(seq), abs=1e-4) diff --git a/tests/test_lab_results.py b/tests/test_lab_results.py new file mode 100644 index 00000000..eb8f95d3 --- /dev/null +++ b/tests/test_lab_results.py @@ -0,0 +1,174 @@ +"""Tests for lab_results.py — active-learning loop ingestion. + +Verifies that: + - Valid lab result JSON loads and validates against schema + - Invalid lab results are rejected with useful errors + - summarise_lab_results produces correct aggregates + - candidate_result_map groups correctly +""" +from __future__ import annotations + +import json + +import pytest + +from openamp_foundry.data.lab_results import ( + candidate_result_map, + load_lab_result, + summarise_lab_results, +) + + +def _valid_result(**overrides) -> dict: + base = { + "result_id": "RES-001", + "candidate_id": "SEED-005_VAR_049", + "assay_type": "MIC", + "organism_or_cell_line": "E. coli ATCC 25922", + "result_value": 4.0, + "result_unit": "µg/mL", + "result_qualitative": "active", + "positive_control_passed": True, + "negative_control_passed": True, + "positive_control_id": "ciprofloxacin 0.25 µg/mL", + "negative_control_id": "PBS", + "assay_date": "2026-07-01", + "replicate_count": 3, + "performed_by_lab": "University Test Lab", + "raw_data_sha256": None, + "computational_candidate_certificate_hash": "abc123def456", + "notes": None, + "disclaimer": ( + "This is an experimental result on a computationally nominated candidate. " + "It does not constitute a drug or clinical claim." + ), + } + base.update(overrides) + return base + + +@pytest.fixture +def valid_result_file(tmp_path): + result = _valid_result() + path = tmp_path / "RES-001.json" + path.write_text(json.dumps(result)) + return path + + +@pytest.fixture +def results_dir(tmp_path): + for i in range(3): + result = _valid_result( + result_id=f"RES-{i:03d}", + candidate_id=f"CAND-{i:03d}", + assay_date=f"2026-07-0{i+1}", + ) + (tmp_path / f"RES-{i:03d}.json").write_text(json.dumps(result)) + return tmp_path + + +class TestLoadLabResult: + def test_valid_result_loads(self, valid_result_file): + result = load_lab_result(valid_result_file) + assert result["result_id"] == "RES-001" + assert result["candidate_id"] == "SEED-005_VAR_049" + assert result["assay_type"] == "MIC" + + def test_invalid_assay_type_rejected(self, tmp_path): + result = _valid_result(assay_type="unknown_assay") + path = tmp_path / "bad.json" + path.write_text(json.dumps(result)) + with pytest.raises(Exception): + load_lab_result(path) + + def test_missing_required_field_rejected(self, tmp_path): + result = _valid_result() + del result["candidate_id"] + path = tmp_path / "missing.json" + path.write_text(json.dumps(result)) + with pytest.raises(Exception): + load_lab_result(path) + + def test_negative_replicate_count_rejected(self, tmp_path): + result = _valid_result(replicate_count=0) + path = tmp_path / "bad_reps.json" + path.write_text(json.dumps(result)) + with pytest.raises(Exception): + load_lab_result(path) + + def test_null_result_value_allowed(self, tmp_path): + result = _valid_result(result_value=None, result_qualitative="inconclusive") + path = tmp_path / "null_val.json" + path.write_text(json.dumps(result)) + loaded = load_lab_result(path) + assert loaded["result_value"] is None + + def test_file_not_found_raises(self, tmp_path): + with pytest.raises(FileNotFoundError): + load_lab_result(tmp_path / "nonexistent.json") + + +class TestSummariseLabResults: + def test_empty_results_handled(self): + summary = summarise_lab_results([]) + assert summary["n_results"] == 0 + assert "disclaimer" in summary + + def test_n_results_correct(self): + results = [_valid_result(result_id=f"R{i}") for i in range(5)] + summary = summarise_lab_results(results) + assert summary["n_results"] == 5 + + def test_by_assay_type_counts(self): + results = [ + _valid_result(result_id="R1", assay_type="MIC"), + _valid_result(result_id="R2", assay_type="MIC"), + _valid_result(result_id="R3", assay_type="hemolysis_RBC"), + ] + summary = summarise_lab_results(results) + assert summary["by_assay_type"]["MIC"] == 2 + assert summary["by_assay_type"]["hemolysis_RBC"] == 1 + + def test_by_qualitative_counts(self): + results = [ + _valid_result(result_id="R1", result_qualitative="active"), + _valid_result(result_id="R2", result_qualitative="inactive"), + _valid_result(result_id="R3", result_qualitative="active"), + ] + summary = summarise_lab_results(results) + assert summary["by_qualitative_result"]["active"] == 2 + assert summary["by_qualitative_result"]["inactive"] == 1 + + def test_valid_controls_counted(self): + results = [ + _valid_result(result_id="R1", positive_control_passed=True, negative_control_passed=True), + _valid_result(result_id="R2", positive_control_passed=True, negative_control_passed=False), + ] + summary = summarise_lab_results(results) + assert summary["n_valid_controls"] == 1 + + def test_disclaimer_present(self): + results = [_valid_result()] + summary = summarise_lab_results(results) + assert "disclaimer" in summary + assert len(summary["disclaimer"]) > 20 + + +class TestCandidateResultMap: + def test_groups_by_candidate_id(self): + results = [ + _valid_result(result_id="R1", candidate_id="CAND-001"), + _valid_result(result_id="R2", candidate_id="CAND-001"), + _valid_result(result_id="R3", candidate_id="CAND-002"), + ] + mapping = candidate_result_map(results) + assert len(mapping["CAND-001"]) == 2 + assert len(mapping["CAND-002"]) == 1 + + def test_unknown_candidate_not_in_map(self): + results = [_valid_result(candidate_id="CAND-001")] + mapping = candidate_result_map(results) + assert "CAND-999" not in mapping + + def test_empty_results_gives_empty_map(self): + assert candidate_result_map([]) == {}