Skip to content

feat: enforce config filters, add run manifest, bench leakage CLI, 37 tests - #1

Merged
cschanhniem merged 1 commit into
mainfrom
feat/pipeline-filters-manifest-benchmarks
Jun 27, 2026
Merged

cschanhniem merged 1 commit into
mainfrom
feat/pipeline-filters-manifest-benchmarks

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

  • Config filters enforced: min_length/max_length from configs/pipeline.yaml now filter sequences in the pipeline; invalid sequences get valid=False and zero activity score with failure modes logged
  • Selection thresholds applied: min_novelty and max_safety_risk from config now gate which candidates are selected for evidence certificates (e.g. near-duplicate reference copies are excluded from selection)
  • Run manifest generated: Every pipeline run now auto-generates run_manifest.json with run_id, pipeline_version, config_hash, per-input SHA-256 hashes, output paths, and timestamp — satisfying the reproducibility requirement
  • bench leakage CLI command: openamp-foundry bench leakage --candidates ... --references ... detects near-duplicate contamination and warns when benchmark results may be inflated; also available as make bench-leakage
  • selected field in JSONL output: Each row now includes a boolean showing whether the candidate passed all filters and was selected for evidence certificates
  • Improved report disclaimer: Clearly states these are baseline heuristics, not validated biological predictors, and no antimicrobial activity has been demonstrated
  • Test expansion: 6 → 37 tests covering pipeline filters, selection thresholds, run manifest, benchmark leakage detection, splits, evaluation utilities, and CLI integration

Test plan

  • make test — all 37 tests pass
  • make demo — demo pipeline runs end-to-end, produces ranked JSONL, report, evidence certificates, and run manifest
  • make bench-leakage — detects 3 near-duplicate candidates in demo dataset with warning
  • All evidence certificates validate against schemas/candidate.schema.json
  • Near-duplicate candidates (novelty=0.0) excluded from selection per min_novelty: 0.20 config
  • Report contains explicit computational-only disclaimer

Safety

No changes to safety policy, allowed amino acids, or optimization objectives. All new code adds filtering (removes unsafe/low-quality candidates) rather than relaxing constraints.

…pand tests

- Enforce min_length/max_length from config in score_candidates()
- Apply min_novelty and max_safety_risk selection thresholds from config
- Add `valid` field to ScoredCandidate; mark invalid sequences with failure reasons
- Add `selected` boolean to JSONL output rows
- Generate run_manifest.json with run_id, input SHA-256 hashes, config hash, pipeline version
- Add `openamp-foundry bench leakage` CLI subcommand and `make bench-leakage` target
- Improve report disclaimer to explicitly state no antimicrobial activity demonstrated
- Expand test suite from 6 to 37 tests covering: pipeline filters, selection thresholds,
  run manifest generation, benchmark leakage detection, splits, evaluation, CLI integration
@cschanhniem
cschanhniem merged commit 3c09a4d into main Jun 27, 2026
2 of 3 checks passed
cschanhniem added a commit that referenced this pull request Jun 28, 2026
* feat: elastase resistance + aggregation propensity scoring

Two new computational features to reduce wet-lab failure rate:

1. Elastase resistance (GAP #3 from audit):
   - physchem.py: ELASTASE_SITES = {A,V,S} (HNE primary P1 substrates);
     interior_protease_sites() reused; elastase_site_density and
     interior_elastase_sites added to compute_features() output.
   - stability.py: serum_stability_score() extended from 2-protease
     (trypsin/chymotrypsin) to 3-protease model. Weighted sum with
     trypsin:2 > chymotrypsin:1 > elastase:0.5. Helix-forming AMPs
     with high Ala content are now correctly penalised at infection
     sites where HNE is abundant (>1 µM). Denominator 3.5 = sum of
     weights; backward-compatible (missing elastase key → 0.0).
   - Literature: Bieth (1986); Doherty et al. (1991 Biochemistry).
   - 12 new tests in test_elastase_stability.py.

2. Aggregation propensity (GAP #1 from audit):
   - physchem.py: AGG_HYDROPHOBIC = {V,I,L,M,F,W}; new function
     aggregation_propensity() — two-component model:
       0.7 × interior_run_risk (run ≥ 4 → ramp 0→1 over 5 residues)
       + 0.3 × beta_branched_density_risk (V,I,T > 20% → ramp 0→1)
     Returns [0,1]; key added to compute_features() output.
   - synthesis.py: synthesis_feasibility_score() now includes:
       if agg > 0: score -= min(agg * 0.25, 0.20)
     Max penalty = 0.20 (capped). Backward-compat (missing key → 0).
   - Literature: Quittot et al. (2017 Protein Sci);
                 Wurth et al. (2006 J Mol Biol).
   - 22 new tests in test_aggregation_propensity.py.

Impact: AUROC=0.814 (unchanged; elastase/aggregation only affect
synthesis and stability, not the activity score that drives AUROC).
Total test count: 1122.

* fix: address code review HIGH issues for PR #49

- physchem.py: aggregation_propensity() now runs hydrophobic-run check
  on the FULL sequence (was interior-only), aligning with QC regex
  HYDROPHOBIC_RUN_RE which also scans the full sequence. Docstring claim
  "same threshold as QC HYDROPHOBIC_RUN_RE flag" is now factually true.
  Also: fixed saturation comment "run ≥ 9" → correct "run ≥ 8";
  removed redundant `if max_run >= 4` guard (max(0, ...) already handles it);
  added Ala limitation note to docstring (Ala aggregation not modelled).
- test_aggregation_propensity.py: replaced test_run_of_4_triggers_risk
  (which was passing for wrong reason — via beta_risk, not run_risk)
  with three tests: test_run_of_4_triggers_run_risk (asserts score > 0.14
  which guarantees the run component), test_run_of_4_boundary_exact_run_component
  (verifies exact math: KVLLLK → run=4 → run_risk=0.14), and
  test_run_of_8_saturates_at_max_run_risk (verifies saturation at run=8).
- Makefile: update test count to 1124.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant