Skip to content

feat: amphipathicity feature, baseline benchmark, recall@k evaluation - #2

Closed
cschanhniem wants to merge 2 commits into
mainfrom
feat/amphipathicity-benchmark
Closed

cschanhniem wants to merge 2 commits into
mainfrom
feat/amphipathicity-benchmark

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

  • Hydrophobic moment (amphipathicity): Added hydrophobic_moment() to physchem.py using the Eisenberg (1984) consensus hydrophobicity scale at 100°/residue helical projection — a literature-established correlate of AMP membrane disruption activity
  • Activity score update: activity_likeness_score() now incorporates amphipathicity (15% weight) alongside length, charge density, hydrophobic fraction, and aromatic content; weights normalised to preserve [0,1] range
  • Benchmark evaluation: benchmark/evaluate.py now provides recall_at_k(), random_recall_at_k(), enrichment_factor(), and benchmark_summary() — all outputs include an honest disclaimer that results do not prove biological efficacy
  • bench baseline CLI: openamp-foundry bench baseline --candidates ... --positives ... --references ... reports pipeline vs random recall at configurable k cutoffs; also make bench-baseline

Honesty note

The demo bench baseline correctly reports EF=0 because the demo positives CSV uses REF-* IDs while the candidates CSV uses AMPF-* IDs — no ID overlap, so recall is 0. This is honest behaviour: the benchmark infrastructure is in place and correctly measures zero recall when positive IDs don't appear in the candidate pool. A real benchmark requires a labelled dataset where positive IDs match candidates.

Test plan

  • make test — 63 tests pass
  • make demo — pipeline runs end-to-end
  • make bench-leakage — leakage detection works
  • make bench-baseline — baseline command runs with honest EF=0 result on demo mismatch
  • 20 new tests: hydrophobic moment edge cases, recall@k boundary conditions, enrichment factor, benchmark summary structure and disclaimer

…pand tests

- Enforce min_length/max_length from config in score_candidates()
- Apply min_novelty and max_safety_risk selection thresholds from config
- Add `valid` field to ScoredCandidate; mark invalid sequences with failure reasons
- Add `selected` boolean to JSONL output rows
- Generate run_manifest.json with run_id, input SHA-256 hashes, config hash, pipeline version
- Add `openamp-foundry bench leakage` CLI subcommand and `make bench-leakage` target
- Improve report disclaimer to explicitly state no antimicrobial activity demonstrated
- Expand test suite from 6 to 37 tests covering: pipeline filters, selection thresholds,
  run manifest generation, benchmark leakage detection, splits, evaluation, CLI integration
…tion

- Add hydrophobic_moment() to physchem.py using Eisenberg (1984) consensus scale
  at 100°/residue helical projection; literature-cited correlate of AMP activity
- Expand activity_likeness_score() to incorporate amphipathicity (15% weight)
  with reduced charge/hydrophobicity weights to keep total at 1.0
- Add recall_at_k(), random_recall_at_k(), enrichment_factor(), benchmark_summary()
  to benchmark/evaluate.py with honest disclaimer in every output
- Add 'openamp-foundry bench baseline' CLI subcommand for pipeline vs random recall
- Add 'make bench-baseline' Makefile target
- 20 new tests: amphipathicity feature, hydrophobic moment edge cases,
  recall@k boundary conditions, enrichment factor, benchmark summary structure
@cschanhniem

Copy link
Copy Markdown
Collaborator Author

Superseded by PR #11 (feat/integrate-all-phases), which merges all Phase 2 + Phase 3 work into a single consolidation PR with 251 tests passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant