Repository navigation
feat: amphipathicity feature, baseline benchmark, recall@k evaluation - #2
Closed
cschanhniem wants to merge 2 commits into
Closed
cschanhniem wants to merge 2 commits into
cschanhniem wants to merge 2 commits into
Conversation
…pand tests - Enforce min_length/max_length from config in score_candidates() - Apply min_novelty and max_safety_risk selection thresholds from config - Add `valid` field to ScoredCandidate; mark invalid sequences with failure reasons - Add `selected` boolean to JSONL output rows - Generate run_manifest.json with run_id, input SHA-256 hashes, config hash, pipeline version - Add `openamp-foundry bench leakage` CLI subcommand and `make bench-leakage` target - Improve report disclaimer to explicitly state no antimicrobial activity demonstrated - Expand test suite from 6 to 37 tests covering: pipeline filters, selection thresholds, run manifest generation, benchmark leakage detection, splits, evaluation, CLI integration
…tion - Add hydrophobic_moment() to physchem.py using Eisenberg (1984) consensus scale at 100°/residue helical projection; literature-cited correlate of AMP activity - Expand activity_likeness_score() to incorporate amphipathicity (15% weight) with reduced charge/hydrophobicity weights to keep total at 1.0 - Add recall_at_k(), random_recall_at_k(), enrichment_factor(), benchmark_summary() to benchmark/evaluate.py with honest disclaimer in every output - Add 'openamp-foundry bench baseline' CLI subcommand for pipeline vs random recall - Add 'make bench-baseline' Makefile target - 20 new tests: amphipathicity feature, hydrophobic moment edge cases, recall@k boundary conditions, enrichment factor, benchmark summary structure
Merged
7 tasks done
Collaborator
Author
|
Superseded by PR #11 (feat/integrate-all-phases), which merges all Phase 2 + Phase 3 work into a single consolidation PR with 251 tests passing. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
hydrophobic_moment()tophyschem.pyusing the Eisenberg (1984) consensus hydrophobicity scale at 100°/residue helical projection — a literature-established correlate of AMP membrane disruption activityactivity_likeness_score()now incorporates amphipathicity (15% weight) alongside length, charge density, hydrophobic fraction, and aromatic content; weights normalised to preserve [0,1] rangebenchmark/evaluate.pynow providesrecall_at_k(),random_recall_at_k(),enrichment_factor(), andbenchmark_summary()— all outputs include an honest disclaimer that results do not prove biological efficacybench baselineCLI:openamp-foundry bench baseline --candidates ... --positives ... --references ...reports pipeline vs random recall at configurable k cutoffs; alsomake bench-baselineHonesty note
The demo
bench baselinecorrectly reports EF=0 because the demo positives CSV uses REF-* IDs while the candidates CSV uses AMPF-* IDs — no ID overlap, so recall is 0. This is honest behaviour: the benchmark infrastructure is in place and correctly measures zero recall when positive IDs don't appear in the candidate pool. A real benchmark requires a labelled dataset where positive IDs match candidates.Test plan
make test— 63 tests passmake demo— pipeline runs end-to-endmake bench-leakage— leakage detection worksmake bench-baseline— baseline command runs with honest EF=0 result on demo mismatch