Skip to content

Commit 431eec1

Browse files
cschanhniemOpenCode
andauthored
feat: Phase W W2 charge-matched challenge schema (CMC-) -- documents pipeline AUROC vs charge-only baseline; VALID_CMC_VERDICTS (4), MEANINGFUL_GAP_THRESHOLD=0.05, auroc_gap auto-computed; blocks performance claims when charge explains the gap; 60 tests (#1004)
Co-authored-by: OpenCode <opencode@example.com>
1 parent ce48497 commit 431eec1

3 files changed

Lines changed: 486 additions & 1 deletion

File tree

‎docs/research/NEXT_100_PR_MAP.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -277,7 +277,7 @@ Make it machine-verifiable that the pipeline produces novel candidates that beat
277277
| ID | Task | Why it matters | Priority |
278278
|----|------|----------------|----------|
279279
| W1 | Add novelty challenge harness schema (NCH-) (complete). — src/openamp_foundry/evidence/novelty_challenge_harness.py: VALID_NCH_VERDICTS (4: novel_batch/mixed_novelty/near_neighbor_dominated/challenge_not_run), VALID_REFERENCE_DATABASES (6), NEAR_NEIGHBOR_IDENTITY_THRESHOLD=0.80, NOVEL_BATCH_CEILING=0.20, NEAR_NEIGHBOR_DOMINATED_FLOOR=0.60; NCHCandidateResult helper; build() auto-computes is_near_neighbor from identity vs threshold, fraction, verdict; dry_lab_only=True enforced; 63 tests in tests/evidence/test_novelty_challenge_harness.py. | Batch-level novelty challenge: documents the fraction of top candidates with ≥80% sequence identity to a known AMP in a reference database (APD3/DRAMP/etc.); blocks novel_batch claim when near-neighbor fraction exceeds 20%; prevents pipeline from advancing near-copies of known AMPs under a novelty label. | C |
280-
| W2 | Add charge-matched challenge schema (CMC-). | Formally documents the charge-matched challenge: compares pipeline AUROC vs a charge-only baseline on the same candidate set; verdict controlled vocabulary (gap_meaningful/gap_marginal/gap_absent/not_run); blocks performance claims when the charge-only baseline explains the gap. | C |
280+
| W2 | Add charge-matched challenge schema (CMC-) (complete). — src/openamp_foundry/evidence/charge_matched_challenge.py: VALID_CMC_VERDICTS (4: gap_meaningful/gap_marginal/gap_absent/challenge_not_run), VALID_CHARGE_BASELINE_METHODS (4), MEANINGFUL_GAP_THRESHOLD=0.05, MARGINAL_GAP_LOWER=0.02; auroc_gap auto-computed; verdict auto-derived; dry_lab_only=True enforced; 60 tests in tests/evidence/test_charge_matched_challenge.py. | Formally documents the charge-matched challenge: compares pipeline AUROC vs a charge-only baseline on the same candidate set; verdict controlled vocabulary (gap_meaningful/gap_marginal/gap_absent/not_run); blocks performance claims when the charge-only baseline explains the gap. | C |
281281
| W3 | Add similarity challenge harness schema (SCH-). | Documents whether pipeline-selected candidates are systematically more similar to known AMPs than random selection from the sequence space; flags selection bias from similarity clustering; prevents "novel panel" claim when selection is proximity-driven. | C |
282282
| W4 | Add benchmark challenge registry schema (BCR-). | Machine-readable registry of which benchmark challenges (NCH/CMC/SCH) have been run and passed for a given pipeline version; aggregates challenge verdicts; overall hardness grade (A: all passed, B: most passed, C: some passed, D: none passed). | C |
283283
| W5 | Add Phase W benchmark gate (WBG-). | Top-level gate asserting NCH + CMC + SCH + BCR all present; overall verdict: hardened/partially_hardened/not_hardened; closes Phase W; no batch-level performance claim is credible without passing this gate. | C |
Lines changed: 154 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,154 @@
1+
"""CMC- charge-matched challenge schema.
2+
3+
Documents the charge-matched challenge run for a batch: compares pipeline
4+
AUROC versus a charge-only baseline on the same candidate set with the same
5+
charge distribution. A meaningful gap is required to claim the pipeline adds
6+
value beyond simply preferring cationic sequences.
7+
"""
8+
9+
from __future__ import annotations
10+
11+
from dataclasses import dataclass
12+
13+
VALID_CMC_VERDICTS: frozenset[str] = frozenset({
14+
"gap_meaningful",
15+
"gap_marginal",
16+
"gap_absent",
17+
"challenge_not_run",
18+
})
19+
20+
VALID_CHARGE_BASELINE_METHODS: frozenset[str] = frozenset({
21+
"charge_only_rank",
22+
"charge_length_rank",
23+
"charge_hydrophobicity_rank",
24+
"logistic_charge_only",
25+
})
26+
27+
MEANINGFUL_GAP_THRESHOLD: float = 0.05
28+
MARGINAL_GAP_LOWER: float = 0.02
29+
30+
MIN_AUROC: float = 0.0
31+
MAX_AUROC: float = 1.0
32+
33+
34+
@dataclass
35+
class ChargeMatchedChallenge:
36+
cmc_id: str
37+
batch_id: str
38+
pipeline_version: str
39+
baseline_method: str
40+
pipeline_auroc: float
41+
baseline_auroc: float
42+
auroc_gap: float
43+
n_candidates: int
44+
mean_charge_pipeline: float
45+
mean_charge_baseline: float
46+
charge_distribution_matched: bool
47+
cmc_verdict: str
48+
dry_lab_only: bool
49+
limitations: list[str]
50+
created_at: str
51+
52+
53+
def validate_charge_matched_challenge(cmc: ChargeMatchedChallenge) -> None:
54+
if not cmc.cmc_id.startswith("CMC-"):
55+
raise ValueError(f"cmc_id must start with 'CMC-': {cmc.cmc_id!r}")
56+
if not cmc.batch_id:
57+
raise ValueError("batch_id must be non-empty")
58+
if not cmc.pipeline_version:
59+
raise ValueError("pipeline_version must be non-empty")
60+
if cmc.baseline_method not in VALID_CHARGE_BASELINE_METHODS:
61+
raise ValueError(
62+
f"baseline_method {cmc.baseline_method!r} not in VALID_CHARGE_BASELINE_METHODS"
63+
)
64+
if not (MIN_AUROC <= cmc.pipeline_auroc <= MAX_AUROC):
65+
raise ValueError(
66+
f"pipeline_auroc must be in [0, 1]: {cmc.pipeline_auroc}"
67+
)
68+
if not (MIN_AUROC <= cmc.baseline_auroc <= MAX_AUROC):
69+
raise ValueError(
70+
f"baseline_auroc must be in [0, 1]: {cmc.baseline_auroc}"
71+
)
72+
expected_gap = round(cmc.pipeline_auroc - cmc.baseline_auroc, 6)
73+
if abs(cmc.auroc_gap - expected_gap) > 1e-4:
74+
raise ValueError(
75+
f"auroc_gap {cmc.auroc_gap} does not match computed "
76+
f"{expected_gap} (pipeline_auroc - baseline_auroc)"
77+
)
78+
if cmc.n_candidates < 0:
79+
raise ValueError("n_candidates must be non-negative")
80+
if cmc.cmc_verdict not in VALID_CMC_VERDICTS:
81+
raise ValueError(
82+
f"cmc_verdict {cmc.cmc_verdict!r} not in VALID_CMC_VERDICTS"
83+
)
84+
if not cmc.dry_lab_only:
85+
raise ValueError("dry_lab_only must be True")
86+
if not cmc.limitations:
87+
raise ValueError("limitations must be non-empty")
88+
if not cmc.created_at:
89+
raise ValueError("created_at must be non-empty")
90+
91+
92+
def _compute_verdict(n_candidates: int, auroc_gap: float) -> str:
93+
if n_candidates == 0:
94+
return "challenge_not_run"
95+
if auroc_gap >= MEANINGFUL_GAP_THRESHOLD:
96+
return "gap_meaningful"
97+
if auroc_gap >= MARGINAL_GAP_LOWER:
98+
return "gap_marginal"
99+
return "gap_absent"
100+
101+
102+
def build_charge_matched_challenge(
103+
*,
104+
cmc_id: str,
105+
batch_id: str,
106+
pipeline_version: str,
107+
baseline_method: str,
108+
pipeline_auroc: float,
109+
baseline_auroc: float,
110+
n_candidates: int,
111+
mean_charge_pipeline: float,
112+
mean_charge_baseline: float,
113+
charge_distribution_matched: bool,
114+
limitations: list[str],
115+
created_at: str,
116+
) -> ChargeMatchedChallenge:
117+
auroc_gap = round(pipeline_auroc - baseline_auroc, 6)
118+
verdict = _compute_verdict(n_candidates, auroc_gap)
119+
cmc = ChargeMatchedChallenge(
120+
cmc_id=cmc_id,
121+
batch_id=batch_id,
122+
pipeline_version=pipeline_version,
123+
baseline_method=baseline_method,
124+
pipeline_auroc=pipeline_auroc,
125+
baseline_auroc=baseline_auroc,
126+
auroc_gap=auroc_gap,
127+
n_candidates=n_candidates,
128+
mean_charge_pipeline=mean_charge_pipeline,
129+
mean_charge_baseline=mean_charge_baseline,
130+
charge_distribution_matched=charge_distribution_matched,
131+
cmc_verdict=verdict,
132+
dry_lab_only=True,
133+
limitations=limitations,
134+
created_at=created_at,
135+
)
136+
validate_charge_matched_challenge(cmc)
137+
return cmc
138+
139+
140+
def format_charge_matched_challenge(cmc: ChargeMatchedChallenge) -> str:
141+
lines = [
142+
f"Charge-Matched Challenge — {cmc.cmc_id}",
143+
f"Batch: {cmc.batch_id} | Pipeline: {cmc.pipeline_version}",
144+
f"Baseline method: {cmc.baseline_method}",
145+
f"Verdict: {cmc.cmc_verdict}",
146+
f"Pipeline AUROC: {cmc.pipeline_auroc:.4f} | Baseline AUROC: {cmc.baseline_auroc:.4f} | Gap: {cmc.auroc_gap:+.4f}",
147+
f"N candidates: {cmc.n_candidates}",
148+
f"Mean charge — pipeline: {cmc.mean_charge_pipeline:.2f} | baseline: {cmc.mean_charge_baseline:.2f}",
149+
f"Charge distribution matched: {cmc.charge_distribution_matched}",
150+
f"Created: {cmc.created_at}",
151+
f"Limitations: {'; '.join(cmc.limitations)}",
152+
f"dry_lab_only: {cmc.dry_lab_only}",
153+
]
154+
return "\n".join(lines)

0 commit comments

Comments
 (0)