Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/research/NEXT_100_PR_MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -279,5 +279,5 @@ Make it machine-verifiable that the pipeline produces novel candidates that beat
| W1 | Add novelty challenge harness schema (NCH-) (complete). — src/openamp_foundry/evidence/novelty_challenge_harness.py: VALID_NCH_VERDICTS (4: novel_batch/mixed_novelty/near_neighbor_dominated/challenge_not_run), VALID_REFERENCE_DATABASES (6), NEAR_NEIGHBOR_IDENTITY_THRESHOLD=0.80, NOVEL_BATCH_CEILING=0.20, NEAR_NEIGHBOR_DOMINATED_FLOOR=0.60; NCHCandidateResult helper; build() auto-computes is_near_neighbor from identity vs threshold, fraction, verdict; dry_lab_only=True enforced; 63 tests in tests/evidence/test_novelty_challenge_harness.py. | Batch-level novelty challenge: documents the fraction of top candidates with ≥80% sequence identity to a known AMP in a reference database (APD3/DRAMP/etc.); blocks novel_batch claim when near-neighbor fraction exceeds 20%; prevents pipeline from advancing near-copies of known AMPs under a novelty label. | C |
| W2 | Add charge-matched challenge schema (CMC-) (complete). — src/openamp_foundry/evidence/charge_matched_challenge.py: VALID_CMC_VERDICTS (4: gap_meaningful/gap_marginal/gap_absent/challenge_not_run), VALID_CHARGE_BASELINE_METHODS (4), MEANINGFUL_GAP_THRESHOLD=0.05, MARGINAL_GAP_LOWER=0.02; auroc_gap auto-computed; verdict auto-derived; dry_lab_only=True enforced; 60 tests in tests/evidence/test_charge_matched_challenge.py. | Formally documents the charge-matched challenge: compares pipeline AUROC vs a charge-only baseline on the same candidate set; verdict controlled vocabulary (gap_meaningful/gap_marginal/gap_absent/not_run); blocks performance claims when the charge-only baseline explains the gap. | C |
| W3 | Add similarity challenge harness schema (SCH-) (complete). — src/openamp_foundry/evidence/similarity_challenge_harness.py: VALID_SCH_VERDICTS (4: selection_adds_value/marginal_improvement/proximity_driven/challenge_not_run), VALID_SIMILARITY_METRICS (4), SELECTION_VALUE_GAP_THRESHOLD=0.10, MARGINAL_IMPROVEMENT_LOWER=0.03; SimilarityGroupStats helper; similarity_gap auto-computed; dry_lab_only=True enforced; 60 tests in tests/evidence/test_similarity_challenge_harness.py. | Documents whether pipeline-selected candidates are systematically more similar to known AMPs than random selection from the sequence space; flags selection bias from similarity clustering; prevents "novel panel" claim when selection is proximity-driven. | C |
| W4 | Add benchmark challenge registry schema (BCR-). | Machine-readable registry of which benchmark challenges (NCH/CMC/SCH) have been run and passed for a given pipeline version; aggregates challenge verdicts; overall hardness grade (A: all passed, B: most passed, C: some passed, D: none passed). | C |
| W4 | Add benchmark challenge registry schema (BCR-) (complete). — src/openamp_foundry/evidence/benchmark_challenge_registry.py: REQUIRED_CHALLENGE_TYPES (3: NCH/CMC/SCH), VALID_CHALLENGE_VERDICTS (4: pass/marginal/fail/not_run), VALID_BCR_HARDNESS_GRADES (A-D); ChallengeEntry helper; verdict mapping (novel_batch→pass, mixed_novelty→marginal, gap_meaningful→pass, etc.); grade A=all pass, B=all pass+marginal, C=some pass/marginal, D=all fail/not_run; dry_lab_only=True; 64 tests in tests/evidence/test_benchmark_challenge_registry.py. | Machine-readable registry of which benchmark challenges (NCH/CMC/SCH) have been run and passed for a given pipeline version; aggregates challenge verdicts; overall hardness grade (A: all passed, B: most passed, C: some passed, D: none passed). | C |
| W5 | Add Phase W benchmark gate (WBG-). | Top-level gate asserting NCH + CMC + SCH + BCR all present; overall verdict: hardened/partially_hardened/not_hardened; closes Phase W; no batch-level performance claim is credible without passing this gate. | C |
235 changes: 235 additions & 0 deletions src/openamp_foundry/evidence/benchmark_challenge_registry.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,235 @@
"""BCR- benchmark challenge registry schema.

Machine-readable registry of which benchmark challenges (NCH, CMC, SCH) have
been run and passed for a given pipeline version and batch. Aggregates the
per-challenge verdicts into an overall hardness grade. No batch-level
performance claim is credible without a passing BCR- registry entry.
"""

from __future__ import annotations

from dataclasses import dataclass

REQUIRED_CHALLENGE_TYPES: tuple[str, ...] = ("NCH", "CMC", "SCH")

VALID_CHALLENGE_VERDICTS: frozenset[str] = frozenset({
"pass",
"marginal",
"fail",
"not_run",
})

VALID_BCR_HARDNESS_GRADES: frozenset[str] = frozenset({
"A",
"B",
"C",
"D",
})

_PASSING_NCH_VERDICTS: frozenset[str] = frozenset({
"novel_batch",
})
_MARGINAL_NCH_VERDICTS: frozenset[str] = frozenset({
"mixed_novelty",
})

_PASSING_CMC_VERDICTS: frozenset[str] = frozenset({
"gap_meaningful",
})
_MARGINAL_CMC_VERDICTS: frozenset[str] = frozenset({
"gap_marginal",
})

_PASSING_SCH_VERDICTS: frozenset[str] = frozenset({
"selection_adds_value",
})
_MARGINAL_SCH_VERDICTS: frozenset[str] = frozenset({
"marginal_improvement",
})


@dataclass
class ChallengeEntry:
challenge_type: str
artifact_id: str
raw_verdict: str
challenge_verdict: str


@dataclass
class BenchmarkChallengeRegistry:
bcr_id: str
batch_id: str
pipeline_version: str
challenge_entries: list[ChallengeEntry]
n_challenges_required: int
n_challenges_passed: int
n_challenges_marginal: int
n_challenges_failed: int
hardness_grade: str
dry_lab_only: bool
limitations: list[str]
created_at: str


def _map_verdict(challenge_type: str, raw_verdict: str) -> str:
if challenge_type == "NCH":
if raw_verdict in _PASSING_NCH_VERDICTS:
return "pass"
if raw_verdict in _MARGINAL_NCH_VERDICTS:
return "marginal"
if raw_verdict == "challenge_not_run":
return "not_run"
return "fail"
if challenge_type == "CMC":
if raw_verdict in _PASSING_CMC_VERDICTS:
return "pass"
if raw_verdict in _MARGINAL_CMC_VERDICTS:
return "marginal"
if raw_verdict == "challenge_not_run":
return "not_run"
return "fail"
if challenge_type == "SCH":
if raw_verdict in _PASSING_SCH_VERDICTS:
return "pass"
if raw_verdict in _MARGINAL_SCH_VERDICTS:
return "marginal"
if raw_verdict == "challenge_not_run":
return "not_run"
return "fail"
return "not_run"


def _compute_hardness_grade(
n_pass: int,
n_marginal: int,
n_required: int,
) -> str:
if n_pass == n_required:
return "A"
if n_pass + n_marginal == n_required:
return "B"
if n_pass > 0 or n_marginal > 0:
return "C"
return "D"


def validate_benchmark_challenge_registry(bcr: BenchmarkChallengeRegistry) -> None:
if not bcr.bcr_id.startswith("BCR-"):
raise ValueError(f"bcr_id must start with 'BCR-': {bcr.bcr_id!r}")
if not bcr.batch_id:
raise ValueError("batch_id must be non-empty")
if not bcr.pipeline_version:
raise ValueError("pipeline_version must be non-empty")
seen_types: set[str] = set()
for entry in bcr.challenge_entries:
if entry.challenge_type not in REQUIRED_CHALLENGE_TYPES:
raise ValueError(
f"challenge_type {entry.challenge_type!r} not in REQUIRED_CHALLENGE_TYPES"
)
if entry.challenge_type in seen_types:
raise ValueError(
f"duplicate challenge_type: {entry.challenge_type!r}"
)
seen_types.add(entry.challenge_type)
if entry.challenge_verdict not in VALID_CHALLENGE_VERDICTS:
raise ValueError(
f"challenge_verdict {entry.challenge_verdict!r} not in VALID_CHALLENGE_VERDICTS"
)
if bcr.n_challenges_required != len(REQUIRED_CHALLENGE_TYPES):
raise ValueError(
f"n_challenges_required must be {len(REQUIRED_CHALLENGE_TYPES)}"
)
n_pass = sum(1 for e in bcr.challenge_entries if e.challenge_verdict == "pass")
n_marginal = sum(1 for e in bcr.challenge_entries if e.challenge_verdict == "marginal")
n_fail = sum(
1 for e in bcr.challenge_entries
if e.challenge_verdict in ("fail", "not_run")
)
if bcr.n_challenges_passed != n_pass:
raise ValueError("n_challenges_passed mismatch")
if bcr.n_challenges_marginal != n_marginal:
raise ValueError("n_challenges_marginal mismatch")
if bcr.n_challenges_failed != n_fail:
raise ValueError("n_challenges_failed mismatch")
if bcr.hardness_grade not in VALID_BCR_HARDNESS_GRADES:
raise ValueError(
f"hardness_grade {bcr.hardness_grade!r} not in VALID_BCR_HARDNESS_GRADES"
)
if not bcr.dry_lab_only:
raise ValueError("dry_lab_only must be True")
if not bcr.limitations:
raise ValueError("limitations must be non-empty")
if not bcr.created_at:
raise ValueError("created_at must be non-empty")


def build_benchmark_challenge_registry(
*,
bcr_id: str,
batch_id: str,
pipeline_version: str,
nch_artifact_id: str = "",
nch_raw_verdict: str = "challenge_not_run",
cmc_artifact_id: str = "",
cmc_raw_verdict: str = "challenge_not_run",
sch_artifact_id: str = "",
sch_raw_verdict: str = "challenge_not_run",
limitations: list[str],
created_at: str,
) -> BenchmarkChallengeRegistry:
raw = {
"NCH": (nch_artifact_id, nch_raw_verdict),
"CMC": (cmc_artifact_id, cmc_raw_verdict),
"SCH": (sch_artifact_id, sch_raw_verdict),
}
entries = [
ChallengeEntry(
challenge_type=ctype,
artifact_id=raw[ctype][0],
raw_verdict=raw[ctype][1],
challenge_verdict=_map_verdict(ctype, raw[ctype][1]),
)
for ctype in REQUIRED_CHALLENGE_TYPES
]
n_pass = sum(1 for e in entries if e.challenge_verdict == "pass")
n_marginal = sum(1 for e in entries if e.challenge_verdict == "marginal")
n_fail = sum(1 for e in entries if e.challenge_verdict in ("fail", "not_run"))
grade = _compute_hardness_grade(n_pass, n_marginal, len(REQUIRED_CHALLENGE_TYPES))
bcr = BenchmarkChallengeRegistry(
bcr_id=bcr_id,
batch_id=batch_id,
pipeline_version=pipeline_version,
challenge_entries=entries,
n_challenges_required=len(REQUIRED_CHALLENGE_TYPES),
n_challenges_passed=n_pass,
n_challenges_marginal=n_marginal,
n_challenges_failed=n_fail,
hardness_grade=grade,
dry_lab_only=True,
limitations=limitations,
created_at=created_at,
)
validate_benchmark_challenge_registry(bcr)
return bcr


def format_benchmark_challenge_registry(bcr: BenchmarkChallengeRegistry) -> str:
lines = [
f"Benchmark Challenge Registry — {bcr.bcr_id}",
f"Batch: {bcr.batch_id} | Pipeline: {bcr.pipeline_version}",
f"Hardness grade: {bcr.hardness_grade} | "
f"Passed: {bcr.n_challenges_passed}/{bcr.n_challenges_required} "
f"Marginal: {bcr.n_challenges_marginal} Failed: {bcr.n_challenges_failed}",
"Challenges:",
]
for entry in bcr.challenge_entries:
aid = entry.artifact_id if entry.artifact_id else "(none)"
lines.append(
f" {entry.challenge_type}: {entry.challenge_verdict.upper()}"
f" [{entry.raw_verdict}] {aid}"
)
lines.append(f"Created: {bcr.created_at}")
lines.append(f"Limitations: {'; '.join(bcr.limitations)}")
lines.append(f"dry_lab_only: {bcr.dry_lab_only}")
return "\n".join(lines)
Loading
Loading