Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 13 additions & 1 deletion docs/research/NEXT_100_PR_MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,7 +134,7 @@ Make learning controlled rather than self-serving.
| G6 | Add calibration-overfit warning when cohort is too small (complete). — calibration/overfit_warning.py: detects cohort-too-small condition and emits OverfitWarning; make calibration-overfit-check target; tests/calibration/test_overfit_warning.py. | Prevents false learning. | C |
| G7 | Add result-quality flag propagation into calibration engine (complete). — calibration/result_quality.py: flags low-quality outcomes before they enter calibration; make result-quality-filter target; tests/calibration/test_result_quality.py (27 tests). | Low-quality outcomes cannot drive updates. | C |
| G8 | Add policy that synthetic results cannot raise proof-ladder level (complete). | SBR- schema: 14 fields, 16 validation rules, synthetic-only evidence cannot propose level ≥4 without violations recorded, policy_enforced=True enforced, violation rate consistency check (tol 0.01); anti-overclaim boundary is now auditable artifact. | B/C |
| G9 | Add calibration decision review checklist. | Human review stronger. | C/D |
| G9 | Add calibration decision review checklist (complete). — calibration/decision_checklist.py: CHECKLIST_ITEMS (12 items, 11 required), CalibrationDecisionChecklist (8 fields), build_checklist(), write_checklist_json(), write_checklist_markdown(); 14 tests in tests/calibration/test_decision_checklist.py. | Human review stronger. | C/D |
| G10 | Add recalibration rollback plan (complete). — calibration/rollback_plan.py: structured plan for rolling back a calibration update if quality degrades; make calibration-rollback-plan target; tests/calibration/test_rollback_plan.py. | Safer updates. | C |

## Phase H — Virtual assay discipline
Expand Down Expand Up @@ -269,3 +269,15 @@ Gate external sharing on all evidence schemas passing; seal the evidence trail f
| V3 | Add scientific reproducibility seal schema (SRS-) (complete). — src/openamp_foundry/evidence/scientific_reproducibility_seal.py: sealed/provisional/invalidated statuses; human_reviewed cross-check; PENDING hash placeholder; 50 tests. | Immutable record asserting that a batch's evidence trail is complete and auditable; includes pipeline version, schema hash placeholder, and human-reviewed flag; enables preprint data availability statements. | C |
| V4 | Add external review packet schema (ERP-) (complete). — src/openamp_foundry/evidence/external_review_packet.py: 5 components (BRC/ECI/FET/PTR/SRS); ready/incomplete/draft status; 45 tests. | Assembles BRC + ECI + FET + PTR + SRS into a single record listing everything a scientist needs to review the batch's computational evidence; dry-lab-only constraint explicit; closes the "what do I send to a reviewer?" question. | C |
| V5 | Add Phase V completeness gate schema (V5G-) (complete). — src/openamp_foundry/evidence/phase_v_completeness_gate.py: 4 components (PRG/EBM/SRS/ERP); prefix-validated artifact IDs; ready/blocked verdict; 63 tests. | Top-level gate asserting PRG + EBM + SRS + ERP all present; closes Phase V and signals the batch is ready for external scientific review. | C |

## Phase W — Batch-level benchmark hardness and novelty challenges

Make it machine-verifiable that the pipeline produces novel candidates that beat cheap baselines. Directly addresses the strategic bottleneck: "Can the system help choose real experiments better than cheap baselines?" Every item produces a named artifact that gates further claims.

| ID | Task | Why it matters | Priority |
|----|------|----------------|----------|
| W1 | Add novelty challenge harness schema (NCH-) (complete). — src/openamp_foundry/evidence/novelty_challenge_harness.py: VALID_NCH_VERDICTS (4: novel_batch/mixed_novelty/near_neighbor_dominated/challenge_not_run), VALID_REFERENCE_DATABASES (6), NEAR_NEIGHBOR_IDENTITY_THRESHOLD=0.80, NOVEL_BATCH_CEILING=0.20, NEAR_NEIGHBOR_DOMINATED_FLOOR=0.60; NCHCandidateResult helper; build() auto-computes is_near_neighbor from identity vs threshold, fraction, verdict; dry_lab_only=True enforced; 63 tests in tests/evidence/test_novelty_challenge_harness.py. | Batch-level novelty challenge: documents the fraction of top candidates with ≥80% sequence identity to a known AMP in a reference database (APD3/DRAMP/etc.); blocks novel_batch claim when near-neighbor fraction exceeds 20%; prevents pipeline from advancing near-copies of known AMPs under a novelty label. | C |
| W2 | Add charge-matched challenge schema (CMC-). | Formally documents the charge-matched challenge: compares pipeline AUROC vs a charge-only baseline on the same candidate set; verdict controlled vocabulary (gap_meaningful/gap_marginal/gap_absent/not_run); blocks performance claims when the charge-only baseline explains the gap. | C |
| W3 | Add similarity challenge harness schema (SCH-). | Documents whether pipeline-selected candidates are systematically more similar to known AMPs than random selection from the sequence space; flags selection bias from similarity clustering; prevents "novel panel" claim when selection is proximity-driven. | C |
| W4 | Add benchmark challenge registry schema (BCR-). | Machine-readable registry of which benchmark challenges (NCH/CMC/SCH) have been run and passed for a given pipeline version; aggregates challenge verdicts; overall hardness grade (A: all passed, B: most passed, C: some passed, D: none passed). | C |
| W5 | Add Phase W benchmark gate (WBG-). | Top-level gate asserting NCH + CMC + SCH + BCR all present; overall verdict: hardened/partially_hardened/not_hardened; closes Phase W; no batch-level performance claim is credible without passing this gate. | C |
209 changes: 209 additions & 0 deletions src/openamp_foundry/evidence/novelty_challenge_harness.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,209 @@
"""NCH- novelty challenge harness schema.

Batch-level novelty challenge record: documents what fraction of top
candidates have high sequence identity (≥ threshold) to known AMPs in a
reference database. Blocks 'novel family' claims at the batch level when
the near-neighbor fraction is above the allowed ceiling.

This schema operates at batch level; per-family novelty is captured in
the FNR- (family novelty report) schema. NCH- provides the aggregate
machine-verifiable gate.
"""

from __future__ import annotations

from dataclasses import dataclass

VALID_NCH_VERDICTS: frozenset[str] = frozenset({
"novel_batch",
"mixed_novelty",
"near_neighbor_dominated",
"challenge_not_run",
})

VALID_REFERENCE_DATABASES: frozenset[str] = frozenset({
"APD3",
"DRAMP",
"DBAASP",
"CAMP",
"LAMP",
"custom",
})

NEAR_NEIGHBOR_IDENTITY_THRESHOLD: float = 0.80
NOVEL_BATCH_CEILING: float = 0.20
NEAR_NEIGHBOR_DOMINATED_FLOOR: float = 0.60


@dataclass
class NCHCandidateResult:
candidate_id: str
max_identity_to_known: float
is_near_neighbor: bool
closest_known_amp_id: str


@dataclass
class NoveltyChallengeHarness:
nch_id: str
batch_id: str
pipeline_version: str
reference_database: str
identity_threshold: float
n_candidates_checked: int
n_near_neighbors: int
near_neighbor_fraction: float
candidate_results: list[NCHCandidateResult]
nch_verdict: str
dry_lab_only: bool
limitations: list[str]
created_at: str


def validate_novelty_challenge_harness(nch: NoveltyChallengeHarness) -> None:
if not nch.nch_id.startswith("NCH-"):
raise ValueError(f"nch_id must start with 'NCH-': {nch.nch_id!r}")
if not nch.batch_id:
raise ValueError("batch_id must be non-empty")
if not nch.pipeline_version:
raise ValueError("pipeline_version must be non-empty")
if nch.reference_database not in VALID_REFERENCE_DATABASES:
raise ValueError(
f"reference_database {nch.reference_database!r} not in VALID_REFERENCE_DATABASES"
)
if not (0.0 < nch.identity_threshold <= 1.0):
raise ValueError(
f"identity_threshold must be in (0, 1]: {nch.identity_threshold}"
)
if nch.n_candidates_checked < 0:
raise ValueError("n_candidates_checked must be non-negative")
if nch.n_near_neighbors < 0:
raise ValueError("n_near_neighbors must be non-negative")
if nch.n_near_neighbors > nch.n_candidates_checked:
raise ValueError("n_near_neighbors cannot exceed n_candidates_checked")
for cr in nch.candidate_results:
if not (0.0 <= cr.max_identity_to_known <= 1.0):
raise ValueError(
f"max_identity_to_known must be in [0, 1]: {cr.max_identity_to_known}"
)
expected_nn = cr.max_identity_to_known >= nch.identity_threshold
if cr.is_near_neighbor != expected_nn:
raise ValueError(
f"is_near_neighbor mismatch for {cr.candidate_id!r}: "
f"identity={cr.max_identity_to_known}, threshold={nch.identity_threshold}"
)
if nch.n_candidates_checked != len(nch.candidate_results):
raise ValueError("n_candidates_checked must equal len(candidate_results)")
expected_nn_count = sum(1 for cr in nch.candidate_results if cr.is_near_neighbor)
if nch.n_near_neighbors != expected_nn_count:
raise ValueError("n_near_neighbors mismatch")
if nch.n_candidates_checked == 0:
expected_fraction = 0.0
else:
expected_fraction = round(
nch.n_near_neighbors / nch.n_candidates_checked, 6
)
if abs(nch.near_neighbor_fraction - expected_fraction) > 1e-4:
raise ValueError(
f"near_neighbor_fraction {nch.near_neighbor_fraction} does not match "
f"computed {expected_fraction}"
)
if nch.nch_verdict not in VALID_NCH_VERDICTS:
raise ValueError(
f"nch_verdict {nch.nch_verdict!r} not in VALID_NCH_VERDICTS"
)
if not nch.dry_lab_only:
raise ValueError("dry_lab_only must be True")
if not nch.limitations:
raise ValueError("limitations must be non-empty")
if not nch.created_at:
raise ValueError("created_at must be non-empty")


def _compute_verdict(
n_candidates: int,
near_neighbor_fraction: float,
) -> str:
if n_candidates == 0:
return "challenge_not_run"
if near_neighbor_fraction <= NOVEL_BATCH_CEILING:
return "novel_batch"
if near_neighbor_fraction >= NEAR_NEIGHBOR_DOMINATED_FLOOR:
return "near_neighbor_dominated"
return "mixed_novelty"


def build_novelty_challenge_harness(
*,
nch_id: str,
batch_id: str,
pipeline_version: str,
reference_database: str,
identity_threshold: float = NEAR_NEIGHBOR_IDENTITY_THRESHOLD,
candidate_result_dicts: list[dict],
limitations: list[str],
created_at: str,
) -> NoveltyChallengeHarness:
"""Build a NoveltyChallengeHarness.

candidate_result_dicts: list of dicts with keys:
candidate_id (str), max_identity_to_known (float),
closest_known_amp_id (str, optional, default "")
"""
candidate_results = []
for d in candidate_result_dicts:
identity = float(d["max_identity_to_known"])
is_nn = identity >= identity_threshold
candidate_results.append(
NCHCandidateResult(
candidate_id=d["candidate_id"],
max_identity_to_known=identity,
is_near_neighbor=is_nn,
closest_known_amp_id=d.get("closest_known_amp_id", ""),
)
)
n = len(candidate_results)
n_nn = sum(1 for cr in candidate_results if cr.is_near_neighbor)
fraction = round(n_nn / n, 6) if n > 0 else 0.0
verdict = _compute_verdict(n, fraction)
nch = NoveltyChallengeHarness(
nch_id=nch_id,
batch_id=batch_id,
pipeline_version=pipeline_version,
reference_database=reference_database,
identity_threshold=identity_threshold,
n_candidates_checked=n,
n_near_neighbors=n_nn,
near_neighbor_fraction=fraction,
candidate_results=candidate_results,
nch_verdict=verdict,
dry_lab_only=True,
limitations=limitations,
created_at=created_at,
)
validate_novelty_challenge_harness(nch)
return nch


def format_novelty_challenge_harness(nch: NoveltyChallengeHarness) -> str:
lines = [
f"Novelty Challenge Harness — {nch.nch_id}",
f"Batch: {nch.batch_id} | Pipeline: {nch.pipeline_version}",
f"Reference DB: {nch.reference_database} | Identity threshold: {nch.identity_threshold:.0%}",
f"Verdict: {nch.nch_verdict}",
f"Near-neighbors: {nch.n_near_neighbors}/{nch.n_candidates_checked} "
f"({nch.near_neighbor_fraction:.1%})",
]
if nch.candidate_results:
lines.append("Top candidates:")
for cr in nch.candidate_results[:5]:
nn_label = "NEAR-NEIGHBOR" if cr.is_near_neighbor else "novel"
lines.append(
f" {cr.candidate_id}: {cr.max_identity_to_known:.1%} identity [{nn_label}]"
)
if len(nch.candidate_results) > 5:
lines.append(f" ... ({len(nch.candidate_results) - 5} more)")
lines.append(f"Created: {nch.created_at}")
lines.append(f"Limitations: {'; '.join(nch.limitations)}")
lines.append(f"dry_lab_only: {nch.dry_lab_only}")
return "\n".join(lines)
Loading
Loading