diff --git a/docs/research/NEXT_100_PR_MAP.md b/docs/research/NEXT_100_PR_MAP.md index ae48055e..700ecc6c 100644 --- a/docs/research/NEXT_100_PR_MAP.md +++ b/docs/research/NEXT_100_PR_MAP.md @@ -134,7 +134,7 @@ Make learning controlled rather than self-serving. | G6 | Add calibration-overfit warning when cohort is too small (complete). — calibration/overfit_warning.py: detects cohort-too-small condition and emits OverfitWarning; make calibration-overfit-check target; tests/calibration/test_overfit_warning.py. | Prevents false learning. | C | | G7 | Add result-quality flag propagation into calibration engine (complete). — calibration/result_quality.py: flags low-quality outcomes before they enter calibration; make result-quality-filter target; tests/calibration/test_result_quality.py (27 tests). | Low-quality outcomes cannot drive updates. | C | | G8 | Add policy that synthetic results cannot raise proof-ladder level (complete). | SBR- schema: 14 fields, 16 validation rules, synthetic-only evidence cannot propose level ≥4 without violations recorded, policy_enforced=True enforced, violation rate consistency check (tol 0.01); anti-overclaim boundary is now auditable artifact. | B/C | -| G9 | Add calibration decision review checklist. | Human review stronger. | C/D | +| G9 | Add calibration decision review checklist (complete). — calibration/decision_checklist.py: CHECKLIST_ITEMS (12 items, 11 required), CalibrationDecisionChecklist (8 fields), build_checklist(), write_checklist_json(), write_checklist_markdown(); 14 tests in tests/calibration/test_decision_checklist.py. | Human review stronger. | C/D | | G10 | Add recalibration rollback plan (complete). — calibration/rollback_plan.py: structured plan for rolling back a calibration update if quality degrades; make calibration-rollback-plan target; tests/calibration/test_rollback_plan.py. | Safer updates. | C | ## Phase H — Virtual assay discipline @@ -269,3 +269,15 @@ Gate external sharing on all evidence schemas passing; seal the evidence trail f | V3 | Add scientific reproducibility seal schema (SRS-) (complete). — src/openamp_foundry/evidence/scientific_reproducibility_seal.py: sealed/provisional/invalidated statuses; human_reviewed cross-check; PENDING hash placeholder; 50 tests. | Immutable record asserting that a batch's evidence trail is complete and auditable; includes pipeline version, schema hash placeholder, and human-reviewed flag; enables preprint data availability statements. | C | | V4 | Add external review packet schema (ERP-) (complete). — src/openamp_foundry/evidence/external_review_packet.py: 5 components (BRC/ECI/FET/PTR/SRS); ready/incomplete/draft status; 45 tests. | Assembles BRC + ECI + FET + PTR + SRS into a single record listing everything a scientist needs to review the batch's computational evidence; dry-lab-only constraint explicit; closes the "what do I send to a reviewer?" question. | C | | V5 | Add Phase V completeness gate schema (V5G-) (complete). — src/openamp_foundry/evidence/phase_v_completeness_gate.py: 4 components (PRG/EBM/SRS/ERP); prefix-validated artifact IDs; ready/blocked verdict; 63 tests. | Top-level gate asserting PRG + EBM + SRS + ERP all present; closes Phase V and signals the batch is ready for external scientific review. | C | + +## Phase W — Batch-level benchmark hardness and novelty challenges + +Make it machine-verifiable that the pipeline produces novel candidates that beat cheap baselines. Directly addresses the strategic bottleneck: "Can the system help choose real experiments better than cheap baselines?" Every item produces a named artifact that gates further claims. + +| ID | Task | Why it matters | Priority | +|----|------|----------------|----------| +| W1 | Add novelty challenge harness schema (NCH-) (complete). — src/openamp_foundry/evidence/novelty_challenge_harness.py: VALID_NCH_VERDICTS (4: novel_batch/mixed_novelty/near_neighbor_dominated/challenge_not_run), VALID_REFERENCE_DATABASES (6), NEAR_NEIGHBOR_IDENTITY_THRESHOLD=0.80, NOVEL_BATCH_CEILING=0.20, NEAR_NEIGHBOR_DOMINATED_FLOOR=0.60; NCHCandidateResult helper; build() auto-computes is_near_neighbor from identity vs threshold, fraction, verdict; dry_lab_only=True enforced; 63 tests in tests/evidence/test_novelty_challenge_harness.py. | Batch-level novelty challenge: documents the fraction of top candidates with ≥80% sequence identity to a known AMP in a reference database (APD3/DRAMP/etc.); blocks novel_batch claim when near-neighbor fraction exceeds 20%; prevents pipeline from advancing near-copies of known AMPs under a novelty label. | C | +| W2 | Add charge-matched challenge schema (CMC-). | Formally documents the charge-matched challenge: compares pipeline AUROC vs a charge-only baseline on the same candidate set; verdict controlled vocabulary (gap_meaningful/gap_marginal/gap_absent/not_run); blocks performance claims when the charge-only baseline explains the gap. | C | +| W3 | Add similarity challenge harness schema (SCH-). | Documents whether pipeline-selected candidates are systematically more similar to known AMPs than random selection from the sequence space; flags selection bias from similarity clustering; prevents "novel panel" claim when selection is proximity-driven. | C | +| W4 | Add benchmark challenge registry schema (BCR-). | Machine-readable registry of which benchmark challenges (NCH/CMC/SCH) have been run and passed for a given pipeline version; aggregates challenge verdicts; overall hardness grade (A: all passed, B: most passed, C: some passed, D: none passed). | C | +| W5 | Add Phase W benchmark gate (WBG-). | Top-level gate asserting NCH + CMC + SCH + BCR all present; overall verdict: hardened/partially_hardened/not_hardened; closes Phase W; no batch-level performance claim is credible without passing this gate. | C | diff --git a/src/openamp_foundry/evidence/novelty_challenge_harness.py b/src/openamp_foundry/evidence/novelty_challenge_harness.py new file mode 100644 index 00000000..f21aade1 --- /dev/null +++ b/src/openamp_foundry/evidence/novelty_challenge_harness.py @@ -0,0 +1,209 @@ +"""NCH- novelty challenge harness schema. + +Batch-level novelty challenge record: documents what fraction of top +candidates have high sequence identity (≥ threshold) to known AMPs in a +reference database. Blocks 'novel family' claims at the batch level when +the near-neighbor fraction is above the allowed ceiling. + +This schema operates at batch level; per-family novelty is captured in +the FNR- (family novelty report) schema. NCH- provides the aggregate +machine-verifiable gate. +""" + +from __future__ import annotations + +from dataclasses import dataclass + +VALID_NCH_VERDICTS: frozenset[str] = frozenset({ + "novel_batch", + "mixed_novelty", + "near_neighbor_dominated", + "challenge_not_run", +}) + +VALID_REFERENCE_DATABASES: frozenset[str] = frozenset({ + "APD3", + "DRAMP", + "DBAASP", + "CAMP", + "LAMP", + "custom", +}) + +NEAR_NEIGHBOR_IDENTITY_THRESHOLD: float = 0.80 +NOVEL_BATCH_CEILING: float = 0.20 +NEAR_NEIGHBOR_DOMINATED_FLOOR: float = 0.60 + + +@dataclass +class NCHCandidateResult: + candidate_id: str + max_identity_to_known: float + is_near_neighbor: bool + closest_known_amp_id: str + + +@dataclass +class NoveltyChallengeHarness: + nch_id: str + batch_id: str + pipeline_version: str + reference_database: str + identity_threshold: float + n_candidates_checked: int + n_near_neighbors: int + near_neighbor_fraction: float + candidate_results: list[NCHCandidateResult] + nch_verdict: str + dry_lab_only: bool + limitations: list[str] + created_at: str + + +def validate_novelty_challenge_harness(nch: NoveltyChallengeHarness) -> None: + if not nch.nch_id.startswith("NCH-"): + raise ValueError(f"nch_id must start with 'NCH-': {nch.nch_id!r}") + if not nch.batch_id: + raise ValueError("batch_id must be non-empty") + if not nch.pipeline_version: + raise ValueError("pipeline_version must be non-empty") + if nch.reference_database not in VALID_REFERENCE_DATABASES: + raise ValueError( + f"reference_database {nch.reference_database!r} not in VALID_REFERENCE_DATABASES" + ) + if not (0.0 < nch.identity_threshold <= 1.0): + raise ValueError( + f"identity_threshold must be in (0, 1]: {nch.identity_threshold}" + ) + if nch.n_candidates_checked < 0: + raise ValueError("n_candidates_checked must be non-negative") + if nch.n_near_neighbors < 0: + raise ValueError("n_near_neighbors must be non-negative") + if nch.n_near_neighbors > nch.n_candidates_checked: + raise ValueError("n_near_neighbors cannot exceed n_candidates_checked") + for cr in nch.candidate_results: + if not (0.0 <= cr.max_identity_to_known <= 1.0): + raise ValueError( + f"max_identity_to_known must be in [0, 1]: {cr.max_identity_to_known}" + ) + expected_nn = cr.max_identity_to_known >= nch.identity_threshold + if cr.is_near_neighbor != expected_nn: + raise ValueError( + f"is_near_neighbor mismatch for {cr.candidate_id!r}: " + f"identity={cr.max_identity_to_known}, threshold={nch.identity_threshold}" + ) + if nch.n_candidates_checked != len(nch.candidate_results): + raise ValueError("n_candidates_checked must equal len(candidate_results)") + expected_nn_count = sum(1 for cr in nch.candidate_results if cr.is_near_neighbor) + if nch.n_near_neighbors != expected_nn_count: + raise ValueError("n_near_neighbors mismatch") + if nch.n_candidates_checked == 0: + expected_fraction = 0.0 + else: + expected_fraction = round( + nch.n_near_neighbors / nch.n_candidates_checked, 6 + ) + if abs(nch.near_neighbor_fraction - expected_fraction) > 1e-4: + raise ValueError( + f"near_neighbor_fraction {nch.near_neighbor_fraction} does not match " + f"computed {expected_fraction}" + ) + if nch.nch_verdict not in VALID_NCH_VERDICTS: + raise ValueError( + f"nch_verdict {nch.nch_verdict!r} not in VALID_NCH_VERDICTS" + ) + if not nch.dry_lab_only: + raise ValueError("dry_lab_only must be True") + if not nch.limitations: + raise ValueError("limitations must be non-empty") + if not nch.created_at: + raise ValueError("created_at must be non-empty") + + +def _compute_verdict( + n_candidates: int, + near_neighbor_fraction: float, +) -> str: + if n_candidates == 0: + return "challenge_not_run" + if near_neighbor_fraction <= NOVEL_BATCH_CEILING: + return "novel_batch" + if near_neighbor_fraction >= NEAR_NEIGHBOR_DOMINATED_FLOOR: + return "near_neighbor_dominated" + return "mixed_novelty" + + +def build_novelty_challenge_harness( + *, + nch_id: str, + batch_id: str, + pipeline_version: str, + reference_database: str, + identity_threshold: float = NEAR_NEIGHBOR_IDENTITY_THRESHOLD, + candidate_result_dicts: list[dict], + limitations: list[str], + created_at: str, +) -> NoveltyChallengeHarness: + """Build a NoveltyChallengeHarness. + + candidate_result_dicts: list of dicts with keys: + candidate_id (str), max_identity_to_known (float), + closest_known_amp_id (str, optional, default "") + """ + candidate_results = [] + for d in candidate_result_dicts: + identity = float(d["max_identity_to_known"]) + is_nn = identity >= identity_threshold + candidate_results.append( + NCHCandidateResult( + candidate_id=d["candidate_id"], + max_identity_to_known=identity, + is_near_neighbor=is_nn, + closest_known_amp_id=d.get("closest_known_amp_id", ""), + ) + ) + n = len(candidate_results) + n_nn = sum(1 for cr in candidate_results if cr.is_near_neighbor) + fraction = round(n_nn / n, 6) if n > 0 else 0.0 + verdict = _compute_verdict(n, fraction) + nch = NoveltyChallengeHarness( + nch_id=nch_id, + batch_id=batch_id, + pipeline_version=pipeline_version, + reference_database=reference_database, + identity_threshold=identity_threshold, + n_candidates_checked=n, + n_near_neighbors=n_nn, + near_neighbor_fraction=fraction, + candidate_results=candidate_results, + nch_verdict=verdict, + dry_lab_only=True, + limitations=limitations, + created_at=created_at, + ) + validate_novelty_challenge_harness(nch) + return nch + + +def format_novelty_challenge_harness(nch: NoveltyChallengeHarness) -> str: + lines = [ + f"Novelty Challenge Harness — {nch.nch_id}", + f"Batch: {nch.batch_id} | Pipeline: {nch.pipeline_version}", + f"Reference DB: {nch.reference_database} | Identity threshold: {nch.identity_threshold:.0%}", + f"Verdict: {nch.nch_verdict}", + f"Near-neighbors: {nch.n_near_neighbors}/{nch.n_candidates_checked} " + f"({nch.near_neighbor_fraction:.1%})", + ] + if nch.candidate_results: + lines.append("Top candidates:") + for cr in nch.candidate_results[:5]: + nn_label = "NEAR-NEIGHBOR" if cr.is_near_neighbor else "novel" + lines.append( + f" {cr.candidate_id}: {cr.max_identity_to_known:.1%} identity [{nn_label}]" + ) + if len(nch.candidate_results) > 5: + lines.append(f" ... ({len(nch.candidate_results) - 5} more)") + lines.append(f"Created: {nch.created_at}") + lines.append(f"Limitations: {'; '.join(nch.limitations)}") + lines.append(f"dry_lab_only: {nch.dry_lab_only}") + return "\n".join(lines) diff --git a/tests/evidence/test_novelty_challenge_harness.py b/tests/evidence/test_novelty_challenge_harness.py new file mode 100644 index 00000000..93d6e044 --- /dev/null +++ b/tests/evidence/test_novelty_challenge_harness.py @@ -0,0 +1,380 @@ +"""Tests for NCH- novelty challenge harness schema.""" + +import pytest +from openamp_foundry.evidence.novelty_challenge_harness import ( + NoveltyChallengeHarness, + NCHCandidateResult, + VALID_NCH_VERDICTS, + VALID_REFERENCE_DATABASES, + NEAR_NEIGHBOR_IDENTITY_THRESHOLD, + NOVEL_BATCH_CEILING, + NEAR_NEIGHBOR_DOMINATED_FLOOR, + build_novelty_challenge_harness, + format_novelty_challenge_harness, + validate_novelty_challenge_harness, +) + +# --------------------------------------------------------------------------- +# Helpers +# --------------------------------------------------------------------------- + +_NOVEL_CANDIDATES = [ + {"candidate_id": f"FAM-{i:03d}", "max_identity_to_known": 0.30, "closest_known_amp_id": "AMP-X"} + for i in range(10) +] + +_MIXED_CANDIDATES = [ + {"candidate_id": "FAM-001", "max_identity_to_known": 0.90, "closest_known_amp_id": "AMP-A"}, + {"candidate_id": "FAM-002", "max_identity_to_known": 0.30, "closest_known_amp_id": "AMP-B"}, + {"candidate_id": "FAM-003", "max_identity_to_known": 0.30, "closest_known_amp_id": "AMP-C"}, + {"candidate_id": "FAM-004", "max_identity_to_known": 0.30, "closest_known_amp_id": "AMP-D"}, +] + +_NN_CANDIDATES = [ + {"candidate_id": f"FAM-{i:03d}", "max_identity_to_known": 0.95, "closest_known_amp_id": "AMP-Z"} + for i in range(10) +] + + +def _build(**kwargs): + defaults = dict( + nch_id="NCH-001", + batch_id="BATCH-01", + pipeline_version="v1.0", + reference_database="APD3", + candidate_result_dicts=_NOVEL_CANDIDATES, + limitations=["dry-lab only", "identity search is approximate"], + created_at="2026-07-10", + ) + defaults.update(kwargs) + return build_novelty_challenge_harness(**defaults) + + +# --------------------------------------------------------------------------- +# 1. Constants +# --------------------------------------------------------------------------- + + +def test_valid_nch_verdicts_is_frozenset(): + assert isinstance(VALID_NCH_VERDICTS, frozenset) + + +def test_valid_nch_verdicts_contains_novel_batch(): + assert "novel_batch" in VALID_NCH_VERDICTS + + +def test_valid_nch_verdicts_contains_mixed_novelty(): + assert "mixed_novelty" in VALID_NCH_VERDICTS + + +def test_valid_nch_verdicts_contains_near_neighbor_dominated(): + assert "near_neighbor_dominated" in VALID_NCH_VERDICTS + + +def test_valid_nch_verdicts_contains_challenge_not_run(): + assert "challenge_not_run" in VALID_NCH_VERDICTS + + +def test_valid_reference_databases_is_frozenset(): + assert isinstance(VALID_REFERENCE_DATABASES, frozenset) + + +def test_valid_reference_databases_contains_apd3(): + assert "APD3" in VALID_REFERENCE_DATABASES + + +def test_valid_reference_databases_contains_dramp(): + assert "DRAMP" in VALID_REFERENCE_DATABASES + + +def test_valid_reference_databases_contains_custom(): + assert "custom" in VALID_REFERENCE_DATABASES + + +def test_near_neighbor_identity_threshold(): + assert NEAR_NEIGHBOR_IDENTITY_THRESHOLD == 0.80 + + +def test_novel_batch_ceiling(): + assert NOVEL_BATCH_CEILING == 0.20 + + +def test_near_neighbor_dominated_floor(): + assert NEAR_NEIGHBOR_DOMINATED_FLOOR == 0.60 + + +# --------------------------------------------------------------------------- +# 2. build – happy paths +# --------------------------------------------------------------------------- + + +def test_build_returns_novelty_challenge_harness(): + assert isinstance(_build(), NoveltyChallengeHarness) + + +def test_build_nch_id_stored(): + assert _build().nch_id == "NCH-001" + + +def test_build_batch_id_stored(): + assert _build().batch_id == "BATCH-01" + + +def test_build_pipeline_version_stored(): + assert _build().pipeline_version == "v1.0" + + +def test_build_dry_lab_only_true(): + assert _build().dry_lab_only is True + + +def test_build_reference_database_stored(): + assert _build().reference_database == "APD3" + + +def test_build_identity_threshold_default(): + assert _build().identity_threshold == NEAR_NEIGHBOR_IDENTITY_THRESHOLD + + +def test_build_identity_threshold_custom(): + r = _build(identity_threshold=0.70) + assert r.identity_threshold == 0.70 + + +def test_build_novel_batch_verdict_when_all_novel(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + assert r.nch_verdict == "novel_batch" + + +def test_build_n_candidates_checked_matches_input(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + assert r.n_candidates_checked == 10 + + +def test_build_n_near_neighbors_zero_when_all_novel(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + assert r.n_near_neighbors == 0 + + +def test_build_near_neighbor_fraction_zero_when_all_novel(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + assert r.near_neighbor_fraction == 0.0 + + +def test_build_mixed_novelty_verdict(): + r = _build(candidate_result_dicts=_MIXED_CANDIDATES) + assert r.nch_verdict == "mixed_novelty" + + +def test_build_mixed_n_near_neighbors(): + r = _build(candidate_result_dicts=_MIXED_CANDIDATES) + assert r.n_near_neighbors == 1 + + +def test_build_mixed_near_neighbor_fraction(): + r = _build(candidate_result_dicts=_MIXED_CANDIDATES) + assert abs(r.near_neighbor_fraction - 0.25) < 1e-4 + + +def test_build_near_neighbor_dominated_verdict(): + r = _build(candidate_result_dicts=_NN_CANDIDATES) + assert r.nch_verdict == "near_neighbor_dominated" + + +def test_build_near_neighbor_dominated_n_nn(): + r = _build(candidate_result_dicts=_NN_CANDIDATES) + assert r.n_near_neighbors == 10 + + +def test_build_challenge_not_run_verdict_when_empty(): + r = _build(candidate_result_dicts=[]) + assert r.nch_verdict == "challenge_not_run" + + +def test_build_empty_candidates_fraction_zero(): + r = _build(candidate_result_dicts=[]) + assert r.near_neighbor_fraction == 0.0 + + +def test_build_candidate_results_are_nch_candidate_result(): + for cr in _build().candidate_results: + assert isinstance(cr, NCHCandidateResult) + + +def test_build_candidate_result_is_near_neighbor_false_for_novel(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + for cr in r.candidate_results: + assert cr.is_near_neighbor is False + + +def test_build_candidate_result_is_near_neighbor_true_for_nn(): + r = _build(candidate_result_dicts=_NN_CANDIDATES) + for cr in r.candidate_results: + assert cr.is_near_neighbor is True + + +def test_build_closest_known_amp_id_stored(): + r = _build(candidate_result_dicts=_NOVEL_CANDIDATES) + assert r.candidate_results[0].closest_known_amp_id == "AMP-X" + + +def test_build_closest_known_amp_id_defaults_empty(): + candidates = [{"candidate_id": "FAM-001", "max_identity_to_known": 0.3}] + r = _build(candidate_result_dicts=candidates) + assert r.candidate_results[0].closest_known_amp_id == "" + + +def test_build_limitations_stored(): + r = _build() + assert "dry-lab only" in r.limitations + + +def test_build_created_at_stored(): + assert _build().created_at == "2026-07-10" + + +def test_build_boundary_novel_batch_at_ceiling(): + candidates = [ + {"candidate_id": f"FAM-{i:03d}", "max_identity_to_known": 0.85 if i == 0 else 0.30} + for i in range(10) + ] + r = _build(candidate_result_dicts=candidates) + assert r.nch_verdict == "novel_batch" + assert r.n_near_neighbors == 1 + assert abs(r.near_neighbor_fraction - 0.10) < 1e-4 + + +def test_build_boundary_mixed_above_ceiling(): + candidates = [ + {"candidate_id": f"FAM-{i:03d}", "max_identity_to_known": 0.85 if i < 3 else 0.30} + for i in range(10) + ] + r = _build(candidate_result_dicts=candidates) + assert r.nch_verdict == "mixed_novelty" + + +def test_build_dramp_database(): + r = _build(reference_database="DRAMP") + assert r.reference_database == "DRAMP" + + +def test_build_custom_database(): + r = _build(reference_database="custom") + assert r.reference_database == "custom" + + +# --------------------------------------------------------------------------- +# 3. validate – rejection cases +# --------------------------------------------------------------------------- + + +def test_validate_rejects_bad_nch_id_prefix(): + with pytest.raises(ValueError, match="NCH-"): + _build(nch_id="BAD-001") + + +def test_validate_rejects_empty_batch_id(): + with pytest.raises(ValueError): + _build(batch_id="") + + +def test_validate_rejects_empty_pipeline_version(): + with pytest.raises(ValueError): + _build(pipeline_version="") + + +def test_validate_rejects_invalid_reference_database(): + with pytest.raises(ValueError, match="VALID_REFERENCE_DATABASES"): + _build(reference_database="UNKNOWN_DB") + + +def test_validate_rejects_zero_identity_threshold(): + with pytest.raises(ValueError, match="identity_threshold"): + _build(identity_threshold=0.0) + + +def test_validate_rejects_identity_threshold_above_one(): + with pytest.raises(ValueError, match="identity_threshold"): + _build(identity_threshold=1.1) + + +def test_validate_rejects_empty_limitations(): + with pytest.raises(ValueError, match="limitations"): + _build(limitations=[]) + + +def test_validate_rejects_empty_created_at(): + with pytest.raises(ValueError): + _build(created_at="") + + +def test_validate_rejects_invalid_nch_verdict(): + nch = _build() + nch.nch_verdict = "UNKNOWN" + with pytest.raises(ValueError, match="nch_verdict"): + validate_novelty_challenge_harness(nch) + + +def test_validate_rejects_dry_lab_only_false(): + nch = _build() + nch.dry_lab_only = False + with pytest.raises(ValueError, match="dry_lab_only"): + validate_novelty_challenge_harness(nch) + + +def test_validate_rejects_n_near_neighbors_exceeds_checked(): + nch = _build() + nch.n_near_neighbors = nch.n_candidates_checked + 1 + with pytest.raises(ValueError, match="n_near_neighbors"): + validate_novelty_challenge_harness(nch) + + +def test_validate_rejects_n_candidates_mismatch(): + nch = _build() + nch.n_candidates_checked = 999 + with pytest.raises(ValueError, match="n_candidates_checked"): + validate_novelty_challenge_harness(nch) + + +def test_validate_rejects_identity_above_one_in_candidate(): + candidates = [{"candidate_id": "FAM-001", "max_identity_to_known": 1.5}] + with pytest.raises(ValueError, match="max_identity_to_known"): + _build(candidate_result_dicts=candidates) + + +# --------------------------------------------------------------------------- +# 4. format +# --------------------------------------------------------------------------- + + +def test_format_contains_nch_id(): + assert "NCH-001" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_batch_id(): + assert "BATCH-01" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_reference_database(): + assert "APD3" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_verdict(): + assert "novel_batch" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_limitations(): + assert "dry-lab only" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_dry_lab_only(): + assert "dry_lab_only: True" in format_novelty_challenge_harness(_build()) + + +def test_format_contains_near_neighbor_label_for_nn(): + r = _build(candidate_result_dicts=_NN_CANDIDATES) + assert "NEAR-NEIGHBOR" in format_novelty_challenge_harness(r) + + +def test_format_is_string(): + assert isinstance(format_novelty_challenge_harness(_build()), str)