Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions docs/research/NEXT_100_PR_MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -293,3 +293,15 @@ Track whether the pipeline actually improves across batches by capturing per-bat
| X3 | Add learning progress report schema (LPR-) (complete). — src/openamp_foundry/evidence/learning_progress_report.py: VALID_LPR_VERDICTS (4: learning_confirmed/learning_inconclusive/no_learning_signal/insufficient_data), VALID_FEATURE_PREDICTIVITY (3: predictive/not_predictive/uncertain), VALID_FEATURE_CATEGORIES (8); FeatureLearningEntry helper; verdict auto-computed: learning_confirmed (n_pred>n_non), learning_inconclusive (equal+>0), no_learning_signal (n_pred<n_non), insufficient_data (no batches or no definitive features); dry_lab_only=True; 58 tests in tests/evidence/test_learning_progress_report.py. | Human-readable summary of what the pipeline has learned from all batches to date; references CIT- for trend data; includes which candidate features proved predictive vs not; links to calibration decision logs. | C |
| X4 | Add recalibration confidence certificate schema (RCC-) (complete). — src/openamp_foundry/evidence/recalibration_confidence_certificate.py: VALID_RCC_GRADES (A-D), VALID_RCC_VERDICTS (4: high_confidence/moderate_confidence/low_confidence/insufficient_data), VALID_CONSISTENCY_RATINGS (4); BatchCohortEntry helper; grade A (≥4 batches+consistent), B (≥2 batches+consistent/moderately_consistent), D (insufficient_data); cross_batch_consistency auto-computed; total_cohort_size auto-summed; dry_lab_only=True; cit_id/lpr_id prefix-validated; 76 tests. | Asserts with what confidence the current calibration weights are reliable based on cohort size, quality, and consistency across batches; A/B/C/D grade; prevents overconfident calibration claims. | C |
| X5 | Add Phase X learning gate (XLG-) (complete). — src/openamp_foundry/evidence/phase_x_learning_gate.py: REQUIRED_X_COMPONENTS=(MBL,CIT,LPR,RCC), VALID_XLG_VERDICTS (3: learning_verified/learning_in_progress/learning_not_started); XComponentCheck helper; verdict auto-computed: learning_verified (all 4 present), learning_in_progress (2-3), learning_not_started (0-1); artifact_id prefix-validated per component; dry_lab_only=True; 55 tests. | Top-level gate asserting MBL + CIT + LPR + RCC all present; overall verdict: learning_verified/learning_in_progress/learning_not_started; closes Phase X; no calibration improvement claim is credible without passing this gate. | C |

## Phase Y — Baseline-vs-pipeline accountability

Track and publish structured comparisons between pipeline selections and cheap baselines. Every pilot claim must reference a pre-registered baseline comparison record. Makes it impossible to claim "the pipeline works" without a machine-verifiable comparison against the cheapest plausible alternative. Directly addresses the bottleneck: can the pipeline choose real experiments better than charge/length heuristics?

| ID | Task | Why it matters | Priority |
|----|------|----------------|----------|
| Y1 | Add cheap baseline comparison record schema (CBR-) (complete). — src/openamp_foundry/evidence/cheap_baseline_comparison_record.py: VALID_CBR_VERDICTS (4: pipeline_superior/tied/baseline_superior/insufficient_data), VALID_BASELINE_METHODS (5: charge_only_rank/length_only_rank/random_selection/charge_length_combined/hydrophobicity_only_rank), VALID_CBR_METRICS (4: auroc/hit_rate/top_k_precision/ndcg), SUPERIORITY_THRESHOLD=0.05, MIN_SAMPLE_SIZE=5; metric_delta auto-computed; verdict auto-derived; dry_lab_only=True; 62 tests. | Structured record: pipeline metric vs charge-only/random/length-only baseline; pre-registered threshold; verdict (pipeline_superior/tied/baseline_superior/insufficient_data). Forces every performance claim to cite the baseline it beat. | C |
| Y2 | Add feature importance audit schema (FIA-). | Documents which features drove selections and whether charge/length alone explains the result; anti-cheap-explanation gate; rejects candidate panels where charge-only ordering recovers the same top-k. | C |
| Y3 | Add selection diversity audit schema (SDA-). | Tracks sequence diversity of selected panel vs random draw; detects proximity-driven selection masquerading as discovery; required before any novelty claim. | C |
| Y4 | Add pipeline maturity certificate schema (PMC-). | Aggregates CBR/FIA/SDA results into A/B/C/D maturity grade; anchors pre-registration; prevents retroactive interpretation of results. | C |
| Y5 | Add Phase Y accountability gate (YAG-). | Top-level gate asserting CBR+FIA+SDA+PMC all present; verdict: accountability_verified/accountability_partial/accountability_not_established; closes Phase Y; no external pilot claim is credible without passing this gate. | C |
175 changes: 175 additions & 0 deletions src/openamp_foundry/evidence/cheap_baseline_comparison_record.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,175 @@
"""CBR- cheap baseline comparison record schema.

Structured record comparing the pipeline's performance metric against a
cheap baseline (charge-only, length-only, random) on the same candidate set.
Forces every performance claim to name the specific baseline it beat and
provide a machine-readable verdict. No performance claim is credible without
this record.
"""

from __future__ import annotations

from dataclasses import dataclass

VALID_CBR_VERDICTS: frozenset[str] = frozenset({
"pipeline_superior",
"tied",
"baseline_superior",
"insufficient_data",
})

VALID_BASELINE_METHODS: frozenset[str] = frozenset({
"charge_only_rank",
"length_only_rank",
"random_selection",
"charge_length_combined",
"hydrophobicity_only_rank",
})

VALID_CBR_METRICS: frozenset[str] = frozenset({
"auroc",
"hit_rate",
"top_k_precision",
"ndcg",
})

SUPERIORITY_THRESHOLD: float = 0.05
INFERIORITY_THRESHOLD: float = -0.05
MIN_SAMPLE_SIZE: int = 5


@dataclass
class CheapBaselineComparisonRecord:
cbr_id: str
pipeline_version: str
baseline_method: str
metric_name: str
pipeline_metric_value: float
baseline_metric_value: float
metric_delta: float
n_candidates_evaluated: int
pre_registered_threshold: float
cbr_verdict: str
comparison_notes: str
dry_lab_only: bool
limitations: list[str]
created_at: str


def validate_cheap_baseline_comparison_record(
cbr: CheapBaselineComparisonRecord,
) -> None:
if not cbr.cbr_id.startswith("CBR-"):
raise ValueError(f"cbr_id must start with 'CBR-': {cbr.cbr_id!r}")
if not cbr.pipeline_version:
raise ValueError("pipeline_version must be non-empty")
if cbr.baseline_method not in VALID_BASELINE_METHODS:
raise ValueError(
f"baseline_method {cbr.baseline_method!r} not in VALID_BASELINE_METHODS"
)
if cbr.metric_name not in VALID_CBR_METRICS:
raise ValueError(
f"metric_name {cbr.metric_name!r} not in VALID_CBR_METRICS"
)
if not (0.0 <= cbr.pipeline_metric_value <= 1.0):
raise ValueError(
f"pipeline_metric_value must be in [0, 1]: {cbr.pipeline_metric_value}"
)
if not (0.0 <= cbr.baseline_metric_value <= 1.0):
raise ValueError(
f"baseline_metric_value must be in [0, 1]: {cbr.baseline_metric_value}"
)
expected_delta = round(cbr.pipeline_metric_value - cbr.baseline_metric_value, 6)
if abs(cbr.metric_delta - expected_delta) > 1e-5:
raise ValueError(
f"metric_delta mismatch: expected {expected_delta}, got {cbr.metric_delta}"
)
if cbr.n_candidates_evaluated < 0:
raise ValueError("n_candidates_evaluated must be non-negative")
if cbr.cbr_verdict not in VALID_CBR_VERDICTS:
raise ValueError(
f"cbr_verdict {cbr.cbr_verdict!r} not in VALID_CBR_VERDICTS"
)
if not cbr.dry_lab_only:
raise ValueError("dry_lab_only must be True")
if not cbr.limitations:
raise ValueError("limitations must be non-empty")
if not cbr.created_at:
raise ValueError("created_at must be non-empty")


def _compute_verdict(
delta: float,
n_candidates: int,
superiority_threshold: float,
) -> str:
if n_candidates < MIN_SAMPLE_SIZE:
return "insufficient_data"
if delta >= superiority_threshold:
return "pipeline_superior"
if delta <= INFERIORITY_THRESHOLD:
return "baseline_superior"
return "tied"


def build_cheap_baseline_comparison_record(
*,
cbr_id: str,
pipeline_version: str,
baseline_method: str,
metric_name: str,
pipeline_metric_value: float,
baseline_metric_value: float,
n_candidates_evaluated: int,
pre_registered_threshold: float = SUPERIORITY_THRESHOLD,
comparison_notes: str = "",
limitations: list[str],
created_at: str,
) -> CheapBaselineComparisonRecord:
"""Build a CheapBaselineComparisonRecord.

metric_delta = pipeline_metric_value - baseline_metric_value (auto-computed).
cbr_verdict is auto-derived from delta vs pre_registered_threshold and n_candidates.
"""
delta = round(pipeline_metric_value - baseline_metric_value, 6)
verdict = _compute_verdict(delta, n_candidates_evaluated, pre_registered_threshold)
cbr = CheapBaselineComparisonRecord(
cbr_id=cbr_id,
pipeline_version=pipeline_version,
baseline_method=baseline_method,
metric_name=metric_name,
pipeline_metric_value=float(pipeline_metric_value),
baseline_metric_value=float(baseline_metric_value),
metric_delta=delta,
n_candidates_evaluated=n_candidates_evaluated,
pre_registered_threshold=float(pre_registered_threshold),
cbr_verdict=verdict,
comparison_notes=comparison_notes,
dry_lab_only=True,
limitations=limitations,
created_at=created_at,
)
validate_cheap_baseline_comparison_record(cbr)
return cbr


def format_cheap_baseline_comparison_record(
cbr: CheapBaselineComparisonRecord,
) -> str:
lines = [
f"Cheap Baseline Comparison Record — {cbr.cbr_id}",
f"Pipeline: {cbr.pipeline_version}",
f"Baseline: {cbr.baseline_method} | Metric: {cbr.metric_name}",
f"Verdict: {cbr.cbr_verdict}",
f"Pipeline {cbr.metric_name}: {cbr.pipeline_metric_value:.4f}",
f"Baseline {cbr.metric_name}: {cbr.baseline_metric_value:.4f}",
f"Delta: {cbr.metric_delta:+.4f} "
f"(threshold: {cbr.pre_registered_threshold:+.4f})",
f"Candidates evaluated: {cbr.n_candidates_evaluated}",
]
if cbr.comparison_notes:
lines.append(f"Notes: {cbr.comparison_notes}")
lines.append(f"Created: {cbr.created_at}")
lines.append(f"Limitations: {'; '.join(cbr.limitations)}")
lines.append(f"dry_lab_only: {cbr.dry_lab_only}")
return "\n".join(lines)
Loading
Loading