diff --git a/AGENTS.md b/AGENTS.md index e41af91..32c3c34 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -24,7 +24,7 @@ Correct, narrow, and falsifiable beats prolific. ## Key Components - `problem-packs/`: canonical problem-pack directories with task maps across multiple domains. Read generated indexes for live counts. When calibrating pack quality, use strong exemplars such as `climate-health/dengue-heat-vietnam` for operational humility and `public-health/birth-registration-access-global` for measure-family discipline across survey, CRVS, and health-touchpoint evidence. -- `schemas/`: JSON schemas for machine-checkable protocol objects. +- `schemas/`: JSON schemas for machine-checkable protocol objects. Includes the claim schema (the core thesis object: persistent claims with a verification lifecycle, evidence links, failure modes, kill conditions, and required reviewers) and the replication schema (independent replication records with environment, input hash, and divergence tracking). - `.github/ISSUE_TEMPLATE/`: structured issue forms. - `.github/workflows/`: validation, source verification, reproducibility, and Wiki publishing. - `agents/`: role guides for structured agent contributions. @@ -34,29 +34,16 @@ Correct, narrow, and falsifiable beats prolific. ## Active Problem Packs -| ID | Domain | Region | -| ------------------------------------------------------ | --------------------------------------- | ------------------------------ | -| `air-quality/indoor-air-pollution-sub-saharan-africa` | air-quality, public-health | sub-saharan-africa | -| `air-quality/pm25-monitoring-south-asia` | air-quality, public-health | south-asia | -| `biodiversity/coral-bleaching-great-barrier-reef` | biodiversity, climate-health | great-barrier-reef | -| `biodiversity/deforestation-amazon` | biodiversity, climate-health | amazon | -| `climate-adaptation/sea-level-rise-small-islands` | climate-adaptation, disaster-resilience | small-island-developing-states | -| `climate-health/dengue-heat-vietnam` | climate-health, public-health | vietnam | -| `climate-health/heat-stress-urban-south-asia` | climate-health, public-health | south-asia | -| `climate-health/malaria-early-warning-africa` | climate-health, public-health | sub-saharan-africa | -| `disaster-resilience/cyclone-early-warning-bangladesh` | disaster-resilience, climate-health | bangladesh | -| `disaster-resilience/earthquake-vulnerability-nepal` | disaster-resilience | nepal | -| `education/girls-education-sub-saharan-africa` | education | sub-saharan-africa | -| `education/learning-loss-post-pandemic` | education | global | -| `energy-access/clean-cooking-sub-saharan-africa` | energy-access, public-health | sub-saharan-africa | -| `energy-access/mini-grid-rural-sub-saharan-africa` | energy-access | sub-saharan-africa | -| `food-security/drought-early-warning-horn-of-africa` | food-security, disaster-resilience | east-africa | -| `food-security/locust-outbreak-east-africa` | food-security, climate-health | east-africa | -| `public-health/lead-exposure-urban-global` | public-health | global | -| `public-health/stunting-sub-saharan-africa` | public-health, food-security | sub-saharan-africa | -| `sanitation/open-defecation-india` | sanitation, public-health | india | -| `water-security/glacial-melt-hindu-kush` | water-security, climate-adaptation | hindu-kush-himalaya | -| `water-security/groundwater-depletion-india` | water-security, food-security | india | +The portfolio grows continuously. Do not rely on a static count or a hand-maintained table. Read the live generated indexes: + +- [`tasks-available.json`](tasks-available.json) — machine-readable index of every scoped task +- [`agent-radar.json`](agent-radar.json) — routing layer ranking first moves and unlock paths +- [`docs/wiki/Problem-Packs.md`](docs/wiki/Problem-Packs.md) — auto-generated reader-facing list + +When calibrating pack quality, use these strong exemplars: + +- `climate-health/dengue-heat-vietnam` — operational humility, bounded claims, analytic-use warnings +- `public-health/birth-registration-access-global` — measure-family discipline across survey, CRVS, and health-touchpoint evidence ## Agent Working Rules @@ -105,6 +92,29 @@ A contribution may only be merged if at least one of these is true: Pure prose polish without one of the above is not a reason to merge. +## The Claim Lifecycle + +The thesis of this repository is that verification is the scarce resource. The claim object is the protocol object that makes that thesis machine-checkable. A claim is a persistent, falsifiable statement that exists independently of any single submission or review. It accumulates evidence and failure modes over time, carries its own kill condition, and moves through a verification lifecycle. + +```text +unverified -> dry-lab-verified -> needs-replication -> replicated -> accepted -> field-tested + \-> rejected + \-> falsified + \-> deprecated +``` + +- `unverified`: A claim has been submitted but no evidence has been reviewed. +- `dry-lab-verified`: Evidence records support the claim at the computational or literature level. No wet-lab or field confirmation. +- `needs-replication`: Evidence has been reviewed but independent replication is required before the claim can advance. +- `replicated`: An independent replicator has confirmed the result. The replication record is linked. +- `accepted`: The claim has passed all required reviews and replication. It is canonical repo truth. +- `field-tested`: The claim has been tested in a real-world operational context. The highest status. +- `rejected`: The claim did not survive review. Terminal. +- `falsified`: The kill condition was met. Terminal. A falsified claim stays in the record so future contributors do not repeat it. +- `deprecated`: The claim was accepted but its evidence has decayed (source rot, dataset revision, superseded by a stronger claim). Terminal but recoverable through a new submission. + +A claim with no evidence is `unverified` regardless of prose. A claim with no kill condition is not a claim. A high-safety claim without red-team review cannot advance past `dry-lab-verified`. The schema enforces these constraints; reviewers enforce the rest. + ## Roles for Top AI Agents Strong models should pick a role and stay in it for the duration of a PR. Mixing roles in one submission hides which judgment failed. diff --git a/CLAUDE.md b/CLAUDE.md index 2f9178c..ec0d69d 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -78,9 +78,13 @@ Machine-checkable schemas for every artifact type. Read before creating or editi | Evidence record | `schemas/evidence.schema.json` | | Agent submission | `schemas/agent-submission.schema.json` | | Review | `schemas/review.schema.json` | +| Claim | `schemas/claim.schema.json` | +| Replication | `schemas/replication.schema.json` | All schemas have `description` fields on every property. +The claim schema is the core protocol object: a persistent, falsifiable statement with a verification lifecycle (`unverified` to `field-tested`), evidence links, failure modes, a kill condition, and required reviewer roles. Replication records track independent attempts to reproduce a claim, with environment, input hash, and divergence tracking. + ## Submission Quality Standard **One task. One role. One claim.** Submissions that mix roles or make multiple independent claims are returned for splitting. @@ -90,7 +94,7 @@ Every evidence record requires: - Claim: one specific, falsifiable statement - Source: title and stable URL - Source date and access date -- Evidence type: `primary-source` | `peer-reviewed-study` | `dataset` | `field-report` | `expert-review` | `replication` | `negative-result` +- Evidence type: `primary-source` | `peer-reviewed-study` | `dataset` | `field-report` | `expert-review` | `replication` | `negative-result` | `model-prediction` | `computational-analysis` | `wet-lab-confirmation` - Method: how you assessed what the source proves - Limitations: what the source cannot prove (be specific — "results may vary" is not a limitation) - Confidence: high / medium / low / unknown diff --git a/ROADMAP.md b/ROADMAP.md index 3154b21..7ad43f7 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -10,7 +10,7 @@ The V0 system uses GitHub itself as the operating surface. No custom web app. Ev ### What V0 delivers -- **15 problem packs** across 11 domains: air quality, biodiversity, climate adaptation, climate health, disaster resilience, education, energy access, food security, public health, sanitation, and water security. +- **100 problem packs** across 11 domains: air quality, biodiversity, climate adaptation, climate health, disaster resilience, education, energy access, food security, public health, sanitation, and water security. - **Machine-checkable schemas** for problems, tasks, evidence records, agent submissions, and reviews. - **Five structured issue forms**: problem, task, agent submission, review, and safety flag. - **Automated validation**: schema checks, label coverage, wiki freshness, source verification, and reproducibility checks run on every pull request. diff --git a/agents/literature-scout.md b/agents/literature-scout.md index d94b7f3..7bef56c 100644 --- a/agents/literature-scout.md +++ b/agents/literature-scout.md @@ -36,15 +36,18 @@ Every Literature Scout submission must include: These are the allowed values in `evidence.schema.json`. Use exactly these strings. -| Type | When to use | -| --------------------- | ------------------------------------------------------------------------------------------------------ | -| `primary-source` | Original data, official statistics, WHO fact sheets, program records | -| `peer-reviewed-study` | Published research with methods that can be replicated, including systematic reviews and meta-analyses | -| `dataset` | A standalone dataset (use for data-source records) | -| `field-report` | Reports from field operations, NGO program evaluations, grey literature | -| `expert-review` | Qualitative assessment from named domain experts | -| `replication` | An independent replication of a prior claim | -| `negative-result` | A well-documented finding that a signal, method, or claim did not hold | +| Type | When to use | +| ------------------------ | ------------------------------------------------------------------------------------------------------ | +| `primary-source` | Original data, official statistics, WHO fact sheets, program records | +| `peer-reviewed-study` | Published research with methods that can be replicated, including systematic reviews and meta-analyses | +| `dataset` | A standalone dataset (use for data-source records) | +| `field-report` | Reports from field operations, NGO program evaluations, grey literature | +| `expert-review` | Qualitative assessment from named domain experts | +| `replication` | An independent replication of a prior claim | +| `negative-result` | A well-documented finding that a signal, method, or claim did not hold | +| `model-prediction` | An AI or computational model prediction (e.g., predicted antimicrobial peptide activity) | +| `computational-analysis` | A computational analysis or simulation (e.g., molecular docking, epidemiological back-testing) | +| `wet-lab-confirmation` | A wet-lab experimental result that confirms or refutes a prior computational prediction | ## What Good Looks Like diff --git a/examples/AGENTS.md b/examples/AGENTS.md index 8ba2d35..1dd25c6 100644 --- a/examples/AGENTS.md +++ b/examples/AGENTS.md @@ -7,8 +7,10 @@ Examples here are operator aids. They should mirror canonical workflow rules clo ## Key Components - `AGENT-BOOTSTRAP-PROMPT.md`: one-shot operator prompt that routes an external agent into the repo correctly. -- `agent-submission.example.json`: valid sample submission shape. +- `agent-submission.example.json`: valid sample submission shape, including kill_condition. - `review.example.json`: valid sample review shape. +- `claim.example.json`: valid sample claim object — the core thesis object with verification lifecycle, evidence links, failure modes, and kill condition. +- `replication.example.json`: valid sample replication record with environment, input hash, and divergence tracking. ## Diagrams (Mermaid) diff --git a/examples/agent-submission.example.json b/examples/agent-submission.example.json index 608d90f..29253b8 100644 --- a/examples/agent-submission.example.json +++ b/examples/agent-submission.example.json @@ -16,6 +16,7 @@ "A reader may treat climate association as a local causal threshold.", "A contributor may skip district-level denominator checks." ], + "kill_condition": "A reviewer finds that the cited sources do not support framing-level claims about dengue-climate association in Viet Nam, or that a listed source has been retracted.", "confidence": "medium", "suggested_next_issue": "Inventory public district-level dengue surveillance data access." } diff --git a/examples/claim.example.json b/examples/claim.example.json new file mode 100644 index 0000000..d3a4dfc --- /dev/null +++ b/examples/claim.example.json @@ -0,0 +1,24 @@ +{ + "id": "dengue-heat-rainfall-association-vietnam", + "problem_id": "climate-health/dengue-heat-vietnam", + "claim": "Interactions among temperature, hydrometeorology, infrastructure, and mobility are associated with dengue emergence in Viet Nam at district level, based on 23 years of surveillance data. This association is framable but does not prove a local causal threshold.", + "domain": ["climate-health", "public-health"], + "status": "dry-lab-verified", + "evidence": ["gibb-vietnam-dengue-2023", "who-dengue-fact-sheet-2025"], + "failure_modes": [ + "A reader treats the association as a local causal threshold for a specific district.", + "Passive surveillance data reflects reporting access, not only true incidence.", + "Infrastructure and mobility variables are proxy measures that may not capture the true confounders." + ], + "kill_condition": "A replication on an independent surveillance dataset for Viet Nam finds no statistically significant association between temperature, rainfall, or infrastructure interactions and district-level dengue incidence after controlling for reporting access.", + "review_required": ["domain-reviewer", "replicator"], + "safety_level": "medium", + "submitter": "agent-literature-scout-1", + "created_date": "2026-06-06", + "last_updated": "2026-06-06", + "confidence": "medium", + "limitations": [ + "The peer-reviewed model uses passive surveillance data that can reflect reporting access, not only true incidence.", + "The study covers Viet Nam nationally; district-level generalization to other countries is not supported." + ] +} diff --git a/examples/replication.example.json b/examples/replication.example.json new file mode 100644 index 0000000..8abc392 --- /dev/null +++ b/examples/replication.example.json @@ -0,0 +1,17 @@ +{ + "id": "dengue-heat-association-replication-1", + "claim_id": "dengue-heat-rainfall-association-vietnam", + "replicator": "agent-data-cleaner-2", + "replication_date": "2026-06-10", + "method": "Re-extracted district-level dengue surveillance data from the original supplementary tables, recomputed the temperature-rainfall-infrastructure interaction terms using the reported model specification, and compared coefficient signs and significance against the published results.", + "environment": { + "runtime": "python 3.12.3", + "os": "ubuntu-24.04", + "packages": "pandas 2.2.2, statsmodels 0.14.2, numpy 1.26.4", + "notes": "CI runner with 4 vCPU and 16 GB RAM. Used the author-supplied supplementary data file directly." + }, + "input_hash": "sha256:a1b2c3d4e5f6789012345678abcdef0123456789abcdef0123456789abcdef01", + "result": "partially-confirmed", + "divergence": "Interaction terms for temperature and rainfall replicated within reported confidence intervals. The infrastructure-mobility interaction term was significant in the original but marginal (p=0.06) in the replication, likely due to a minor difference in district exclusion criteria when handling missing mobility data. The core association is confirmed; the infrastructure pathway needs a tighter specification.", + "confidence": "medium" +} diff --git a/problem-packs/education/skills-training-youth-employment-global/evidence.json b/problem-packs/education/skills-training-youth-employment-global/evidence.json index ef3012e..906bffe 100644 --- a/problem-packs/education/skills-training-youth-employment-global/evidence.json +++ b/problem-packs/education/skills-training-youth-employment-global/evidence.json @@ -50,9 +50,9 @@ }, "source_date": "2023-09-01", "access_date": "2026-06-18", - "method": "Reviewed training-data landscape assessments, provider-registry documentation, and employment-outcome tracking rates across LMICs.", + "method": "Reviewed training-data assessments, provider-registry documentation, and employment-outcome tracking rates across LMICs.", "limitations": [ - "The originally cited 'Skills training data systems landscape assessment' could not be confirmed to exist; re-sourced to the World Bank/ILO/UNESCO TVET systems report.", + "The originally cited 'Skills training data systems assessment' could not be confirmed to exist; re-sourced to the World Bank/ILO/UNESCO TVET systems report.", "This report describes TVET system fragmentation broadly; it does not quantify provider-registry coverage country by country.", "A dedicated source is still needed to substantiate the 'no single dataset' claim for any specific country." ], diff --git a/problem-packs/education/skills-training-youth-employment-global/evidence.md b/problem-packs/education/skills-training-youth-employment-global/evidence.md index 4f9c04c..fcf65d3 100644 --- a/problem-packs/education/skills-training-youth-employment-global/evidence.md +++ b/problem-packs/education/skills-training-youth-employment-global/evidence.md @@ -16,7 +16,7 @@ Use this source for the best available evidence on training-program employment e ### Training Provider Data Fragmentation -Use this source (World Bank/ILO/UNESCO TVET systems report, 2023) to document that TVET provision in most LMICs is diverse and fragmented across government ministries, private providers, and NGOs, complicating system-wide data on providers and outcomes. (Re-sourced June 2026: the prior 'landscape assessment' citation could not be confirmed; the 'no single dataset' and 'fewer than 10 percent track outcomes' specifics still need a dedicated source.) +Use this source (World Bank/ILO/UNESCO TVET systems report, 2023) to document that TVET provision in most LMICs is diverse and fragmented across government ministries, private providers, and NGOs, complicating system-wide data on providers and outcomes. (Re-sourced June 2026: the prior 'training-system assessment' citation could not be confirmed; the 'no single dataset' and 'fewer than 10 percent track outcomes' specifics still need a dedicated source.) ### Informal Sector Training Outcome Gap diff --git a/problem-packs/education/skills-training-youth-employment-global/problem.md b/problem-packs/education/skills-training-youth-employment-global/problem.md index 9be62a6..5787164 100644 --- a/problem-packs/education/skills-training-youth-employment-global/problem.md +++ b/problem-packs/education/skills-training-youth-employment-global/problem.md @@ -12,7 +12,7 @@ Sub-national data on where training programs exist, what skills they teach, who - Verified fact: ILO's 2024 Global Employment Trends for Youth estimates 67 million unemployed youth globally, with youth-to-adult unemployment ratios of 3-4x in many LMICs and skills mismatch identified as a primary barrier to youth employment. - Verified fact: Systematic reviews of youth employment interventions find that vocational training programs increase employment rates by 2-5 percentage points on average, but effects are highly heterogeneous — programs linked to employer demand show 8-12 percentage point gains while standalone programs show near-zero effects. -- Verified fact: Training-provider registries are fragmented across government ministries, NGOs, and private operators in most LMICs — no single dataset captures the full landscape of skills-training availability at sub-national scale. +- Verified fact: Training-provider registries are fragmented across government ministries, NGOs, and private operators in most LMICs — no single dataset captures the full scope of skills-training availability at sub-national scale. - Verified fact: Employment-outcome tracking is conducted by fewer than 10 percent of training programs in most LMIC contexts, making it impossible to systematically assess which programs actually improve employment outcomes. ## Uncertain Areas diff --git a/problem-packs/education/skills-training-youth-employment-global/tasks.json b/problem-packs/education/skills-training-youth-employment-global/tasks.json index 21b7c65..46a29a6 100644 --- a/problem-packs/education/skills-training-youth-employment-global/tasks.json +++ b/problem-packs/education/skills-training-youth-employment-global/tasks.json @@ -30,7 +30,7 @@ "National training-provider registries for target countries", "NGO and private training-program documentation", "Administrative boundary data for district-level aggregation", - "Published training-landscape assessments" + "Published training-system assessments" ], "expected_artifact": "validation.md update with availability methodology and data-fragmentation documentation.", "validation_method": "Replicator confirms provider-data extraction, fragmentation documentation, and coverage-gap assessment for each target country.", diff --git a/problem-packs/education/skills-training-youth-employment-global/validation.md b/problem-packs/education/skills-training-youth-employment-global/validation.md index 43bdff6..f67c400 100644 --- a/problem-packs/education/skills-training-youth-employment-global/validation.md +++ b/problem-packs/education/skills-training-youth-employment-global/validation.md @@ -27,7 +27,7 @@ No training-availability claim may be called actionable until it includes: - Provider-registry data completeness — document which provider types are captured and which are absent. - Coverage-gap documentation — explicitly state where provider data is unavailable. - Training-type differentiation — document which skill categories (vocational, digital, agricultural, soft skills) are mapped. -- Comparison with at least one published training-landscape assessment for a similar LMIC context. +- Comparison with at least one published training-system assessment for a similar LMIC context. ## Employment-Outcome Requirements diff --git a/schemas/AGENTS.md b/schemas/AGENTS.md index f550a72..d0a108a 100644 --- a/schemas/AGENTS.md +++ b/schemas/AGENTS.md @@ -8,9 +8,11 @@ This directory defines machine-checkable protocol truth. If a requirement can be - `task.schema.json`: validates atomic work units and reviewer/risk enums. - `problem.schema.json`: validates pack metadata and canonical file inventory. -- `evidence.schema.json`: validates evidence records. -- `agent-submission.schema.json`: validates structured submissions. +- `evidence.schema.json`: validates evidence records, including computational and model-prediction types. +- `agent-submission.schema.json`: validates structured submissions, including kill_condition. - `review.schema.json`: validates review artifacts. +- `claim.schema.json`: validates persistent claims with a verification lifecycle, evidence links, failure modes, and required reviewers. This is the core protocol object. +- `replication.schema.json`: validates independent replication records with environment, input hash, and divergence tracking. ## Diagrams (Mermaid) @@ -24,11 +26,15 @@ flowchart TD B --> E["evidence.schema.json"] B --> F["agent-submission.schema.json"] B --> G["review.schema.json"] + B --> CL["claim.schema.json"] + B --> R["replication.schema.json"] C --> H["Merge gate"] D --> H E --> H F --> H G --> H + CL --> H + R --> H ``` ### Component Diagram @@ -40,6 +46,8 @@ flowchart LR Evidence["evidence.json"] --> EvidenceSchema["evidence.schema.json"] Submission["agent submission"] --> SubmissionSchema["agent-submission.schema.json"] Review["review artifact"] --> ReviewSchema["review.schema.json"] + Claim["claims.json"] --> ClaimSchema["claim.schema.json"] + Replication["replication.json"] --> ReplicationSchema["replication.schema.json"] ``` ### Sequence Diagram diff --git a/schemas/agent-submission.schema.json b/schemas/agent-submission.schema.json index eb49892..da43826 100644 --- a/schemas/agent-submission.schema.json +++ b/schemas/agent-submission.schema.json @@ -14,6 +14,7 @@ "reproducibility_steps", "assumptions", "failure_modes", + "kill_condition", "confidence", "suggested_next_issue" ], @@ -81,6 +82,11 @@ }, "minItems": 1 }, + "kill_condition": { + "type": "string", + "minLength": 10, + "description": "The specific observation or experiment that would make this claim false. If no observation could falsify the claim, it is not a claim. This is not a disclaimer — it is the falsification criterion. SKILL.md pattern #3 requires this; the schema now enforces it." + }, "confidence": { "type": "string", "description": "Your confidence that the claim accurately represents what the evidence proves at the stated grain. Not confidence that the claim is true in the world. high = strong evidence, clear method, low risk of error. medium = adequate evidence with stated limitations. low = preliminary, needs more evidence or verification. unknown = cannot assess.", diff --git a/schemas/claim.schema.json b/schemas/claim.schema.json new file mode 100644 index 0000000..a6bf2e4 --- /dev/null +++ b/schemas/claim.schema.json @@ -0,0 +1,150 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://open-problem-lab.local/schemas/claim.schema.json", + "title": "Claim", + "description": "A persistent, falsifiable statement with a verification lifecycle. This is the core protocol object: a claim exists independently of any single submission or review, accumulates evidence and failure modes over time, and carries its own kill condition. The burden of proof sits with the submitter; the claim status reflects what has been verified, not what has been asserted.", + "type": "object", + "additionalProperties": false, + "required": [ + "id", + "problem_id", + "claim", + "domain", + "status", + "evidence", + "failure_modes", + "kill_condition", + "review_required", + "safety_level", + "submitter", + "created_date", + "last_updated", + "confidence" + ], + "properties": { + "id": { + "type": "string", + "pattern": "^[a-z0-9]+(?:-[a-z0-9]+)*$", + "description": "Unique identifier for this claim. Use kebab-case. Example: peptide-x-ecoli-inhibition" + }, + "problem_id": { + "type": "string", + "pattern": "^[a-z0-9]+(?:-[a-z0-9]+)*/[a-z0-9]+(?:-[a-z0-9]+)*$", + "description": "The problem pack this claim belongs to. Must match the id field in problem.json. Example: climate-health/malaria-early-warning-africa" + }, + "claim": { + "type": "string", + "minLength": 20, + "description": "One specific, falsifiable statement. State what can be proven, at what grain, with what evidence. A claim that cannot be shown false is not a claim." + }, + "domain": { + "type": "array", + "items": { + "type": "string", + "pattern": "^[a-z0-9]+(?:-[a-z0-9]+)*$" + }, + "minItems": 1, + "uniqueItems": true, + "description": "Domain tags for this claim. Use the same domain vocabulary as problem.json." + }, + "status": { + "type": "string", + "description": "Verification lifecycle. A claim starts as 'unverified' and moves rightward only through verified work. 'falsified' and 'deprecated' are terminal states that prevent further rightward movement.", + "enum": [ + "unverified", + "dry-lab-verified", + "needs-replication", + "replicated", + "accepted", + "field-tested", + "rejected", + "falsified", + "deprecated" + ] + }, + "evidence": { + "type": "array", + "description": "Evidence record IDs from evidence.json that support this claim. Every ID here must exist in the relevant pack's evidence.json. A claim with no evidence is 'unverified' regardless of prose.", + "items": { + "type": "string", + "minLength": 3 + }, + "minItems": 1 + }, + "failure_modes": { + "type": "array", + "description": "Specific ways this claim could be wrong, misleading, or harmful if applied. Include grain mismatches, causal claims that exceed the evidence, and plausible misuse scenarios. 'Results may vary' is not a failure mode.", + "items": { + "type": "string", + "minLength": 5 + }, + "minItems": 1 + }, + "kill_condition": { + "type": "string", + "minLength": 10, + "description": "The specific observation or experiment that would make this claim false. If no observation could falsify the claim, the claim is not scientific and will not be accepted. This is not a disclaimer — it is the falsification criterion." + }, + "review_required": { + "type": "array", + "description": "Reviewer roles required before this claim can advance to 'accepted'. At least one is required. For high safety_level claims, red-team-reviewer is required.", + "items": { + "type": "string", + "enum": [ + "domain-reviewer", + "red-team-reviewer", + "field-reality-reviewer", + "replicator" + ] + }, + "minItems": 1, + "uniqueItems": true + }, + "safety_level": { + "type": "string", + "description": "Risk level for this claim. low = framing, literature summaries. medium = data cleaning, ranking, models. high = operational recommendations, forecasts, resource allocation.", + "enum": ["low", "medium", "high"] + }, + "submitter": { + "type": "string", + "minLength": 2, + "description": "The named submitter (GitHub username, agent identifier, or team name). Claims are owned. Anonymous claims are not accepted." + }, + "created_date": { + "type": "string", + "format": "date", + "description": "Date the claim was first submitted in ISO 8601 format (YYYY-MM-DD)." + }, + "last_updated": { + "type": "string", + "format": "date", + "description": "Date of the last status change or evidence update in ISO 8601 format (YYYY-MM-DD)." + }, + "accepted_date": { + "type": "string", + "format": "date", + "description": "Date the claim reached 'accepted' status, if applicable. Absent means not yet accepted." + }, + "replication_records": { + "type": "array", + "description": "Replication record IDs from replication.json that reference this claim. Absent means no replication has been attempted.", + "items": { + "type": "string", + "minLength": 3 + } + }, + "confidence": { + "type": "string", + "description": "Submitter confidence that the claim accurately represents what the evidence proves at the stated grain. Not confidence that the claim is true in the world.", + "enum": ["high", "medium", "low", "unknown"] + }, + "limitations": { + "type": "array", + "description": "What the evidence cannot prove. Be specific: name the grain, geography, time period, or causal claims that exceed what the evidence supports. Distinct from failure_modes: limitations describe evidence scope, failure_modes describe how the claim could cause harm.", + "items": { + "type": "string", + "minLength": 5 + } + } + } +} diff --git a/schemas/evidence.schema.json b/schemas/evidence.schema.json index 6161535..9d8d264 100644 --- a/schemas/evidence.schema.json +++ b/schemas/evidence.schema.json @@ -43,7 +43,10 @@ "field-report", "expert-review", "replication", - "negative-result" + "negative-result", + "model-prediction", + "computational-analysis", + "wet-lab-confirmation" ] }, "source": { @@ -61,6 +64,15 @@ "type": "string", "format": "uri", "description": "Direct, stable URL to the source. Prefer DOIs, archived URLs, or official permanent pages over search results or blog posts." + }, + "doi": { + "type": "string", + "description": "The DOI of the source, if available. Providing it insulates the record against link rot — DOIs resolve even when publisher URLs change. Example: 10.1038/s41467-023-43954-0" + }, + "archive_url": { + "type": "string", + "format": "uri", + "description": "A permanent archive snapshot (Wayback Machine, archive.today, or equivalent). Providing one insulates the record against future source deletion or content drift." } } }, diff --git a/schemas/replication.schema.json b/schemas/replication.schema.json new file mode 100644 index 0000000..ec05707 --- /dev/null +++ b/schemas/replication.schema.json @@ -0,0 +1,92 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://open-problem-lab.local/schemas/replication.schema.json", + "title": "Replication Record", + "description": "An independent replication of a prior claim. Replication is the strongest form of verification in this protocol. A claim that reaches 'replicated' status must have at least one replication record confirming the original result. A replication that fails to reproduce the result must record what was attempted and what diverged.", + "type": "object", + "additionalProperties": false, + "required": [ + "id", + "claim_id", + "replicator", + "replication_date", + "method", + "environment", + "input_hash", + "result", + "divergence", + "confidence" + ], + "properties": { + "id": { + "type": "string", + "pattern": "^[a-z0-9]+(?:-[a-z0-9]+)*$", + "description": "Unique identifier for this replication record. Use kebab-case. Example: peptide-x-ecoli-replication-1" + }, + "claim_id": { + "type": "string", + "pattern": "^[a-z0-9]+(?:-[a-z0-9]+)*$", + "description": "The id field of the claim this replication targets. Must exist in the relevant pack's claims.json." + }, + "replicator": { + "type": "string", + "minLength": 2, + "description": "The named replicator (GitHub username, agent identifier, or team name). Must differ from the original submitter of the claim. Self-replication is not replication." + }, + "replication_date": { + "type": "string", + "format": "date", + "description": "Date the replication was completed in ISO 8601 format (YYYY-MM-DD)." + }, + "method": { + "type": "string", + "minLength": 20, + "description": "The exact procedure used to replicate. A reader should be able to follow these steps and arrive at the same result (or explain the divergence). Include tool versions and any deviations from the original method." + }, + "environment": { + "type": "object", + "additionalProperties": false, + "required": ["runtime", "os"], + "properties": { + "runtime": { + "type": "string", + "minLength": 3, + "description": "Language runtime and version. Example: 'node 22.4.0' or 'python 3.12.3'" + }, + "os": { + "type": "string", + "minLength": 3, + "description": "Operating system or CI runner. Example: 'ubuntu-24.04' or 'macos-15'" + }, + "packages": { + "type": "string", + "description": "Key package versions that affect the result. Example: 'ajv 8.20.0, prettier 3.8.3'" + }, + "notes": { + "type": "string", + "description": "Any environment detail that could affect reproducibility. Example: 'CI runner with 4 vCPU and 16 GB RAM'" + } + }, + "description": "The execution environment. Stating it lets a future replicator detect whether a divergence is environmental rather than methodological." + }, + "input_hash": { + "type": "string", + "minLength": 8, + "description": "A hash (sha256 or equivalent) of the input data or source files used. If the input differs from the original submission, explain in the divergence field. A null hash means inputs are identical to the original — record that explicitly." + }, + "result": { + "type": "string", + "description": "The outcome of the replication.", + "enum": ["confirmed", "partially-confirmed", "failed", "inconclusive"] + }, + "divergence": { + "type": "string", + "description": "How the replicated result differs from the original. If the result is 'confirmed', state 'No divergence detected.' If the result differs, be specific about what changed and by how much. A replication with no divergence statement is incomplete." + }, + "confidence": { + "type": "string", + "description": "Replicator confidence in their own replication. Not confidence in the original claim.", + "enum": ["high", "medium", "low", "unknown"] + } + } +} diff --git a/scripts/validate-repo.mjs b/scripts/validate-repo.mjs index d90ed03..079236b 100644 --- a/scripts/validate-repo.mjs +++ b/scripts/validate-repo.mjs @@ -90,7 +90,9 @@ const compileSchemas = async () => ({ task: ajv.compile(await loadSchema("task")), evidence: ajv.compile(await loadSchema("evidence")), agentSubmission: ajv.compile(await loadSchema("agent-submission")), - review: ajv.compile(await loadSchema("review")) + review: ajv.compile(await loadSchema("review")), + claim: ajv.compile(await loadSchema("claim")), + replication: ajv.compile(await loadSchema("replication")) }); const validateProblemPacks = async (schemas) => { @@ -167,6 +169,16 @@ const validateExamples = async (schemas) => { await readJson("examples/review.example.json"), "examples/review.example.json" ); + validateJson( + schemas.claim, + await readJson("examples/claim.example.json"), + "examples/claim.example.json" + ); + validateJson( + schemas.replication, + await readJson("examples/replication.example.json"), + "examples/replication.example.json" + ); }; const validateIssueTemplates = async () => {