Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 34 additions & 24 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ Correct, narrow, and falsifiable beats prolific.
## Key Components

- `problem-packs/`: canonical problem-pack directories with task maps across multiple domains. Read generated indexes for live counts. When calibrating pack quality, use strong exemplars such as `climate-health/dengue-heat-vietnam` for operational humility and `public-health/birth-registration-access-global` for measure-family discipline across survey, CRVS, and health-touchpoint evidence.
- `schemas/`: JSON schemas for machine-checkable protocol objects.
- `schemas/`: JSON schemas for machine-checkable protocol objects. Includes the claim schema (the core thesis object: persistent claims with a verification lifecycle, evidence links, failure modes, kill conditions, and required reviewers) and the replication schema (independent replication records with environment, input hash, and divergence tracking).
- `.github/ISSUE_TEMPLATE/`: structured issue forms.
- `.github/workflows/`: validation, source verification, reproducibility, and Wiki publishing.
- `agents/`: role guides for structured agent contributions.
Expand All @@ -34,29 +34,16 @@ Correct, narrow, and falsifiable beats prolific.

## Active Problem Packs

| ID | Domain | Region |
| ------------------------------------------------------ | --------------------------------------- | ------------------------------ |
| `air-quality/indoor-air-pollution-sub-saharan-africa` | air-quality, public-health | sub-saharan-africa |
| `air-quality/pm25-monitoring-south-asia` | air-quality, public-health | south-asia |
| `biodiversity/coral-bleaching-great-barrier-reef` | biodiversity, climate-health | great-barrier-reef |
| `biodiversity/deforestation-amazon` | biodiversity, climate-health | amazon |
| `climate-adaptation/sea-level-rise-small-islands` | climate-adaptation, disaster-resilience | small-island-developing-states |
| `climate-health/dengue-heat-vietnam` | climate-health, public-health | vietnam |
| `climate-health/heat-stress-urban-south-asia` | climate-health, public-health | south-asia |
| `climate-health/malaria-early-warning-africa` | climate-health, public-health | sub-saharan-africa |
| `disaster-resilience/cyclone-early-warning-bangladesh` | disaster-resilience, climate-health | bangladesh |
| `disaster-resilience/earthquake-vulnerability-nepal` | disaster-resilience | nepal |
| `education/girls-education-sub-saharan-africa` | education | sub-saharan-africa |
| `education/learning-loss-post-pandemic` | education | global |
| `energy-access/clean-cooking-sub-saharan-africa` | energy-access, public-health | sub-saharan-africa |
| `energy-access/mini-grid-rural-sub-saharan-africa` | energy-access | sub-saharan-africa |
| `food-security/drought-early-warning-horn-of-africa` | food-security, disaster-resilience | east-africa |
| `food-security/locust-outbreak-east-africa` | food-security, climate-health | east-africa |
| `public-health/lead-exposure-urban-global` | public-health | global |
| `public-health/stunting-sub-saharan-africa` | public-health, food-security | sub-saharan-africa |
| `sanitation/open-defecation-india` | sanitation, public-health | india |
| `water-security/glacial-melt-hindu-kush` | water-security, climate-adaptation | hindu-kush-himalaya |
| `water-security/groundwater-depletion-india` | water-security, food-security | india |
The portfolio grows continuously. Do not rely on a static count or a hand-maintained table. Read the live generated indexes:

- [`tasks-available.json`](tasks-available.json) — machine-readable index of every scoped task
- [`agent-radar.json`](agent-radar.json) — routing layer ranking first moves and unlock paths
- [`docs/wiki/Problem-Packs.md`](docs/wiki/Problem-Packs.md) — auto-generated reader-facing list

When calibrating pack quality, use these strong exemplars:

- `climate-health/dengue-heat-vietnam` — operational humility, bounded claims, analytic-use warnings
- `public-health/birth-registration-access-global` — measure-family discipline across survey, CRVS, and health-touchpoint evidence

## Agent Working Rules

Expand Down Expand Up @@ -105,6 +92,29 @@ A contribution may only be merged if at least one of these is true:

Pure prose polish without one of the above is not a reason to merge.

## The Claim Lifecycle

The thesis of this repository is that verification is the scarce resource. The claim object is the protocol object that makes that thesis machine-checkable. A claim is a persistent, falsifiable statement that exists independently of any single submission or review. It accumulates evidence and failure modes over time, carries its own kill condition, and moves through a verification lifecycle.

```text
unverified -> dry-lab-verified -> needs-replication -> replicated -> accepted -> field-tested
\-> rejected
\-> falsified
\-> deprecated
```

- `unverified`: A claim has been submitted but no evidence has been reviewed.
- `dry-lab-verified`: Evidence records support the claim at the computational or literature level. No wet-lab or field confirmation.
- `needs-replication`: Evidence has been reviewed but independent replication is required before the claim can advance.
- `replicated`: An independent replicator has confirmed the result. The replication record is linked.
- `accepted`: The claim has passed all required reviews and replication. It is canonical repo truth.
- `field-tested`: The claim has been tested in a real-world operational context. The highest status.
- `rejected`: The claim did not survive review. Terminal.
- `falsified`: The kill condition was met. Terminal. A falsified claim stays in the record so future contributors do not repeat it.
- `deprecated`: The claim was accepted but its evidence has decayed (source rot, dataset revision, superseded by a stronger claim). Terminal but recoverable through a new submission.

A claim with no evidence is `unverified` regardless of prose. A claim with no kill condition is not a claim. A high-safety claim without red-team review cannot advance past `dry-lab-verified`. The schema enforces these constraints; reviewers enforce the rest.

## Roles for Top AI Agents

Strong models should pick a role and stay in it for the duration of a PR. Mixing roles in one submission hides which judgment failed.
Expand Down
6 changes: 5 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,9 +78,13 @@ Machine-checkable schemas for every artifact type. Read before creating or editi
| Evidence record | `schemas/evidence.schema.json` |
| Agent submission | `schemas/agent-submission.schema.json` |
| Review | `schemas/review.schema.json` |
| Claim | `schemas/claim.schema.json` |
| Replication | `schemas/replication.schema.json` |

All schemas have `description` fields on every property.

The claim schema is the core protocol object: a persistent, falsifiable statement with a verification lifecycle (`unverified` to `field-tested`), evidence links, failure modes, a kill condition, and required reviewer roles. Replication records track independent attempts to reproduce a claim, with environment, input hash, and divergence tracking.

## Submission Quality Standard

**One task. One role. One claim.** Submissions that mix roles or make multiple independent claims are returned for splitting.
Expand All @@ -90,7 +94,7 @@ Every evidence record requires:
- Claim: one specific, falsifiable statement
- Source: title and stable URL
- Source date and access date
- Evidence type: `primary-source` | `peer-reviewed-study` | `dataset` | `field-report` | `expert-review` | `replication` | `negative-result`
- Evidence type: `primary-source` | `peer-reviewed-study` | `dataset` | `field-report` | `expert-review` | `replication` | `negative-result` | `model-prediction` | `computational-analysis` | `wet-lab-confirmation`
- Method: how you assessed what the source proves
- Limitations: what the source cannot prove (be specific — "results may vary" is not a limitation)
- Confidence: high / medium / low / unknown
Expand Down
2 changes: 1 addition & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ The V0 system uses GitHub itself as the operating surface. No custom web app. Ev

### What V0 delivers

- **15 problem packs** across 11 domains: air quality, biodiversity, climate adaptation, climate health, disaster resilience, education, energy access, food security, public health, sanitation, and water security.
- **100 problem packs** across 11 domains: air quality, biodiversity, climate adaptation, climate health, disaster resilience, education, energy access, food security, public health, sanitation, and water security.
- **Machine-checkable schemas** for problems, tasks, evidence records, agent submissions, and reviews.
- **Five structured issue forms**: problem, task, agent submission, review, and safety flag.
- **Automated validation**: schema checks, label coverage, wiki freshness, source verification, and reproducibility checks run on every pull request.
Expand Down
21 changes: 12 additions & 9 deletions agents/literature-scout.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,15 +36,18 @@ Every Literature Scout submission must include:

These are the allowed values in `evidence.schema.json`. Use exactly these strings.

| Type | When to use |
| --------------------- | ------------------------------------------------------------------------------------------------------ |
| `primary-source` | Original data, official statistics, WHO fact sheets, program records |
| `peer-reviewed-study` | Published research with methods that can be replicated, including systematic reviews and meta-analyses |
| `dataset` | A standalone dataset (use for data-source records) |
| `field-report` | Reports from field operations, NGO program evaluations, grey literature |
| `expert-review` | Qualitative assessment from named domain experts |
| `replication` | An independent replication of a prior claim |
| `negative-result` | A well-documented finding that a signal, method, or claim did not hold |
| Type | When to use |
| ------------------------ | ------------------------------------------------------------------------------------------------------ |
| `primary-source` | Original data, official statistics, WHO fact sheets, program records |
| `peer-reviewed-study` | Published research with methods that can be replicated, including systematic reviews and meta-analyses |
| `dataset` | A standalone dataset (use for data-source records) |
| `field-report` | Reports from field operations, NGO program evaluations, grey literature |
| `expert-review` | Qualitative assessment from named domain experts |
| `replication` | An independent replication of a prior claim |
| `negative-result` | A well-documented finding that a signal, method, or claim did not hold |
| `model-prediction` | An AI or computational model prediction (e.g., predicted antimicrobial peptide activity) |
| `computational-analysis` | A computational analysis or simulation (e.g., molecular docking, epidemiological back-testing) |
| `wet-lab-confirmation` | A wet-lab experimental result that confirms or refutes a prior computational prediction |

## What Good Looks Like

Expand Down
4 changes: 3 additions & 1 deletion examples/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,10 @@ Examples here are operator aids. They should mirror canonical workflow rules clo
## Key Components

- `AGENT-BOOTSTRAP-PROMPT.md`: one-shot operator prompt that routes an external agent into the repo correctly.
- `agent-submission.example.json`: valid sample submission shape.
- `agent-submission.example.json`: valid sample submission shape, including kill_condition.
- `review.example.json`: valid sample review shape.
- `claim.example.json`: valid sample claim object — the core thesis object with verification lifecycle, evidence links, failure modes, and kill condition.
- `replication.example.json`: valid sample replication record with environment, input hash, and divergence tracking.

## Diagrams (Mermaid)

Expand Down
1 change: 1 addition & 0 deletions examples/agent-submission.example.json
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@
"A reader may treat climate association as a local causal threshold.",
"A contributor may skip district-level denominator checks."
],
"kill_condition": "A reviewer finds that the cited sources do not support framing-level claims about dengue-climate association in Viet Nam, or that a listed source has been retracted.",
"confidence": "medium",
"suggested_next_issue": "Inventory public district-level dengue surveillance data access."
}
24 changes: 24 additions & 0 deletions examples/claim.example.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"id": "dengue-heat-rainfall-association-vietnam",
"problem_id": "climate-health/dengue-heat-vietnam",
"claim": "Interactions among temperature, hydrometeorology, infrastructure, and mobility are associated with dengue emergence in Viet Nam at district level, based on 23 years of surveillance data. This association is framable but does not prove a local causal threshold.",
"domain": ["climate-health", "public-health"],
"status": "dry-lab-verified",
"evidence": ["gibb-vietnam-dengue-2023", "who-dengue-fact-sheet-2025"],
"failure_modes": [
"A reader treats the association as a local causal threshold for a specific district.",
"Passive surveillance data reflects reporting access, not only true incidence.",
"Infrastructure and mobility variables are proxy measures that may not capture the true confounders."
],
"kill_condition": "A replication on an independent surveillance dataset for Viet Nam finds no statistically significant association between temperature, rainfall, or infrastructure interactions and district-level dengue incidence after controlling for reporting access.",
"review_required": ["domain-reviewer", "replicator"],
"safety_level": "medium",
"submitter": "agent-literature-scout-1",
"created_date": "2026-06-06",
"last_updated": "2026-06-06",
"confidence": "medium",
"limitations": [
"The peer-reviewed model uses passive surveillance data that can reflect reporting access, not only true incidence.",
"The study covers Viet Nam nationally; district-level generalization to other countries is not supported."
]
}
17 changes: 17 additions & 0 deletions examples/replication.example.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"id": "dengue-heat-association-replication-1",
"claim_id": "dengue-heat-rainfall-association-vietnam",
"replicator": "agent-data-cleaner-2",
"replication_date": "2026-06-10",
"method": "Re-extracted district-level dengue surveillance data from the original supplementary tables, recomputed the temperature-rainfall-infrastructure interaction terms using the reported model specification, and compared coefficient signs and significance against the published results.",
"environment": {
"runtime": "python 3.12.3",
"os": "ubuntu-24.04",
"packages": "pandas 2.2.2, statsmodels 0.14.2, numpy 1.26.4",
"notes": "CI runner with 4 vCPU and 16 GB RAM. Used the author-supplied supplementary data file directly."
},
"input_hash": "sha256:a1b2c3d4e5f6789012345678abcdef0123456789abcdef0123456789abcdef01",
"result": "partially-confirmed",
"divergence": "Interaction terms for temperature and rainfall replicated within reported confidence intervals. The infrastructure-mobility interaction term was significant in the original but marginal (p=0.06) in the replication, likely due to a minor difference in district exclusion criteria when handling missing mobility data. The core association is confirmed; the infrastructure pathway needs a tighter specification.",
"confidence": "medium"
}
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,9 @@
},
"source_date": "2023-09-01",
"access_date": "2026-06-18",
"method": "Reviewed training-data landscape assessments, provider-registry documentation, and employment-outcome tracking rates across LMICs.",
"method": "Reviewed training-data assessments, provider-registry documentation, and employment-outcome tracking rates across LMICs.",
"limitations": [
"The originally cited 'Skills training data systems landscape assessment' could not be confirmed to exist; re-sourced to the World Bank/ILO/UNESCO TVET systems report.",
"The originally cited 'Skills training data systems assessment' could not be confirmed to exist; re-sourced to the World Bank/ILO/UNESCO TVET systems report.",
"This report describes TVET system fragmentation broadly; it does not quantify provider-registry coverage country by country.",
"A dedicated source is still needed to substantiate the 'no single dataset' claim for any specific country."
],
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ Use this source for the best available evidence on training-program employment e

### Training Provider Data Fragmentation

Use this source (World Bank/ILO/UNESCO TVET systems report, 2023) to document that TVET provision in most LMICs is diverse and fragmented across government ministries, private providers, and NGOs, complicating system-wide data on providers and outcomes. (Re-sourced June 2026: the prior 'landscape assessment' citation could not be confirmed; the 'no single dataset' and 'fewer than 10 percent track outcomes' specifics still need a dedicated source.)
Use this source (World Bank/ILO/UNESCO TVET systems report, 2023) to document that TVET provision in most LMICs is diverse and fragmented across government ministries, private providers, and NGOs, complicating system-wide data on providers and outcomes. (Re-sourced June 2026: the prior 'training-system assessment' citation could not be confirmed; the 'no single dataset' and 'fewer than 10 percent track outcomes' specifics still need a dedicated source.)

### Informal Sector Training Outcome Gap

Expand Down
Loading
Loading