Skip to content

feat: reproducibility tests — Phase 2 benchmark honesty - #9

Closed
cschanhniem wants to merge 1 commit into
mainfrom
feat/reproducibility-manifest
Closed

cschanhniem wants to merge 1 commit into
mainfrom
feat/reproducibility-manifest

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

Implements Phase 2: Reproducibility from AGENTS.md:

"Another machine can reproduce rankings from the same inputs."

This PR verifies that the pipeline produces identical results across independent runs and that every run generates a cryptographically traceable manifest.

Changes

schemas/run_manifest.schema.json

Added generated_at and input_hashes as required fields (they were being generated but not schema-validated).

tests/test_reproducibility.py (18 tests)

Category Tests
Deterministic ranking Same order, same scores, same selected candidates across 2 independent runs
Manifest generation Generated alongside output; explicit path respected; validates against schema
Manifest contents All required fields present; version matches; UUID format; unique per run
Input hash integrity SHA-256 in manifest matches actual files; candidate and reference paths recorded
Hashing correctness Deterministic, changes with content, matches stdlib, 64-char lowercase hex

What the tests prove

  • Rankings are deterministic (no hidden random state)
  • Manifest is always generated and validates against schema
  • Input file integrity is verifiable via SHA-256
  • Config changes are detectable via config hash
  • Each run has a unique traceability ID

Test plan

  • make test — 55 tests pass
  • ruff check src tests — clean

Phase 2: Reproducibility — rankings from same inputs must be reproducible.

- Update schemas/run_manifest.schema.json: add generated_at and input_hashes
  as required fields (previously absent from schema but present in output)
- Add test_reproducibility.py: 18 tests verifying:
  - Two independent runs produce identical JSONL ranking order
  - Two runs produce identical scores for every candidate
  - Two runs produce identical selected candidate set
  - run_manifest.json is generated alongside ranked.jsonl
  - Explicit manifest_path argument is respected
  - Manifest validates against updated JSON Schema
  - All required fields present (run_id, pipeline_version, config_hash,
    generated_at, inputs, input_hashes, outputs)
  - pipeline_version in manifest matches installed package __version__
  - run_id is valid UUID format
  - Two runs produce different run_ids (unique per run)
  - SHA-256 hashes in manifest match actual files
  - Candidate and reference paths recorded in manifest
  - stable_json_hash() is deterministic for same config
  - stable_json_hash() changes when config changes
  - build_run_manifest() produces correct structure
  - SHA-256 output is 64-char lowercase hex
  - Different file content → different SHA-256
  - file_sha256() matches stdlib hashlib computation
- Add recall_at_k, enrichment_factor, benchmark_summary to evaluate.py
- Fix ruff F401/E741 issues across pipeline.py, test_cli.py, test_pipeline_filters.py
- 55 tests passing, ruff clean
@cschanhniem

Copy link
Copy Markdown
Collaborator Author

Superseded by PR #11 (feat/integrate-all-phases), which merges all Phase 2 + Phase 3 work into a single consolidation PR with 251 tests passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant