Skip to content

Fix the Ouroboros CUDA benchmark protocol and budget accounting - #5

Draft
ouroboros-agent wants to merge 10 commits into
AlexWortega:rsifrom
ouroboros-agent:fix/report-budget-exposure
Draft

Fix the Ouroboros CUDA benchmark protocol and budget accounting#5
ouroboros-agent wants to merge 10 commits into
AlexWortega:rsifrom
ouroboros-agent:fix/report-budget-exposure

Conversation

@ouroboros-agent

Copy link
Copy Markdown

This fixes the Ouroboros arm of the CUDA benchmark and defines a separately labeled H200 cohort. Every acting and review slot is pinned to the arm model, task acceptance remains automatic with advisory enforcement, and the benchmark runs as a strict single root with schedule_subagent and plan_task disabled. Each run records the exact Ouroboros source, image, applied settings, provider routing, raw usage ledger, and shared-GPU evidence.

The evaluator runs each submitted candidate twice from a read-only mount and binds both CUDA receipts to the candidate and evaluator process. Non-final provider usage keeps the independent capability score while charging the full $50 cell cap. The report shows reported cost, cost finality, and charged budget exposure separately, and fails closed when exposure authority is missing or malformed.

A local corrected H200 campaign completed all 20 preregistered cells. GPT and Kimi each passed 9/10, for 18/20 overall, with $72.696679 of charged exposure. These are diagnostic results and are kept separate from the published RTX A6000 table because the hardware, prompts, evaluator, agent configuration, budget, runtime version, and replicate count differ.

Validation on this branch: 95 harness tests pass. Ruff F checks for harnessblog and its tests, Python compileall, and git diff --check also pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants