Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
234 changes: 234 additions & 0 deletions .github/workflows/golden-pr-evals.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,234 @@
# Golden PR evals (products/tasks/evals/golden_prs): run a coding agent at the commit before a
# human-written PR, prompt it with the PR description, and score its diff against the merged PR.
#
# Manual dispatch only. Every case runs an agent for up to half an hour and spends LLM credits,
# and the score is a directional signal for comparing models, not a merge gate.
# If a required secret is missing the job skips green with a warning, so a fork or a
# not-yet-configured environment is unaffected.
name: Golden PR Evals

on:
workflow_dispatch:
inputs:
Comment on lines +11 to +12

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review

[must_fix] Restrict provider secrets to trusted workflow refs

must_fix security

Issue description

This workflow can be dispatched against a non-default branch, and GitHub runs the workflow version from that ref. A modified workflow on that branch could read and exfiltrate the repository secrets used by the evaluation, even before it starts the agent. The workflow_dispatch branch selector and ref behavior make this a separate exposure path from the agent prompt. (docs.github.com)

Why we think it's a valid issue
  • Checked: Reviewed .github/workflows/golden-pr-evals.yml from dispatch through the evaluation job, including its secret checks and provider-key use.
  • Found: The workflow allows workflow_dispatch at .github/workflows/golden-pr-evals.yml:11. The evaluate job reads repository secrets at .github/workflows/golden-pr-evals.yml:92-93 and passes both provider keys to the eval step at .github/workflows/golden-pr-evals.yml:153-154. The job does not declare a protected environment.
  • Impact: A modified workflow on a dispatched ref can change its steps and access the repository secrets. persist-credentials: false does not restrict access to these explicitly injected provider keys. Keeping the keys only as environment secrets, with deployment limited to trusted refs, prevents a branch from accessing them without the required environment binding and approval.
  • Priority: must_fix is appropriate because this exposes paid provider credentials to workflow code on an untrusted ref.
Suggested fix

Store the provider keys in a protected GitHub environment that only allows the trusted default branch, and bind the evaluate job to that environment. Add required approval if appropriate. A guard in this workflow alone is not sufficient because a modified branch can remove it.

Prompt to fix with AI (copy-paste)
## Context
@.github/workflows/golden-pr-evals.yml#L11-12

<issue_description>
This workflow can be dispatched against a non-default branch, and GitHub runs the workflow version from that ref. A modified workflow on that branch could read and exfiltrate the repository secrets used by the evaluation, even before it starts the agent. The `workflow_dispatch` branch selector and ref behavior make this a separate exposure path from the agent prompt. ([docs.github.com](https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows?utm_source=openai))
</issue_description>

<issue_validation>
- **Checked:** Reviewed `.github/workflows/golden-pr-evals.yml` from dispatch through the evaluation job, including its secret checks and provider-key use.
- **Found:** The workflow allows `workflow_dispatch` at `.github/workflows/golden-pr-evals.yml:11`. The `evaluate` job reads repository secrets at `.github/workflows/golden-pr-evals.yml:92-93` and passes both provider keys to the eval step at `.github/workflows/golden-pr-evals.yml:153-154`. The job does not declare a protected environment.
- **Impact:** A modified workflow on a dispatched ref can change its steps and access the repository secrets. `persist-credentials: false` does not restrict access to these explicitly injected provider keys. Keeping the keys only as environment secrets, with deployment limited to trusted refs, prevents a branch from accessing them without the required environment binding and approval.
- **Priority:** `must_fix` is appropriate because this exposes paid provider credentials to workflow code on an untrusted ref.
</issue_validation>

## Task
Investigate the issue and solve it

<potential_solution>
Store the provider keys in a protected GitHub environment that only allows the trusted default branch, and bind the `evaluate` job to that environment. Add required approval if appropriate. A guard in this workflow alone is not sufficient because a modified branch can remove it.
</potential_solution>

runtime:
description: 'Agent CLI that does the work'
type: choice
options:
- claude
- codex
default: claude
model:
description: 'Agent model. Leave empty for the runtime default.'
type: string
default: ''
prs:
description: 'PR numbers to run, comma separated. "all" runs the whole golden set.'
type: string
default: 'all'
judge_model:
description: 'Anthropic model that grades each result'
type: string
default: 'claude-opus-5'
case_timeout_minutes:
description: 'Minutes the agent gets per PR'
type: string
default: '30'

permissions:
contents: read

env:
RESULTS_DIR: /tmp/golden-pr-results

jobs:
plan:
name: Select golden PRs
runs-on: ubuntu-24.04
timeout-minutes: 5
outputs:
prs: ${{ steps.select.outputs.prs }}
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
sparse-checkout: products/tasks/evals/golden_prs/golden_prs.json
sparse-checkout-cone-mode: false

- name: Resolve the PR list
id: select
env:
# Through env, not inline ${{ }}, so a dispatch input can never inject into the shell.
PRS: ${{ inputs.prs }}
run: |
prs=$(python3 - <<'EOF'
import json, os
golden = [entry["number"] for entry in json.load(open("products/tasks/evals/golden_prs/golden_prs.json"))]
wanted = os.environ["PRS"].strip()
selected = golden if wanted == "all" else [int(n) for n in wanted.split(",") if n.strip()]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review

[should_fix] Reject an empty PR selection

should_fix bug

Issue description

A blank or comma-only prs input produces an empty list. The evaluate job then skips, but the report job still runs and has no results to summarize. This can make an invalid dispatch look like a completed run with no scores.

Why we think it's a valid issue
  • Checked: .github/workflows/golden-pr-evals.yml:63-74 filters blank PR tokens but does not reject an empty selection. The plan step can therefore output [].
  • Found: .github/workflows/golden-pr-evals.yml:79-80 skips evaluate for [], but report at .github/workflows/golden-pr-evals.yml:191-195 has no corresponding selection guard. products/tasks/evals/golden_prs/__main__.py:92-94 returns No results found.\n for empty results, and products/tasks/evals/golden_prs/__main__.py:142-144 exits successfully after printing it.
  • Impact: A blank or comma-only dispatch can finish successfully without scores, so the run does not clearly signal that no PRs were evaluated. Rejecting the empty selection or skipping the report addresses this reachable input case.
Suggested fix

Fail the plan step with a clear error when selected is empty, or explicitly skip the report job when no PRs were selected.

Prompt to fix with AI (copy-paste)
## Context
@.github/workflows/golden-pr-evals.yml#L67

<issue_description>
A blank or comma-only `prs` input produces an empty list. The evaluate job then skips, but the report job still runs and has no results to summarize. This can make an invalid dispatch look like a completed run with no scores.
</issue_description>

<issue_validation>
- **Checked:** `.github/workflows/golden-pr-evals.yml:63-74` filters blank PR tokens but does not reject an empty selection. The plan step can therefore output `[]`.
- **Found:** `.github/workflows/golden-pr-evals.yml:79-80` skips `evaluate` for `[]`, but `report` at `.github/workflows/golden-pr-evals.yml:191-195` has no corresponding selection guard. `products/tasks/evals/golden_prs/__main__.py:92-94` returns `No results found.\n` for empty results, and `products/tasks/evals/golden_prs/__main__.py:142-144` exits successfully after printing it.
- **Impact:** A blank or comma-only dispatch can finish successfully without scores, so the run does not clearly signal that no PRs were evaluated. Rejecting the empty selection or skipping the report addresses this reachable input case.
</issue_validation>

## Task
Investigate the issue and solve it

<potential_solution>
Fail the plan step with a clear error when `selected` is empty, or explicitly skip the report job when no PRs were selected.
</potential_solution>

unknown = sorted(set(selected) - set(golden))
Comment thread
pauldambra marked this conversation as resolved.
if unknown:
raise SystemExit(f"Not in the golden set: {unknown}")
print(json.dumps(selected))
EOF
)
echo "prs=$prs" >> "$GITHUB_OUTPUT"

evaluate:
name: PR ${{ matrix.pr }}
needs: plan
# A matrix that expands to zero cells fails its dependents and posts no check run.
if: needs.plan.outputs.prs != '[]'
runs-on: ubuntu-24.04
timeout-minutes: 60
Comment thread
pauldambra marked this conversation as resolved.
strategy:
fail-fast: false
max-parallel: 4
matrix:
pr: ${{ fromJSON(needs.plan.outputs.prs) }}
steps:
- name: Check required secrets
id: gate
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
RUNTIME: ${{ inputs.runtime }}
run: |
missing=""
[ -n "$ANTHROPIC_API_KEY" ] || missing="$missing ANTHROPIC_API_KEY"
if [ "$RUNTIME" = "codex" ] && [ -z "$OPENAI_API_KEY" ]; then
missing="$missing OPENAI_API_KEY"
fi
if [ -n "$missing" ]; then
echo "enabled=false" >> "$GITHUB_OUTPUT"
echo "::warning::Missing secrets:$missing; skipping the golden PR eval"
else
echo "enabled=true" >> "$GITHUB_OUTPUT"
fi

- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
Comment thread
pauldambra marked this conversation as resolved.
if: steps.gate.outputs.enabled == 'true'
with:
# The agent runs unsandboxed in this checkout; it has no reason to read a persisted token.
persist-credentials: false
# The runner fetches each golden commit itself; the checkout only needs the eval package
# and the files the toolchain setup steps read.
sparse-checkout: |
.nvmrc
pyproject.toml
products/__init__.py
products/tasks/__init__.py
products/tasks/evals
sparse-checkout-cone-mode: false
Comment thread
pauldambra marked this conversation as resolved.

- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
if: steps.gate.outputs.enabled == 'true'
with:
node-version-file: .nvmrc

- name: Install the agent CLI
if: steps.gate.outputs.enabled == 'true'
env:
RUNTIME: ${{ inputs.runtime }}
run: |
if [ "$RUNTIME" = "codex" ]; then
npm install -g @openai/codex
else
npm install -g @anthropic-ai/claude-code
Comment on lines +134 to +136

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review

[should_fix] Pin the agent CLI before exposing credentials

should_fix security

Issue description

These commands install the latest CLI version on every run. The eval later gives the selected provider key to the agent process. If a registry release is compromised or changes unexpectedly, the installed CLI can read or send that key. This also makes runs less reproducible.

Why we think it's a valid issue
  • Checked: Traced the CLI installation at .github/workflows/golden-pr-evals.yml:128-138 to the later agent launch and environment filtering in products/tasks/evals/golden_prs/agents.py:39-45.
  • Found: Both npm install -g commands omit a version, so each run installs the package version currently selected by the registry. The launched CLI receives the selected provider key through agent_environment; only the other provider key is removed.
  • Impact: A compromised or malicious newly published CLI version can access and send the provider key when the agent starts. Pinning the CLI version makes upgrades intentional and reduces this exposure. The issue also affects reproducibility because runs can use different CLI versions.
Suggested fix

Install an explicitly reviewed CLI version instead of the unqualified latest version. Keep version updates intentional and controlled.

Prompt to fix with AI (copy-paste)
## Context
@.github/workflows/golden-pr-evals.yml#L134-136

<issue_description>
These commands install the latest CLI version on every run. The eval later gives the selected provider key to the agent process. If a registry release is compromised or changes unexpectedly, the installed CLI can read or send that key. This also makes runs less reproducible.
</issue_description>

<issue_validation>
- **Checked:** Traced the CLI installation at `.github/workflows/golden-pr-evals.yml:128-138` to the later agent launch and environment filtering in `products/tasks/evals/golden_prs/agents.py:39-45`.
- **Found:** Both `npm install -g` commands omit a version, so each run installs the package version currently selected by the registry. The launched CLI receives the selected provider key through `agent_environment`; only the other provider key is removed.
- **Impact:** A compromised or malicious newly published CLI version can access and send the provider key when the agent starts. Pinning the CLI version makes upgrades intentional and reduces this exposure. The issue also affects reproducibility because runs can use different CLI versions.
</issue_validation>

## Task
Investigate the issue and solve it

<potential_solution>
Install an explicitly reviewed CLI version instead of the unqualified latest version. Keep version updates intentional and controlled.
</potential_solution>

fi
"$RUNTIME" --version

- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
if: steps.gate.outputs.enabled == 'true'
with:
python-version-file: 'pyproject.toml'

- name: Install the judge SDK
if: steps.gate.outputs.enabled == 'true'
# Only what the runner imports; a full `uv sync` of the repo is minutes of work the eval never uses.
run: python -m pip install --quiet 'anthropic>=0.80,<1' 'pydantic>=2,<3'

- name: Run the eval
if: steps.gate.outputs.enabled == 'true'
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
Comment thread
pauldambra marked this conversation as resolved.
Comment thread
pauldambra marked this conversation as resolved.
PR: ${{ matrix.pr }}
RUNTIME: ${{ inputs.runtime }}
MODEL: ${{ inputs.model }}
JUDGE_MODEL: ${{ inputs.judge_model }}
CASE_TIMEOUT_MINUTES: ${{ inputs.case_timeout_minutes }}
run: |
# Validated before the arithmetic expansion below: Bash evaluates a command
# substitution embedded in an unchecked arithmetic string, and this step's
# environment holds both provider API keys. The upper bound also keeps setup,
# judging and artifact upload inside the job's 60-minute timeout.
case "$CASE_TIMEOUT_MINUTES" in
''|*[!0-9]*)
echo "::error::case_timeout_minutes must be a positive integer, got '$CASE_TIMEOUT_MINUTES'"
exit 1
;;
esac
if [ "$CASE_TIMEOUT_MINUTES" -lt 1 ] || [ "$CASE_TIMEOUT_MINUTES" -gt 45 ]; then
echo "::error::case_timeout_minutes must be between 1 and 45"
exit 1
fi
python -m products.tasks.evals.golden_prs run \
--pr "$PR" \
--runtime "$RUNTIME" \
${MODEL:+--model "$MODEL"} \
--judge-model "$JUDGE_MODEL" \
--case-timeout "$((10#$CASE_TIMEOUT_MINUTES * 60))" \
--results-dir "$RESULTS_DIR"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review

[should_fix] Keep agent-written files out of score results

should_fix bug

Issue description

The agent inherits RESULTS_DIR and runs without a filesystem sandbox. It can write a fabricated .json file under this directory while it works. The runner later loads every JSON file recursively from the same directory, so that file can add fake scores to the report and distort the mean.

Why we think it's a valid issue
  • Checked: Traced RESULTS_DIR from .github/workflows/golden-pr-evals.yml through run_agent and the artifact/report steps.
  • Found: The workflow passes RESULTS_DIR to the eval command at .github/workflows/golden-pr-evals.yml:181. agent_environment removes only GitHub credentials and the other provider key, so the agent inherits RESULTS_DIR (products/tasks/evals/golden_prs/agents.py:39-41). The workflow uploads that directory (.github/workflows/golden-pr-evals.yml:183-189), then downloads the artifacts into it and runs the report (.github/workflows/golden-pr-evals.yml:212-229). load_results reads every *.json recursively (products/tasks/evals/golden_prs/__main__.py:88-89), and report includes every loaded result in the rows and means (products/tasks/evals/golden_prs/__main__.py:92-108).
  • Impact: An agent-written JSON file can be included in the uploaded artifacts and loaded by the report job, adding a fabricated score to the table and changing its means. This is a reachable score-integrity bug, so the finding meets the bar.
Suggested fix

Isolate the agent from the evaluator's result directory, then have the report read only the expected result files for the selected PRs. A separate container or user for the agent provides a stronger boundary than hiding the path from its environment.

Prompt to fix with AI (copy-paste)
## Context
@.github/workflows/golden-pr-evals.yml#L181

<issue_description>
The agent inherits `RESULTS_DIR` and runs without a filesystem sandbox. It can write a fabricated `.json` file under this directory while it works. The runner later loads every JSON file recursively from the same directory, so that file can add fake scores to the report and distort the mean.
</issue_description>

<issue_validation>
- **Checked:** Traced `RESULTS_DIR` from `.github/workflows/golden-pr-evals.yml` through `run_agent` and the artifact/report steps.
- **Found:** The workflow passes `RESULTS_DIR` to the eval command at `.github/workflows/golden-pr-evals.yml:181`. `agent_environment` removes only GitHub credentials and the other provider key, so the agent inherits `RESULTS_DIR` (`products/tasks/evals/golden_prs/agents.py:39-41`). The workflow uploads that directory (`.github/workflows/golden-pr-evals.yml:183-189`), then downloads the artifacts into it and runs the report (`.github/workflows/golden-pr-evals.yml:212-229`). `load_results` reads every `*.json` recursively (`products/tasks/evals/golden_prs/__main__.py:88-89`), and `report` includes every loaded result in the rows and means (`products/tasks/evals/golden_prs/__main__.py:92-108`).
- **Impact:** An agent-written JSON file can be included in the uploaded artifacts and loaded by the report job, adding a fabricated score to the table and changing its means. This is a reachable score-integrity bug, so the finding meets the bar.
</issue_validation>

## Task
Investigate the issue and solve it

<potential_solution>
Isolate the agent from the evaluator's result directory, then have the report read only the expected result files for the selected PRs. A separate container or user for the agent provides a stronger boundary than hiding the path from its environment.
</potential_solution>


- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
# Keep the diff and agent log of a failed case too; they are what explains the failure.
if: always() && steps.gate.outputs.enabled == 'true'
with:
name: golden-pr-${{ matrix.pr }}
path: ${{ env.RESULTS_DIR }}
if-no-files-found: warn

report:
name: Score summary
needs: evaluate
if: ${{ !cancelled() }}
runs-on: ubuntu-24.04
timeout-minutes: 5
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false
sparse-checkout: |
pyproject.toml
products/__init__.py
products/tasks/__init__.py
products/tasks/evals
sparse-checkout-cone-mode: false

- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
with:
python-version-file: 'pyproject.toml'

- uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7.0.0
with:
pattern: golden-pr-*
path: ${{ env.RESULTS_DIR }}
merge-multiple: true
Comment on lines +212 to +216

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review

[should_fix] Mark incomplete score summaries

should_fix bug

Issue description

If an eval fails before it writes a result file, its artifact can be empty while other PR artifacts still download. The report then calculates scores and the mean from only the PRs that produced results, without marking the table as incomplete. This can make model comparisons misleading.

Why we think it's a valid issue
  • Checked: Traced PR selection, matrix execution, artifact upload and report generation in .github/workflows/golden-pr-evals.yml, plus result loading and mean calculation in products/tasks/evals/golden_prs/__main__.py.
  • Found: .github/workflows/golden-pr-evals.yml:48-49 exposes the selected PR list from plan, but the report job at .github/workflows/golden-pr-evals.yml:191-195 depends only on evaluate. The upload step at .github/workflows/golden-pr-evals.yml:183-189 runs after failure and warns when no result files exist.
  • Found: products/tasks/evals/golden_prs/__main__.py:88-89 loads only JSON files that exist. At products/tasks/evals/golden_prs/__main__.py:103-105, report calculates means across those loaded results without checking them against the selected PRs.
  • Impact: When at least one PR produces a result and another does not, the run summary presents a mean over only the available results without indicating that the selected set is incomplete. This can mislead model comparisons, so the finding meets the bar for a real reporting correctness issue.
Suggested fix

Pass the selected PR list to the report job and compare it with the downloaded results. Mark missing PRs as failed or label the summary as incomplete so readers do not treat a partial mean as a complete run.

Prompt to fix with AI (copy-paste)
## Context
@.github/workflows/golden-pr-evals.yml#L212-216

<issue_description>
If an eval fails before it writes a result file, its artifact can be empty while other PR artifacts still download. The report then calculates scores and the mean from only the PRs that produced results, without marking the table as incomplete. This can make model comparisons misleading.
</issue_description>

<issue_validation>
- **Checked:** Traced PR selection, matrix execution, artifact upload and report generation in `.github/workflows/golden-pr-evals.yml`, plus result loading and mean calculation in `products/tasks/evals/golden_prs/__main__.py`.
- **Found:** `.github/workflows/golden-pr-evals.yml:48-49` exposes the selected PR list from `plan`, but the report job at `.github/workflows/golden-pr-evals.yml:191-195` depends only on `evaluate`. The upload step at `.github/workflows/golden-pr-evals.yml:183-189` runs after failure and warns when no result files exist.
- **Found:** `products/tasks/evals/golden_prs/__main__.py:88-89` loads only JSON files that exist. At `products/tasks/evals/golden_prs/__main__.py:103-105`, `report` calculates means across those loaded results without checking them against the selected PRs.
- **Impact:** When at least one PR produces a result and another does not, the run summary presents a mean over only the available results without indicating that the selected set is incomplete. This can mislead model comparisons, so the finding meets the bar for a real reporting correctness issue.
</issue_validation>

## Task
Investigate the issue and solve it

<potential_solution>
Pass the selected PR list to the report job and compare it with the downloaded results. Mark missing PRs as failed or label the summary as incomplete so readers do not treat a partial mean as a complete run.
</potential_solution>


- name: Publish the score table
env:
RUNTIME: ${{ inputs.runtime }}
MODEL: ${{ inputs.model }}
run: |
set -o pipefail
python3 -m pip install --quiet 'anthropic>=0.80,<1' 'pydantic>=2,<3'
Comment thread
pauldambra marked this conversation as resolved.
{
echo "## Golden PR evals: $RUNTIME ${MODEL:-default model}"
echo ""
python3 -m products.tasks.evals.golden_prs report --results-dir "$RESULTS_DIR"
} | tee -a "$GITHUB_STEP_SUMMARY"
Comment thread
pauldambra marked this conversation as resolved.

- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: golden-pr-results
path: ${{ env.RESULTS_DIR }}
Empty file.
1 change: 1 addition & 0 deletions products/tasks/evals/golden_prs/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
results/
52 changes: 52 additions & 0 deletions products/tasks/evals/golden_prs/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# Golden PR evals

Score how well a coding agent can one-shot a real PostHog pull request.

The golden set is a list of pull requests from 2023 and 2024, written and reviewed by people before coding agents were common.
For each PR, the eval:

1. Checks out the repository at the commit before the PR merged. The agent gets a fresh git repository with only that tree, so it cannot read the merged PR from git history.
2. Sends the PR title and description to the agent as its only prompt.
3. Compares the agent's diff with the merged PR's diff.

## Scores

| Score | What it measures |
| --------- | ----------------------------------------------------------------------------------------------- |
| Files hit | Share of the golden PR's files that the agent also changed. |
| Line F1 | Overlap between the added lines of the two diffs. |
| Judge | An Anthropic model reads the task and both diffs and scores behavioral equivalence from 0 to 1. |

Snapshot files and images are ignored, because an agent cannot regenerate them without the test suite.
The deterministic scores are strict, so a correct change written in a different way scores low on them. Read the judge score and its reasoning together with them.

## Known limits

The agent runs unsandboxed (`--dangerously-skip-permissions` / `--dangerously-bypass-approvals-and-sandbox`), with its working directory as a soft boundary rather than a hard one. It is not run inside a container or namespace, so it could, in principle, read the original checkout that fetched the golden merge commit, or other host state, instead of relying on the PR description alone. Treat a score as a signal for comparing agents and models, not as proof the agent only ever saw the task description.

## Run it

From GitHub, open the **Golden PR Evals** workflow and choose **Run workflow**. Pick the runtime, the model, and the PR numbers. Each PR runs as its own job. The run summary shows the score table, and the artifacts hold every diff, agent log, and score file.

From a devbox, with `claude` or `codex` on `PATH`.
The judge uses `ANTHROPIC_API_KEY` when it is set, and the signed-in `claude` CLI otherwise:

```bash
python -m products.tasks.evals.golden_prs list
python -m products.tasks.evals.golden_prs run --pr 25832 --runtime claude --model claude-opus-5
python -m products.tasks.evals.golden_prs run --runtime codex --model gpt-5.5
python -m products.tasks.evals.golden_prs report --results-dir products/tasks/evals/golden_prs/results/<run>
```

Results land in `products/tasks/evals/golden_prs/results/<timestamp>-<runtime>-<model>/`, which git ignores.

## Add a golden PR

Pick a merged PR that a person wrote, with a description that says what should change. Small, self-contained fixes and features work best. Then append it to `golden_prs.json`:

```bash
gh pr view <number> --json number,title,body,mergedAt,mergeCommit,author \
--jq '{number, author: .author.login, title, merged_at: .mergedAt, merge_commit_sha: .mergeCommit.oid, body}'
```

The tests in `test_golden_prs.py` check that every entry is complete.
Empty file.
Loading
Loading