-
Notifications
You must be signed in to change notification settings - Fork 3.4k
feat(tasks): add golden PR evals for coding agents #107390
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
f8f203d
77c1895
f7dece8
3530644
874641f
4e39a69
f6eceeb
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,234 @@ | ||
| # Golden PR evals (products/tasks/evals/golden_prs): run a coding agent at the commit before a | ||
| # human-written PR, prompt it with the PR description, and score its diff against the merged PR. | ||
| # | ||
| # Manual dispatch only. Every case runs an agent for up to half an hour and spends LLM credits, | ||
| # and the score is a directional signal for comparing models, not a merge gate. | ||
| # If a required secret is missing the job skips green with a warning, so a fork or a | ||
| # not-yet-configured environment is unaffected. | ||
| name: Golden PR Evals | ||
|
|
||
| on: | ||
| workflow_dispatch: | ||
| inputs: | ||
| runtime: | ||
| description: 'Agent CLI that does the work' | ||
| type: choice | ||
| options: | ||
| - claude | ||
| - codex | ||
| default: claude | ||
| model: | ||
| description: 'Agent model. Leave empty for the runtime default.' | ||
| type: string | ||
| default: '' | ||
| prs: | ||
| description: 'PR numbers to run, comma separated. "all" runs the whole golden set.' | ||
| type: string | ||
| default: 'all' | ||
| judge_model: | ||
| description: 'Anthropic model that grades each result' | ||
| type: string | ||
| default: 'claude-opus-5' | ||
| case_timeout_minutes: | ||
| description: 'Minutes the agent gets per PR' | ||
| type: string | ||
| default: '30' | ||
|
|
||
| permissions: | ||
| contents: read | ||
|
|
||
| env: | ||
| RESULTS_DIR: /tmp/golden-pr-results | ||
|
|
||
| jobs: | ||
| plan: | ||
| name: Select golden PRs | ||
| runs-on: ubuntu-24.04 | ||
| timeout-minutes: 5 | ||
| outputs: | ||
| prs: ${{ steps.select.outputs.prs }} | ||
| steps: | ||
| - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 | ||
| with: | ||
| persist-credentials: false | ||
| sparse-checkout: products/tasks/evals/golden_prs/golden_prs.json | ||
| sparse-checkout-cone-mode: false | ||
|
|
||
| - name: Resolve the PR list | ||
| id: select | ||
| env: | ||
| # Through env, not inline ${{ }}, so a dispatch input can never inject into the shell. | ||
| PRS: ${{ inputs.prs }} | ||
| run: | | ||
| prs=$(python3 - <<'EOF' | ||
| import json, os | ||
| golden = [entry["number"] for entry in json.load(open("products/tasks/evals/golden_prs/golden_prs.json"))] | ||
| wanted = os.environ["PRS"].strip() | ||
| selected = golden if wanted == "all" else [int(n) for n in wanted.split(",") if n.strip()] | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review [should_fix] Reject an empty PR selectionIssue descriptionA blank or comma-only Why we think it's a valid issue
Suggested fixFail the plan step with a clear error when Prompt to fix with AI (copy-paste) |
||
| unknown = sorted(set(selected) - set(golden)) | ||
|
pauldambra marked this conversation as resolved.
|
||
| if unknown: | ||
| raise SystemExit(f"Not in the golden set: {unknown}") | ||
| print(json.dumps(selected)) | ||
| EOF | ||
| ) | ||
| echo "prs=$prs" >> "$GITHUB_OUTPUT" | ||
|
|
||
| evaluate: | ||
| name: PR ${{ matrix.pr }} | ||
| needs: plan | ||
| # A matrix that expands to zero cells fails its dependents and posts no check run. | ||
| if: needs.plan.outputs.prs != '[]' | ||
| runs-on: ubuntu-24.04 | ||
| timeout-minutes: 60 | ||
|
pauldambra marked this conversation as resolved.
|
||
| strategy: | ||
| fail-fast: false | ||
| max-parallel: 4 | ||
| matrix: | ||
| pr: ${{ fromJSON(needs.plan.outputs.prs) }} | ||
| steps: | ||
| - name: Check required secrets | ||
| id: gate | ||
| env: | ||
| ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} | ||
| OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} | ||
| RUNTIME: ${{ inputs.runtime }} | ||
| run: | | ||
| missing="" | ||
| [ -n "$ANTHROPIC_API_KEY" ] || missing="$missing ANTHROPIC_API_KEY" | ||
| if [ "$RUNTIME" = "codex" ] && [ -z "$OPENAI_API_KEY" ]; then | ||
| missing="$missing OPENAI_API_KEY" | ||
| fi | ||
| if [ -n "$missing" ]; then | ||
| echo "enabled=false" >> "$GITHUB_OUTPUT" | ||
| echo "::warning::Missing secrets:$missing; skipping the golden PR eval" | ||
| else | ||
| echo "enabled=true" >> "$GITHUB_OUTPUT" | ||
| fi | ||
|
|
||
| - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 | ||
|
pauldambra marked this conversation as resolved.
|
||
| if: steps.gate.outputs.enabled == 'true' | ||
| with: | ||
| # The agent runs unsandboxed in this checkout; it has no reason to read a persisted token. | ||
| persist-credentials: false | ||
| # The runner fetches each golden commit itself; the checkout only needs the eval package | ||
| # and the files the toolchain setup steps read. | ||
| sparse-checkout: | | ||
| .nvmrc | ||
| pyproject.toml | ||
| products/__init__.py | ||
| products/tasks/__init__.py | ||
| products/tasks/evals | ||
| sparse-checkout-cone-mode: false | ||
|
pauldambra marked this conversation as resolved.
|
||
|
|
||
| - uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0 | ||
| if: steps.gate.outputs.enabled == 'true' | ||
| with: | ||
| node-version-file: .nvmrc | ||
|
|
||
| - name: Install the agent CLI | ||
| if: steps.gate.outputs.enabled == 'true' | ||
| env: | ||
| RUNTIME: ${{ inputs.runtime }} | ||
| run: | | ||
| if [ "$RUNTIME" = "codex" ]; then | ||
| npm install -g @openai/codex | ||
| else | ||
| npm install -g @anthropic-ai/claude-code | ||
|
Comment on lines
+134
to
+136
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review [should_fix] Pin the agent CLI before exposing credentialsIssue descriptionThese commands install the latest CLI version on every run. The eval later gives the selected provider key to the agent process. If a registry release is compromised or changes unexpectedly, the installed CLI can read or send that key. This also makes runs less reproducible. Why we think it's a valid issue
Suggested fixInstall an explicitly reviewed CLI version instead of the unqualified latest version. Keep version updates intentional and controlled. Prompt to fix with AI (copy-paste) |
||
| fi | ||
| "$RUNTIME" --version | ||
|
|
||
| - uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0 | ||
| if: steps.gate.outputs.enabled == 'true' | ||
| with: | ||
| python-version-file: 'pyproject.toml' | ||
|
|
||
| - name: Install the judge SDK | ||
| if: steps.gate.outputs.enabled == 'true' | ||
| # Only what the runner imports; a full `uv sync` of the repo is minutes of work the eval never uses. | ||
| run: python -m pip install --quiet 'anthropic>=0.80,<1' 'pydantic>=2,<3' | ||
|
|
||
| - name: Run the eval | ||
| if: steps.gate.outputs.enabled == 'true' | ||
| env: | ||
| ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} | ||
| OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} | ||
|
pauldambra marked this conversation as resolved.
pauldambra marked this conversation as resolved.
|
||
| PR: ${{ matrix.pr }} | ||
| RUNTIME: ${{ inputs.runtime }} | ||
| MODEL: ${{ inputs.model }} | ||
| JUDGE_MODEL: ${{ inputs.judge_model }} | ||
| CASE_TIMEOUT_MINUTES: ${{ inputs.case_timeout_minutes }} | ||
| run: | | ||
| # Validated before the arithmetic expansion below: Bash evaluates a command | ||
| # substitution embedded in an unchecked arithmetic string, and this step's | ||
| # environment holds both provider API keys. The upper bound also keeps setup, | ||
| # judging and artifact upload inside the job's 60-minute timeout. | ||
| case "$CASE_TIMEOUT_MINUTES" in | ||
| ''|*[!0-9]*) | ||
| echo "::error::case_timeout_minutes must be a positive integer, got '$CASE_TIMEOUT_MINUTES'" | ||
| exit 1 | ||
| ;; | ||
| esac | ||
| if [ "$CASE_TIMEOUT_MINUTES" -lt 1 ] || [ "$CASE_TIMEOUT_MINUTES" -gt 45 ]; then | ||
| echo "::error::case_timeout_minutes must be between 1 and 45" | ||
| exit 1 | ||
| fi | ||
| python -m products.tasks.evals.golden_prs run \ | ||
| --pr "$PR" \ | ||
| --runtime "$RUNTIME" \ | ||
| ${MODEL:+--model "$MODEL"} \ | ||
| --judge-model "$JUDGE_MODEL" \ | ||
| --case-timeout "$((10#$CASE_TIMEOUT_MINUTES * 60))" \ | ||
| --results-dir "$RESULTS_DIR" | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review [should_fix] Keep agent-written files out of score resultsIssue descriptionThe agent inherits Why we think it's a valid issue
Suggested fixIsolate the agent from the evaluator's result directory, then have the report read only the expected result files for the selected PRs. A separate container or user for the agent provides a stronger boundary than hiding the path from its environment. Prompt to fix with AI (copy-paste) |
||
|
|
||
| - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | ||
| # Keep the diff and agent log of a failed case too; they are what explains the failure. | ||
| if: always() && steps.gate.outputs.enabled == 'true' | ||
| with: | ||
| name: golden-pr-${{ matrix.pr }} | ||
| path: ${{ env.RESULTS_DIR }} | ||
| if-no-files-found: warn | ||
|
|
||
| report: | ||
| name: Score summary | ||
| needs: evaluate | ||
| if: ${{ !cancelled() }} | ||
| runs-on: ubuntu-24.04 | ||
| timeout-minutes: 5 | ||
| steps: | ||
| - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 | ||
| with: | ||
| persist-credentials: false | ||
| sparse-checkout: | | ||
| pyproject.toml | ||
| products/__init__.py | ||
| products/tasks/__init__.py | ||
| products/tasks/evals | ||
| sparse-checkout-cone-mode: false | ||
|
|
||
| - uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0 | ||
| with: | ||
| python-version-file: 'pyproject.toml' | ||
|
|
||
| - uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7.0.0 | ||
| with: | ||
| pattern: golden-pr-* | ||
| path: ${{ env.RESULTS_DIR }} | ||
| merge-multiple: true | ||
|
Comment on lines
+212
to
+216
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review [should_fix] Mark incomplete score summariesIssue descriptionIf an eval fails before it writes a result file, its artifact can be empty while other PR artifacts still download. The report then calculates scores and the mean from only the PRs that produced results, without marking the table as incomplete. This can make model comparisons misleading. Why we think it's a valid issue
Suggested fixPass the selected PR list to the report job and compare it with the downloaded results. Mark missing PRs as failed or label the summary as incomplete so readers do not treat a partial mean as a complete run. Prompt to fix with AI (copy-paste) |
||
|
|
||
| - name: Publish the score table | ||
| env: | ||
| RUNTIME: ${{ inputs.runtime }} | ||
| MODEL: ${{ inputs.model }} | ||
| run: | | ||
| set -o pipefail | ||
| python3 -m pip install --quiet 'anthropic>=0.80,<1' 'pydantic>=2,<3' | ||
|
pauldambra marked this conversation as resolved.
|
||
| { | ||
| echo "## Golden PR evals: $RUNTIME ${MODEL:-default model}" | ||
| echo "" | ||
| python3 -m products.tasks.evals.golden_prs report --results-dir "$RESULTS_DIR" | ||
| } | tee -a "$GITHUB_STEP_SUMMARY" | ||
|
pauldambra marked this conversation as resolved.
|
||
|
|
||
| - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 | ||
| with: | ||
| name: golden-pr-results | ||
| path: ${{ env.RESULTS_DIR }} | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| results/ |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,52 @@ | ||
| # Golden PR evals | ||
|
|
||
| Score how well a coding agent can one-shot a real PostHog pull request. | ||
|
|
||
| The golden set is a list of pull requests from 2023 and 2024, written and reviewed by people before coding agents were common. | ||
| For each PR, the eval: | ||
|
|
||
| 1. Checks out the repository at the commit before the PR merged. The agent gets a fresh git repository with only that tree, so it cannot read the merged PR from git history. | ||
| 2. Sends the PR title and description to the agent as its only prompt. | ||
| 3. Compares the agent's diff with the merged PR's diff. | ||
|
|
||
| ## Scores | ||
|
|
||
| | Score | What it measures | | ||
| | --------- | ----------------------------------------------------------------------------------------------- | | ||
| | Files hit | Share of the golden PR's files that the agent also changed. | | ||
| | Line F1 | Overlap between the added lines of the two diffs. | | ||
| | Judge | An Anthropic model reads the task and both diffs and scores behavioral equivalence from 0 to 1. | | ||
|
|
||
| Snapshot files and images are ignored, because an agent cannot regenerate them without the test suite. | ||
| The deterministic scores are strict, so a correct change written in a different way scores low on them. Read the judge score and its reasoning together with them. | ||
|
|
||
| ## Known limits | ||
|
|
||
| The agent runs unsandboxed (`--dangerously-skip-permissions` / `--dangerously-bypass-approvals-and-sandbox`), with its working directory as a soft boundary rather than a hard one. It is not run inside a container or namespace, so it could, in principle, read the original checkout that fetched the golden merge commit, or other host state, instead of relying on the PR description alone. Treat a score as a signal for comparing agents and models, not as proof the agent only ever saw the task description. | ||
|
|
||
| ## Run it | ||
|
|
||
| From GitHub, open the **Golden PR Evals** workflow and choose **Run workflow**. Pick the runtime, the model, and the PR numbers. Each PR runs as its own job. The run summary shows the score table, and the artifacts hold every diff, agent log, and score file. | ||
|
|
||
| From a devbox, with `claude` or `codex` on `PATH`. | ||
| The judge uses `ANTHROPIC_API_KEY` when it is set, and the signed-in `claude` CLI otherwise: | ||
|
|
||
| ```bash | ||
| python -m products.tasks.evals.golden_prs list | ||
| python -m products.tasks.evals.golden_prs run --pr 25832 --runtime claude --model claude-opus-5 | ||
| python -m products.tasks.evals.golden_prs run --runtime codex --model gpt-5.5 | ||
| python -m products.tasks.evals.golden_prs report --results-dir products/tasks/evals/golden_prs/results/<run> | ||
| ``` | ||
|
|
||
| Results land in `products/tasks/evals/golden_prs/results/<timestamp>-<runtime>-<model>/`, which git ignores. | ||
|
|
||
| ## Add a golden PR | ||
|
|
||
| Pick a merged PR that a person wrote, with a description that says what should change. Small, self-contained fixes and features work best. Then append it to `golden_prs.json`: | ||
|
|
||
| ```bash | ||
| gh pr view <number> --json number,title,body,mergedAt,mergeCommit,author \ | ||
| --jq '{number, author: .author.login, title, merged_at: .mergedAt, merge_commit_sha: .mergeCommit.oid, body}' | ||
| ``` | ||
|
|
||
| The tests in `test_golden_prs.py` check that every entry is complete. |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
FLASH MODE - Faster, but stupid, use regular ReviewHog for a heavy review
[must_fix] Restrict provider secrets to trusted workflow refs
Issue description
This workflow can be dispatched against a non-default branch, and GitHub runs the workflow version from that ref. A modified workflow on that branch could read and exfiltrate the repository secrets used by the evaluation, even before it starts the agent. The
workflow_dispatchbranch selector and ref behavior make this a separate exposure path from the agent prompt. (docs.github.com)Why we think it's a valid issue
.github/workflows/golden-pr-evals.ymlfrom dispatch through the evaluation job, including its secret checks and provider-key use.workflow_dispatchat.github/workflows/golden-pr-evals.yml:11. Theevaluatejob reads repository secrets at.github/workflows/golden-pr-evals.yml:92-93and passes both provider keys to the eval step at.github/workflows/golden-pr-evals.yml:153-154. The job does not declare a protected environment.persist-credentials: falsedoes not restrict access to these explicitly injected provider keys. Keeping the keys only as environment secrets, with deployment limited to trusted refs, prevents a branch from accessing them without the required environment binding and approval.must_fixis appropriate because this exposes paid provider credentials to workflow code on an untrusted ref.Suggested fix
Store the provider keys in a protected GitHub environment that only allows the trusted default branch, and bind the
evaluatejob to that environment. Add required approval if appropriate. A guard in this workflow alone is not sufficient because a modified branch can remove it.Prompt to fix with AI (copy-paste)