Skip to content

feat(signals): add scout rubric editor and generator - #106580

Merged
trunk-io[bot] merged 36 commits into
masterfrom
signals/scout-rubric-generator
Sep 29, 2026
Merged

trunk-io[bot] merged 36 commits into
masterfrom
signals/scout-rubric-generator

Conversation

@sortafreel

@sortafreel sortafreel commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Scout owners need a place to save what good work looks like before reviewing or comparing runs.

Changes

  • Staff in project 2 can open Rubrics to write, edit and save a scout's evaluation criteria.
  • Generate suggestions drafts criteria with short titles, passing conditions and explanations from the scout's description, instructions, references and recent run summaries. It also works before the scout has any runs.
  • Rubric generation defaults to GPT-6 Sol at high effort. Existing model overrides still apply.
  • The generator checks the saved rubric before offering additions. Owners review, select and edit suggestions before clicking Save rubrics.
  • Defaults start enabled and remain in every save. Owners can edit or disable them, and remove custom criteria.
  • Saved criteria and generated suggestions survive closing the page. Stale saves and late generation results cannot overwrite newer work.
  • This adds criteria for future evaluations; it does not score runs or change how scouts execute.

The two-step generator reuses the existing background worker. Its merged prerequisite #107828 must deploy first.
API client updates are generated.

The example uses the public Product analytics scout.
GPT-6 Sol at high effort generated these suggestions without prior runs. The text is unedited.
These replace the earlier screenshots with output from the final readability prompt. Storybook displays the saved result using replayed API responses.

Unedited generated Product analytics criteria

Editing a criterion

Expanded Product analytics criterion

Before and after

Before:

flowchart LR
    Scout[Scout] --> Review[Define criteria elsewhere]
    classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
    class Scout,Review phGray;
Loading

After:

flowchart LR
    Generate[Generate suggestions] --> Agent{{Background agent}}
    Agent --> Review[Review and edit]
    Review --> Save[Save rubric]
    classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
    class Agent phBlue;
    class Generate,Review phYellow;
    class Save phGray;
Loading

How did you test this code?

  • Backend and frontend tests cover access, retained defaults, conflicting saves, polling, invalid output, cancellation, late results and requested permissions.
  • Runtime tests check the Sol/high default, retained overrides, and unchanged defaults for other steps. The rubric tests verify the model and effort passed to the sandbox.
  • Browser checks covered generation, editing, saving, reopening, failure states and narrow layouts. Native API checks confirmed generation preserves saved criteria.
  • The quality report records checks across run histories, saved rubrics and description-only inputs. One native output needed a duplicate removed: suggestions still require review.

The final readability confirmation covered twelve scouts: six with history and six without.
All twelve preserved the intended behavior; 59 of 60 criteria were clear without edits. One feedback criterion needs a sentence repair.
The report preserves all 54 generations from this expanded-field round, including rejected revisions. These were familiar development cases, not a blind benchmark.
Every final analytics criterion was opened and visually inspected; the screenshots show unedited generated text.
This pass did not repeat browser-triggered backend generation or the full earlier quality matrix. This diff does not read event properties or change ingestion.
Some worker dispatch and early exits are missing from the coverage report.
Native runs exercise dispatch; exits for revoked access and removed or replaced records remain untested.

👉 Stay up-to-date with PostHog coding conventions for a smoother review.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Access is restricted to staff in project 2.
Existing generation records remain visible to their creator while they retain project access, even if their staff status changes.
The backend prepares the source context; the generator requests no additional project-read permissions.
The sandbox still has internal credentials and tool access. This remains an accepted limitation of the internal v0.

Automatic notifications

  • Publish to changelog?

Docs update

Updated the sandboxed-agent guide, Signals architecture and harness directory reference.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Codex, GPT-6

Tools: terminal, Playwright and GitHub CLI. No shareable session link is available.
CodeRabbit CLI was signed out; the local pass was skipped. No duplicate generator PR was found.
Review fixes cover timeout handling, default retention, Temporal scoping and reduced permissions.
A later readability pass added final writing instructions and selected Sol/high as the default; generation still uses the same two steps. Expanded-field review later replaced dense rule lists with plain outcomes and short policy references.
Private research informed the design; committed fixtures are invented, and screenshots use only the public scout source.

Skills used

/improving-drf-endpoints, /django-migrations, /writing-tests, /maintaining-python-tests, /fixing-flaky-tests, /writing-dataclasses, /adding-activity-logging, /writing-ui-components, /writing-kea-logics, /using-kea-disposables, /writing-user-facing-copy, /writing-code-comments, /run-posthog, /hogli, /debugging-local-task-agent-runs, /querying-local-postgres, /editing-agents-md, /stacking-prs, /running-ci-preflight, /debugging-ci-failures, /writing-pr-descriptions, /reviewing-with-coderabbit.

@sortafreel sortafreel self-assigned this Sep 25, 2026
@trunk-io

trunk-io Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 5 functions above the limit (max 36)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
ScoutRubricsModal products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx:17 36 10
saveRubrics products/signals/frontend/inbox/logics/scoutRubricsLogic.ts:388 16 10
ScoutDetailHeader products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx:102 14 10
generateSuggestions products/signals/frontend/inbox/logics/scoutRubricsLogic.ts:365 14 10
ScoutAttentionBanner products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx:240 11 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Bundle size — 🔺 +13.6 KiB (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 68.90 MiB · 🔺 +13.6 KiB (+0.0%)

File Size Δ vs base
posthog-app/_parent/products/signals/frontend/inbox/InboxScene.js 488.4 KiB 🔺 +13.6 KiB (+2.9%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.57 MiB · 22 files no change █████████░ 85.5% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.51 MiB · 629 files no change █████████░ 87.2% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.34 MiB · 2,339 files no change █████████░ 88.0% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
88.4 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.4 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
271.7 KiB src/taxonomy/core-filter-definitions-by-group.json
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
98.5 KiB ../packages/quill/packages/quill/dist/index.js
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
88.4 KiB src/products.tsx

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.16 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.16 MiB · 19 files no change ████░░░░░░ 37.7% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
800.5 KiB dist/toolbar/toolbar-app-QUJ43CJ4.css
651.6 KiB dist/toolbar/chunk-chunk-ZPCK2O6G.js
259.4 KiB dist/toolbar/chunk-chunk-CV2VU6SQ.js
138.3 KiB dist/toolbar/chunk-chunk-DYPTRYMF.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-4HYNQ5KU.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-HOKNU4ZL.js
21.0 KiB dist/toolbar/chunk-chunk-Z5ELNJKM.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +154.9 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 946.46 MiB · 🔺 +154.9 KiB (+0.0%)

ℹ️ MCP UI apps size — 33 app(s), 17631.0 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 196.2 KB
action 454.1 KB 196.2 KB
action-list 564.2 KB 196.2 KB
cohort 453.1 KB 196.2 KB
cohort-list 563.2 KB 196.2 KB
email-template 452.9 KB 196.2 KB
error-details 469.6 KB 196.2 KB
error-issue 453.8 KB 196.2 KB
error-issue-list 564.1 KB 196.2 KB
experiment 561.3 KB 196.2 KB
experiment-list 564.9 KB 196.2 KB
experiment-results 566.3 KB 196.2 KB
feature-flag 566.8 KB 196.2 KB
feature-flag-list 570.5 KB 196.2 KB
feature-flag-testing 457.3 KB 196.2 KB
inline-scan 453.6 KB 196.2 KB
insight-actors 562.3 KB 196.2 KB
invite-email-preview 452.3 KB 196.2 KB
llm-costs 559.3 KB 196.2 KB
session-recording 455.3 KB 196.2 KB
survey 454.7 KB 196.2 KB
survey-global-stats 561.9 KB 196.2 KB
survey-list 564.9 KB 196.2 KB
survey-stats 561.9 KB 196.2 KB
trace-span 453.5 KB 196.2 KB
trace-span-list 564.1 KB 196.2 KB
vision-observation-list 563.3 KB 196.2 KB
workflow 453.4 KB 196.2 KB
workflow-list 563.5 KB 196.2 KB
loops-review 457.8 KB 196.2 KB
query-results 774.1 KB 196.2 KB
render-ui 857.4 KB 196.2 KB
visual-review-snapshots 457.9 KB 196.2 KB
✅ Playwright — all passed

All tests passed.

View test results →

⚠️ Backend coverage — 98.0% of changed backend lines covered — 17 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ████████████████████ 98.0% (849 / 866)

File Patch Uncovered changed lines
products/signals/backend/temporal/agentic/scout_rubrics.py 81.0% 33, 37–38, 45, 54–55, 63–64
products/signals/backend/scout_harness/rubrics_runner.py 96.8% 343, 349, 359, 390
products/signals/backend/scout_harness/rubrics.py 97.9% 234, 259, 283, 286
products/signals/backend/presentation/scout_rubrics.py 99.0% 175

🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 36548668960 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
warehouse_sources_queue ██████░░░░░░░░░░░░░░ 29.1% 92 / 316
demo ███████████░░░░░░░░░ 52.9% 1,413 / 2,673
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
aeo ███████████████░░░░░ 76.3% 617 / 809
batch_exports ████████████████░░░░ 81.2% 21,473 / 26,446
apm █████████████████░░░ 84.1% 1,306 / 1,553
cdp ██████████████████░░ 88.2% 4,545 / 5,155
ml_inference ██████████████████░░ 88.4% 509 / 576
mcp_analytics ██████████████████░░ 88.9% 4,927 / 5,540
product_tours ██████████████████░░ 89.3% 1,331 / 1,491
dashboards ██████████████████░░ 89.5% 6,855 / 7,657
data_warehouse ██████████████████░░ 90.0% 14,019 / 15,581
signals ██████████████████░░ 90.2% 57,444 / 63,698
notebooks ██████████████████░░ 90.2% 15,287 / 16,945
cohorts ██████████████████░░ 90.4% 8,420 / 9,316
streamlit_apps ██████████████████░░ 90.6% 2,623 / 2,895
managed_warehouse ██████████████████░░ 91.0% 10,252 / 11,263
data_modeling ██████████████████░░ 91.4% 10,504 / 11,498
exports ██████████████████░░ 91.6% 9,680 / 10,562
engineering_analytics ██████████████████░░ 91.7% 11,032 / 12,030
tasks ██████████████████░░ 91.9% 74,866 / 81,470
ai_training ██████████████████░░ 92.2% 356 / 386
business_knowledge ██████████████████░░ 92.2% 7,684 / 8,330
conversations ███████████████████░ 92.5% 28,734 / 31,062
early_access_features ███████████████████░ 92.6% 1,332 / 1,439
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
visual_review ███████████████████░ 92.8% 9,247 / 9,969
canvas ███████████████████░ 92.8% 6,877 / 7,409
approvals ███████████████████░ 93.0% 3,974 / 4,271
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
error_tracking ███████████████████░ 93.2% 15,928 / 17,097
notifications ███████████████████░ 93.2% 1,145 / 1,229
slack_app ███████████████████░ 93.2% 13,677 / 14,674
stamphog ███████████████████░ 93.2% 7,885 / 8,456
surveys ███████████████████░ 93.3% 6,571 / 7,040
autoresearch ███████████████████░ 93.6% 8,481 / 9,061
context_layer ███████████████████░ 93.9% 3,415 / 3,638
web_analytics ███████████████████░ 94.0% 21,815 / 23,218
alerts ███████████████████░ 94.0% 8,569 / 9,114
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.4% 8,940 / 9,472
ai_observability ███████████████████░ 94.5% 24,259 / 25,676
wizard ███████████████████░ 94.7% 6,150 / 6,496
reminders ███████████████████░ 94.8% 760 / 802
workflows ███████████████████░ 94.8% 14,379 / 15,160
review_hog ███████████████████░ 95.0% 11,507 / 12,119
endpoints ███████████████████░ 95.1% 9,206 / 9,681
annotations ███████████████████░ 95.1% 817 / 859
customer_analytics ███████████████████░ 95.2% 25,078 / 26,349
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
marketing_analytics ███████████████████░ 95.3% 19,322 / 20,274
posthog_ai ███████████████████░ 95.3% 2,488 / 2,610
experiments ███████████████████░ 95.4% 32,458 / 34,023
actions ███████████████████░ 95.5% 756 / 792
logs ███████████████████░ 95.5% 15,399 / 16,130
data_catalog ███████████████████░ 95.5% 4,401 / 4,606
tracing ███████████████████░ 95.6% 3,536 / 3,699
replay_vision ███████████████████░ 95.6% 27,675 / 28,939
growth ███████████████████░ 95.7% 11,228 / 11,734
messaging ███████████████████░ 95.8% 3,798 / 3,963
skills ███████████████████░ 95.8% 6,972 / 7,274
product_analytics ███████████████████░ 96.0% 28,470 / 29,647
access_control ███████████████████░ 96.3% 7,112 / 7,386
revenue_analytics ███████████████████░ 96.4% 1,876 / 1,946
user_interviews ███████████████████░ 96.5% 2,867 / 2,971
feature_flags ███████████████████░ 96.5% 25,588 / 26,509
warehouse_sources ███████████████████░ 97.2% 456,705 / 469,652
data_quality ████████████████████ 97.6% 7,587 / 7,774
security ████████████████████ 97.9% 1,202 / 1,228
links ████████████████████ 97.9% 234 / 239
metrics ████████████████████ 98.0% 4,084 / 4,166
analytics_platform ████████████████████ 98.3% 2,784 / 2,833
pulse ████████████████████ 98.5% 2,046 / 2,078
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

⚠️ Django migration SQL — 1 new migration to review

We've detected new migrations on this PR. Review the SQL output for each migration:

products/signals/backend/migrations/0155_signalscoutconfig_rubrics.py

/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/anyio/from_thread.py:119: SyntaxWarning: 'return' in a 'finally' block
  return result
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/structlog/stdlib.py:1166: UserWarning: Remove `format_exc_info` from your processor chain if you want pretty exceptions.
  ed = p(logger, meth_name, ed)  # type: ignore[arg-type]
2026-09-29T09:24:45.185188Z [error    ] Path must be a valid database or directory containing databases. [posthog.exceptions_capture] pid=7653 tid=140555622476672
Traceback (most recent call last):
  File "/home/runner/work/posthog/posthog/posthog/geoip.py", line 15, in <module>
    geoip: Optional[GeoIP2] = GeoIP2(cache=8)
                              ~~~~~~^^^^^^^^^
  File "/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/django/contrib/gis/geoip2.py", line 116, in __init__
    raise GeoIP2Exception(
        "Path must be a valid database or directory containing databases."
    )
django.contrib.gis.geoip2.GeoIP2Exception: Path must be a valid database or directory containing databases.
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/sshtunnel.py:1040: SyntaxWarning: 'return' in a 'finally' block
  return (ssh_host,
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/langchain_core/_api/deprecation.py:27: UserWarning: Core Pydantic V1 functionality isn't compatible with Python 3.14 or greater.
  from pydantic.v1.fields import FieldInfo as FieldInfoV1
System check identified some issues:

WARNINGS:
?: (axes.W001) You are using the django-axes cache handler for login attempt tracking. Your cache configuration is however invalid and will not work correctly with django-axes. This can leave security holes in your login systems as attempts are not tracked correctly. Reconfigure settings.AXES_CACHE and settings.CACHES per django-axes configuration documentation.
?: (staticfiles.W004) The directory '/home/runner/work/posthog/posthog/frontend/dist' in the STATICFILES_DIRS setting does not exist.
BEGIN;
--
-- Add field rubrics to signalscoutconfig
--
ALTER TABLE "signals_signalscoutconfig" ADD COLUMN "rubrics" jsonb NULL;
COMMIT;

Last updated: 2026-09-29 09:25 UTC (1bc5dab)

✅ Django migration risk — migration analysis complete

We've analyzed your migrations for potential risks.

Summary: 1 Safe | 0 Needs Review | 0 Blocked

✅ Safe

Brief or no lock, backwards compatible

signals.0155_signalscoutconfig_rubrics
  └─ #1 ✅ AddField
     Adding nullable field requires brief lock
     model: signalscoutconfig, field: rubrics

📚 How to Deploy These Changes Safely

AddField:

This operation acquires a brief lock but doesn't rewrite the table.

Deployment uses lock timeouts with automatic retries, so lock contention will cause retries rather than connection pile-up.

Last updated: 2026-09-29 09:25 UTC (1bc5dab)

@greptile-apps

greptile-apps Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[High risk] Adds database schema, API endpoints, and background job system for scout rubrics.

The reviewed rubric changes appear safe to merge.

Reviews (2) · Last reviewed commit: "chore(signals): separate the scout harne..."

@coderabbitai

coderabbitai Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 00ec8ea4-bfa7-4cdf-8201-abe72f0a04bd

📥 Commits

Reviewing files that changed from the base of the PR and between e21e1d9 and d020bea.

📒 Files selected for processing (1)
  • frontend/snapshots.yml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 2 remain after this review.


📝 Walkthrough

Walkthrough

Adds persistent, revisioned scout rubrics with default criteria and API endpoints for retrieval, updates, and generation requests. A Temporal workflow runs a background agent that stores generated suggestions separately from saved criteria. The scout interface adds controls to edit criteria, review suggestions, and select criteria before saving. The change also excludes scout suggestion background runs from finish-tool availability and updates report-channel guidance selection.

Priority: ⬇️ Low

Merge Risk: 🟠 High · up to d020b

Core rubric generation and persistence still have unresolved correctness and integration risks, including possible missing workflow execution and loss or corruption of generated results. Resolve these before merging.

Security Architecture Review

Security architecture risk: 🔵 Low · up to d020b

Rubric editing and generation are restricted to authorized staff in one project, and suggestions do not change saved criteria automatically. A late generation could nevertheless change a status users have already seen as failed.

Retained concerns

  • Low · reliability · inferred: A generation shown as timed out can be completed later under the same generation ID: visible expiry is not persisted, and terminal writes do not reject an expired or already terminal generation. This is conditional on sufficiently late work; the shorter workflow limit is counterevidence for the normal path.
Security review details

Security Blast Radius

  • inferred — The intended generation authority is limited to an authorized staff request for project 2 and read access to scout-related project data. The source does not establish broader tenant exposure or prove enforcement inside the shared task runtime.

Trust Boundaries and Controls

  • observed — Scout-authored instructions and run summaries enter an agent context as untrusted material. The runner requests read-scoped credentials, no repository or GitHub access, and no external MCP gateways; generated output is schema-validated and remains a suggestion until saved.

Resilience and Maintainability Implications

  • inferred — A disagreement between visible timeout state and later same-ID writes weakens failure containment for suggestions, though it does not itself save those suggestions as criteria.

Hardening Proposals

  • proposed — Make expiry or terminal status durable, or reject same-ID completion and failure writes once the generation is terminal or expired.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description follows the repository template and clearly documents the problem, user-visible changes, testing scope and limitations, release status, documentation updates, agent context, architectu…
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (3)
products/signals/backend/scout_harness/rubrics.py-217-221 (1)

217-221: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Do not let fail_generation overwrite a terminal generation.

fail_generation matches only on generation.id. It sets FAILED whatever the current status is. In run_rubric_generation, the final update_generation runs in a worker thread through database_sync_to_async. If the activity is cancelled after that write commits, the CancelledError goes to the except branch, which calls fail_generation. The stored generation then shows failed with its suggestions still present. The UI tells the user to retry, and the retry discards a valid proposal. Any Temporal-side failure handler that calls fail_generation after the activity has finished has the same problem.

Fail only generations that are still in progress.

Proposed fix
-        if state.generation is None or state.generation.id != generation_id:
+        if (
+            state.generation is None
+            or state.generation.id != generation_id
+            or state.generation.status
+            not in (ScoutRubricGenerationStatus.QUEUED, ScoutRubricGenerationStatus.RUNNING)
+        ):
             return
products/signals/backend/models.py-2298-2303 (1)

2298-2303: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Exclude rubric-generation state from activity logging.

The generation path saves rubrics for the queued, running, and terminal states. Early quota or start failures can save only queued and failed states, so the exact count varies. Each save uses config.save(update_fields=["rubrics", "updated_at"]).

rubrics is absent from both SignalScoutConfig exclusion lists. These machine-state saves can therefore create activity entries and include the full JSON value in their diffs.

Suggested fix
@@ signal_exclusions
     "SignalScoutConfig": [
         "last_run_at",
         "consecutive_failure_count",
+        "rubrics",
     ],
@@ field_exclusions
     "SignalScoutConfig": [
         "last_run_at",
         "consecutive_failure_count",
+        "rubrics",
         "status_changed_at",

Alternatively, store generation state in a separate excluded field.

products/signals/backend/scout_harness/rubrics_runner.py-100-101 (1)

100-101: 🔒 Security & Privacy | 🛡️ Detected with Advanced Tier | 🟡 Minor | ⚡ Quick win

LLM Security

Reachability: External
CWE: CWE-1427

Treat skill instructions and run summaries as untrusted data. json.dumps(context) does not stop embedded text from influencing the model. A malicious skill or poisoned run summary can redirect reads within the granted project scopes or inject unrelated content into the persisted summary and suggestions. Delimit each field clearly and enforce the selected skill and run IDs in the MCP handlers or an equivalent server-side allowlist. Do not rely on the prompt warning alone.

Source: Learnings

🧹 Nitpick comments (2)
products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.tsx (1)

17-17: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Keep modal-open state in Kea.

isOpen uses React-local state beside Kea-backed scene state. Move open and close actions to the owning logic. As per coding guidelines, “Don't use useState or useEffect to store local state.”

Source: Coding guidelines

products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx (1)

96-99: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Give Close a stable interaction attribute.

Add a kebab-case data-attr to the Close button so autocapture and Playwright can identify this exit action. As per coding guidelines, “New buttons and key interactive elements get a kebab-case data-attr.”

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 2c4fffb3-dca6-4a18-a5e7-00501c0d7610

📥 Commits

Reviewing files that changed from the base of the PR and between 7cf5437 and d7c6c37.

⛔ Files ignored due to path filters (3)
  • products/signals/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.zod.ts is excluded by !**/generated/**
📒 Files selected for processing (23)
  • docs/published/handbook/engineering/ai/sandboxed-agents.md
  • posthog/temporal/tests/ai/test_module_integrity.py
  • products/signals/ARCHITECTURE.md
  • products/signals/backend/migrations/0155_signalscoutconfig_rubrics.py
  • products/signals/backend/migrations/max_migration.txt
  • products/signals/backend/models.py
  • products/signals/backend/routes.py
  • products/signals/backend/scout_harness/AGENTS.md
  • products/signals/backend/scout_harness/rubrics.py
  • products/signals/backend/scout_harness/rubrics_runner.py
  • products/signals/backend/scout_rubrics_api.py
  • products/signals/backend/temporal/__init__.py
  • products/signals/backend/temporal/agentic/scout_rubrics.py
  • products/signals/backend/test/test_scout_rubrics.py
  • products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricCriterionEditor.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.test.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.stories.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.test.ts
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.ts
  • services/mcp/src/api/generated.ts

Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review.

Comment thread products/signals/backend/routes.py Outdated
Comment thread products/signals/backend/scout_rubrics_api.py Outdated
@trunk-io

trunk-io Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
Scenes-App/SidePanels SidePanelNotebooks smoke-test The test timed out while waiting for a loading indicator or spinner to disappear. Logs ↗︎
HogFlowPropertyFilters search shows workflow variables in the dedicated tab The test exceeded the 5000 ms timeout and was terminated. Logs ↗︎

View Full Report ↗︎ ⋅ Docs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/signals/backend/temporal/agentic/scout_rubrics.py-28-44 (1)

28-44: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Add @scoped_temporal() to both activities.

The repository requires @scoped_temporal() from posthog/temporal/common/scoped.py on every async Signals Temporal activity. generate_scout_rubrics_activity and fail_scout_rubrics_activity do not have it. Without the decorator, these activities run outside the standard analytics scoping that other Signals activities use.

As per coding guidelines: "Every async Signals Temporal activity is decorated with @scoped_temporal() from posthog/temporal/common/scoped.py (not upstream @posthoganalytics.scoped())."

Proposed fix
+from posthog.temporal.common.scoped import scoped_temporal
 ...
 `@activity.defn`
+@scoped_temporal()
 `@close_db_connections`
 async def generate_scout_rubrics_activity(input: ScoutRubricGenerationInput) -> None:
 ...
 `@activity.defn`
+@scoped_temporal()
 `@close_db_connections`
 async def fail_scout_rubrics_activity(input: ScoutRubricGenerationInput) -> None:

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 290b7ca4-95e8-47db-9612-a15855b86c13

📥 Commits

Reviewing files that changed from the base of the PR and between d7c6c37 and 6f259a1.

⛔ Files ignored due to path filters (3)
  • products/signals/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.zod.ts is excluded by !**/generated/**
📒 Files selected for processing (26)
  • docs/published/handbook/engineering/ai/sandboxed-agents.md
  • posthog/temporal/tests/ai/test_module_integrity.py
  • products/signals/ARCHITECTURE.md
  • products/signals/backend/facade/rubrics.py
  • products/signals/backend/migrations/0155_signalscoutconfig_rubrics.py
  • products/signals/backend/migrations/max_migration.txt
  • products/signals/backend/models.py
  • products/signals/backend/presentation/__init__.py
  • products/signals/backend/presentation/scout_rubrics.py
  • products/signals/backend/routes.py
  • products/signals/backend/scout_harness/AGENTS.md
  • products/signals/backend/scout_harness/rubrics.py
  • products/signals/backend/scout_harness/rubrics_runner.py
  • products/signals/backend/temporal/__init__.py
  • products/signals/backend/temporal/agentic/scout_rubrics.py
  • products/signals/backend/test/test_scout_rubrics.py
  • products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricCriterionEditor.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.test.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.stories.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.test.ts
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.ts
  • products/signals/mcp/tools.yaml
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (1)
  • products/signals/backend/presentation/init.py

Included review availability: Your plan provides up to 12 included reviews per hour; 10 remain after this review.

Comment thread products/signals/backend/scout_harness/rubrics.py
Include descriptions in rubric context, review the draft in the same
session, and allow one bounded format correction without replacing saved
criteria on failure. Extend the runner regression coverage and guide.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
products/signals/backend/scout_harness/rubrics.py (1)

285-291: 🎯 Functional Correctness | 🔴 Critical | ⚡ Quick win

Await the Temporal dispatch.

sync_connect() and start_scout_rubric_generation() are both async def. This code creates coroutines and never awaits them. No workflow starts. The generation stays queued until GENERATION_TIMEOUT. The except block also cannot catch connection failures, because the coroutine never runs.

Proposed fix
-        start_scout_rubric_generation(
-            sync_connect(),
-            team_id=team_id,
-            config_id=config_id,
-            generation_id=generation.id,
-            user_id=user_id,
-        )
+        async def _dispatch() -> None:
+            await start_scout_rubric_generation(
+                await sync_connect(),
+                team_id=team_id,
+                config_id=config_id,
+                generation_id=generation.id,
+                user_id=user_id,
+            )
+
+        async_to_sync(_dispatch)()

Import async_to_sync from asgiref.sync.


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 33179ab6-405a-4a67-afe8-32c160b9b4ef

📥 Commits

Reviewing files that changed from the base of the PR and between 6f259a1 and 90bfa08.

⛔ Files ignored due to path filters (3)
  • products/signals/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.zod.ts is excluded by !**/generated/**
📒 Files selected for processing (26)
  • docs/published/handbook/engineering/ai/sandboxed-agents.md
  • posthog/temporal/tests/ai/test_module_integrity.py
  • products/signals/ARCHITECTURE.md
  • products/signals/backend/facade/rubrics.py
  • products/signals/backend/migrations/0155_signalscoutconfig_rubrics.py
  • products/signals/backend/migrations/max_migration.txt
  • products/signals/backend/models.py
  • products/signals/backend/presentation/__init__.py
  • products/signals/backend/presentation/scout_rubrics.py
  • products/signals/backend/routes.py
  • products/signals/backend/scout_harness/AGENTS.md
  • products/signals/backend/scout_harness/rubrics.py
  • products/signals/backend/scout_harness/rubrics_runner.py
  • products/signals/backend/temporal/__init__.py
  • products/signals/backend/temporal/agentic/scout_rubrics.py
  • products/signals/backend/test/test_scout_rubrics.py
  • products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricCriterionEditor.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.test.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.stories.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.test.ts
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.ts
  • products/signals/mcp/tools.yaml
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (1)
  • products/signals/backend/presentation/init.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

Draft complete criteria from bounded scout context, then select additions
against the saved rubric without rewriting them. Preserve caller-owned
session completion and require saving pending edits before generation.

Add regression coverage and document the quality iterations, native UI
validation, and remaining need to review generated suggestions.
Bring the rubric branch up to date for the current generated API inputs
and CI workflows before publishing the completed quality changes.
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

React Doctor found 1 issue in 1 file · 1 warning.

1 warning

packages/ui/src/features/canvas/components/NavRail.tsx

Reviewed by React Doctor for commit 3f4ab85.

@hosthog

hosthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

HostHog preview — posthog-desktop-web

The previews for this PR have been torn down and no longer serve.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

♻️ Duplicate comments (1)
products/signals/backend/scout_harness/rubrics.py (1)

285-291: 🎯 Functional Correctness | 🔴 Critical | ⚡ Quick win

Await the Temporal dispatch. The workflow never starts.

sync_connect() and start_scout_rubric_generation() are both async def. The code calls them without await, so it creates two coroutines and discards them without running them. Creating a coroutine does not raise an exception, so the except branch never runs. The endpoint returns 202 with a queued generation that stays queued until GENERATION_TIMEOUT marks it failed. The endpoint also consumes a daily attempt and does not refund it.

Proposed fix
-        start_scout_rubric_generation(
-            sync_connect(),
-            team_id=team_id,
-            config_id=config_id,
-            generation_id=generation.id,
-            user_id=user_id,
-        )
+        async def _dispatch() -> None:
+            await start_scout_rubric_generation(
+                await sync_connect(),
+                team_id=team_id,
+                config_id=config_id,
+                generation_id=generation.id,
+                user_id=user_id,
+            )
+
+        async_to_sync(_dispatch)()

Add from asgiref.sync import async_to_sync at module level.

🟡 Other comments (2)
products/signals/eval/experiments/2026-09-scout-rubrics/FINAL_REPORT.md-152-152 (1)

152-152: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Fix the missing spaces in "copied33" and "full34".

Change them to "copied 33" and "full 34".

Also applies to: 171-171

Source: Linters/SAST tools

products/signals/backend/scout_harness/rubrics_runner.py-408-410 (1)

408-410: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

The selection prompt reads saved criteria from a stale snapshot.

build_selection_prompt(read_rubric_state(config).criteria, ...) uses the config loaded before the sandbox started. A generation can run for up to 15 minutes. If a user saves during that time, the draft is compared against criteria that no longer exist. The selection can then keep duplicates or drop useful additions. The docs state that "generation uses the saved criteria". Reload the config before the selection turn.

Proposed fix
+            fresh_config = await database_sync_to_async(
+                lambda: SignalScoutConfig.objects.for_team(team_id).get(id=config_id), thread_sensitive=True
+            )()
             selected_output = await session.send_followup_raw(
-                build_selection_prompt(read_rubric_state(config).criteria, draft), label="rubric_saved_selection"
+                build_selection_prompt(read_rubric_state(fresh_config).criteria, draft), label="rubric_saved_selection"
             )

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 731cd678-478f-461a-a929-5d80ef4dafe8

📥 Commits

Reviewing files that changed from the base of the PR and between 90bfa08 and e21e1d9.

⛔ Files ignored due to path filters (3)
  • products/signals/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.ts is excluded by !**/generated/**
  • products/signals/frontend/generated/api.zod.ts is excluded by !**/generated/**
📒 Files selected for processing (30)
  • docs/published/handbook/engineering/ai/sandboxed-agents.md
  • posthog/temporal/tests/ai/test_module_integrity.py
  • products/desktop/packages/harness/src/extensions/local-tools/tools/finish.test.ts
  • products/desktop/packages/harness/src/extensions/local-tools/tools/finish.ts
  • products/signals/ARCHITECTURE.md
  • products/signals/backend/facade/rubrics.py
  • products/signals/backend/migrations/0155_signalscoutconfig_rubrics.py
  • products/signals/backend/migrations/max_migration.txt
  • products/signals/backend/models.py
  • products/signals/backend/presentation/__init__.py
  • products/signals/backend/presentation/scout_rubrics.py
  • products/signals/backend/routes.py
  • products/signals/backend/scout_harness/AGENTS.md
  • products/signals/backend/scout_harness/prompt.py
  • products/signals/backend/scout_harness/rubrics.py
  • products/signals/backend/scout_harness/rubrics_runner.py
  • products/signals/backend/temporal/__init__.py
  • products/signals/backend/temporal/agentic/scout_rubrics.py
  • products/signals/backend/test/test_scout_rubrics.py
  • products/signals/eval/experiments/2026-09-scout-rubrics/FINAL_REPORT.md
  • products/signals/frontend/inbox/components/config/scouts/ScoutDetailHeader.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricCriterionEditor.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.test.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsButton.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.stories.tsx
  • products/signals/frontend/inbox/components/config/scouts/ScoutRubricsModal.tsx
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.test.ts
  • products/signals/frontend/inbox/logics/scoutRubricsLogic.ts
  • products/signals/mcp/tools.yaml
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (1)
  • products/signals/backend/presentation/init.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@posthog

posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

✅ Visual changes approved by @sortafreel — baseline updated in d020bea.

View this run in PostHog

12 new.

12 updated
Run: 2ed0c011-4d67-4614-9e09-f4eb5f8e12d7

Co-authored-by: sortafreel <354488+sortafreel@users.noreply.github.com>
@sortafreel sortafreel added the reviewhog ($$$) Reviews pull requests before humans do label Sep 28, 2026
@posthog

posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🦔 PostHog Review reviewed this pull request

Found 0 must fix, 3 should fix, 3 consider.

Published 6 findings (view the review).

Resolved comments: 1 already settled

@posthog

posthog Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏

@posthog posthog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PostHog Review

Found 3 should fix, 3 consider.

Comment thread products/signals/backend/scout_harness/rubrics.py
Comment thread products/signals/backend/models.py
Comment thread products/signals/backend/scout_harness/rubrics_runner.py
Comment thread products/signals/backend/presentation/scout_rubrics.py
Comment thread products/signals/backend/presentation/scout_rubrics.py
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Sep 28, 2026
@veria-ai

veria-ai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 2 · PR risk: 0/10

@hosthog

hosthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

HostHog preview — storybook-quill

The previews for this PR have been torn down and no longer serve.

Reject updates to expired or terminal generation requests. Persist expired
requests as failed when a late worker returns, without changing saved criteria.

Extend the existing timeout and completion tests to cover late updates.
Require every shared default in rubric API saves while preserving edits and
disabled choices. Extend validation coverage and retain complete criteria in
the existing stale-save and concurrent-generation tests.
Comment thread products/signals/backend/scout_harness/rubrics_runner.py
Supply the existing prepared context without requesting project-read MCP
scopes. Preserve the tested sandbox flow and document its remaining access.
Cover empty scope lists through dispatch and token minting.

Sync master, including the merged standalone harness prerequisite.
@sortafreel sortafreel added the stamphog Request AI approval (no full review) label Sep 28, 2026
@stamphog stamphog Bot removed the stamphog Request AI approval (no full review) label Sep 28, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

This pull request was refused automatically, not reviewed. The deny-list gate matched auth (this PR touches authentication/authorization-related code, e.g. posthog/permissions, OAuthAccessTokenAuthentication, APIScopePermission in products/signals/backend/presentation/scout_rubrics.py), and the size gate failed because the change is far over the 800-line ceiling (2269 substantive lines across 20 files, 4143 lines across 36 files including docs and generated snapshots). The tier gate also failed, classifying this as T2-never due to its large, cross-cutting feature footprint spanning backend models/migrations, temporal workflows, frontend logic, and generated API clients.

Given these deterministic gate failures, this pull request cannot be auto-reviewed. Consider splitting it into smaller, self-contained pieces (e.g. separate the migration/model change, the backend API/temporal workflow, and the frontend components), and route the auth-related changes to a human reviewer directly.

  • 👍 on the PR from greptile-apps[bot].
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: auth
size ✗ too large for auto-review (2269L substantive in global — ceiling is 800L; 2269L, 20F total, 1 binary; 4143L/36F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (4143L, 36F, cross-cutting, feat)
stamphog 2.2.0 .stamphog/policy.yml @ unknown · reviewed head 470b1df

@github-actions
github-actions Bot requested a deployment to preview-pr-106580 September 29, 2026 06:56 In progress
Describe observable outcomes before short source-policy references. Keep
conditional duties and required deliverables explicit without copying
procedural checklists into the editor.

Record the expanded-field recheck and its remaining wording limitations.
Explain checks in everyday language and preserve complete source rules.
Keep selection from rewriting criteria and simplify its visible summary.

Record the expanded-field iterations and twelve-scout confirmation,
including one remaining sentence that needs an owner edit.
@trunk-io
trunk-io Bot merged commit 328f6d6 into master Sep 29, 2026
298 checks passed
@trunk-io
trunk-io Bot deleted the signals/scout-rubric-generator branch September 29, 2026 10:43
@deployment-status-posthog

deployment-status-posthog Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-09-29 11:03 UTC Run
prod-us ✅ Deployed 2026-09-29 11:15 UTC Run
prod-eu ✅ Deployed 2026-09-29 11:16 UTC Run

This branch was successfully deployed

1 active deployment
preview-pr-106580 — 1bc5dab0 Deployed Sep 29, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants