Skip to content

feat(replay-vision): generate search suggestions up front for every scanner and team - #107428

Merged
trunk-io[bot] merged 9 commits into
masterfrom
tue/replay-vision-suggestion-quality
Sep 28, 2026
Merged

trunk-io[bot] merged 9 commits into
masterfrom
tue/replay-vision-suggestion-quality

Conversation

@TueHaulund

Copy link
Copy Markdown
Contributor

Problem

  • The first person to open a scanner's Search tab usually sees the fixed example phrases, not phrases from their own data.
    • Suggestions were generated only for scanners someone had already viewed, and only at the next hourly run.
  • The "All scanners" view merged phrases from the team's most recently swept scanners, which usually had none stored.
  • Monitor phrases often described routine sessions ("browsing job listings on careers page") instead of what the scanner looks for.
    • The model saw only observation text: no scanner purpose, and no outcome. Most monitor observations are no verdicts that describe normal sessions.

Changes

  • Phrases now exist before anyone opens the Search tab.
    • Every enabled scanner with new observations is refreshed, whether or not anyone viewed it. Watched scanners go first.
    • The refresher runs every 10 minutes, not hourly. A scope that never had a model call retries after 10 minutes, and its first set needs 2 observations.
  • The "All scanners" view gets its own phrases, drawn from the team's most recently active scanners.
    • They live on a new TeamReplayVisionConfig team extension (migration 0100, a plain CREATE TABLE with no lock on posthog_team).
    • A viewer sees them only if they can read every scanner that fed them. Otherwise the view falls back to the per-scanner merge.
    • Experiment-targeted scanners, and rows captured under an experiment, never feed them. Deleting a source scanner clears them.
  • Phrases name the thing each scanner looks for.
    • The model now reads the scanner's name, type and prompt, and each observation's outcome label, e.g. [verdict=no].
    • The sample takes one row from each outcome in turn, so a rare outcome keeps its share. The model decides from the scanner's instructions which outcome matters, because that depends on how the question is phrased.
  • The refresh is more robust.
    • An empty model answer keeps the stored phrases.
    • The watermark uses completed_at, so a row that finishes after a refresh still counts as new.
    • Phrases that break the prompt's shape, such as a URL, an email, or more than 10 words, are dropped.
  • search_viewed warms the phrases the view actually shows.
  • Budget: the daily cap is now 40,000 model calls on gemini-3.5-flash-lite, with 200 per run at a concurrency of 16.
  • Mechanical: the workflow input takes an optional scanner_id (none means the team scope), and its result keys are strings.

Note

#106990 adds a 10-minute jitter to this refresher's schedule. With this PR's 10-minute interval, whichever lands second should shrink that jitter.

No screenshot: the empty state looks the same, and only which phrases it shows changes.

How did you test this code?

  • test_search_suggestions.py:
    • scopes are listed without a view, retry soon when they never had a model call, and are skipped when disabled or below 2 new observations;
    • a rare outcome keeps its share of the labeled sample;
    • an empty model answer keeps stored phrases and waits a full interval;
    • a row that completes after a refresh still counts;
    • team phrases skip targeted scanners and experiment rows, show only to viewers of every source, and clear when a source is deleted;
    • the "All scanners" view serves and warms team phrases;
    • scanner text and model output with markup, URLs or emails are defanged or dropped.
  • A read-only comparison on 20 production scanners, run on a replay-vision worker with the same model and data for both prompts:
    • Among each phrase's top matches, the share of monitor yes findings rose from 24% to 33%.
    • The average score of frustration-scorer matches rose from 1.2 to 2.2 out of 10.
  • Not checked: the load of the candidate queries at production scale, and the first-deploy backlog, where every active scope becomes due at once.

👉 Stay up-to-date with PostHog coding conventions for a smoother review.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

None. The product README covers the new model and module.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5)

  • Skills: /django-migrations, /writing-tests, /writing-code-comments, /reviewing-with-coderabbit, /writing-pr-descriptions, /simplify.
  • An earlier version filtered the sample to findings (yes verdicts, top scores). It moved to outcome labels because "the interesting outcome" depends on each scanner's question.
  • An earlier version also checked each phrase with a ClickHouse search before storing it. The production comparison showed it dropped 1 of 102 phrases, so it was removed.
  • CodeRabbit CLI (--deep): one finding, fixed. Experiment rows could feed team phrases after a scanner's targeting was removed.
  • Shepherd review (Claude and Codex), all fixed:
    • disabled scanners fed team phrases;
    • an empty answer wiped phrases;
    • the created_at watermark missed late rows;
    • stuck scopes were relisted every run;
    • team phrases outlived a deleted source;
    • model output had no shape check;
    • a full run of slow calls could exceed the execution timeout.
  • Duplicate search: no open PR covers this.
  • Public artifact: the test data is invented. The production comparison output is not committed.

…ctive scanner and team

Generate phrases for every active scanner and for each team's
cross-scanner view on a 10-minute schedule, without waiting for a view.
Feed the model an outcome-labeled sample balanced across outcomes, plus
the scanner's name and prompt, and keep only candidate phrases that a
real search shows to find something. Scopes without phrases retry after
10 minutes and get a first set from 2 observations.
…estions

A comparison on real scanners showed nearly every phrase clears the
search distance ceiling, so the check dropped almost nothing for a
ClickHouse scan per refresh.
Keep stored phrases when the model finds no theme and back off a full
interval, watermark on completed_at so late rows still count, list only
enabled scopes with enough new rows, drop phrases that break the prompt's
shape, clear team phrases when a source scanner is deleted, and fit a full
run of slow calls inside the execution timeout.
@TueHaulund TueHaulund self-assigned this Sep 27, 2026
@trunk-io

trunk-io Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@TueHaulund TueHaulund added the stamphog Request AI approval (no full review) label Sep 27, 2026
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

✅ Complexity (TypeScript) — clean

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Comment density — 5% of added code lines are comments (23 of 490)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/replay_vision/backend/tests/test_search_suggestions.py 8 143
products/replay_vision/backend/search_suggestions.py 7 219
products/replay_vision/backend/temporal/constants.py 4 8
products/replay_vision/backend/models/team_replay_vision_config.py 2 26
products/replay_vision/backend/temporal/activities/refresh_search_suggestions.py 1 37
products/replay_vision/backend/temporal/search_suggestions_types.py 1 8

This check does not block merging. It updates on every push and clears when the share drops.

✅ Bundle size — 🟢 -41 B (-0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 68.88 MiB · 🟢 -41 B (-0.0%)

No file changed by more than 1000 B.

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.57 MiB · 22 files no change █████████░ 85.2% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.58 MiB · 628 files no change █████████░ 88.7% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.37 MiB · 2,326 files 🟢 -41 B (-0.0%) █████████░ 88.4% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
267.6 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.4 KiB src/lib/api.ts
85.5 KiB src/products.tsx
69.1 KiB src/lib/lemon-ui/icons/icons.tsx
63.9 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.3 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
271.7 KiB src/taxonomy/core-filter-definitions-by-group.json
267.6 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.4 KiB src/lib/api.ts
98.5 KiB ../packages/quill/packages/quill/dist/index.js
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
85.5 KiB src/products.tsx

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.37 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.37 MiB · 19 files no change ████░░░░░░ 41.4% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
791.7 KiB dist/toolbar/toolbar-app-XDB3CY2H.css
650.8 KiB dist/toolbar/chunk-chunk-OLZXKW3U.js
483.6 KiB dist/toolbar/chunk-chunk-LP5DDLVQ.js
138.3 KiB dist/toolbar/chunk-chunk-FVYKO6VU.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-M6TDBA3J.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-GL4SRUHV.js
21.0 KiB dist/toolbar/chunk-chunk-Z4YQYAC3.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🟢 -437 B (-0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 944.90 MiB · 🟢 -437 B (-0.0%)

ℹ️ MCP UI apps size — 33 app(s), 17630.1 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 196.2 KB
action 454.1 KB 196.2 KB
action-list 564.2 KB 196.2 KB
cohort 453.1 KB 196.2 KB
cohort-list 563.2 KB 196.2 KB
email-template 452.9 KB 196.2 KB
error-details 469.1 KB 196.2 KB
error-issue 453.8 KB 196.2 KB
error-issue-list 564.1 KB 196.2 KB
experiment 561.3 KB 196.2 KB
experiment-list 564.9 KB 196.2 KB
experiment-results 566.3 KB 196.2 KB
feature-flag 566.8 KB 196.2 KB
feature-flag-list 570.5 KB 196.2 KB
feature-flag-testing 457.3 KB 196.2 KB
inline-scan 453.6 KB 196.2 KB
insight-actors 562.3 KB 196.2 KB
invite-email-preview 452.3 KB 196.2 KB
llm-costs 559.3 KB 196.2 KB
session-recording 455.3 KB 196.2 KB
survey 454.7 KB 196.2 KB
survey-global-stats 561.9 KB 196.2 KB
survey-list 564.9 KB 196.2 KB
survey-stats 561.9 KB 196.2 KB
trace-span 453.5 KB 196.2 KB
trace-span-list 564.1 KB 196.2 KB
vision-observation-list 563.3 KB 196.2 KB
workflow 453.4 KB 196.2 KB
workflow-list 563.5 KB 196.2 KB
loops-review 457.8 KB 196.2 KB
query-results 774.1 KB 196.2 KB
render-ui 857.0 KB 196.2 KB
visual-review-snapshots 457.9 KB 196.2 KB
✅ Playwright — all passed

All tests passed.

View test results →

⚠️ Backend coverage — 91.0% of changed backend lines covered — 20 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ██████████████████░░ 91.0% (228 / 248)

File Patch Uncovered changed lines
products/replay_vision/backend/temporal/activities/refresh_search_suggestions.py 42.1% 46, 75–80, 82–85
products/replay_vision/backend/search_suggestions.py 90.6% 151, 316–317, 380–383, 385–386

🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 36335838068 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
warehouse_sources_queue ██████░░░░░░░░░░░░░░ 29.1% 92 / 316
demo ████████████░░░░░░░░ 57.8% 1,545 / 2,673
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
aeo ██████████████░░░░░░ 70.5% 467 / 662
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
batch_exports ████████████████░░░░ 81.2% 21,452 / 26,424
apm █████████████████░░░ 84.1% 1,306 / 1,553
ml_inference █████████████████░░░ 87.2% 482 / 553
cdp ██████████████████░░ 88.2% 4,548 / 5,155
mcp_analytics ██████████████████░░ 88.9% 4,910 / 5,523
product_tours ██████████████████░░ 89.3% 1,331 / 1,491
dashboards ██████████████████░░ 89.5% 6,839 / 7,641
signals ██████████████████░░ 89.9% 54,456 / 60,606
data_warehouse ██████████████████░░ 89.9% 13,912 / 15,470
notebooks ██████████████████░░ 90.2% 15,287 / 16,945
cohorts ██████████████████░░ 90.4% 8,420 / 9,316
streamlit_apps ██████████████████░░ 90.7% 2,625 / 2,895
managed_warehouse ██████████████████░░ 90.9% 10,215 / 11,234
tasks ██████████████████░░ 91.1% 73,950 / 81,169
data_modeling ██████████████████░░ 91.5% 10,525 / 11,498
business_knowledge ██████████████████░░ 91.6% 6,899 / 7,528
engineering_analytics ██████████████████░░ 91.7% 11,002 / 11,999
exports ██████████████████░░ 91.8% 9,685 / 10,555
ai_training ██████████████████░░ 92.2% 356 / 386
conversations ███████████████████░ 92.5% 28,726 / 31,047
early_access_features ███████████████████░ 92.6% 1,341 / 1,448
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
visual_review ███████████████████░ 92.8% 9,244 / 9,966
canvas ███████████████████░ 92.8% 6,877 / 7,409
approvals ███████████████████░ 93.0% 3,919 / 4,214
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
error_tracking ███████████████████░ 93.1% 15,783 / 16,950
notifications ███████████████████░ 93.2% 1,145 / 1,229
slack_app ███████████████████░ 93.2% 13,677 / 14,674
stamphog ███████████████████░ 93.2% 7,885 / 8,456
surveys ███████████████████░ 93.3% 6,571 / 7,040
context_layer ███████████████████░ 93.8% 3,373 / 3,595
web_analytics ███████████████████░ 93.9% 21,653 / 23,051
alerts ███████████████████░ 94.0% 8,541 / 9,082
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.4% 8,940 / 9,472
ai_observability ███████████████████░ 94.4% 22,139 / 23,454
wizard ███████████████████░ 94.7% 6,151 / 6,496
reminders ███████████████████░ 94.8% 760 / 802
workflows ███████████████████░ 94.8% 13,615 / 14,360
review_hog ███████████████████░ 94.9% 11,490 / 12,109
annotations ███████████████████░ 95.1% 817 / 859
customer_analytics ███████████████████░ 95.1% 24,756 / 26,028
endpoints ███████████████████░ 95.1% 9,211 / 9,681
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
marketing_analytics ███████████████████░ 95.3% 19,216 / 20,161
posthog_ai ███████████████████░ 95.4% 2,489 / 2,610
tracing ███████████████████░ 95.4% 3,483 / 3,650
growth ███████████████████░ 95.4% 9,812 / 10,282
logs ███████████████████░ 95.4% 15,290 / 16,022
experiments ███████████████████░ 95.5% 32,956 / 34,526
actions ███████████████████░ 95.5% 756 / 792
data_catalog ███████████████████░ 95.6% 4,402 / 4,606
skills ███████████████████░ 95.8% 6,972 / 7,274
messaging ███████████████████░ 95.9% 3,766 / 3,927
replay_vision ███████████████████░ 95.9% 27,154 / 28,310
product_analytics ███████████████████░ 96.2% 28,495 / 29,617
autoresearch ███████████████████░ 96.2% 8,115 / 8,434
revenue_analytics ███████████████████░ 96.4% 1,876 / 1,946
access_control ███████████████████░ 96.4% 7,122 / 7,386
user_interviews ███████████████████░ 96.5% 2,859 / 2,963
feature_flags ███████████████████░ 96.5% 25,499 / 26,416
warehouse_sources ███████████████████░ 97.2% 448,749 / 461,565
data_quality ████████████████████ 97.7% 7,592 / 7,774
links ████████████████████ 97.9% 234 / 239
security ████████████████████ 98.1% 1,258 / 1,283
metrics ████████████████████ 98.1% 4,085 / 4,166
analytics_platform ████████████████████ 98.3% 2,783 / 2,832
pulse ████████████████████ 98.5% 2,043 / 2,075
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

⚠️ Django migration SQL — 1 new migration to review

We've detected new migrations on this PR. Review the SQL output for each migration:

products/replay_vision/backend/migrations/0100_team_replay_vision_config.py

/opt/hostedtoolcache/Python/3.13.13/x64/lib/python3.13/site-packages/infi/clickhouse_orm/__init__.py:1: UserWarning: pkg_resources is deprecated as an API. See https://setuptools.pypa.io/en/latest/pkg_resources.html. The pkg_resources package is slated for removal as early as 2025-11-30. Refrain from using this package or pin to Setuptools<81.
  __import__("pkg_resources").declare_namespace(__name__)
System check identified some issues:

WARNINGS:
?: (axes.W001) You are using the django-axes cache handler for login attempt tracking. Your cache configuration is however invalid and will not work correctly with django-axes. This can leave security holes in your login systems as attempts are not tracked correctly. Reconfigure settings.AXES_CACHE and settings.CACHES per django-axes configuration documentation.
?: (staticfiles.W004) The directory '/home/runner/work/posthog/posthog/frontend/dist' in the STATICFILES_DIRS setting does not exist.
BEGIN;
--
-- Create model TeamReplayVisionConfig
--
CREATE TABLE "replay_vision_teamreplayvisionconfig" ("team_id" integer NOT NULL PRIMARY KEY, "search_suggestions" jsonb NOT NULL, "search_suggestions_sources" jsonb NOT NULL, "search_suggestions_watermark" timestamp with time zone NULL, "search_suggestions_generated_at" timestamp with time zone NULL);
COMMIT;

Last updated: 2026-09-27 17:12 UTC (d76b4f8)

✅ Django migration risk — migration analysis complete

We've analyzed your migrations for potential risks.

Summary: 1 Safe | 0 Needs Review | 0 Blocked

✅ Safe

Brief or no lock, backwards compatible

replay_vision.0100_team_replay_vision_config
  └─ #1 ✅ CreateModel
     Creating new table is safe
     model: TeamReplayVisionConfig
  │
  └──> ℹ️  INFO:
       ℹ️  Skipped operations on newly created tables (empty tables
       don't cause lock contention).

Last updated: 2026-09-27 17:12 UTC (d76b4f8)

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

The Migration risk check has not finished for this commit, so stamphog cannot tell a safe migration from a risky one yet. The review runs again on the next push, or you can re-request it once the check reports.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✓ 517L, 11F substantive, 677L/14F incl. docs/generated/snapshots — within ceiling
tier ✗ classified as T2-never: T2-never (677L, 14F, single-area, feat)
stamphog 2.2.0 .stamphog/policy.yml @ unknown · reviewed head 9d547c8

@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested review from a team, arnohillen, fasyy612 and ksvat and removed request for a team September 27, 2026 15:49
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

🦔 Hogbox preview · ✅ ready

▶ Open the preview

🔑 Login test@posthog.com / 12345678 (demo data)
🧩 Running this PR's backend and frontend, on the PostHog :master base
🔗 Link stable across rebuilds — a re-push swaps the box underneath, the URL stays
🔒 Access tailnet only (PostHog VPN)
🛠️ Admin inspect & debug state in hogland
💤 Idle sleeps after ~30 min idle (snapshot to S3, zero node cost) and wakes on your next visit in ~30s, behind a brief "waking up" screen

commit d76b4f8 · box box-8c6383631c4f · ready in 578s (push → usable) · build log · rebuilds on every push, torn down on close

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

The Migration risk check has not finished for this commit, so stamphog cannot tell a safe migration from a risky one yet. The review runs again on the next push, or you can re-request it once the check reports.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✓ 517L, 11F substantive, 685L/16F incl. docs/generated/snapshots — within ceiling
tier ✗ classified as T2-never: T2-never (685L, 16F, single-area, feat)
stamphog 2.2.0 .stamphog/policy.yml @ unknown · reviewed head b6e00e0

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[High risk] Adds database model and scheduled workflow for team-wide search suggestions.

The PR is not ready to merge because teams with observations spread across scanners may never receive their first team phrases.

Reviews (1) · Last reviewed commit: "fix(replay-vision): harden search sugges..."

Comment thread products/replay_vision/backend/search_suggestions.py
Comment thread products/replay_vision/backend/search_suggestions.py
Comment thread products/replay_vision/backend/search_suggestions.py Outdated
@trunk-io

trunk-io Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
ErrorTrackingWidget loads the declared renderer and falls back for the MCP issues-list response The test exceeded the maximum allowed time of 5000 milliseconds and timed out. Logs ↗︎
the activity log logic humanizing insights can handle change of insight query as a query wrapped in an InsightVizNode The test exceeded the maximum allowed time of 15 seconds and timed out. Logs ↗︎
the activity log logic humanizing insights can handle change of a SQL insight query The test failed because the code could not find the 'models.cohortsModel' path in the store. Logs ↗︎

View Full Report ↗︎ ⋅ Docs

@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

📝 Walkthrough

Walkthrough

Adds per-team storage and scheduled generation of cross-scanner search suggestions. Updates scanner suggestion eligibility, observation sampling, and prompt construction. The API selects team suggestions for all-scanner requests when the source scanners are readable, and otherwise falls back to merged scanner suggestions. View tracking warms vectors for the displayed suggestions when AI data processing is approved.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to d76b4

Deleting a scanner during refresh can temporarily leave the All scanners view using per-scanner suggestions instead of team suggestions. The fallback keeps Search usable, but the refresh race should be fixed or explicitly accepted before merging.

Security Architecture Review

Security architecture risk: 🔵 Low · up to d76b4

The new team suggestions retain scanner-level access checks, limiting who can see them. A deletion race could nevertheless leave recording-derived suggestions stored after their source scanner is removed. The available evidence does not establish a broader access bypass.

Retained concerns

  • Low · security · inferred: A refresh can restore deleted-scanner-derived team suggestions after the deletion cleanup has run, leaving recording-derived data and stale provenance in team storage.
Security review details

Security Blast Radius

  • inferred — Team-derived phrases may combine observations from several scanners in one team, but serving them is bounded by the viewer's access to every recorded source. The deletion race primarily affects retention in that team's extension state.

Security Findings and Attack Paths

  • inferred — If a source is deleted after refresh's existence check but before its final update, refresh can repopulate phrases that the deletion receiver cleared. A subsequent request that recomputes readable scanner IDs will normally reject those team phrases because the deleted source is absent.

Trust Boundaries and Controls

  • observed — Team phrases are withheld unless all stored sources are in the viewer-readable scanner set. Vector warming uses the same displayed set and is conditional on AI-processing approval.

Resilience and Maintainability Implications

  • inferred — Deletion cleanup and refresh both update the same team-owned fields without a shared version or atomic source-validation-and-write step. Failure backoff and schedule overlap limits do not close that source-lifecycle race.

Hardening Proposals

  • proposed — Make final source validation and team-state replacement conditional on an unchanged source lifecycle, or serialize them with deletion cleanup, so a completed deletion cannot be followed by a stale refresh write.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description follows the required structure and clearly explains the problem, user-visible changes, testing, release status, documentation, and agent decisions. It omits a session link and explicit…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/replay_vision/backend/temporal/activities/refresh_search_suggestions.py-29-36 (1)

29-36: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Allow one team slot when the remaining budget is positive.

When the daily remainder is 1–3, remaining // 4 is zero, so stale_team_candidates cannot schedule a due team. If no scanner is eligible, the activity returns no work despite unused budget. The 10-minute workflow can repeat this until the daily budget resets. Each team refresh consumes one model call, not four.

Suggested fix
-    teams = [RefreshScannerSuggestionsInputs(team_id=team_id) for team_id in stale_team_candidates(remaining // 4)]
+    teams = [
+        RefreshScannerSuggestionsInputs(team_id=team_id)
+        for team_id in stale_team_candidates(max(1, remaining // 4))
+    ]

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 6151c97f-21ea-49c9-b462-a9808134c88e

📥 Commits

Reviewing files that changed from the base of the PR and between 49b83c3 and b6e00e0.

⛔ Files ignored due to path filters (2)
  • products/replay_vision/frontend/generated/api.ts is excluded by !**/generated/**
  • products/replay_vision/frontend/generated/api.zod.ts is excluded by !**/generated/**
📒 Files selected for processing (14)
  • products/replay_vision/README.md
  • products/replay_vision/backend/api/observations.py
  • products/replay_vision/backend/apps.py
  • products/replay_vision/backend/migrations/0100_team_replay_vision_config.py
  • products/replay_vision/backend/migrations/max_migration.txt
  • products/replay_vision/backend/models/__init__.py
  • products/replay_vision/backend/models/team_replay_vision_config.py
  • products/replay_vision/backend/search_suggestion_cleanup.py
  • products/replay_vision/backend/search_suggestions.py
  • products/replay_vision/backend/temporal/activities/refresh_search_suggestions.py
  • products/replay_vision/backend/temporal/constants.py
  • products/replay_vision/backend/temporal/search_suggestions.py
  • products/replay_vision/backend/temporal/search_suggestions_types.py
  • products/replay_vision/backend/tests/test_search_suggestions.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

…tisfy repo checks

Count new rows across a team's scanners when listing team suggestions, skip
a team refresh whose source scanner was deleted mid-call, register the
team extension and delete receiver with the repo checks, and regenerate
the API types for the updated view docstring.
…lity' into tue/replay-vision-suggestion-quality

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

The Migration risk check has not finished for this commit, so stamphog cannot tell a safe migration from a risky one yet. The review runs again on the next push, or you can re-request it once the check reports.

  • 👍 on the PR from greptile-apps[bot].
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✓ 526L, 12F substantive, 708L/18F incl. docs/generated/snapshots — within ceiling
tier ✗ classified as T2-never: T2-never (708L, 18F, cross-cutting, feat)
stamphog 2.2.0 .stamphog/policy.yml @ unknown · reviewed head f92a674

@github-actions
github-actions Bot requested a deployment to preview-pr-107428 September 27, 2026 17:05 In progress

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

The Migration risk check has not finished for this commit, so stamphog cannot tell a safe migration from a risky one yet. The review runs again on the next push, or you can re-request it once the check reports.

  • 👍 on the PR from greptile-apps[bot].
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✓ 526L, 12F substantive, 708L/18F incl. docs/generated/snapshots — within ceiling
tier ✗ classified as T2-never: T2-never (708L, 18F, cross-cutting, feat)
stamphog 2.2.0 .stamphog/policy.yml @ unknown · reviewed head d76b4f8

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/replay_vision/backend/search_suggestions.py-327-327 (1)

327-327: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Make source validation and the config write atomic.

refresh_team_suggestions validates source existence with a separate count() before updating TeamReplayVisionConfig. A deletion can run the post_delete cleanup between those statements. When generated phrases are non-empty, the refresh can then restore the deleted source ID and watermark.

The display helper rejects those team phrases when the deleted ID is not readable, so the caller falls back to per-scanner suggestions. The restored watermark and timestamp also delay another team refresh until the due interval and enough newer observations exist.

Lock the source rows immediately before validation and update the config in the same short transaction. Do not hold the locks during the model call.

Suggested fix
+from django.db import transaction
 from django.db.models import DateTimeField, Exists, F, OuterRef, Q, QuerySet, Subquery, Value
@@
-    if ReplayScanner.objects.filter(team_id=team.id, id__in=source_ids).count() < len(source_ids):
-        # A source was deleted during the model call, and its delete already cleared the team's phrases.
-        TeamReplayVisionConfig.objects.filter(pk=team.id).update(search_suggestions_generated_at=timezone.now())
-        return False
-    TeamReplayVisionConfig.objects.filter(pk=team.id).update(
-        search_suggestions_watermark=newest,
-        search_suggestions_generated_at=timezone.now(),
-        **({"search_suggestions": phrases, "search_suggestions_sources": source_ids} if phrases else {}),
-    )
+    with transaction.atomic():
+        existing_source_ids = set(
+            ReplayScanner.objects.select_for_update()
+            .filter(team_id=team.id, id__in=source_ids)
+            .values_list("id", flat=True)
+        )
+        if len(existing_source_ids) < len(source_ids):
+            TeamReplayVisionConfig.objects.filter(pk=team.id).update(
+                search_suggestions_generated_at=timezone.now()
+            )
+            return False
+        TeamReplayVisionConfig.objects.filter(pk=team.id).update(
+            search_suggestions_watermark=newest,
+            search_suggestions_generated_at=timezone.now(),
+            **({"search_suggestions": phrases, "search_suggestions_sources": source_ids} if phrases else {}),
+        )

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: aacaec9c-3e10-4d44-82cc-ad7e183a094a

📥 Commits

Reviewing files that changed from the base of the PR and between b6e00e0 and d76b4f8.

📒 Files selected for processing (4)
  • .github/scripts/check-idor-model-coverage.py
  • posthog/test/repo_invariants/setup_receivers_baseline.txt
  • products/replay_vision/backend/search_suggestions.py
  • products/replay_vision/backend/tests/test_search_suggestions.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@fasyy612 fasyy612 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Amazing stuff 💯

@trunk-io
trunk-io Bot merged commit 192c60a into master Sep 28, 2026
282 checks passed
@trunk-io
trunk-io Bot deleted the tue/replay-vision-suggestion-quality branch September 28, 2026 10:11
@deployment-status-posthog

deployment-status-posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-09-28 10:40 UTC Run
prod-us ✅ Deployed 2026-09-28 10:51 UTC Run
prod-eu ✅ Deployed 2026-09-28 10:51 UTC Run

This branch was successfully deployed

1 active deployment
preview-pr-107428 — d76b4f81 Deployed Sep 27, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stamphog Request AI approval (no full review)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants