Skip to content

feat(replay-vision): lead observations with a condensed prompt question - #108268

Open
TueHaulund wants to merge 8 commits into
masterfrom
tue/observation-stories-long-prompt
Open

TueHaulund wants to merge 8 commits into
masterfrom
tue/observation-stories-long-prompt

Conversation

@TueHaulund

@TueHaulund TueHaulund commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • On an observation, the scanner's prompt sits above the answer as a raw multi-paragraph block, so the question gets lost.
  • The result styles clash: green verdict text in a different green than elsewhere, heavy orange category pills, and a confidence tag parked in the facts list.
  • The player dock, the What to watch feed and the scanner list show raw prompt text, or only the scanner name, where a reader needs to know what was asked.

Changes

Condensed prompt question

  • Every scanner now carries its prompt condensed into one question, such as "Did the user struggle to complete checkout?".
  • The question is written in the same save as the prompt: on create, prompt edits, duplicate, inline scans and applied recommendations.
  • Gemini 3.5 Flash Lite writes it before the save's transaction opens, with a 10 second timeout.
  • A failed or unusable answer, or no AI consent, falls back to the prompt's first line, so a save never fails.
  • The five app templates map to hand-written questions, so a template copy needs no model call.
  • A hash of the source prompt is stored with the question. The observation API returns prompt_question only when it matches the prompt that observation was scanned with.
  • backfill_replay_scanner_prompt_questions fills existing scanners. It condenses each distinct prompt once and does not bump the scanner version.

Observation page

  • "Question · Show full prompt" leads the card, and the full prompt opens in a modal. Old observations show the prompt's first paragraph.
  • The confidence badge moves next to the result heading. The scorer's scale name joins the heading, capitalized.
  • The Question, result and Reasoning headers are larger and bolder than the other labels.
  • The verdict is colored text with an icon, using the same green and red as LemonTag. Categories are grey filled pills.
  • Reasoning collapses only past about 20 lines, which few scans reach.
  • The "Rate more results for this scanner" link and its flag constant are removed.
  • Removed with their elements: the vision-observation-calibration-entry-point and vision-observation-prompt-toggle data-attrs.

Other surfaces

  • The player dock card uses the same question row.
  • What to watch cards show the question beside the result, with the scanner name as a tooltip.
  • The scanner list shows the question only when a scanner has no description.
  • New Slack alerts get a Question line, and new webhooks get scanner_question. Existing alert destinations are left unchanged.
  • The scanner overview's side column moves to the left and grows to 32rem on wide screens.

Mechanical: generated API types, story fixtures (longer prompts, one-paragraph and long reasoning, an inconclusive monitor), and a new dock card story.

1-observation-page
2-dock-card
3-watch-feed
4-scanner-list
5-scanner-overview

Note

After deploy, run python manage.py backfill_replay_scanner_prompt_questions --dry-run, then without --dry-run. Until then, existing scanners show their prompt text where the question would go.

How did you test this code?

  • New test_prompt_questions.py (15 cases) covers the fallbacks, including a failed client setup, the consent gate, template lookups, the create and edit write paths, question matching on observations and scanners, and the backfill's one call per distinct prompt.
  • The existing replay_vision backend suites for the scanner API, prompt suggestions, scanning, MCP tools and alerts pass locally, as does the product's frontend Jest suite.
  • Screenshots are Storybook renders with mock data, checked in light and dark mode for the pills.
  • Not checked: the live app, and the model's question quality on real prompts.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

The What to watch feed itself stays behind its existing experiment flag.

Automatic notifications

  • Publish to changelog?

Docs update

None.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Claude Opus 5.5

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@TueHaulund TueHaulund self-assigned this Sep 29, 2026
@trunk-io

trunk-io Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

🧪 Running tests on this pull request (testing on PR #108681) - details.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 11 functions above the limit (max 55)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
ObservationPrimaryOutput products/replay_vision/frontend/components/ObservationCard.tsx:116 55 10
watchReasonCopy products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:44 32 10
ReplayObservationSceneComponent products/replay_vision/frontend/observations/ReplayObservation.tsx:53 27 10
ObservationDockCard products/replay_vision/frontend/components/ObservationCard.tsx:393 26 10
readCitations products/replay_vision/frontend/utils/observation.ts:202 23 10
WatchFeedCard products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:205 20 10
ReplayScannersScene products/replay_vision/frontend/replay_scanners/ReplayScannersScene.tsx:133 17 10
observationClipboardText products/replay_vision/frontend/utils/observation.ts:330 17 10
watchCardHeadline products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:162 15 10
ObservationStatusTag products/replay_vision/frontend/components/ObservationCard.tsx:29 12 10
HeadlineValue products/replay_vision/frontend/observations/ObservationHeadline.tsx:91 11 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Comment density — 5% of added code lines are comments (50 of 992)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/replay_vision/backend/prompt_questions.py 12 208
products/replay_vision/frontend/observations/ObservationHeadline.tsx 9 78
products/replay_vision/frontend/replay_scanners/ReplayScannersScene.stories.tsx 4 128
products/replay_vision/backend/api/scanners.py 3 33
products/replay_vision/frontend/components/ObservationPrompt.tsx 3 36
products/replay_vision/backend/api/prompt_suggestions.py 2 14
products/replay_vision/backend/models/replay_scanner.py 2 19
products/replay_vision/frontend/components/LabeledRow.tsx 2 14

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Bundle size — 🔺 +4.5 KiB (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 68.92 MiB · 🔺 +4.5 KiB (+0.0%)

File Size Δ vs base
posthog-app/_parent/products/business_knowledge/frontend/scenes/BusinessKnowledgePlaygroundScene.js removed 🟢 -26.2 KiB (-100.0%)
posthog-app/_parent/products/business_knowledge/frontend/scenes/playground/BusinessKnowledgePlaygroundScene.js 26.1 KiB 🔺 +26.1 KiB (new)
posthog-app/_parent/products/business_knowledge/frontend/scenes/BusinessKnowledgeScene.js removed 🟢 -18.7 KiB (-100.0%)
posthog-app/_parent/products/business_knowledge/frontend/scenes/sources/BusinessKnowledgeScene.js 18.6 KiB 🔺 +18.6 KiB (new)
posthog-app/_parent/products/business_knowledge/frontend/scenes/KnowledgeSourceScene.js removed 🟢 -18.0 KiB (-100.0%)
posthog-app/_parent/products/business_knowledge/frontend/scenes/source/KnowledgeSourceScene.js 18.0 KiB 🔺 +18.0 KiB (new)
posthog-app/_parent/products/business_knowledge/frontend/scenes/BusinessKnowledgeSettingsScene.js removed 🟢 -16.8 KiB (-100.0%)
posthog-app/_parent/products/business_knowledge/frontend/scenes/settings/BusinessKnowledgeSettingsScene.js 16.8 KiB 🔺 +16.8 KiB (new)
render-query/src/render-query/render-query.js 20.16 MiB 🔺 +2.3 KiB (+0.0%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.58 MiB · 22 files 🔺 +564 B (+0.0%) █████████░ 85.8% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.52 MiB · 629 files 🔺 +456 B (+0.0%) █████████░ 87.4% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.35 MiB · 2,339 files 🔺 +2.1 KiB (+0.0%) █████████░ 88.1% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
88.5 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.4 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
272.2 KiB src/taxonomy/core-filter-definitions-by-group.json
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
98.8 KiB ../packages/quill/packages/quill/dist/index.js
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
88.5 KiB src/products.tsx

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.16 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.16 MiB · 19 files 🔺 +436 B (+0.0%) ████░░░░░░ 37.8% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
805.7 KiB dist/toolbar/toolbar-app-6NVL45AI.css
651.7 KiB dist/toolbar/chunk-chunk-GWJK2C7V.js
259.4 KiB dist/toolbar/chunk-chunk-RK3L52NC.js
138.3 KiB dist/toolbar/chunk-chunk-DCSVQMI3.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-5ZRCEWPZ.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-4YDOS2K2.js
21.0 KiB dist/toolbar/chunk-chunk-76J7HSSP.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +313.1 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 947.34 MiB · 🔺 +313.1 KiB (+0.0%)

ℹ️ MCP UI apps size — 33 app(s), 17633.1 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 199.2 KB
action 454.1 KB 199.2 KB
action-list 564.2 KB 199.2 KB
cohort 453.1 KB 199.2 KB
cohort-list 563.2 KB 199.2 KB
email-template 452.9 KB 199.2 KB
error-details 469.6 KB 199.2 KB
error-issue 454.5 KB 199.2 KB
error-issue-list 564.8 KB 199.2 KB
experiment 561.3 KB 199.2 KB
experiment-list 564.9 KB 199.2 KB
experiment-results 566.3 KB 199.2 KB
feature-flag 566.8 KB 199.2 KB
feature-flag-list 570.5 KB 199.2 KB
feature-flag-testing 457.3 KB 199.2 KB
inline-scan 453.6 KB 199.2 KB
insight-actors 562.3 KB 199.2 KB
invite-email-preview 452.3 KB 199.2 KB
llm-costs 559.3 KB 199.2 KB
session-recording 455.3 KB 199.2 KB
survey 454.7 KB 199.2 KB
survey-global-stats 561.9 KB 199.2 KB
survey-list 564.9 KB 199.2 KB
survey-stats 561.9 KB 199.2 KB
trace-span 453.5 KB 199.2 KB
trace-span-list 564.1 KB 199.2 KB
vision-observation-list 563.3 KB 199.2 KB
workflow 453.4 KB 199.2 KB
workflow-list 563.5 KB 199.2 KB
loops-review 457.8 KB 199.2 KB
query-results 774.1 KB 199.2 KB
render-ui 858.1 KB 199.2 KB
visual-review-snapshots 457.9 KB 199.2 KB
⚠️ Playwright — 1 flaky

🎭 Playwright report · View test results →

⚠️ 1 flaky test:

  • Edit mode button enters and exits edit mode (chromium)

These issues are not necessarily caused by your changes.
Annoyed by this section? Help fix flakies and failures and it will go green!

⚠️ Backend coverage — 94.0% of changed backend lines covered — 16 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ███████████████████░ 94.0% (259 / 275)

File Patch Uncovered changed lines
products/replay_vision/backend/management/commands/backfill_replay_scanner_prompt_questions.py 0.0% 1, 3, 5, 8–9, 11–13, 16–17, 19–20, 26–27
products/replay_vision/backend/prompt_questions.py 98.3% 179, 229

🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 36584183449 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
warehouse_sources_queue ██████░░░░░░░░░░░░░░ 29.1% 92 / 316
demo ███████████░░░░░░░░░ 52.8% 1,411 / 2,673
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
aeo ███████████████░░░░░ 76.3% 617 / 809
batch_exports ████████████████░░░░ 81.3% 21,543 / 26,502
apm █████████████████░░░ 84.1% 1,306 / 1,553
cdp ██████████████████░░ 88.2% 4,545 / 5,155
ml_inference ██████████████████░░ 88.4% 509 / 576
mcp_analytics ██████████████████░░ 89.0% 4,948 / 5,561
product_tours ██████████████████░░ 89.3% 1,331 / 1,491
dashboards ██████████████████░░ 89.5% 6,904 / 7,714
data_warehouse ██████████████████░░ 90.0% 14,027 / 15,589
signals ██████████████████░░ 90.2% 57,575 / 63,828
notebooks ██████████████████░░ 90.2% 15,287 / 16,945
cohorts ██████████████████░░ 90.4% 8,420 / 9,316
streamlit_apps ██████████████████░░ 90.6% 2,623 / 2,895
managed_warehouse ██████████████████░░ 91.0% 10,252 / 11,263
tasks ██████████████████░░ 91.1% 75,145 / 82,502
data_modeling ██████████████████░░ 91.4% 10,504 / 11,498
exports ██████████████████░░ 91.6% 9,680 / 10,562
engineering_analytics ██████████████████░░ 91.7% 11,032 / 12,030
ai_training ██████████████████░░ 92.2% 356 / 386
business_knowledge ██████████████████░░ 92.2% 7,684 / 8,330
conversations ███████████████████░ 92.5% 28,734 / 31,062
early_access_features ███████████████████░ 92.6% 1,332 / 1,439
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
visual_review ███████████████████░ 92.8% 9,247 / 9,969
stamphog ███████████████████░ 92.8% 8,109 / 8,742
canvas ███████████████████░ 92.8% 6,877 / 7,409
approvals ███████████████████░ 93.0% 3,974 / 4,271
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
notifications ███████████████████░ 93.2% 1,145 / 1,229
slack_app ███████████████████░ 93.2% 13,670 / 14,667
error_tracking ███████████████████░ 93.2% 16,359 / 17,547
surveys ███████████████████░ 93.3% 6,571 / 7,040
autoresearch ███████████████████░ 93.6% 8,481 / 9,061
context_layer ███████████████████░ 93.9% 3,415 / 3,638
web_analytics ███████████████████░ 94.0% 21,815 / 23,218
alerts ███████████████████░ 94.0% 8,569 / 9,114
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.4% 8,940 / 9,472
ai_observability ███████████████████░ 94.5% 24,473 / 25,896
wizard ███████████████████░ 94.7% 6,150 / 6,496
reminders ███████████████████░ 94.8% 760 / 802
workflows ███████████████████░ 94.9% 14,449 / 15,230
review_hog ███████████████████░ 95.0% 11,507 / 12,119
endpoints ███████████████████░ 95.1% 9,206 / 9,681
annotations ███████████████████░ 95.1% 817 / 859
customer_analytics ███████████████████░ 95.2% 25,078 / 26,349
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
marketing_analytics ███████████████████░ 95.3% 19,339 / 20,299
posthog_ai ███████████████████░ 95.3% 2,488 / 2,610
experiments ███████████████████░ 95.4% 32,452 / 34,015
actions ███████████████████░ 95.5% 756 / 792
logs ███████████████████░ 95.5% 15,399 / 16,130
data_catalog ███████████████████░ 95.5% 4,401 / 4,606
tracing ███████████████████░ 95.6% 3,536 / 3,699
replay_vision ███████████████████░ 95.6% 27,934 / 29,214
growth ███████████████████░ 95.7% 11,228 / 11,734
messaging ███████████████████░ 95.8% 3,798 / 3,963
skills ███████████████████░ 95.8% 6,972 / 7,274
product_analytics ███████████████████░ 96.0% 28,470 / 29,647
access_control ███████████████████░ 96.3% 7,112 / 7,386
revenue_analytics ███████████████████░ 96.4% 1,876 / 1,946
user_interviews ███████████████████░ 96.5% 2,867 / 2,971
feature_flags ███████████████████░ 96.5% 25,588 / 26,509
warehouse_sources ███████████████████░ 97.2% 458,860 / 471,856
data_quality ████████████████████ 97.6% 7,587 / 7,774
links ████████████████████ 97.9% 234 / 239
security ████████████████████ 98.0% 1,286 / 1,312
metrics ████████████████████ 98.0% 4,084 / 4,166
analytics_platform ████████████████████ 98.3% 2,784 / 2,833
pulse ████████████████████ 98.5% 2,046 / 2,078
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

⚠️ Django migration SQL — 1 new migration to review

We've detected new migrations on this PR. Review the SQL output for each migration:

products/replay_vision/backend/migrations/0101_replayscanner_prompt_question.py

/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/anyio/from_thread.py:119: SyntaxWarning: 'return' in a 'finally' block
  return result
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/structlog/stdlib.py:1166: UserWarning: Remove `format_exc_info` from your processor chain if you want pretty exceptions.
  ed = p(logger, meth_name, ed)  # type: ignore[arg-type]
2026-09-29T14:41:19.205040Z [error    ] Path must be a valid database or directory containing databases. [posthog.exceptions_capture] pid=8198 tid=140117135055744
Traceback (most recent call last):
  File "/home/runner/work/posthog/posthog/posthog/geoip.py", line 15, in <module>
    geoip: Optional[GeoIP2] = GeoIP2(cache=8)
                              ~~~~~~^^^^^^^^^
  File "/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/django/contrib/gis/geoip2.py", line 116, in __init__
    raise GeoIP2Exception(
        "Path must be a valid database or directory containing databases."
    )
django.contrib.gis.geoip2.GeoIP2Exception: Path must be a valid database or directory containing databases.
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/sshtunnel.py:1040: SyntaxWarning: 'return' in a 'finally' block
  return (ssh_host,
/opt/hostedtoolcache/Python/3.14.7/x64/lib/python3.14/site-packages/langchain_core/_api/deprecation.py:27: UserWarning: Core Pydantic V1 functionality isn't compatible with Python 3.14 or greater.
  from pydantic.v1.fields import FieldInfo as FieldInfoV1
System check identified some issues:

WARNINGS:
?: (axes.W001) You are using the django-axes cache handler for login attempt tracking. Your cache configuration is however invalid and will not work correctly with django-axes. This can leave security holes in your login systems as attempts are not tracked correctly. Reconfigure settings.AXES_CACHE and settings.CACHES per django-axes configuration documentation.
?: (staticfiles.W004) The directory '/home/runner/work/posthog/posthog/frontend/dist' in the STATICFILES_DIRS setting does not exist.
BEGIN;
--
-- Add field prompt_question to replayscanner
--
ALTER TABLE "replay_vision_replayscanner" ADD COLUMN "prompt_question" text DEFAULT '' NOT NULL;
--
-- Add field prompt_question_source to replayscanner
--
ALTER TABLE "replay_vision_replayscanner" ADD COLUMN "prompt_question_source" varchar(64) DEFAULT '' NOT NULL;
COMMIT;

Last updated: 2026-09-29 14:41 UTC (51f9dc4)

✅ Django migration risk — migration analysis complete

We've analyzed your migrations for potential risks.

Summary: 1 Safe | 0 Needs Review | 0 Blocked

✅ Safe

Brief or no lock, backwards compatible

replay_vision.0101_replayscanner_prompt_question
  └─ #1 ✅ AddField
     Adding NOT NULL field with constant default (safe in PG11+)
     model: replayscanner, field: prompt_question
  └─ #2 ✅ AddField
     Adding NOT NULL field with constant default (safe in PG11+)
     model: replayscanner, field: prompt_question_source

Last updated: 2026-09-29 14:42 UTC (51f9dc4)

@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[Medium risk] Adds AI-condensed prompt questions to scanner observations.

The PR is not safe to merge until scanner saves retain their fallback on client-construction failure and recommendation application cannot overwrite a newer question.

Reviews (1) · Last reviewed commit: "feat(replay-vision): lead observations w..."

Comment thread products/replay_vision/backend/prompt_questions.py Outdated
Comment thread products/replay_vision/backend/api/prompt_suggestions.py Outdated
Comment thread products/replay_vision/backend/prompt_questions.py Outdated
Comment thread products/replay_vision/backend/temporal/vision_alerts/activities.py Outdated
Comment thread products/replay_vision/backend/prompt_questions.py Outdated
@TueHaulund
TueHaulund marked this pull request as ready for review September 29, 2026 11:35
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions
github-actions Bot requested a deployment to preview-pr-108268 September 29, 2026 11:35 In progress
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

🦔 Hogbox preview · ✅ ready

▶ Open the preview

🔑 Login test@posthog.com / 12345678 (demo data)
🧩 Running this PR's backend and frontend, on the PostHog :master base
🔗 Link stable across rebuilds — a re-push swaps the box underneath, the URL stays
🔒 Access tailnet only (PostHog VPN)
🛠️ Admin inspect & debug state in hogland
💤 Idle sleeps after ~30 min idle (snapshot to S3, zero node cost) and wakes on your next visit in ~30s, behind a brief "waking up" screen

commit 51f9dc4 · box box-ef86d2ba990d · ready in 707s (push → usable) · build log · rebuilds on every push, torn down on close

@github-actions
github-actions Bot requested a deployment to preview-pr-108268 September 29, 2026 11:36 In progress
@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested review from a team, arnohillen, fasyy612 and ksvat and removed request for a team September 29, 2026 11:36
@TueHaulund TueHaulund added the stamphog Request AI approval (no full review) label Sep 29, 2026
@pr-assigner-resolver-posthog

Copy link
Copy Markdown

👀 Auto-assigned reviewers

These soft owners were skipped because they only have minor changes here. Nothing blocks merge, so self-assign if you'd like a look:

  • @PostHog/team-platform-features (posthog/models/owners.yaml)

Soft owners come from each directory's owners.yaml and each product's product.yaml (resolved nearest-file-wins). For a skipped owner, the locator is the file that decided it. Generated files and lockfiles are ignored when deciding ownership.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

Stamphog can't auto-review this pull request because three gates refused it. The deny-list gate matched the migrations path: 0101_replayscanner_prompt_question.py and max_migration.txt under products/replay_vision/backend/migrations/ are changed. The size gate refused it as well: 1027 substantive lines is over the 800-line ceiling, and the total is 1282 lines across 35 files. The tier gate classified it as T2-never because it is a cross-cutting feature.

The next step is to ask a human reviewer to look at it. If you want to shrink it, you could split the schema migration and backend question-writing from the frontend changes. A migration will still need a human reviewer, though.

  • 👍 on the PR from greptile-apps[bot].
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✗ too large for auto-review (1027L substantive in global — ceiling is 800L; 1027L, 30F total, 1282L/35F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1282L, 35F, cross-cutting, feat)
stamphog 2.3.0 .stamphog/policy.yml @ unknown · reviewed head 439bf4e

@stamphog stamphog Bot removed the stamphog Request AI approval (no full review) label Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions
github-actions Bot requested a deployment to preview-pr-108268 September 29, 2026 11:39 In progress
Comment thread products/replay_vision/backend/prompt_questions.py
@veria-ai

veria-ai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

PR overview

All previously flagged issues have been addressed. No open security concerns remain on this pull request.

Security review

No open security issues remain on this pull request.

Fixed/addressed: 1 · PR risk: 0/10

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions
github-actions Bot requested a deployment to preview-pr-108268 September 29, 2026 11:44 In progress
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions
github-actions Bot requested a deployment to preview-pr-108268 September 29, 2026 11:48 In progress
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

📝 Walkthrough

Walkthrough

Replay Vision now stores fingerprinted scanner prompt questions, generates or backfills them, and exposes them through scanner and observation APIs and alert webhooks. Frontend views display questions and update observation prompt, reasoning, confidence, and result presentation. The calibration entry point and its feature flag are removed.

Priority: ➖ Normal

Merge Risk: 🟡 Moderate · up to 51f9d

Scanner changes can fail during question-generation setup, and a suggestion apply can report an error after changing the prompt. Fix the fallback and save the question with the apply before merging.

🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description clearly covers the problem, user-visible changes, testing, screenshots, release status, documentation, and agent context. It also documents the backfill procedure and known testing lim…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@trunk-io

trunk-io Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx-323-326 (1)

323-326: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make the tooltip trigger keyboard-focusable.

When observation.prompt_question is present, the tooltip wraps a plain <span>. Base UI requires custom tooltip triggers to be focusable, so keyboard users cannot access scannerName through the current trigger.

Suggested fix
-                            <span className="text-xs truncate">{observation.prompt_question}</span>
+                            <span tabIndex={0} className="text-xs truncate">
+                                {observation.prompt_question}
+                            </span>
🧹 Nitpick comments (1)
products/replay_vision/frontend/observations/ObservationReasoning.tsx (1)

16-16: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Move expansion state to Kea logic.

expanded uses local React state. Store the expansion state in a Kea logic instead. As per coding guidelines: “Don't use useState or useEffect to store local state. It's false convenience. Take the extra 3 minutes and change it to a logic early on in the development.”

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 58c72c56-cca0-4ea4-a9ce-f2f2ba4648e4

📥 Commits

Reviewing files that changed from the base of the PR and between d84cd86 and 9a49920.

⛔ Files ignored due to path filters (1)
  • products/replay_vision/frontend/generated/api.schemas.ts is excluded by !**/generated/**
📒 Files selected for processing (34)
  • frontend/src/lib/constants.tsx
  • posthog/models/activity_logging/activity_log.py
  • products/replay_vision/backend/alert_destinations.py
  • products/replay_vision/backend/api/observations.py
  • products/replay_vision/backend/api/prompt_suggestions.py
  • products/replay_vision/backend/api/scanners.py
  • products/replay_vision/backend/inline_scan.py
  • products/replay_vision/backend/management/commands/backfill_replay_scanner_prompt_questions.py
  • products/replay_vision/backend/migrations/0101_replayscanner_prompt_question.py
  • products/replay_vision/backend/migrations/max_migration.txt
  • products/replay_vision/backend/models/replay_observation.py
  • products/replay_vision/backend/models/replay_scanner.py
  • products/replay_vision/backend/prompt_questions.py
  • products/replay_vision/backend/temporal/vision_alerts/activities.py
  • products/replay_vision/backend/tests/conftest.py
  • products/replay_vision/backend/tests/test_prompt_questions.py
  • products/replay_vision/frontend/components/LabeledRow.tsx
  • products/replay_vision/frontend/components/ObservationCard.tsx
  • products/replay_vision/frontend/components/ObservationDockCard.stories.tsx
  • products/replay_vision/frontend/components/ObservationPrompt.tsx
  • products/replay_vision/frontend/observations/ObservationFacts.tsx
  • products/replay_vision/frontend/observations/ObservationHeadline.tsx
  • products/replay_vision/frontend/observations/ObservationReasoning.tsx
  • products/replay_vision/frontend/observations/ReplayObservation.tsx
  • products/replay_vision/frontend/replay_scanners/ReplayScannersScene.stories.tsx
  • products/replay_vision/frontend/replay_scanners/ReplayScannersScene.tsx
  • products/replay_vision/frontend/replay_scanners/components/ClippedPreview.tsx
  • products/replay_vision/frontend/replay_scanners/components/FullPromptModal.tsx
  • products/replay_vision/frontend/replay_scanners/components/PromptPreview.tsx
  • products/replay_vision/frontend/replay_scanners/components/ScannerOverview.tsx
  • products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx
  • products/replay_vision/frontend/replay_scanners/scannerTemplates.ts
  • products/replay_vision/frontend/utils/observation.ts
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (1)
  • frontend/src/lib/constants.tsx

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

Comment thread products/replay_vision/frontend/components/ObservationPrompt.tsx
…n observations

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@posthog

posthog Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

✅ Visual changes approved by @TueHaulund — baseline updated in 51f9dc4.

View this run in PostHog

40 changed, 6 new, 2 removed.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

46 updated, 2 removed
Run: c5dbbf93-43d5-469c-8048-b43dc80bae83

Co-authored-by: TueHaulund <2675352+TueHaulund@users.noreply.github.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to GitHub limitations.

⚠️ Outside diff range comments (2)

🟠 Major · Catch setup failures inside condense_prompt. · prompt_questions.py:137-165

products/replay_vision/backend/prompt_questions.py:137-165
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Catch setup failures inside condense_prompt.

condense_prompt promises a fallback, but _take_model_call can raise when cache.set fails after cache.incr returns ValueError, and is_ai_data_processing_approved is also outside any catch boundary. These exceptions can abort scanner create/update, inline scanner creation, and backfill before their writes. Handle the consent, metering, and generation branch in the shared helper so all callers receive the fallback.

Suggested fix
     question = None
     # The prompt is the team's own text, but it still only goes to the model under the org's AI consent.
-    if is_ai_data_processing_approved(team_id) and (not metered or _take_model_call(team_id)):
-        question = _generate(prompt=prompt, scanner_type=scanner_type, team_id=team_id)
+    try:
+        if is_ai_data_processing_approved(team_id) and (not metered or _take_model_call(team_id)):
+            question = _generate(prompt=prompt, scanner_type=scanner_type, team_id=team_id)
+    except Exception:
+        logger.exception("replay_vision.prompt_question.generation_failed", team_id=team_id)
     if question is None:
🟡 Minor · Persist the question with the prompt configuration. · prompt_suggestions.py:388-439

products/replay_vision/backend/api/prompt_suggestions.py:388-439
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Persist the question with the prompt configuration.

products/replay_vision/backend/prompt_questions.py requires question fields to be generated before the save transaction and states that the question is written in the same save as its prompt. apply commits scanner_config and APPLIED before calling question_fields_for_save. An exception or persistence failure can therefore return an error after the apply committed, while the question fingerprint remains stale. The suggestion is already applied, so the client cannot retry it.

Resolve the candidate config and question before the write transaction. Revalidate the locked scanner and suggestion, then save scanner_config, prompt_question, prompt_question_source, and the applied status in the same transaction.

Suggested fix
         edited_config = input_serializer.validated_data.get("config")
+        if edited_config is not None:
+            config = dict(edited_config)
+        elif suggestion.suggested_config is not None:
+            config = dict(suggestion.suggested_config)
+        else:
+            config = {**(scanner.scanner_config or {}), "prompt": suggestion.suggested_prompt}
+        message = scanner_config_error(ScannerType(scanner.scanner_type), config)
+        if message:
+            raise ValidationError(f"This recommendation can't be applied: {message}")
+        question = question_fields_for_save(
+            team_id=self.team_id,
+            scanner_type=scanner.scanner_type,
+            scanner_config=config,
+            current_source=scanner.prompt_question_source,
+        )
         # Guards must run on locked rows: unlocked reads let two concurrent applies both pass,
         # and the second silently overwrites the first.
         with transaction.atomic():
             scanner = ReplayScanner.objects.select_for_update().get(team_id=self.team_id, id=scanner.id)
             suggestion = ReplayScannerPromptSuggestion.objects.select_for_update().get(
@@
-            if edited_config is not None:
-                config = dict(edited_config)
-            elif suggestion.suggested_config is not None:
-                config = dict(suggestion.suggested_config)
-            else:
-                config = {**(scanner.scanner_config or {}), "prompt": suggestion.suggested_prompt}
-            message = scanner_config_error(ScannerType(scanner.scanner_type), config)
-            if message:
-                raise ValidationError(f"This recommendation can't be applied: {message}")
             before_config = scanner.scanner_config or {}
             applied_fields = sorted(
                 key
                 for key in {*before_config, *config}
                 if before_config.get(key) != config.get(key) and (key in before_config or config.get(key))
             )
             scanner.scanner_config = config
-            scanner.save(update_fields=["scanner_config"])
+            if question:
+                scanner.prompt_question = question["prompt_question"]
+                scanner.prompt_question_source = question["prompt_question_source"]
+                scanner.save(update_fields=["scanner_config", *question])
+            else:
+                scanner.save(update_fields=["scanner_config"])
             suggestion.status = PromptSuggestionStatus.APPLIED
             suggestion.applied_at = timezone.now()
             suggestion.applied_by = cast(User, request.user)
             suggestion.save(update_fields=["status", "applied_at", "applied_by"])
-        question = question_fields_for_save(
-            team_id=self.team_id,
-            scanner_type=scanner.scanner_type,
-            scanner_config=config,
-            current_source=scanner.prompt_question_source,
-        )
-        if question:
-            ReplayScanner.objects.filter(pk=scanner.pk, scanner_version=scanner.scanner_version).update(
-                prompt_question=question["prompt_question"],
-                prompt_question_source=question["prompt_question_source"],
-            )

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: fae1bdf9-f666-409d-b591-61a86a3ce47a

📥 Commits

Reviewing files that changed from the base of the PR and between 843043a and 51f9dc4.

📒 Files selected for processing (1)
  • frontend/snapshots.yml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review.

@fasyy612 fasyy612 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM 💯

This branch was successfully deployed

1 active deployment
preview-pr-108268 — 51f9dc42 Deployed Sep 29, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants