Skip to content

feat(replay-vision): rank the watch feed with jev behind a flag - #107096

Open
ksvat wants to merge 10 commits into
posthog/vision-jev-watch-rank-sweepfrom
posthog/vision-jev-watch-feed-ranker
Open

ksvat wants to merge 10 commits into
posthog/vision-jev-watch-rank-sweepfrom
posthog/vision-jev-watch-feed-ranker

Conversation

@ksvat

@ksvat ksvat commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Problem

Layer 2 of the Jev watch-feed stack, on top of #107092. The base PR judges scanner windows hourly and caches the probabilities; nothing reads them yet.

Changes

  • Teams whose vision-watch-feed-ranker flag is jev get a What to watch feed ranked on the cached Jev probabilities: highest probability first, viewed rows docked, newest as the tiebreak. Their cards say "The decision model judged this session worth watching."
  • The two rankers are independent branches in the endpoint. weighted-score and jev-shadow teams run the existing rank_watch_feed_candidates exactly as today, and no Jev value enters that blend. Neither arm makes a model call at request time.
  • A row without a cached probability (not yet swept, or the cache went cold) falls to the recency filler tier, so the feed never empties during rollout or a sweep outage.
  • New jev_watchable reason kind: serializer choice with jev_probability, regenerated OpenAPI types, and the frontend copy case.

How did you test this code?

  • New API test: catches a cached probability moving the feed for a weighted-score or jev-shadow team, and, under jev, broken probability ordering or a lost recency fallback for unjudged rows.
  • Ranker unit tests live in the base PR (test_jev_watch_feed.py), including the viewed-row dock and the filler-tier reasons.
  • Ran locally: the watch feed API test class, ruff, and repo-wide mypy. All passed.
  • Not run: the full backend suite and frontend typecheck. CI covers them.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

None. The behavior is flag-gated and not user-visible yet.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code (PostHog Desktop task), Claude Fable 5

  • See the base PR for the design history. This layer holds only the read path and its contract changes, so the feed switch can be reviewed apart from the sweep machinery.
  • Skills invoked: /writing-tests, /writing-user-facing-copy (reason copy), /improving-drf-endpoints, /writing-pr-descriptions.
  • CodeRabbit local pass skipped (cloud task run).

Created with PostHog Desktop

🤖 Generated with Claude Code

@ksvat ksvat self-assigned this Sep 25, 2026
@ksvat
ksvat added this pull request to stack #107097 September 25, 2026 21:02
@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 3 functions above the limit (max 33)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
watchReasonCopy products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:43 33 10
WatchFeedCard products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:206 17 10
watchCardHeadline products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx:163 15 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

🚨 Comment density — 15% of added code lines are comments (18 of 118)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/replay_vision/backend/tests/test_api.py 11 70
products/replay_vision/backend/api/scanners.py 7 46

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Bundle size — 🔺 +83 B (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 68.87 MiB · 🔺 +83 B (+0.0%)

No file changed by more than 1000 B.

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.57 MiB · 22 files no change █████████░ 85.5% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.51 MiB · 629 files no change █████████░ 87.2% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.33 MiB · 2,332 files no change █████████░ 87.9% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
216.0 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.4 KiB src/lib/api.ts
88.4 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.4 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
271.7 KiB src/taxonomy/core-filter-definitions-by-group.json
216.0 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.4 KiB src/lib/api.ts
98.5 KiB ../packages/quill/packages/quill/dist/index.js
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
88.4 KiB src/products.tsx

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.16 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.16 MiB · 19 files no change ████░░░░░░ 37.7% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
800.5 KiB dist/toolbar/toolbar-app-JG6G2U7S.css
651.5 KiB dist/toolbar/chunk-chunk-DJ3DVXIY.js
259.4 KiB dist/toolbar/chunk-chunk-CV2VU6SQ.js
138.3 KiB dist/toolbar/chunk-chunk-TUPNHXOT.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-QSMPSM2L.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-P5FAEBEQ.js
21.0 KiB dist/toolbar/chunk-chunk-W5LHFGNZ.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +1.8 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 946.07 MiB · 🔺 +1.8 KiB (+0.0%)

ℹ️ MCP UI apps size — 33 app(s), 17631.0 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 196.2 KB
action 454.1 KB 196.2 KB
action-list 564.2 KB 196.2 KB
cohort 453.1 KB 196.2 KB
cohort-list 563.2 KB 196.2 KB
email-template 452.9 KB 196.2 KB
error-details 469.6 KB 196.2 KB
error-issue 453.8 KB 196.2 KB
error-issue-list 564.1 KB 196.2 KB
experiment 561.3 KB 196.2 KB
experiment-list 564.9 KB 196.2 KB
experiment-results 566.3 KB 196.2 KB
feature-flag 566.8 KB 196.2 KB
feature-flag-list 570.5 KB 196.2 KB
feature-flag-testing 457.3 KB 196.2 KB
inline-scan 453.6 KB 196.2 KB
insight-actors 562.3 KB 196.2 KB
invite-email-preview 452.3 KB 196.2 KB
llm-costs 559.3 KB 196.2 KB
session-recording 455.3 KB 196.2 KB
survey 454.7 KB 196.2 KB
survey-global-stats 561.9 KB 196.2 KB
survey-list 564.9 KB 196.2 KB
survey-stats 561.9 KB 196.2 KB
trace-span 453.5 KB 196.2 KB
trace-span-list 564.1 KB 196.2 KB
vision-observation-list 563.3 KB 196.2 KB
workflow 453.4 KB 196.2 KB
workflow-list 563.5 KB 196.2 KB
loops-review 457.8 KB 196.2 KB
query-results 774.1 KB 196.2 KB
render-ui 857.4 KB 196.2 KB
visual-review-snapshots 457.9 KB 196.2 KB
✅ Playwright — all passed

All tests passed.

View test results →

⚠️ Backend coverage — 96.0% of changed backend lines covered — 25 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ███████████████████░ 96.0% (667 / 692)

File Patch Uncovered changed lines
products/ml_inference/backend/facade/api.py 50.0% 16
products/replay_vision/backend/temporal/jev_watch_rank/schedule.py 87.5% 19
products/replay_vision/backend/temporal/jev_watch_rank/workflow.py 87.5% 26, 33
products/replay_vision/backend/temporal/jev_watch_rank/activities.py 89.3% 115–116, 141–142, 155–156, 158–159, 163, 221, 223–224, 250
products/replay_vision/backend/jev_watch_feed.py 96.1% 375–377, 399–401, 403, 425

🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 36495114150 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
warehouse_sources_queue ██████░░░░░░░░░░░░░░ 29.1% 92 / 316
demo ███████████░░░░░░░░░ 52.8% 1,411 / 2,673
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
aeo ██████████████░░░░░░ 70.8% 472 / 667
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
batch_exports ████████████████░░░░ 81.2% 21,475 / 26,449
apm █████████████████░░░ 84.1% 1,306 / 1,553
cdp ██████████████████░░ 88.2% 4,545 / 5,155
ml_inference ██████████████████░░ 88.2% 510 / 578
mcp_analytics ██████████████████░░ 88.9% 4,927 / 5,540
product_tours ██████████████████░░ 89.3% 1,331 / 1,491
dashboards ██████████████████░░ 89.5% 6,855 / 7,657
data_warehouse ██████████████████░░ 90.0% 14,019 / 15,581
signals ██████████████████░░ 90.0% 56,092 / 62,306
notebooks ██████████████████░░ 90.2% 15,287 / 16,945
cohorts ██████████████████░░ 90.4% 8,420 / 9,316
streamlit_apps ██████████████████░░ 90.6% 2,623 / 2,895
tasks ██████████████████░░ 91.0% 74,108 / 81,470
managed_warehouse ██████████████████░░ 91.0% 10,252 / 11,263
data_modeling ██████████████████░░ 91.4% 10,504 / 11,498
exports ██████████████████░░ 91.7% 9,676 / 10,556
engineering_analytics ██████████████████░░ 91.7% 11,032 / 12,030
ai_training ██████████████████░░ 92.2% 356 / 386
business_knowledge ██████████████████░░ 92.2% 7,684 / 8,330
conversations ███████████████████░ 92.5% 28,734 / 31,062
early_access_features ███████████████████░ 92.6% 1,341 / 1,448
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
visual_review ███████████████████░ 92.8% 9,247 / 9,969
canvas ███████████████████░ 92.8% 6,877 / 7,409
approvals ███████████████████░ 93.0% 3,974 / 4,271
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
error_tracking ███████████████████░ 93.2% 15,928 / 17,097
notifications ███████████████████░ 93.2% 1,145 / 1,229
slack_app ███████████████████░ 93.2% 13,677 / 14,674
stamphog ███████████████████░ 93.2% 7,885 / 8,456
surveys ███████████████████░ 93.3% 6,571 / 7,040
context_layer ███████████████████░ 93.9% 3,415 / 3,638
web_analytics ███████████████████░ 93.9% 21,687 / 23,096
alerts ███████████████████░ 94.0% 8,569 / 9,114
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.4% 8,940 / 9,472
ai_observability ███████████████████░ 94.6% 24,063 / 25,440
wizard ███████████████████░ 94.7% 6,150 / 6,496
reminders ███████████████████░ 94.8% 760 / 802
workflows ███████████████████░ 94.8% 14,332 / 15,113
review_hog ███████████████████░ 95.0% 11,507 / 12,119
endpoints ███████████████████░ 95.1% 9,206 / 9,681
annotations ███████████████████░ 95.1% 817 / 859
customer_analytics ███████████████████░ 95.2% 25,078 / 26,349
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
marketing_analytics ███████████████████░ 95.3% 19,322 / 20,274
posthog_ai ███████████████████░ 95.3% 2,488 / 2,610
experiments ███████████████████░ 95.4% 32,458 / 34,023
actions ███████████████████░ 95.5% 756 / 792
logs ███████████████████░ 95.5% 15,399 / 16,130
data_catalog ███████████████████░ 95.5% 4,401 / 4,606
tracing ███████████████████░ 95.6% 3,536 / 3,699
autoresearch ███████████████████░ 95.7% 8,481 / 8,865
growth ███████████████████░ 95.7% 11,228 / 11,734
messaging ███████████████████░ 95.8% 3,798 / 3,963
skills ███████████████████░ 95.8% 6,972 / 7,274
replay_vision ███████████████████░ 95.9% 27,818 / 28,999
product_analytics ███████████████████░ 96.0% 28,470 / 29,647
access_control ███████████████████░ 96.3% 7,112 / 7,386
revenue_analytics ███████████████████░ 96.4% 1,876 / 1,946
user_interviews ███████████████████░ 96.5% 2,859 / 2,963
feature_flags ███████████████████░ 96.5% 25,556 / 26,473
warehouse_sources ███████████████████░ 97.2% 456,705 / 469,652
data_quality ████████████████████ 97.6% 7,587 / 7,774
security ████████████████████ 97.9% 1,202 / 1,228
links ████████████████████ 97.9% 234 / 239
metrics ████████████████████ 98.0% 4,084 / 4,166
analytics_platform ████████████████████ 98.3% 2,784 / 2,833
pulse ████████████████████ 98.5% 2,043 / 2,075
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

@coderabbitai

coderabbitai Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: b7cad819-ed8e-49f7-8296-0eb8b16c6bec

📥 Commits

Reviewing files that changed from the base of the PR and between e73fb1d and 721ca3c.

⛔ Files ignored due to path filters (1)
  • products/replay_vision/frontend/generated/api.schemas.ts is excluded by !**/generated/**
📒 Files selected for processing (1)
  • services/mcp/src/api/generated.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The watch-feed API selects JEV ranking when the team’s ranker is jev. Other ranker modes use the existing candidate ranker. Watch-feed reasons include jev_watchable and an optional nullable probability. The frontend adds explanatory copy for the new reason. A test covers behavior across weighted-score, jev-shadow, and jev.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to 721ca

Teams assigned to Jev can see more than the documented three recency-filler rows while judgments are unavailable. The default remains weighted-score, so this is bounded to the Jev rollout; address or accept the cap difference before broadening it.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 721ca

Access checks still apply before ranking, and no cross-tenant exposure was found. When scores are unavailable, however, the feed can return more routine cards than its documented fallback permits.

Retained concerns

  • Low · reliability · observed: On the JEV arm, missing cached scores can cause routine filler rows to occupy the full requested limit, rather than the existing three-item fallback. This changes the feed’s documented behavior during cache failure or rollout.
Security review details

Security Blast Radius

  • inferred — The changed ordering and additional reason metadata can affect authorized watch-feed readers on JEV-enabled teams; the examined path does not expand the candidate set beyond the existing team and access filters.

Trust Boundaries and Controls

  • observed — Caller-provided feed filters are validated, scanner IDs are intersected with readable IDs, and observation access is checked before cached ranks influence the response.

Resilience and Maintainability Implications

  • observed — A cache read failure preserves endpoint availability through recency fallback, but the JEV fallback does not preserve the weighted arm’s three-item filler bound.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and aligned with the template. It explains the problem, user-visible changes, testing performed and not performed, feature-flag release status, documentation status, and ag…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 2a11e840-749e-487e-a3f6-43d0d33a1f96

📥 Commits

Reviewing files that changed from the base of the PR and between 7df6c53 and 7f97ce8.

⛔ Files ignored due to path filters (1)
  • products/replay_vision/frontend/generated/api.schemas.ts is excluded by !**/generated/**
📒 Files selected for processing (4)
  • products/replay_vision/backend/api/scanners.py
  • products/replay_vision/backend/tests/test_api.py
  • products/replay_vision/frontend/replay_scanners/components/WatchFeedCard.tsx
  • services/mcp/src/api/generated.ts

Included review availability: Your plan provides up to 12 included reviews per hour; 10 remain after this review.

Comment thread products/replay_vision/backend/api/scanners.py Outdated
@ksvat
ksvat force-pushed the posthog/vision-jev-watch-feed-ranker branch from 7f97ce8 to 6bb9c83 Compare September 25, 2026 21:28
@posthog

posthog Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

👋 Visual changes detected for this PR.

Review and approve in PostHog Visual Review

If these changes are unexpected, they may be caused by a flaky test or a broken snapshot on master. Don't approve — rerun the job or wait for a fix.

@trunk-io

trunk-io Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

@ksvat
ksvat force-pushed the posthog/vision-jev-watch-feed-ranker branch 5 times, most recently from b6f2f7c to 1bd808e Compare September 26, 2026 01:18
@ksvat
ksvat marked this pull request as ready for review September 26, 2026 01:30
@ksvat ksvat added the stamphog Request AI approval (no full review) label Sep 26, 2026
@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

🦔 Hogbox preview · ✅ ready

▶ Open the preview

🔑 Login test@posthog.com / 12345678 (demo data)
🧩 Running this PR's backend and frontend, on the PostHog :master base
🔗 Link stable across rebuilds — a re-push swaps the box underneath, the URL stays
🔒 Access tailnet only (PostHog VPN)
🛠️ Admin inspect & debug state in hogland
💤 Idle sleeps after ~30 min idle (snapshot to S3, zero node cost) and wakes on your next visit in ~30s, behind a brief "waking up" screen

commit b720826 · box box-565f92c03600 · ready in 771s (push → usable) · build log · rebuilds on every push, torn down on close

@ksvat
ksvat requested review from a team, TueHaulund, arnohillen and fasyy612 and removed request for a team September 26, 2026 01:30
@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team September 26, 2026 01:30

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved — this change needs a human reviewer.

Re-add the stamphog label to request another review once you have addressed this.

CodeRabbit's Major finding on scanners.py — that a cold Jev cache lets the new ranking branch return up to the full page-size limit of recency-filler rows instead of the feed's normal 3-row floor for a window with no findings — is marked outdated but is still present in the current code: rank_watch_feed_by_jev has no equivalent of the weighted ranker's filler trim, and the call site slices straight by params["limit"] with no floor applied. That's a genuine, unaddressed behavioral regression the PR itself says will happen during rollout.

  • Author wrote 75% of the modified lines and has 37 merged PRs in these paths (familiarity STRONG).
  • Unaddressed CodeRabbit comment on products/replay_vision/backend/api/scanners.py: the jev ranking branch can return up to params["limit"] recency-filler rows when the cache is cold, instead of trimming to the feed's normal no-findings floor of 3 rows (see watch_feed.py's _trim_filler / WATCH_FEED_MIN_ITEMS, which rank_watch_feed_by_jev has no equivalent of).
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 52L, 2F substantive, 150L/5F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1c-medium (150L, 5F, two-areas, feat)
stamphog 2.2.0 .stamphog/policy.yml @ 1bd808e · reviewed head 1bd808e

@stamphog stamphog Bot removed the stamphog Request AI approval (no full review) label Sep 26, 2026
.annotate(feed_viewed=viewed)
.values("id", "scanner_id", "created_at", "scanner_result", "feed_viewed")
)
ranked = rank_watch_feed_by_jev(jev_rows, probabilities)[: params["limit"]]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Agent-drafted, reviewed by Tue.

CodeRabbit is right here. The jev path never drops filler rows like the weighted path does. If the cache is empty, a team gets 20 "new since you last looked" cards instead of 3. That also makes the two variants hard to compare. Wrapping this in _trim_filler(...) before the slice should fix it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in aca1032 with exactly that shape — the jev result now runs through _trim_filler, so both arms keep the same 3-row floor for windows without findings and stay comparable in the experiment. Unit test plus a call-site assertion cover it.

"reason kind, and preferred over copy derived from the reason kind. Absent on observations "
"scanned before notability shipped."
"The scan's own sentence naming why the session is worth watching. Present on the `notable` and "
"`jev_watchable` reason kinds when the scan wrote one, and preferred over copy derived from the "

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Agent-drafted, reviewed by Tue.

A jev_watchable card always shows the scan's notability_reason, even when that reason is "nothing stands out". The weighted path only shows it when the scan found something notable. Could the jev path do the same?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in aca1032: the jev card now carries notability_reason only when the scan's own notability clears NOTABLE_MIN_SCORE, the same bar the weighted path applies, so a card Jev rated watchable relative to a dull window can no longer lead with "nothing stands out". Covered by a unit test and the end-to-end reason assertion in the API tests.

@ksvat
ksvat force-pushed the posthog/vision-jev-watch-feed-ranker branch from 1bd808e to 08236ba Compare September 28, 2026 16:28
@github-actions
github-actions Bot requested a deployment to preview-pr-107096 September 28, 2026 16:28 In progress
# The flag selects one of two independent rankers; nothing is blended between them. Shadow
# teams rank on the weighted score too, because only the `jev` arm reads the probabilities
# the hourly sweep cached. Neither arm makes a model call here.
if watch_feed_ranker(self.team_id) == "jev":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

list_watch uses the jev ranker in all regions, but the sweep writes ranks only where decision_api.decisions_available_here() is true (activities.py:110), so an EU team with the flag on gets only recency filler rows. Make watch_feed_ranker return "weighted-score" when decisions_available_here() is false, and add a test for it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deliberately deferring this while the experiment is internal-only: enrollment is explicitly team-targeted (PostHog's own team, US cloud), so no EU team can land on the jev arm, and with the filler floor fixed a misconfigured team would now see 3 recency rows rather than a full page. Agreed the region gate belongs in watch_feed_ranker before the flag opens beyond internal enrollment — it's recorded in the same TODO as the team-2 pin (constants.py, base PR), so both internal-only assumptions get unwound together.

The watch feed endpoint branches on the vision-watch-feed-ranker flag:
weighted-score teams (default and jev-shadow) rank on the unchanged
deterministic blend, and jev teams rank on the probabilities the hourly
sweep cached. Nothing is blended between the two rankers, and neither arm
makes a model call at request time.

Adds the jev_watchable reason kind with its serializer field, generated
types, and frontend copy. A row without a cached probability falls to the
recency filler tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
Follows the base layer's JEV_WATCHABLE_MIN change: a row the model rated
0.2 now carries a filler reason and ranks by recency, so the API test
asserts that instead of a jev_watchable reason.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
…ce in jev mode

The candidate query keeps only each scanner's 100 newest rows, which on a
high-volume scanner covers minutes, so the most interesting sessions of
the last days could never reach the feed. In jev mode the endpoint now
also fetches the rows whose cached probability clears the watchable
threshold, through the same filtered queryset (team, readable scanners,
date window, search), and ranks the union. The weighted-score arm is
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
Follows the base layer's review fixes: the store call takes the split
judged-set and watchable-map arguments, the jev arm's API test asserts the
scan's notability_reason reaches the reason payload, and the serializer's
help_text says the field now appears on jev_watchable reasons too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
Follows the base layer's store_watch_ranks signature change, which added
the judge-attempt counts to the judged-key payload.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
@ksvat
ksvat force-pushed the posthog/vision-jev-watch-feed-ranker branch from ef0048d to 7980812 Compare September 28, 2026 21:08
…ason gate

Follows the base layer's review fixes: the slice-bypass test now asserts
the feed pads its one finding only to the 3-row floor and carries the
scan's sentence when the scan itself found the session notable, and the
serializer help_text says so. The flag-switch test drops its notability
fixture, because a notable score legitimately reorders the weighted arm
and that test's premise is that the cached probability alone moves
nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
@ksvat ksvat added the stamphog Request AI approval (no full review) label Sep 28, 2026 — with PostHog
stamphog[bot]

This comment was marked as outdated.

…n-jev-watch-feed-ranker

Generated-By: PostHog Desktop
Task-Id: ae8c41df-c0e7-45c1-a61a-dec6b0b2c893
@stamphog
stamphog Bot dismissed their stale review September 28, 2026 22:50

A new stamphog review started for this PR — the fresh verdict replaces this approval.

@github-actions
github-actions Bot requested a deployment to preview-pr-107096 September 28, 2026 22:50 In progress
stamphog[bot]

This comment was marked as outdated.

@stamphog
stamphog Bot dismissed their stale review September 28, 2026 22:54

A new stamphog review started for this PR — the fresh verdict replaces this approval.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved.

This is a flag-gated read path in a contained product area. The earlier reviewer concerns (filler trimming, notability reason) are fixed in the tree. The unresolved EU-region point degrades safely to recency rows and is deferred while enrollment is internal-only. The author has STRONG familiarity with the code, and the change is easy to revert.

  • Author wrote 75% of the modified lines and has 37 merged PRs in these paths (familiarity STRONG).
  • Non-blocking: @arnohillen's unresolved thread notes the jev ranker isn't region-gated (EU teams would only get recency filler). Author has deferred it to before the flag widens beyond internal teams; make sure that follow-up lands.
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 52L, 2F substantive, 152L/5F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1c-medium (152L, 5F, two-areas, feat)
stamphog 2.3.0 .stamphog/policy.yml @ b720826 · reviewed head b720826

This branch was successfully deployed

1 active deployment
preview-pr-107096 — b720826f Deployed Sep 28, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stamphog Request AI approval (no full review)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants