Skip to content

feat(today): write the briefing in one llm call over the top reports - #110221

Open
VojtechBartos wants to merge 7 commits into
masterfrom
posthog/today-briefing-single-llm-call
Open

VojtechBartos wants to merge 7 commits into
masterfrom
posthog/today-briefing-single-llm-call

Conversation

@VojtechBartos

@VojtechBartos VojtechBartos commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Problem

  • A person who opens Today waits minutes for a briefing, and a run can take up to 30 minutes before it fails.
  • Each briefing starts a sandbox agent that reads insights, alerts, tickets and GitHub over MCP, then fixes its own text in up to 3 turns.
  • The agent picks the items itself, so the items and the numbers in the text depend on its tool calls.

Changes

  • The briefing now names only the person's top 5 Inbox reports, in the order signals already ranks them.
  • Code builds the items (key, reason, URL, facts) before the LLM call, so the item list is frozen.
  • One LLM call (gpt-6-luna through the AI gateway, with the model's default reasoning) writes only the headline, the paragraphs, and the labels and signals.
  • With no reports, there is no LLM call, and the briefing says "Nothing needs you right now".
  • The writing-rule checks and the fix-up turns are gone. A link to a key that is not in the item list becomes plain text.
  • The code highlights the link to the top item, so the LLM no longer chooses the highlight.
  • The empty state and the MCP tool descriptions now say the briefing covers reports only.
  • Resolved and suppressed reports still show crossed out, because the item state is read live when the page loads.
  • The generate workflow retries an attempt up to 3 times, because a retry no longer starts a second sandbox. The run budget drops from 30 to 10 minutes, and the page polls for 11 minutes.
  • The Temporal activity keeps its type name run_agent_activity, so workflows in flight during the deploy replay without a non-determinism error.
  • Mechanical: old enum values stay, so stored briefings still load. The unused urgency field is gone, and old rows that have it still validate. Comments that named the sandbox are updated.

Before:

flowchart LR
  A[Generate workflow]:::phYellow --> B{{Sandbox agent}}:::phBlue
  B --> C[MCP: reports, insights, alerts, tickets]:::phRed
  B --> D[GitHub via gh]:::phRed
  B --> E[Checks]:::phGray
  E -- broken rules, up to 3 turns --> B
  E --> F[Stored briefing]:::phGray
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
  classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
Loading

After:

flowchart LR
  A[Generate workflow]:::phYellow --> B[Top 5 reports from signals]:::phRed
  B --> C[Fact sheet, frozen]:::phGray
  C --> D{{One LLM call: text only}}:::phBlue
  D --> E[Stored briefing]:::phGray
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
  classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
Loading

Related: #110214 shows resolved and dismissed items in the list.

How did you test this code?

  • Ran products/today/backend/tests locally against the dev stack. The workflow test needed SSL_CERT_FILE set to the certifi bundle to download the Temporal test server.
  • Ran repo-wide mypy, hogli lint:tach and hogli product:lint today.
  • Not checked: a real call through the AI gateway, and the page in a browser.

Test rationale: test_generate.py replaces the sandbox tests with tests at the LLM client boundary. They catch these regressions: items that do not follow the signals order, a link to an invented key that survives, an LLM call with no reports, and an unparsable reply that is stored as an empty briefing. test_workflows.py now asserts the retry count, so a change back to one attempt fails. test_checks.py is deleted with the checks.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

None.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5), in a PostHog Desktop cloud task.


Created with PostHog Desktop

🤖 Generated with Claude Code

Code picks the top 5 signals reports and stores them as the fact sheet. One LLM call writes only the headline, the paragraphs and the labels. The sandbox agent, its MCP and GitHub reads, and the writing-rule checks are removed. The generate workflow now retries the attempt, because a retry no longer starts a second sandbox.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
@VojtechBartos VojtechBartos self-assigned this Oct 1, 2026
@trunk-io

trunk-io Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

🧪 Running tests on this pull request (testing on PR #110325) - details.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

✅ Complexity (TypeScript) — clean

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Comment density — 5% of added code lines are comments (14 of 261)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/today/backend/logic/generate.py 7 62
products/today/backend/temporal/activities.py 2 5
frontend/src/scenes/project-homepage/today/todayLogic.ts 1 2
products/today/backend/logic/fact_sheet.py 1 34
products/today/backend/logic/llm_output.py 1 45
products/today/backend/models.py 1 1
products/today/backend/temporal/workflows.py 1 7

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Bundle size — 🔺 +148.9 KiB (+0.2%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 70.23 MiB · 🔺 +148.9 KiB (+0.2%)

File Size Δ vs base
posthog-app/_parent/products/canvas/frontend/scene/CanvasScene.js 58.7 KiB 🔺 +58.7 KiB (new)
posthog-app/_parent/products/canvas/frontend/sidePanel/CanvasSidePanel.js 50.4 KiB 🔺 +50.4 KiB (new)
posthog-app/src/scenes/views/Views.js 9.7 KiB 🔺 +9.7 KiB (new)
posthog-app/_parent/products/tasks/frontend/spaces/SpaceScene.js 74.6 KiB 🔺 +8.0 KiB (+12.1%)
posthog-app/_parent/products/canvas/frontend/newCanvas/CanvasNewScene.js 4.2 KiB 🔺 +4.2 KiB (new)
posthog-app/_parent/products/visual_review/frontend/scenes/VisualReviewRunScene.js 54.2 KiB 🟢 -3.6 KiB (-6.3%)
posthog-app/src/scenes/AuthenticatedShell.js 319.9 KiB 🔺 +3.5 KiB (+1.1%)
posthog-app/_parent/products/posthog_ai/frontend/scenes/TaskTracker/TaskTracker.js 54.2 KiB 🔺 +3.0 KiB (+5.8%)
exporter/_parent/products/posthog_ai/frontend/scenes/TaskTracker/TaskTracker.js 46.3 KiB 🔺 +2.9 KiB (+6.7%)
render-query/src/render-query/render-query.js 20.22 MiB 🔺 +2.3 KiB (+0.0%)
posthog-app/_parent/products/autoresearch/frontend/AutoresearchPipelineScene.js 54.9 KiB 🔺 +1.7 KiB (+3.1%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.62 MiB · 22 files 🔺 +5.2 KiB (+0.3%) █████████░ 87.8% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.57 MiB · 630 files 🔺 +8.4 KiB (+0.2%) █████████░ 88.7% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.75 MiB · 2,486 files 🔺 +151.2 KiB (+1.9%) █████████░ 92.9% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
91.9 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.7 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
109.9 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
91.9 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.19 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.19 MiB · 19 files 🔺 +4.1 KiB (+0.2%) ████░░░░░░ 38.3% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
828.7 KiB dist/toolbar/toolbar-app-JMWOYNWY.css
656.8 KiB dist/toolbar/chunk-chunk-GGHNBOXG.js
259.4 KiB dist/toolbar/chunk-chunk-7EIYIDXQ.js
138.2 KiB dist/toolbar/chunk-chunk-5LHZX24R.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-MYVEDPG7.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-W4A23KPO.js
21.0 KiB dist/toolbar/chunk-chunk-7CMZVTBI.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +2.69 MiB (+0.3%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 964.78 MiB · 🔺 +2.69 MiB (+0.3%)

ℹ️ MCP UI apps size — 33 app(s), 17633.1 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 199.3 KB
action 454.1 KB 199.3 KB
action-list 564.2 KB 199.3 KB
cohort 453.1 KB 199.3 KB
cohort-list 563.2 KB 199.3 KB
email-template 452.9 KB 199.3 KB
error-details 469.6 KB 199.3 KB
error-issue 454.5 KB 199.3 KB
error-issue-list 564.8 KB 199.3 KB
experiment 561.3 KB 199.3 KB
experiment-list 564.9 KB 199.3 KB
experiment-results 566.3 KB 199.3 KB
feature-flag 566.8 KB 199.3 KB
feature-flag-list 570.5 KB 199.3 KB
feature-flag-testing 457.3 KB 199.3 KB
inline-scan 453.6 KB 199.3 KB
insight-actors 562.3 KB 199.3 KB
invite-email-preview 452.3 KB 199.3 KB
llm-costs 559.3 KB 199.3 KB
session-recording 455.3 KB 199.3 KB
survey 454.7 KB 199.3 KB
survey-global-stats 561.9 KB 199.3 KB
survey-list 564.9 KB 199.3 KB
survey-stats 561.9 KB 199.3 KB
trace-span 453.5 KB 199.3 KB
trace-span-list 564.1 KB 199.3 KB
vision-observation-list 563.3 KB 199.3 KB
workflow 453.4 KB 199.3 KB
workflow-list 563.5 KB 199.3 KB
loops-review 457.8 KB 199.3 KB
query-results 774.1 KB 199.3 KB
render-ui 858.1 KB 199.3 KB
visual-review-snapshots 457.9 KB 199.3 KB
✅ Playwright — all passed

All tests passed.

View test results →

⚠️ Backend coverage — 97.0% of changed backend lines covered — 3 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ███████████████████░ 97.0% (131 / 134)

File Patch Uncovered changed lines
products/today/backend/facade/temporal.py 0.0% 7, 12
products/today/backend/temporal/activities.py 66.7% 45

🤖 Agents: add a test only if an uncovered line exposes a realistic regression that existing tests miss. Otherwise explain why no new test is needed under "How did you test this code?". Gap list: the patch-coverage artifact on this run (gh run download 36911768526 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
demo ███████████░░░░░░░░░ 53.4% 1,445 / 2,707
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
warehouse_sources_queue █████████████░░░░░░░ 65.9% 1,611 / 2,446
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
aeo ███████████████░░░░░ 76.3% 617 / 809
batch_exports ████████████████░░░░ 81.3% 21,573 / 26,551
apm █████████████████░░░ 84.1% 1,306 / 1,553
cdp ██████████████████░░ 88.3% 4,559 / 5,164
ml_inference ██████████████████░░ 88.8% 539 / 607
mcp_analytics ██████████████████░░ 89.2% 5,038 / 5,651
product_tours ██████████████████░░ 89.3% 1,340 / 1,500
dashboards ██████████████████░░ 89.6% 6,924 / 7,727
notebooks ██████████████████░░ 90.2% 15,304 / 16,971
data_warehouse ██████████████████░░ 90.3% 14,298 / 15,839
signals ██████████████████░░ 90.3% 59,285 / 65,644
cohorts ██████████████████░░ 90.5% 8,534 / 9,434
streamlit_apps ██████████████████░░ 90.8% 2,684 / 2,956
managed_warehouse ██████████████████░░ 91.0% 10,252 / 11,263
data_modeling ██████████████████░░ 91.2% 10,562 / 11,584
tasks ██████████████████░░ 91.3% 78,937 / 86,495
today ██████████████████░░ 91.4% 894 / 978
exports ██████████████████░░ 91.7% 9,684 / 10,566
business_knowledge ██████████████████░░ 92.0% 8,446 / 9,180
engineering_analytics ██████████████████░░ 92.1% 11,458 / 12,436
ai_training ██████████████████░░ 92.2% 356 / 386
conversations ███████████████████░ 92.6% 29,137 / 31,466
early_access_features ███████████████████░ 92.6% 1,339 / 1,446
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
stamphog ███████████████████░ 92.8% 8,109 / 8,742
visual_review ███████████████████░ 92.9% 9,534 / 10,265
canvas ███████████████████░ 92.9% 7,155 / 7,703
approvals ███████████████████░ 93.0% 3,974 / 4,271
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
notifications ███████████████████░ 93.2% 1,144 / 1,228
error_tracking ███████████████████░ 93.2% 16,368 / 17,555
surveys ███████████████████░ 93.4% 6,644 / 7,113
slack_app ███████████████████░ 93.6% 14,560 / 15,557
autoresearch ███████████████████░ 93.6% 8,837 / 9,442
context_layer ███████████████████░ 93.8% 3,415 / 3,639
web_analytics ███████████████████░ 93.9% 23,680 / 25,229
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.3% 8,959 / 9,501
ai_observability ███████████████████░ 94.6% 24,974 / 26,409
alerts ███████████████████░ 94.6% 9,267 / 9,796
wizard ███████████████████░ 94.7% 6,150 / 6,496
workflows ███████████████████░ 94.7% 15,159 / 16,006
reminders ███████████████████░ 94.8% 760 / 802
review_hog ███████████████████░ 95.0% 11,623 / 12,235
annotations ███████████████████░ 95.1% 817 / 859
endpoints ███████████████████░ 95.1% 9,222 / 9,694
customer_analytics ███████████████████░ 95.2% 25,991 / 27,301
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
posthog_ai ███████████████████░ 95.3% 2,491 / 2,614
marketing_analytics ███████████████████░ 95.3% 19,596 / 20,559
actions ███████████████████░ 95.5% 756 / 792
logs ███████████████████░ 95.5% 15,468 / 16,200
experiments ███████████████████░ 95.5% 32,931 / 34,481
data_catalog ███████████████████░ 95.6% 4,402 / 4,606
tracing ███████████████████░ 95.6% 3,536 / 3,699
replay_vision ███████████████████░ 95.6% 29,078 / 30,404
growth ███████████████████░ 95.7% 11,381 / 11,888
skills ███████████████████░ 95.8% 6,972 / 7,274
messaging ███████████████████░ 95.9% 3,824 / 3,989
product_analytics ███████████████████░ 96.0% 28,521 / 29,696
revenue_analytics ███████████████████░ 96.4% 1,889 / 1,959
access_control ███████████████████░ 96.5% 7,246 / 7,512
user_interviews ███████████████████░ 96.5% 2,870 / 2,974
feature_flags ███████████████████░ 96.6% 26,748 / 27,687
warehouse_sources ███████████████████░ 97.3% 468,088 / 480,989
data_quality ████████████████████ 97.5% 7,701 / 7,895
links ████████████████████ 97.9% 234 / 239
security ████████████████████ 98.0% 1,283 / 1,309
metrics ████████████████████ 98.0% 4,245 / 4,331
analytics_platform ████████████████████ 98.3% 2,784 / 2,833
pulse ████████████████████ 98.5% 2,046 / 2,078
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 6974811e-d9cc-4e72-8b33-52103c5bd6ed

📥 Commits

Reviewing files that changed from the base of the PR and between e709b07 and b0780d2.

📒 Files selected for processing (3)
  • frontend/src/scenes/project-homepage/today/TodayBriefing.tsx
  • products/today/backend/logic/prompt.py
  • products/today/backend/logic/prompts/briefing.md.j2

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The change replaces sandbox-based briefing generation with direct LLM completion using report facts prepared before the request. It adds structured output models and converts valid responses into stored briefing content. Temporal retries failed attempts, and the frontend poll limit decreases from 190 to 66.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to b0780

The report-driven briefing changes are mergeable with normal checks. Each completion attempt now manages client cleanup explicitly.

Security Architecture Review

Security architecture risk: 🔵 Low · up to b0780

The change reduces automated access and limits briefings to selected reports. The main remaining risk is recovery: a repeated generation attempt can revisit an already published briefing. This affects an individual user’s briefing; no new broader access path was demonstrated.

Retained concerns

  • Low · reliability · inferred: New whole-attempt retries do not preserve a committed terminal state. If an attempt stores READY but its completion acknowledgement is lost, preparation for the next attempt can set that same row back to WRITING, regenerate its contents, or delete it when eligibility has changed. A concurrent refresh can also conflict with the one-pending-row constraint. Workflow deduplication and single-save persistence protect normal execution but do not fence repeated attempts. This is a per-user recovery and state-ownership concern, not demonstrated cross-tenant exposure.
Security review details

Security Blast Radius

  • inferred — The inspected generation request is bounded to up to five selected reports and three earlier briefings for the same team and user. The retry-state concern affects that user’s briefing row and day. No cross-tenant mutation or credential-bearing execution path was demonstrated in the new flow.

Security Findings and Attack Paths

  • inferred — Report-derived titles and summaries enter the writing request and can influence displayed prose. In the inspected output path, generated strings are rendered as text, link keys must match authoritative items, and the request exposes no tools. This bounds prompt-injection consequences to presentation rather than demonstrating executable content, arbitrary destinations, or expanded data access.

Trust Boundaries and Controls

  • observed — The provider controls prose, labels, and signals, but not authoritative report selection or destination URLs. Report URLs are built from team and report identifiers, unknown output keys are discarded or converted to plain text, and report links use the application router. Removed writing-rule checks do not remove these enforced boundaries.

Resilience and Maintainability Implications

  • observed — Provider requests have a two-minute timeout, an output-token limit, and no client-level retries; retry ownership resides in the workflow. The factory-created client is context-managed for each request, so the inspected path does not leave a shared client or sandbox lifecycle to clean up.

Hardening Proposals

  • proposed — Make committed READY rows terminal for repeated generation attempts, and use conditional or attempt-scoped transitions for preparation, success, and failure writes. Validate recovery after a database commit but before activity acknowledgement, including a concurrent refresh.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and matches the template. It explains the problem, user-visible and mechanical changes, testing, rationale, release status, agent context, workflow diagrams, and related PR…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/today/backend/logic/generate.py-74-94 (1)

74-94: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

_write_text never closes the async OpenAI client.

build_async_openai_client creates a new httpx.AsyncClient on every call, and this code never closes it. Each attempt leaks connections until garbage collection runs, and Temporal retries add more attempts. Open the client with async with so it closes when the call returns or fails.

Proposed fix
-    client = build_async_openai_client(
+    async with build_async_openai_client(
         ...
-    ).with_options(timeout=CALL_TIMEOUT.total_seconds(), max_retries=0)
-    response = await client.chat.completions.create(
+    ) as base:
+        client = base.with_options(timeout=CALL_TIMEOUT.total_seconds(), max_retries=0)
+        response = await client.chat.completions.create(

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 56f7df56-6bc8-4b67-9e3c-a0492119b898

📥 Commits

Reviewing files that changed from the base of the PR and between eee7ed0 and f5a80e2.

📒 Files selected for processing (18)
  • frontend/src/scenes/project-homepage/today/todayLogic.ts
  • products/today/backend/facade/temporal.py
  • products/today/backend/feature_flags.py
  • products/today/backend/logic/agent_output.py
  • products/today/backend/logic/briefings.py
  • products/today/backend/logic/checks.py
  • products/today/backend/logic/content.py
  • products/today/backend/logic/fact_sheet.py
  • products/today/backend/logic/generate.py
  • products/today/backend/logic/llm_output.py
  • products/today/backend/logic/prompt.py
  • products/today/backend/logic/prompts/briefing.md.j2
  • products/today/backend/models.py
  • products/today/backend/temporal/activities.py
  • products/today/backend/temporal/workflows.py
  • products/today/backend/tests/test_checks.py
  • products/today/backend/tests/test_generate.py
  • products/today/backend/tests/test_workflows.py
💤 Files with no reviewable changes (3)
  • products/today/backend/tests/test_checks.py
  • products/today/backend/logic/checks.py
  • products/today/backend/logic/agent_output.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.

VojtechBartos and others added 2 commits October 1, 2026 19:59
The code highlights the top item, so the LLM no longer chooses highlight. The prompt rows come from the fact sheet, plus the report summary. The unused urgency field, logger and project id are gone. The LLM client closes after the call, and the workflow enforces the run budget. The MCP descriptions and the empty-state copy now describe report-only briefings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
@trunk-io

trunk-io Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

The Python LLM gateway rejects reasoning_effort for gpt-6-luna on the chat completions route, so every briefing attempt failed with a 400. The call now uses the model's default reasoning, like the other single-call gateway callers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
The prompt now asks for two to four short paragraphs: the top item alone first, then the other items grouped by what the person does with them. Each item says what the person can do. Links sit inside a sentence, so every sentence starts with a capital letter. The word limit goes from 70 to 130, and a short example shows the shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
The prompt now asks for three short paragraphs and one short sentence for each item. Sentences say what is going on and never start with a command such as "Review". Link text names the problem, and the prompt states that segments carry their own spaces. The word limit goes from 130 to 80.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 887c6d68-77e2-4776-81be-a4974d08acea
@VojtechBartos
VojtechBartos requested a review from a team October 1, 2026 19:05
@VojtechBartos VojtechBartos added the stamphog Request AI approval (no full review) label Oct 1, 2026
@VojtechBartos
VojtechBartos marked this pull request as ready for review October 1, 2026 19:05
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

🦔 Hogbox preview · ✅ ready

▶ Open the preview

🔑 Login test@posthog.com / 12345678 (demo data)
🧩 Running this PR's backend and frontend, on the PostHog :master base
🔗 Link stable across rebuilds — a re-push swaps the box underneath, the URL stays
🔒 Access tailnet only (PostHog VPN)
🛠️ Admin inspect & debug state in hogland
💤 Idle sleeps after ~30 min idle (snapshot to S3, zero node cost) and wakes on your next visit in ~30s, behind a brief "waking up" screen

commit b0780d2 · box box-72f92e77aaca · ready in 1435s (push → usable) · build log · rebuilds on every push, torn down on close

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved.

Contained change to a feature-flagged product that swaps a sandbox agent for one gateway LLM call over items code picks. It shrinks the attack surface, keeps the Temporal activity type name so in-flight workflows replay, and is covered by tests. The author has strong familiarity, and the earlier CodeRabbit client-leak concern is fixed with async with.

  • Author wrote 99% of the modified lines and has 27 merged PRs in these paths (familiarity STRONG).
  • Real AI gateway call and in-browser behaviour were not tested, per the PR description. The feature flag limits the impact.
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 726L, 20F substantive, 1054L/23F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1d-complex (1054L, 23F, cross-cutting, feat)
stamphog 2.3.1 .stamphog/policy.yml @ b0780d2 · reviewed head b0780d2

This branch was successfully deployed

1 active deployment
preview-pr-110221 — b0780d2c Deployed Oct 1, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stamphog Request AI approval (no full review)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant