Skip to content

feat(aio): add offline experiment inspection and score trends - #107937

Open
Radu-Raicea wants to merge 36 commits into
masterfrom
feat/aio-offline-evals-ui
Open

Radu-Raicea wants to merge 36 commits into
masterfrom
feat/aio-offline-evals-ui

Conversation

@Radu-Raicea

@Radu-Raicea Radu-Raicea commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Problem

People who upload offline experiments need to inspect results and follow recurring scores in the Offline evals tab.

The read APIs landed in #106977. This adds their browser UI and gives manual-review score definitions a shared home.

Changes

  • Incomplete overview history shows the loaded count and date span, with a notice that earlier experiments and scorer versions may be absent.
  • Recent experiments and score trends share time range, source, and upload-state filters, defaulting to the last 30 days across all upload states.
  • First-time users see onboarding across the offline tab, with setup guidance for reporting their first experiment. Datasets remain optional provenance.
  • Score charts connect points with smooth curves, alternate theme colors, and show one version at a time with navigation arrows.
  • Shared chart hover uses one time range across selected scorers, including all-time views with sparse histories and different versions.
  • Readable time axes show time of day for short ranges, years across year boundaries, and useful elapsed durations for comparisons.
  • Score selection defaults to available scorers when empty and preserves personal chart ordering. Shared URLs do not overwrite saved choices.
  • Result inspection separates scores, reasoning, and payloads into distinct panels. Large text, JSON, and metadata start collapsed.
  • Compact tables show single-line experiment names and coverage with execution time in its own column. Item tables retain every scorer version.
  • Score history keeps upload state, source, date range, comparison, and version filters. Custom comparisons ending now survive sharing and reloading.
  • Scorer summaries aggregate the experiments shown for the selected period and version, with means weighted by scored items and distinct experiment counts.
  • Hiding chart series removes their paths without generating invalid SVG.
  • Scorers moves into Evaluations, with redirects from Human reviews and the former routes. Version creation handles conflicts without silently overwriting another update.
  • Bounded cell reads fetch at most 50 items by 20 scorer versions, enforce existing read permissions, and exclude payloads.
  • Internal cleanup shares experiment status/source presentation and separates history tables and run details while preserving their rendered markup and interactions.
  • Generated API and scene bindings are mechanical. The retired legacy frontend is removed; its backend remains.
Partial history coverage, using invented Storybook data

Before:

Offline overview before partial-history disclosure

After:

Offline overview with partial-history disclosure

Scorer navigation before and after

Before:

flowchart LR
    Reviews["Human reviews"] --> Definitions["Score definitions"]
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
    class Reviews phYellow;
    class Definitions phBlue;
Loading

After:

flowchart LR
    Evaluations["Evaluations"] --> Scorers["Scorers"]
    Reviews["Human reviews / old links"] -->|Redirect| Scorers
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
    class Evaluations,Reviews phYellow;
    class Scorers phBlue;
Loading
Original views before UI refinements, using invented Storybook data

These screenshots document the initial layouts; the former offline tab used the legacy harness.

Recent experiments and score trends

First experiment onboarding

Score tooltip in dark mode

Selected score and item inspector

Scorer history at 520px

Short comparison axes and all-time shared hover, using invented Storybook data

One-hour comparison before:

pr107937-before-comparison

One-hour comparison after:

pr107937-after-comparison

All-time overview before:

pr107937-before-all-time

All-time overview after:

pr107937-after-all-time

Period summaries, using invented Storybook data

Before:

Latest-experiment scorer summaries

After:

Scorer summaries across the displayed period

How did you test this code?

  • Period-summary regressions cover unequal experiment sizes, boolean and category rates, zero-success runs, and switching versions and date ranges.

  • Browser checks verify period scores and experiment counts at 1440px and 520px.

  • URL round-trip tests cover comparison periods ending now and explicit historical timestamps.

  • Shared-domain regressions cover sparse scorers, version changes, selection changes, refreshes, and stale responses.

  • Storybook checks cover short and multi-year axes, hidden legend series, and synchronized all-time hover at desktop and narrow widths in both themes.

  • The frontend TypeScript check, focused regression suites, formatter/linter, and CI preflight pass for these fixes.

  • Storybook checks verify existing controls at desktop/520px in light/dark themes. Partial-history fixtures verify the disclosure.

  • Browser checks cover history expansion, score selection/save, item inspection, payloads, and all scorer summaries.

  • The cleanup passes the full frontend TypeScript check, focused frontend suites, formatting/linting, and CI preflight.

  • A component regression verifies that incomplete history is disclosed and complete-history cards retain their existing presentation.

  • Backend regressions cover cell-query limits, experiment membership, scorer authorization, and serializers.

  • Frontend regressions cover stale responses, failed pagination, local preference isolation, redirects, version conflicts, and missing/zero scores.

  • Automated browser checks exercise score selection and order, 25 scorer columns, lazy payload reads, keyboard focus, and period comparison. Storybook renders cover narrow layouts, dark mode, and both flag states.

  • A real-view database roundtrip uploads numeric, boolean, and categorical results without a dataset, completes the experiment, and reads its history.

  • Feedback regressions cover automatic scorer selection, version navigation, and filter pagination. Browser checks cover collapsible payloads and compact tables at normal and narrow widths.

  • Hover regression stories verify timestamp alignment across sparse charts, switching the active chart, and clearing guides on mouse leave in light/dark and narrow layouts.

  • Browser checks confirm shared upload-state filtering, unchanged checkbox positions, onboarding at 520px, and readable light/dark tooltips.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Offline views, history links, and offline-specific copy use ai-observability-offline-evaluations. Scorer relocation and manual-review management remain available without that flag.

Automatic notifications

  • Publish to changelog?

Docs update

Updated docs/internal/ai-offline-evaluation-reporting.md with the UI behavior and bounded cell-read contract.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Codex, gpt-6-astra

Tools included shell commands, Playwright, and collaborating agents. Fixtures and screenshots contain invented data. The local CodeRabbit review was skipped at the user's request.

Skills applied: UI components, frontend placement, Kea logic, generated API types, user-facing copy, charts, Storybook feature flags, test design, efficient validation, local development, CI preflight, PR descriptions, and GitHub operations. Feedback refinements add version navigation, compact tables, and collapsible payloads. The update keeps the existing screenshots at the author's request.

The cleanup following the code audit preserves existing UI behavior, adds partial-history disclosure, and verifies the rendered result against the previous PR head.

@trunk-io

trunk-io Bot commented Sep 28, 2026

Copy link
Copy Markdown

Merging to master in this repository is managed by Trunk.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

@Radu-Raicea Radu-Raicea self-assigned this Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 27 functions above the limit (max 37)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
onFailure frontend/src/initKea.ts:190 37 10
openScene frontend/src/scenes/sceneLogic.tsx:871 36 10
OfflineItemInspector products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsx:23 34 10
applySearchParams products/ai_observability/frontend/aiObservabilitySharedLogic.ts:572 32 10
buildOfflineTrendPanels products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts:192 32 10
OfflineExperimentsOverview products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx:28 26 10
OfflineScorerHistory products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx:21 25 10
OfflineExperimentContent products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx:22 23 10
pluralizeResource frontend/src/lib/utils/accessControlUtils.ts:97 22 10
setScene frontend/src/scenes/sceneLogic.tsx:787 22 10
AIObservabilityEvaluationsContent products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx:133 20 10
OfflineItemMatrix products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx:12 17 10
<anonymous> products/ai_observability/frontend/offline-evaluations/OfflinePayload.tsx:31 17 10
offlineScoreConfigurationLabel products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts:168 17 10
resourceTypeToString frontend/src/lib/utils/accessControlUtils.ts:172 16 10
OfflineOverviewTrend products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx:13 16 10
cleanFilters products/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.ts:55 16 10
<anonymous> products/ai_observability/frontend/aiObservabilitySharedLogic.ts:486 15 10
offlineCompletionError products/ai_observability/frontend/offline-evaluations/offlineResultPresentation.ts:52 15 10
<anonymous> frontend/src/scenes/sceneLogic.tsx:558 14 10
loadScene frontend/src/scenes/sceneLogic.tsx:985 14 10
render products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx:333 14 10
TraceReviewButton products/ai_observability/frontend/traceReviews/TraceReviewButton.tsx:322 14 10
ScoreDefinitionModal products/ai_observability/frontend/scoreDefinitions/AIObservabilityScoreDefinitions.tsx:263 12 10
submit products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.ts:110 12 10
<anonymous> frontend/src/scenes/sceneLogic.tsx:688 11 10
applyDashboard products/ai_observability/frontend/aiObservabilitySharedLogic.ts:673 11 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Duplication (TypeScript) — 4 new duplicated blocks (worst 1320 tokens)

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

First copy Second copy Lines Tokens
frontend/src/products.tsx:326 products/ai_observability/manifest.tsx:212 117 1320
frontend/src/products.tsx:357 products/ai_observability/manifest.tsx:246 87 1022
frontend/src/products.tsx:1206 products/ai_observability/manifest.tsx:344 77 596
frontend/src/products.tsx:69 products/ai_observability/manifest.tsx:173 38 237
⚠️ Bundle size — 🔺 +55.7 KiB (+0.1%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 69.04 MiB · 🔺 +55.7 KiB (+0.1%)

File Size Δ vs base
posthog-app/_parent/products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.js removed 🟢 -48.7 KiB (-100.0%)
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.js 31.4 KiB 🔺 +31.4 KiB (new)
posthog-app/_parent/products/ai_observability/frontend/evaluations/EvaluationsScene.js 30.9 KiB 🔺 +30.9 KiB (new)
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineExperimentsScene.js 26.8 KiB 🔺 +26.8 KiB (new)
posthog-app/_parent/products/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.js 21.2 KiB 🔺 +21.2 KiB (new)
posthog-app/_parent/products/ai_observability/frontend/AIObservabilityScene.js 168.5 KiB 🟢 -18.4 KiB (-9.9%)
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.js 15.7 KiB 🔺 +15.7 KiB (new)
posthog-app/src/lib/components/ProductEmptyState/ProductEmptyStateGate.js 4.4 KiB 🟢 -9.5 KiB (-68.3%)
render-query/src/render-query/render-query.js 20.18 MiB 🔺 +2.4 KiB (+0.0%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.58 MiB · 22 files 🔺 +1.4 KiB (+0.1%) █████████░ 85.9% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.53 MiB · 630 files 🔺 +4.5 KiB (+0.1%) █████████░ 87.5% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.35 MiB · 2,340 files 🔺 +3.7 KiB (+0.0%) █████████░ 88.2% of 8.34 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
90.4 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
28.4 KiB src/scenes/scenes.ts
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
301.8 KiB ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
272.3 KiB src/taxonomy/core-filter-definitions-by-group.json
216.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
98.8 KiB ../packages/quill/packages/quill/dist/index.js
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
90.4 KiB src/products.tsx

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.17 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.17 MiB · 19 files 🔺 +1.4 KiB (+0.1%) ████░░░░░░ 37.8% of 5.72 MiB
Deferred (lazy) 2.10 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
807.2 KiB dist/toolbar/toolbar-app-NXU2N5Z6.css
651.7 KiB dist/toolbar/chunk-chunk-WOUFMKXM.js
259.4 KiB dist/toolbar/chunk-chunk-OMXHH7OQ.js
138.3 KiB dist/toolbar/chunk-chunk-35ZFDAZZ.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-OMAJT7IB.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-XKFUP4SJ.js
21.0 KiB dist/toolbar/chunk-chunk-NVDDRIT3.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +842.8 KiB (+0.1%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 948.89 MiB · 🔺 +842.8 KiB (+0.1%)

ℹ️ MCP UI apps size — 33 app(s), 17633.1 KB JS

Built size of each MCP UI app (main.js + styles.css).

App JS CSS
debug 597.9 KB 199.2 KB
action 454.1 KB 199.2 KB
action-list 564.2 KB 199.2 KB
cohort 453.1 KB 199.2 KB
cohort-list 563.2 KB 199.2 KB
email-template 452.9 KB 199.2 KB
error-details 469.6 KB 199.2 KB
error-issue 454.5 KB 199.2 KB
error-issue-list 564.8 KB 199.2 KB
experiment 561.3 KB 199.2 KB
experiment-list 564.9 KB 199.2 KB
experiment-results 566.3 KB 199.2 KB
feature-flag 566.8 KB 199.2 KB
feature-flag-list 570.5 KB 199.2 KB
feature-flag-testing 457.3 KB 199.2 KB
inline-scan 453.6 KB 199.2 KB
insight-actors 562.3 KB 199.2 KB
invite-email-preview 452.3 KB 199.2 KB
llm-costs 559.3 KB 199.2 KB
session-recording 455.3 KB 199.2 KB
survey 454.7 KB 199.2 KB
survey-global-stats 561.9 KB 199.2 KB
survey-list 564.9 KB 199.2 KB
survey-stats 561.9 KB 199.2 KB
trace-span 453.5 KB 199.2 KB
trace-span-list 564.1 KB 199.2 KB
vision-observation-list 563.3 KB 199.2 KB
workflow 453.4 KB 199.2 KB
workflow-list 563.5 KB 199.2 KB
loops-review 457.8 KB 199.2 KB
query-results 774.1 KB 199.2 KB
render-ui 858.1 KB 199.2 KB
visual-review-snapshots 457.9 KB 199.2 KB
✅ Playwright — all passed

All tests passed.

View test results →

⚠️ MCP snapshots — 1 updated (1 modified, 0 added, 0 deleted)

Snapshots: MCP unit test snapshots updated

Changes: 1 snapshots (1 modified, 0 added, 0 deleted)

What this means:

  • Snapshots have been automatically updated to match current output

Next steps:

  • Review the changes to ensure they're intentional
  • If unexpected, investigate what caused the output to change

Review snapshot changes →

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[Medium risk] Adds offline experiment inspection UI and backend read endpoints.

The PR appears safe to merge, though all-time hover alignment should be corrected for sparse scorer charts.

Reviews (2) · Last reviewed commit: "fix(aio): clarify offline history covera..."

Comment thread products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx Outdated
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The change adds a bounded result-cell read endpoint and dedicated routes for online evaluations, offline experiments, scorers, and scorer history. New frontend workflows provide experiment overview and detail views, trend charts, item inspection, scorer history, and scorer version creation. Access checks now support scenes mapped to multiple resource types. Documentation describes the read API and offline evaluation views.

Priority: ➖ Normal

Merge Risk: 🟡 Moderate · up to 6a35c

Scorer-history links can change a chosen comparison range, and experiment details can show dependent panels when the experiment cannot load. Previously reported scorer-version and result-cell issues also remain open. Resolve these issues before merging unless their effects are explicitly accepted.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to 6a35c

The new result read and dedicated pages warrant design review because they expose more ways to inspect evaluation data. The inspected server-side path enforces read permissions and limits the response; no introduced security issue was established. Coverage of the broader change remains incomplete.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — A direct caller can request up to 1,000 selected result cells per request, but inspected querysets and identity checks constrain the read to the caller’s team and authorized scorer definitions.

Trust Boundaries and Controls

  • observed — Client-selected UUIDs cross the HTTP boundary through validated query fields. Server-side scope checks and exact item and authorized-version resolution govern the response rather than frontend visibility alone.
  • observed — The result-cell serializer exposes scores and payload lifecycle metadata, not payload contents; payload retrieval remains on distinct actions.

Resilience and Maintainability Implications

  • observed — Completion remains a server-controlled state transition: a concurrent or repeated request is handled by the existing locked transition and terminal-state behavior, not by the new button’s status check.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and follows the repository template. It explains the problem, user-visible changes, feature-flagged release status, screenshots, documentation updates, testing, and agent c…
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (2)
products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.ts-135-139 (1)

135-139: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

On a 409, the loader cancels itself, so finishSubmit never runs and submitting stays true.

In the catch branch, breakpoint() runs first. The finally block then calls actions.finishSubmit(). If the modal unmounts or submit runs again while the request is still pending, breakpoint() throws. The finally block still runs because breakpoint() is inside the try. So that case is fine.

The 409 path has a different problem. await asyncActions.loadCurrentDefinition() dispatches loadCurrentDefinition. The currentDefinition reducer then resets the value to null. If the reload fails, currentDefinition stays null. The loader-failure reducer sets error, but the code on line 139 then replaces it with "confirm again". This happens even though no new definition loaded, so the user sees the wrong message.

Set the conflict error only when the reload returns a definition.

Proposed fix
-                    await asyncActions.loadCurrentDefinition()
-                    actions.setError('The current version changed. Review the new source version and confirm again.')
+                    await asyncActions.loadCurrentDefinition()
+                    if (values.currentDefinition) {
+                        actions.setError('The current version changed. Review the new source version and confirm again.')
+                    }
products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts-585-592 (1)

585-592: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Filter changes from replaceFilters never reach the URL.

actionToUrl handles only setFilters. This is the normal flow: on mount, replaceFilters runs, then loadOfflineHistoryDefinitionSuccess calls setFilters({version}), which writes the URL. But readFilters takes all 15 keys from search, and the URL is built from every filter except empty or null ones. The defaults date_to: null and compare_to: '-30d' are left out of the URL, so a reload applies them again. This part is consistent.

The real gap is on line 588. It drops null values, but readFilters accepts null for date_to only when the key is present. Suppose the user picks a bounded date_to and then clears it back to null. The URL loses the key, and urlToAction reads the default, which is also null. That is fine. For compare_to, the default is '-30d'. If the user clears it to null, the URL drops the key, and a reload restores '-30d'. The comparison window then changes after a reload or when someone opens a shared link.

Keep null values in the URL and compare them against the defaults. Only drop a key when its value equals DEFAULT_FILTERS.

Proposed fix
-            Object.fromEntries(Object.entries(values.filters).filter(([, value]) => value !== null && value !== '')),
+            Object.fromEntries(
+                Object.entries(values.filters).filter(
+                    ([key, value]) => value !== DEFAULT_FILTERS[key as keyof OfflineHistoryFilters]
+                )
+            ),
🧹 Nitpick comments (2)
products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts (1)

482-486: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Share the matrix column widths and batch size between the logic and the matrix.

The viewport batch math in the logic and the rendered column widths in the matrix use separate hardcoded values: 220, 180, and 20. If one file changes and the other does not, the logic requests the wrong column batches. Visible columns then stay at "Scroll to load". Export shared constants from offlineExperimentLogic.ts, for example OFFLINE_ITEM_COLUMN_WIDTH, OFFLINE_SCORER_COLUMN_WIDTH, and OFFLINE_CELL_BATCH_SIZE.

  • products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts#L482-L486: replace the literals in requestedBatches with the constants. Also replace index * 20 in pumpCells.
  • products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L41-L42: use OFFLINE_ITEM_COLUMN_WIDTH for width. Add a comment that explains why the arbitrary min-w-[220px]/max-w-[220px] classes are needed.
  • products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L73-L74: use OFFLINE_SCORER_COLUMN_WIDTH for width. Add the same comment to the 180px classes.
  • products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L83-L83: use Math.floor(index / OFFLINE_CELL_BATCH_SIZE).

As per coding guidelines: "an arbitrary value (w-[347px]) needs a comment explaining the constraint, or it should be a scale value."

Source: Coding guidelines

products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx (1)

455-461: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use offlineRunSourceLabel for every run-source display.

offlineRunSourceLabel in offlineOverviewState.ts is the single mapping from run source to label. Two new views skip it. One view shows raw enum values to users. The other view duplicates the mapping.

  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx#L455-L461: replace the inline { ci, local, scheduled } map with offlineRunSourceLabel(experiment.run_source).
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsx#L82-L82: replace experiment.run_source || 'Source not specified' with offlineRunSourceLabel(experiment.run_source). Without this change, the tooltip shows ci and local where the table shows CI and Local.

The run-source select options also appear three times: OfflineScorerHistory.tsx lines 239-245 and OfflineExperimentsOverview.tsx lines 159-165 and 273-279. Export one options constant next to offlineRunSourceLabel.

This follows the path instruction that code "says everything once and only once."

Source: Path instructions


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: cec31315-b1c0-43ef-87a7-609f1b67d1a0

📥 Commits

Reviewing files that changed from the base of the PR and between 40eeb71 and e467fff.

⛔ Files ignored due to path filters (2)
  • products/ai_observability/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/ai_observability/frontend/generated/api.ts is excluded by !**/generated/**
📒 Files selected for processing (80)
  • docs/internal/ai-offline-evaluation-reporting.md
  • frontend/src/initKea.ts
  • frontend/src/lib/constants.tsx
  • frontend/src/lib/utils/accessControlUtils.ts
  • frontend/src/productScenes.tsx
  • frontend/src/products.tsx
  • frontend/src/scenes/sceneLogic.test.tsx
  • frontend/src/scenes/sceneLogic.tsx
  • frontend/src/scenes/sceneTypes.ts
  • products/ai_observability/backend/api/offline_experiment_read_serializers.py
  • products/ai_observability/backend/api/offline_experiment_reads.py
  • products/ai_observability/backend/api/test/test_offline_experiment_reads.py
  • products/ai_observability/backend/api/test/test_offline_experiment_serializers.py
  • products/ai_observability/backend/offline_evaluation_read_service.py
  • products/ai_observability/backend/offline_evaluation_read_types.py
  • products/ai_observability/frontend/aiObservabilityLogic.test.ts
  • products/ai_observability/frontend/aiObservabilitySharedLogic.ts
  • products/ai_observability/frontend/emptyState/evaluationsEmptyState.tsx
  • products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx
  • products/ai_observability/frontend/evaluations/EvaluationsOnlineNav.tsx
  • products/ai_observability/frontend/evaluations/EvaluationsScene.tsx
  • products/ai_observability/frontend/evaluations/EvaluationsTabs.tsx
  • products/ai_observability/frontend/evaluations/components/OfflineEvaluationsTab.tsx
  • products/ai_observability/frontend/evaluations/evaluationsEntryLogic.ts
  • products/ai_observability/frontend/evaluations/evaluationsEntryRedirect.test.ts
  • products/ai_observability/frontend/evaluations/evaluationsEntryRedirect.ts
  • products/ai_observability/frontend/evaluations/llmEvaluationsLogic.ts
  • products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.test.ts
  • products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.ts
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsScene.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.test.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflinePayload.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreChooser.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.tsx
  • products/ai_observability/frontend/offline-evaluations/offlineDetailFixtures.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineItemInspectorLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineItemInspectorLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewFixtures.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewState.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewState.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineResultPresentation.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.fixtures.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts
  • products/ai_observability/frontend/scoreDefinitions/AIObservabilityScoreDefinitions.tsx
  • products/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.stories.tsx
  • products/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.tsx
  • products/ai_observability/frontend/scoreDefinitions/ScoreDefinitionActions.tsx
  • products/ai_observability/frontend/scoreDefinitions/ScoreDefinitionVersionButton.tsx
  • products/ai_observability/frontend/scoreDefinitions/ScoreDefinitionVersionModal.tsx
  • products/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.test.ts
  • products/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.ts
  • products/ai_observability/frontend/scoreDefinitions/scoreDefinitionNavigation.ts
  • products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.test.ts
  • products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.ts
  • products/ai_observability/frontend/traceReviews/AIObservabilityHumanReviews.tsx
  • products/ai_observability/frontend/traceReviews/TraceReviewButton.tsx
  • products/ai_observability/manifest.tsx
  • products/ai_observability/mcp/tools.yaml
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (3)
  • products/ai_observability/frontend/evaluations/components/OfflineEvaluationsTab.tsx
  • products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.test.ts
  • products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.

@trunk-io

trunk-io Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
AI observability/Offline experiments/Score trends HiddenSeries play-test The story remount on retry failed because 'act' is not supported in production React builds. Logs ↗︎
Scenes-App/SidePanels SidePanelNotebooks smoke-test The test timed out while waiting for a loading indicator or spinner to disappear. Logs ↗︎

View Full Report ↗︎ ⋅ Docs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts-496-498 (1)

496-498: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Cell batches go stale when summaries reload after item pages load.

refresh calls loadOfflineSummaries and loadOfflineItems together. loadOfflineSummaries resets summaries to []. addSummaryPage then appends the new scorers and changes the scorer slices for each batch index. However, cellBatches is reset only by loadOfflineItems.

addSummaryPage fires pumpCells. If items load first, pumpCells requests batch 0 with only the first summary page's scorers. The later summary pages can add scorers to batch 0's index range. Those scorers do not get requests, because cellBatches[0] already exists. Their cells show "No result" when a result may exist.

With 100-scorer summary pages this is uncommon. It still happens on the first load if items resolve before the summaries finish paginating. The first summary page already fills batches 0–4, so the gap appears when a later page adds scorers to a batch that was requested from a partial slice.

Fix: store the requested versionIds in each batch. In pumpCells, request the batch again when the current slice differs from the stored IDs. A simpler alternative is to wait for loadOfflineSummariesSuccess before pumping batches that were only partially covered.

🧹 Nitpick comments (2)
products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts (2)

528-549: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Merge the duplicate setFilters and replaceFilters listeners.

Lines 528-538 and Lines 539-549 are the same. This breaks the "says everything once and only once" rule. Move the shared body into one helper.

♻️ Proposed refactor
-    listeners(({ actions, values }) => ({
+    listeners(({ actions, values }) => {
+        const reloadForFilters = (): void => {
+            if (values.filters.version) {
+                if (values.selectedVersion?.id !== values.filters.version) {
+                    actions.loadOfflineHistoryVersion()
+                }
+                if (!values.windows.error) {
+                    actions.loadOfflineHistoryPrimaryPage()
+                    actions.loadOfflineHistoryComparisonPage()
+                }
+            }
+        }
+        return {
 ...
-        setFilters: () => { ... },
-        replaceFilters: () => { ... },
+        setFilters: reloadForFilters,
+        replaceFilters: reloadForFilters,

As per path instructions: "says everything once and only once".

Source: Path instructions


556-558: 🎯 Functional Correctness | 🔵 Trivial | 💤 Low value

Guard changePage against an invalid date window.

The setFilters and refreshHistory listeners skip loading when values.windows.error is set. The changePage listener does not. If a page change happens while the window is invalid, loadOfflineHistoryPrimaryPage sends a request with date_from and date_to set to undefined. The backend then returns unbounded history. The UI hides the table in this state, so this is unlikely to happen today. Use the same guard for consistency.

         changePage: ({ period }) => {
+            if (!values.filters.version || values.windows.error) {
+                return
+            }
             period === 'primary' ? actions.loadOfflineHistoryPrimaryPage() : actions.loadOfflineHistoryComparisonPage()
         },

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 79c9045b-6755-451f-a3f2-6c946305c4a6

📥 Commits

Reviewing files that changed from the base of the PR and between beaa7db and af44918.

⛔ Files ignored due to path filters (3)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
  • products/ai_observability/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/ai_observability/frontend/generated/api.ts is excluded by !**/generated/**
📒 Files selected for processing (30)
  • docs/internal/ai-offline-evaluation-reporting.md
  • frontend/src/lib/constants.tsx
  • frontend/src/products.tsx
  • frontend/src/scenes/sceneLogic.test.tsx
  • frontend/src/scenes/sceneLogic.tsx
  • products/ai_observability/frontend/evaluations/EvaluationsTabs.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.test.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflinePayload.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendLines.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.tsx
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts
  • products/ai_observability/manifest.tsx
  • products/ai_observability/package.json
  • services/mcp/src/api/generated.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts-76-76 (1)

76-76: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Compare resolved date bounds by instant.

URL filters accept arbitrary date strings, so an encoded offset ISO value can reach the dayjs(value) branch. toISOString() can emit expanded years, whose lexical order is not chronological. The current check can reject a valid range or accept reversed bounds.

Suggested fix
-    if (range.dateFrom && range.dateFrom >= range.dateTo) {
+    if (range.dateFrom && Date.parse(range.dateFrom) >= Date.parse(range.dateTo)) {

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: b4f796e0-ec23-46c4-a287-030251b4fd76

📥 Commits

Reviewing files that changed from the base of the PR and between af44918 and 2558f0b.

⛔ Files ignored due to path filters (3)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
  • products/ai_observability/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/ai_observability/frontend/generated/api.ts is excluded by !**/generated/**
📒 Files selected for processing (11)
  • docs/internal/ai-offline-evaluation-reporting.md
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendCrosshair.tsx
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.test.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts
  • services/mcp/src/api/generated.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx-153-153 (1)

153-153: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Render dependent panels only after the experiment resolves.

If the initial experiment request fails, the error banner appears while scorer summaries and OfflineItemMatrix still render. Keep those panels behind the experiment loading and error states so the page does not show dependent content for an unavailable experiment. As per coding guidelines, “Branch in resolution order: loading → error → empty → content.”

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: d8197e3f-4169-469b-96c3-6977ef1ea21b

📥 Commits

Reviewing files that changed from the base of the PR and between 2558f0b and 6a35c10.

⛔ Files ignored due to path filters (3)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
  • products/ai_observability/frontend/generated/api.schemas.ts is excluded by !**/generated/**
  • products/ai_observability/frontend/generated/api.ts is excluded by !**/generated/**
📒 Files selected for processing (15)
  • docs/internal/ai-offline-evaluation-reporting.md
  • frontend/src/lib/constants.tsx
  • products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentStatus.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.test.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryRunDetails.tsx
  • products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryTable.tsx
  • products/ai_observability/frontend/offline-evaluations/offlineExperimentPresentation.ts
  • products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts
  • services/mcp/src/api/generated.ts
💤 Files with no reviewable changes (2)
  • frontend/src/lib/constants.tsx
  • products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

🦔 Hogbox preview · ✅ ready

▶ Open the preview

🔑 Login test@posthog.com / 12345678 (demo data)
🧩 Running this PR's backend and frontend, on the PostHog :master base
🔗 Link stable across rebuilds — a re-push swaps the box underneath, the URL stays
🔒 Access tailnet only (PostHog VPN)
🛠️ Admin inspect & debug state in hogland
💤 Idle sleeps after ~30 min idle (snapshot to S3, zero node cost) and wakes on your next visit in ~30s, behind a brief "waking up" screen

commit e953215 · box box-f965cb07a6ec · ready in 968s (push → usable) · build log · rebuilds on every push, torn down on close

@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team September 29, 2026 16:33
A search change in the offline overview reset the list clock. The trend date window therefore moved on each settled keystroke, and every selected score chart reloaded with the same data.

A search-only filter change now keeps the current clock. Date, source, upload-state, refresh, and URL changes still reset it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:27 In progress
The Scorers scene export had no productKey, unlike the other AI observability scenes. Its pageviews therefore lost product_key, and Quick Start did not select AI observability on that page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:29 In progress
Refresh clears the scorer summaries, which removes every score column. The browser then moves the table scroll back to 0. The matrix ignores scroll events while the table fits its container, so the stored viewport kept the old position. After a refresh far to the right, the first visible score batch therefore never loaded.

Loading summaries now resets the stored scroll position to 0 and keeps the measured width.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:31 In progress
The router decodes URL text such as 00123 or true as a number or boolean. The offline overview accepted only string filters, so a reload, back navigation, or shared link dropped those searches without notice.

The overview now writes the search as a one-element array, which the router keeps exactly. The reader accepts that array and still accepts a plain string search.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:35 In progress
With all time selected, each overview chart spans only its own results, so a chart hides the shared guide when the hovered time is outside that span. The reporting doc now says so and keeps the unconditional claim for bounded date ranges only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:37 In progress
LemonTabs blocks onChange for a disabled tab but still renders its link. A user without access to online evals, offline evals, or scorers could therefore click the greyed tab and land on the Access denied page.

A tab with an access disabled reason now has no link. Access checks are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:39 In progress
The panel's open prop came from the loaded summary list. Refresh empties that list first, so for more than 10 scorers the prop went from false to true and back to false. React then wrote open = false and collapsed a panel the user had opened.

The panel now reads the last loaded total, which stays set during a refresh. The prop stays the same, so the user's choice remains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 18:42 In progress
ApiError.formattedRetryAfter already reads as "in 30s" or "later" for a 429. The offline evaluation error message added its own "in", so users saw "Try again in in 30s" or "Try again in later".

The message now uses the formatted value as returned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 91388465-09c8-4846-8672-437ea66bed96
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 19:32 In progress
@github-actions
github-actions Bot requested a deployment to preview-pr-107937 September 29, 2026 19:43 In progress

This branch was successfully deployed

1 active deployment
preview-pr-107937 — e953215a Deployed Sep 29, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant