feat(aio): add offline experiment inspection and score trends - #107937
Radu-Raicea wants to merge 36 commits into
Conversation
|
Merging to
After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here |
🤖 CI report
|
| Function | Location | Complexity | Limit |
|---|---|---|---|
onFailure |
frontend/src/initKea.ts:190 |
37 | 10 |
openScene |
frontend/src/scenes/sceneLogic.tsx:871 |
36 | 10 |
OfflineItemInspector |
products/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsx:23 |
34 | 10 |
applySearchParams |
products/ai_observability/frontend/aiObservabilitySharedLogic.ts:572 |
32 | 10 |
buildOfflineTrendPanels |
products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts:192 |
32 | 10 |
OfflineExperimentsOverview |
products/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsx:28 |
26 | 10 |
OfflineScorerHistory |
products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx:21 |
25 | 10 |
OfflineExperimentContent |
products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx:22 |
23 | 10 |
pluralizeResource |
frontend/src/lib/utils/accessControlUtils.ts:97 |
22 | 10 |
setScene |
frontend/src/scenes/sceneLogic.tsx:787 |
22 | 10 |
AIObservabilityEvaluationsContent |
products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx:133 |
20 | 10 |
OfflineItemMatrix |
products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx:12 |
17 | 10 |
<anonymous> |
products/ai_observability/frontend/offline-evaluations/OfflinePayload.tsx:31 |
17 | 10 |
offlineScoreConfigurationLabel |
products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts:168 |
17 | 10 |
resourceTypeToString |
frontend/src/lib/utils/accessControlUtils.ts:172 |
16 | 10 |
OfflineOverviewTrend |
products/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsx:13 |
16 | 10 |
cleanFilters |
products/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.ts:55 |
16 | 10 |
<anonymous> |
products/ai_observability/frontend/aiObservabilitySharedLogic.ts:486 |
15 | 10 |
offlineCompletionError |
products/ai_observability/frontend/offline-evaluations/offlineResultPresentation.ts:52 |
15 | 10 |
<anonymous> |
frontend/src/scenes/sceneLogic.tsx:558 |
14 | 10 |
loadScene |
frontend/src/scenes/sceneLogic.tsx:985 |
14 | 10 |
render |
products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx:333 |
14 | 10 |
TraceReviewButton |
products/ai_observability/frontend/traceReviews/TraceReviewButton.tsx:322 |
14 | 10 |
ScoreDefinitionModal |
products/ai_observability/frontend/scoreDefinitions/AIObservabilityScoreDefinitions.tsx:263 |
12 | 10 |
submit |
products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.ts:110 |
12 | 10 |
<anonymous> |
frontend/src/scenes/sceneLogic.tsx:688 |
11 | 10 |
applyDashboard |
products/ai_observability/frontend/aiObservabilitySharedLogic.ts:673 |
11 | 10 |
✅ Duplication (Python) — clean
New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
⚠️ Duplication (TypeScript) — 4 new duplicated blocks (worst 1320 tokens)
New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
| First copy | Second copy | Lines | Tokens |
|---|---|---|---|
frontend/src/products.tsx:326 |
products/ai_observability/manifest.tsx:212 |
117 | 1320 |
frontend/src/products.tsx:357 |
products/ai_observability/manifest.tsx:246 |
87 | 1022 |
frontend/src/products.tsx:1206 |
products/ai_observability/manifest.tsx:344 |
77 | 596 |
frontend/src/products.tsx:69 |
products/ai_observability/manifest.tsx:173 |
38 | 237 |
⚠️ Bundle size — 🔺 +55.7 KiB (+0.1%)
Uncompressed size of every built .js bundle, compared against the base branch.
Total: 69.04 MiB · 🔺 +55.7 KiB (+0.1%)
| File | Size | Δ vs base |
|---|---|---|
posthog-app/_parent/products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.js |
removed | 🟢 -48.7 KiB (-100.0%) |
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.js |
31.4 KiB | 🔺 +31.4 KiB (new) |
posthog-app/_parent/products/ai_observability/frontend/evaluations/EvaluationsScene.js |
30.9 KiB | 🔺 +30.9 KiB (new) |
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineExperimentsScene.js |
26.8 KiB | 🔺 +26.8 KiB (new) |
posthog-app/_parent/products/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.js |
21.2 KiB | 🔺 +21.2 KiB (new) |
posthog-app/_parent/products/ai_observability/frontend/AIObservabilityScene.js |
168.5 KiB | 🟢 -18.4 KiB (-9.9%) |
posthog-app/_parent/products/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.js |
15.7 KiB | 🔺 +15.7 KiB (new) |
posthog-app/src/lib/components/ProductEmptyState/ProductEmptyStateGate.js |
4.4 KiB | 🟢 -9.5 KiB (-68.3%) |
render-query/src/render-query/render-query.js |
20.18 MiB | 🔺 +2.4 KiB (+0.0%) |
Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report
✅ Eager graph — within budget
How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.
| Root | Eager (shipped) | Δ vs base | Budget |
|---|---|---|---|
entry (logged-out pages, app bootstrap)src/index.tsx |
1.58 MiB · 22 files | 🔺 +1.4 KiB (+0.1%) | █████████░ 85.9% of 1.84 MiB |
logged-out boot: index + App + bootApp (preloaded by every page, including /login)src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts |
3.53 MiB · 630 files | 🔺 +4.5 KiB (+0.1%) | █████████░ 87.5% of 4.03 MiB |
authenticated shell (every logged-in page)src/scenes/AuthenticatedShell.tsx |
7.35 MiB · 2,340 files | 🔺 +3.7 KiB (+0.0%) | █████████░ 88.2% of 8.34 MiB |
🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
Largest files eagerly shipped from src/index.tsx
| Size | File |
|---|---|
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 24.6 KiB | ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js |
| 6.3 KiB | ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js |
| 4.5 KiB | ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js |
| 3.9 KiB | ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js |
| 1.4 KiB | ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js |
| 1.3 KiB | src/index.tsx |
| 1.3 KiB | src/RootErrorBoundary.tsx |
| 912 B | ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js |
| 854 B | src/scenes/ChunkLoadErrorBoundary.tsx |
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
| Size | File |
|---|---|
| 301.8 KiB | ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 216.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 100.5 KiB | src/lib/api.ts |
| 90.4 KiB | src/products.tsx |
| 69.4 KiB | src/lib/lemon-ui/icons/icons.tsx |
| 40.1 KiB | src/lib/utils/eventUsageLogic.ts |
| 38.7 KiB | ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js |
| 33.9 KiB | ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js |
| 28.4 KiB | src/scenes/scenes.ts |
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
| Size | File |
|---|---|
| 301.8 KiB | ../node_modules/.pnpm/posthog-js@1.434.14_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 272.3 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 216.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 100.5 KiB | src/lib/api.ts |
| 98.8 KiB | ../packages/quill/packages/quill/dist/index.js |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 90.6 KiB | ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js |
| 90.4 KiB | src/products.tsx |
Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479
✅ Toolbar bundle — eager 2.17 MiB within budget
What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.
| Metric | Size | Δ vs base | Budget |
|---|---|---|---|
| Eager (shipped) entry + static imports |
2.17 MiB · 19 files | 🔺 +1.4 KiB (+0.1%) | ████░░░░░░ 37.8% of 5.72 MiB |
| Deferred (lazy) | 2.10 MiB · 44 files | no change | n/a — loads on demand |
Loader dist/toolbar.js |
1.2 KiB | no change | █░░░░░░░░░ 6.0% of 19.5 KiB |
Largest eagerly-shipped chunks
| Size | File |
|---|---|
| 807.2 KiB | dist/toolbar/toolbar-app-NXU2N5Z6.css |
| 651.7 KiB | dist/toolbar/chunk-chunk-WOUFMKXM.js |
| 259.4 KiB | dist/toolbar/chunk-chunk-OMXHH7OQ.js |
| 138.3 KiB | dist/toolbar/chunk-chunk-35ZFDAZZ.js |
| 131.8 KiB | dist/toolbar/chunk-chunk-FDH2IBXT.js |
| 75.2 KiB | dist/toolbar/toolbar-app-OMAJT7IB.js |
| 69.0 KiB | dist/toolbar/chunk-chunk-TSAL54PB.js |
| 35.6 KiB | dist/toolbar/chunk-chunk-XKFUP4SJ.js |
| 21.0 KiB | dist/toolbar/chunk-chunk-NVDDRIT3.js |
| 6.8 KiB | dist/toolbar/chunk-chunk-DV7IWQNF.js |
Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile
✅ Dist folder size — 🔺 +842.8 KiB (+0.1%)
Total size of the built frontend/dist folder (all assets), compared against the base branch.
Total: 948.89 MiB · 🔺 +842.8 KiB (+0.1%)
ℹ️ MCP UI apps size — 33 app(s), 17633.1 KB JS
Built size of each MCP UI app (main.js + styles.css).
| App | JS | CSS |
|---|---|---|
| debug | 597.9 KB | 199.2 KB |
| action | 454.1 KB | 199.2 KB |
| action-list | 564.2 KB | 199.2 KB |
| cohort | 453.1 KB | 199.2 KB |
| cohort-list | 563.2 KB | 199.2 KB |
| email-template | 452.9 KB | 199.2 KB |
| error-details | 469.6 KB | 199.2 KB |
| error-issue | 454.5 KB | 199.2 KB |
| error-issue-list | 564.8 KB | 199.2 KB |
| experiment | 561.3 KB | 199.2 KB |
| experiment-list | 564.9 KB | 199.2 KB |
| experiment-results | 566.3 KB | 199.2 KB |
| feature-flag | 566.8 KB | 199.2 KB |
| feature-flag-list | 570.5 KB | 199.2 KB |
| feature-flag-testing | 457.3 KB | 199.2 KB |
| inline-scan | 453.6 KB | 199.2 KB |
| insight-actors | 562.3 KB | 199.2 KB |
| invite-email-preview | 452.3 KB | 199.2 KB |
| llm-costs | 559.3 KB | 199.2 KB |
| session-recording | 455.3 KB | 199.2 KB |
| survey | 454.7 KB | 199.2 KB |
| survey-global-stats | 561.9 KB | 199.2 KB |
| survey-list | 564.9 KB | 199.2 KB |
| survey-stats | 561.9 KB | 199.2 KB |
| trace-span | 453.5 KB | 199.2 KB |
| trace-span-list | 564.1 KB | 199.2 KB |
| vision-observation-list | 563.3 KB | 199.2 KB |
| workflow | 453.4 KB | 199.2 KB |
| workflow-list | 563.5 KB | 199.2 KB |
| loops-review | 457.8 KB | 199.2 KB |
| query-results | 774.1 KB | 199.2 KB |
| render-ui | 858.1 KB | 199.2 KB |
| visual-review-snapshots | 457.9 KB | 199.2 KB |
⚠️ MCP snapshots — 1 updated (1 modified, 0 added, 0 deleted)
Snapshots: MCP unit test snapshots updated
Changes: 1 snapshots (1 modified, 0 added, 0 deleted)
What this means:
- Snapshots have been automatically updated to match current output
Next steps:
- Review the changes to ensure they're intentional
- If unexpected, investigate what caused the output to change
|
[Medium risk] Adds offline experiment inspection UI and backend read endpoints. The PR appears safe to merge, though all-time hover alignment should be corrected for sparse scorer charts. Reviews (2) · Last reviewed commit: "fix(aio): clarify offline history covera..." |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe change adds a bounded result-cell read endpoint and dedicated routes for online evaluations, offline experiments, scorers, and scorer history. New frontend workflows provide experiment overview and detail views, trend charts, item inspection, scorer history, and scorer version creation. Access checks now support scenes mapped to multiple resource types. Documentation describes the read API and offline evaluation views. Priority: ➖ Normal Merge Risk: 🟡 Moderate · up to Scorer-history links can change a chosen comparison range, and experiment details can show dependent panels when the experiment cannot load. Previously reported scorer-version and result-cell issues also remain open. Resolve these issues before merging unless their effects are explicitly accepted. Security Architecture ReviewSecurity architecture risk: 🟡 Moderate · up to The new result read and dedicated pages warrant design review because they expose more ways to inspect evaluation data. The inspected server-side path enforces read permissions and limits the response; no introduced security issue was established. Coverage of the broader change remains incomplete. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 1✅ Passed checks (1 passed)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (2)
products/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.ts-135-139 (1)
135-139: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winOn a 409, the loader cancels itself, so
finishSubmitnever runs andsubmittingstaystrue.In the
catchbranch,breakpoint()runs first. Thefinallyblock then callsactions.finishSubmit(). If the modal unmounts orsubmitruns again while the request is still pending,breakpoint()throws. Thefinallyblock still runs becausebreakpoint()is inside thetry. So that case is fine.The 409 path has a different problem.
await asyncActions.loadCurrentDefinition()dispatchesloadCurrentDefinition. ThecurrentDefinitionreducer then resets the value tonull. If the reload fails,currentDefinitionstaysnull. The loader-failure reducer setserror, but the code on line 139 then replaces it with "confirm again". This happens even though no new definition loaded, so the user sees the wrong message.Set the conflict error only when the reload returns a definition.
Proposed fix
- await asyncActions.loadCurrentDefinition() - actions.setError('The current version changed. Review the new source version and confirm again.') + await asyncActions.loadCurrentDefinition() + if (values.currentDefinition) { + actions.setError('The current version changed. Review the new source version and confirm again.') + }products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts-585-592 (1)
585-592: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winFilter changes from
replaceFiltersnever reach the URL.
actionToUrlhandles onlysetFilters. This is the normal flow: on mount,replaceFiltersruns, thenloadOfflineHistoryDefinitionSuccesscallssetFilters({version}), which writes the URL. ButreadFilterstakes all 15 keys fromsearch, and the URL is built from every filter except empty or null ones. The defaultsdate_to: nullandcompare_to: '-30d'are left out of the URL, so a reload applies them again. This part is consistent.The real gap is on line 588. It drops
nullvalues, butreadFiltersacceptsnullfordate_toonly when the key is present. Suppose the user picks a boundeddate_toand then clears it back tonull. The URL loses the key, andurlToActionreads the default, which is alsonull. That is fine. Forcompare_to, the default is'-30d'. If the user clears it tonull, the URL drops the key, and a reload restores'-30d'. The comparison window then changes after a reload or when someone opens a shared link.Keep
nullvalues in the URL and compare them against the defaults. Only drop a key when its value equalsDEFAULT_FILTERS.Proposed fix
- Object.fromEntries(Object.entries(values.filters).filter(([, value]) => value !== null && value !== '')), + Object.fromEntries( + Object.entries(values.filters).filter( + ([key, value]) => value !== DEFAULT_FILTERS[key as keyof OfflineHistoryFilters] + ) + ),
🧹 Nitpick comments (2)
products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts (1)
482-486: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winShare the matrix column widths and batch size between the logic and the matrix.
The viewport batch math in the logic and the rendered column widths in the matrix use separate hardcoded values:
220,180, and20. If one file changes and the other does not, the logic requests the wrong column batches. Visible columns then stay at "Scroll to load". Export shared constants fromofflineExperimentLogic.ts, for exampleOFFLINE_ITEM_COLUMN_WIDTH,OFFLINE_SCORER_COLUMN_WIDTH, andOFFLINE_CELL_BATCH_SIZE.
products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts#L482-L486: replace the literals inrequestedBatcheswith the constants. Also replaceindex * 20inpumpCells.products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L41-L42: useOFFLINE_ITEM_COLUMN_WIDTHforwidth. Add a comment that explains why the arbitrarymin-w-[220px]/max-w-[220px]classes are needed.products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L73-L74: useOFFLINE_SCORER_COLUMN_WIDTHforwidth. Add the same comment to the180pxclasses.products/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsx#L83-L83: useMath.floor(index / OFFLINE_CELL_BATCH_SIZE).As per coding guidelines: "an arbitrary value (
w-[347px]) needs a comment explaining the constraint, or it should be a scale value."Source: Coding guidelines
products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx (1)
455-461: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winUse
offlineRunSourceLabelfor every run-source display.
offlineRunSourceLabelinofflineOverviewState.tsis the single mapping from run source to label. Two new views skip it. One view shows raw enum values to users. The other view duplicates the mapping.
products/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsx#L455-L461: replace the inline{ ci, local, scheduled }map withofflineRunSourceLabel(experiment.run_source).products/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsx#L82-L82: replaceexperiment.run_source || 'Source not specified'withofflineRunSourceLabel(experiment.run_source). Without this change, the tooltip showsciandlocalwhere the table showsCIandLocal.The run-source select options also appear three times:
OfflineScorerHistory.tsxlines 239-245 andOfflineExperimentsOverview.tsxlines 159-165 and 273-279. Export one options constant next toofflineRunSourceLabel.This follows the path instruction that code "says everything once and only once."
Source: Path instructions
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: cec31315-b1c0-43ef-87a7-609f1b67d1a0
⛔ Files ignored due to path filters (2)
products/ai_observability/frontend/generated/api.schemas.tsis excluded by!**/generated/**products/ai_observability/frontend/generated/api.tsis excluded by!**/generated/**
📒 Files selected for processing (80)
docs/internal/ai-offline-evaluation-reporting.mdfrontend/src/initKea.tsfrontend/src/lib/constants.tsxfrontend/src/lib/utils/accessControlUtils.tsfrontend/src/productScenes.tsxfrontend/src/products.tsxfrontend/src/scenes/sceneLogic.test.tsxfrontend/src/scenes/sceneLogic.tsxfrontend/src/scenes/sceneTypes.tsproducts/ai_observability/backend/api/offline_experiment_read_serializers.pyproducts/ai_observability/backend/api/offline_experiment_reads.pyproducts/ai_observability/backend/api/test/test_offline_experiment_reads.pyproducts/ai_observability/backend/api/test/test_offline_experiment_serializers.pyproducts/ai_observability/backend/offline_evaluation_read_service.pyproducts/ai_observability/backend/offline_evaluation_read_types.pyproducts/ai_observability/frontend/aiObservabilityLogic.test.tsproducts/ai_observability/frontend/aiObservabilitySharedLogic.tsproducts/ai_observability/frontend/emptyState/evaluationsEmptyState.tsxproducts/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsxproducts/ai_observability/frontend/evaluations/EvaluationsOnlineNav.tsxproducts/ai_observability/frontend/evaluations/EvaluationsScene.tsxproducts/ai_observability/frontend/evaluations/EvaluationsTabs.tsxproducts/ai_observability/frontend/evaluations/components/OfflineEvaluationsTab.tsxproducts/ai_observability/frontend/evaluations/evaluationsEntryLogic.tsproducts/ai_observability/frontend/evaluations/evaluationsEntryRedirect.test.tsproducts/ai_observability/frontend/evaluations/evaluationsEntryRedirect.tsproducts/ai_observability/frontend/evaluations/llmEvaluationsLogic.tsproducts/ai_observability/frontend/evaluations/offlineEvaluationsLogic.test.tsproducts/ai_observability/frontend/evaluations/offlineEvaluationsLogic.tsproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsScene.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemInspector.test.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsxproducts/ai_observability/frontend/offline-evaluations/OfflinePayload.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreChooser.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.tsxproducts/ai_observability/frontend/offline-evaluations/offlineDetailFixtures.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineItemInspectorLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineItemInspectorLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewFixtures.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewState.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewState.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineResultPresentation.tsproducts/ai_observability/frontend/offline-evaluations/offlineScoreTrends.fixtures.tsproducts/ai_observability/frontend/offline-evaluations/offlineScoreTrends.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineScoreTrends.tsproducts/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.tsproducts/ai_observability/frontend/scoreDefinitions/AIObservabilityScoreDefinitions.tsxproducts/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.stories.tsxproducts/ai_observability/frontend/scoreDefinitions/AIObservabilityScorersScene.tsxproducts/ai_observability/frontend/scoreDefinitions/ScoreDefinitionActions.tsxproducts/ai_observability/frontend/scoreDefinitions/ScoreDefinitionVersionButton.tsxproducts/ai_observability/frontend/scoreDefinitions/ScoreDefinitionVersionModal.tsxproducts/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.test.tsproducts/ai_observability/frontend/scoreDefinitions/aiObservabilityScoreDefinitionsLogic.tsproducts/ai_observability/frontend/scoreDefinitions/scoreDefinitionNavigation.tsproducts/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.test.tsproducts/ai_observability/frontend/scoreDefinitions/scoreDefinitionVersionLogic.tsproducts/ai_observability/frontend/traceReviews/AIObservabilityHumanReviews.tsxproducts/ai_observability/frontend/traceReviews/TraceReviewButton.tsxproducts/ai_observability/manifest.tsxproducts/ai_observability/mcp/tools.yamlservices/mcp/src/api/generated.ts
💤 Files with no reviewable changes (3)
- products/ai_observability/frontend/evaluations/components/OfflineEvaluationsTab.tsx
- products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.test.ts
- products/ai_observability/frontend/evaluations/offlineEvaluationsLogic.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.
|
…feat/aio-offline-evals-ui
There was a problem hiding this comment.
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.ts-496-498 (1)
496-498: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winCell batches go stale when summaries reload after item pages load.
refreshcallsloadOfflineSummariesandloadOfflineItemstogether.loadOfflineSummariesresetssummariesto[].addSummaryPagethen appends the new scorers and changes the scorer slices for each batch index. However,cellBatchesis reset only byloadOfflineItems.
addSummaryPagefirespumpCells. If items load first,pumpCellsrequests batch 0 with only the first summary page's scorers. The later summary pages can add scorers to batch 0's index range. Those scorers do not get requests, becausecellBatches[0]already exists. Their cells show "No result" when a result may exist.With 100-scorer summary pages this is uncommon. It still happens on the first load if items resolve before the summaries finish paginating. The first summary page already fills batches 0–4, so the gap appears when a later page adds scorers to a batch that was requested from a partial slice.
Fix: store the requested
versionIdsin each batch. InpumpCells, request the batch again when the current slice differs from the stored IDs. A simpler alternative is to wait forloadOfflineSummariesSuccessbefore pumping batches that were only partially covered.
🧹 Nitpick comments (2)
products/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.ts (2)
528-549: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueMerge the duplicate
setFiltersandreplaceFilterslisteners.Lines 528-538 and Lines 539-549 are the same. This breaks the "says everything once and only once" rule. Move the shared body into one helper.
♻️ Proposed refactor
- listeners(({ actions, values }) => ({ + listeners(({ actions, values }) => { + const reloadForFilters = (): void => { + if (values.filters.version) { + if (values.selectedVersion?.id !== values.filters.version) { + actions.loadOfflineHistoryVersion() + } + if (!values.windows.error) { + actions.loadOfflineHistoryPrimaryPage() + actions.loadOfflineHistoryComparisonPage() + } + } + } + return { ... - setFilters: () => { ... }, - replaceFilters: () => { ... }, + setFilters: reloadForFilters, + replaceFilters: reloadForFilters,As per path instructions: "says everything once and only once".
Source: Path instructions
556-558: 🎯 Functional Correctness | 🔵 Trivial | 💤 Low valueGuard
changePageagainst an invalid date window.The
setFiltersandrefreshHistorylisteners skip loading whenvalues.windows.erroris set. ThechangePagelistener does not. If a page change happens while the window is invalid,loadOfflineHistoryPrimaryPagesends a request withdate_fromanddate_toset toundefined. The backend then returns unbounded history. The UI hides the table in this state, so this is unlikely to happen today. Use the same guard for consistency.changePage: ({ period }) => { + if (!values.filters.version || values.windows.error) { + return + } period === 'primary' ? actions.loadOfflineHistoryPrimaryPage() : actions.loadOfflineHistoryComparisonPage() },
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: 79c9045b-6755-451f-a3f2-6c946305c4a6
⛔ Files ignored due to path filters (3)
pnpm-lock.yamlis excluded by!**/pnpm-lock.yamlproducts/ai_observability/frontend/generated/api.schemas.tsis excluded by!**/generated/**products/ai_observability/frontend/generated/api.tsis excluded by!**/generated/**
📒 Files selected for processing (30)
docs/internal/ai-offline-evaluation-reporting.mdfrontend/src/lib/constants.tsxfrontend/src/products.tsxfrontend/src/scenes/sceneLogic.test.tsxfrontend/src/scenes/sceneLogic.tsxproducts/ai_observability/frontend/evaluations/EvaluationsTabs.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentScene.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemInspector.test.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemInspector.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineItemMatrix.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsxproducts/ai_observability/frontend/offline-evaluations/OfflinePayload.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendLines.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryScene.tsxproducts/ai_observability/frontend/offline-evaluations/offlineExperimentLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineOverviewTrendLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.tsproducts/ai_observability/manifest.tsxproducts/ai_observability/package.jsonservices/mcp/src/api/generated.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.
There was a problem hiding this comment.
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/offlineScoreTrends.ts-76-76 (1)
76-76: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winCompare resolved date bounds by instant.
URL filters accept arbitrary date strings, so an encoded offset ISO value can reach the
dayjs(value)branch.toISOString()can emit expanded years, whose lexical order is not chronological. The current check can reject a valid range or accept reversed bounds.Suggested fix
- if (range.dateFrom && range.dateFrom >= range.dateTo) { + if (range.dateFrom && Date.parse(range.dateFrom) >= Date.parse(range.dateTo)) {
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: b4f796e0-ec23-46c4-a287-030251b4fd76
⛔ Files ignored due to path filters (3)
pnpm-lock.yamlis excluded by!**/pnpm-lock.yamlproducts/ai_observability/frontend/generated/api.schemas.tsis excluded by!**/generated/**products/ai_observability/frontend/generated/api.tsis excluded by!**/generated/**
📒 Files selected for processing (11)
docs/internal/ai-offline-evaluation-reporting.mdproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendChart.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScoreTrendCrosshair.tsxproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineExperimentsLogic.tsproducts/ai_observability/frontend/offline-evaluations/offlineScoreTrends.test.tsproducts/ai_observability/frontend/offline-evaluations/offlineScoreTrends.tsservices/mcp/src/api/generated.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.
There was a problem hiding this comment.
Actionable comments posted: 1
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (1)
products/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsx-153-153 (1)
153-153: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winRender dependent panels only after the experiment resolves.
If the initial experiment request fails, the error banner appears while scorer summaries and
OfflineItemMatrixstill render. Keep those panels behind the experiment loading and error states so the page does not show dependent content for an unavailable experiment. As per coding guidelines, “Branch in resolution order: loading → error → empty → content.”Source: Coding guidelines
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: d8197e3f-4169-469b-96c3-6977ef1ea21b
⛔ Files ignored due to path filters (3)
pnpm-lock.yamlis excluded by!**/pnpm-lock.yamlproducts/ai_observability/frontend/generated/api.schemas.tsis excluded by!**/generated/**products/ai_observability/frontend/generated/api.tsis excluded by!**/generated/**
📒 Files selected for processing (15)
docs/internal/ai-offline-evaluation-reporting.mdfrontend/src/lib/constants.tsxproducts/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentContent.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentStatus.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.stories.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineExperimentsOverview.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.test.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineOverviewTrend.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistory.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryRunDetails.tsxproducts/ai_observability/frontend/offline-evaluations/OfflineScorerHistoryTable.tsxproducts/ai_observability/frontend/offline-evaluations/offlineExperimentPresentation.tsproducts/ai_observability/frontend/offline-evaluations/offlineScorerHistoryLogic.tsservices/mcp/src/api/generated.ts
💤 Files with no reviewable changes (2)
- frontend/src/lib/constants.tsx
- products/ai_observability/frontend/evaluations/AIObservabilityEvaluationsScene.tsx
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.
🦔 Hogbox preview · ✅ ready▶ Open the preview
commit |
A search change in the offline overview reset the list clock. The trend date window therefore moved on each settled keystroke, and every selected score chart reloaded with the same data. A search-only filter change now keeps the current clock. Date, source, upload-state, refresh, and URL changes still reset it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
The Scorers scene export had no productKey, unlike the other AI observability scenes. Its pageviews therefore lost product_key, and Quick Start did not select AI observability on that page. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
Refresh clears the scorer summaries, which removes every score column. The browser then moves the table scroll back to 0. The matrix ignores scroll events while the table fits its container, so the stored viewport kept the old position. After a refresh far to the right, the first visible score batch therefore never loaded. Loading summaries now resets the stored scroll position to 0 and keeps the measured width. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
The router decodes URL text such as 00123 or true as a number or boolean. The offline overview accepted only string filters, so a reload, back navigation, or shared link dropped those searches without notice. The overview now writes the search as a one-element array, which the router keeps exactly. The reader accepts that array and still accepts a plain string search. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
With all time selected, each overview chart spans only its own results, so a chart hides the shared guide when the hovered time is outside that span. The reporting doc now says so and keeps the unconditional claim for bounded date ranges only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
LemonTabs blocks onChange for a disabled tab but still renders its link. A user without access to online evals, offline evals, or scorers could therefore click the greyed tab and land on the Access denied page. A tab with an access disabled reason now has no link. Access checks are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
The panel's open prop came from the loaded summary list. Refresh empties that list first, so for more than 10 scorers the prop went from false to true and back to false. React then wrote open = false and collapsed a panel the user had opened. The panel now reads the last loaded total, which stays set during a refresh. The prop stays the same, so the user's choice remains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
ApiError.formattedRetryAfter already reads as "in 30s" or "later" for a 429. The offline evaluation error message added its own "in", so users saw "Try again in in 30s" or "Try again in later". The message now uses the formatted value as returned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: 91388465-09c8-4846-8672-437ea66bed96
Problem
People who upload offline experiments need to inspect results and follow recurring scores in the Offline evals tab.
The read APIs landed in #106977. This adds their browser UI and gives manual-review score definitions a shared home.
Changes
Partial history coverage, using invented Storybook data
Before:
After:
Scorer navigation before and after
Before:
flowchart LR Reviews["Human reviews"] --> Definitions["Score definitions"] classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000; classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff; class Reviews phYellow; class Definitions phBlue;After:
flowchart LR Evaluations["Evaluations"] --> Scorers["Scorers"] Reviews["Human reviews / old links"] -->|Redirect| Scorers classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000; classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff; class Evaluations,Reviews phYellow; class Scorers phBlue;Original views before UI refinements, using invented Storybook data
These screenshots document the initial layouts; the former offline tab used the legacy harness.
Short comparison axes and all-time shared hover, using invented Storybook data
One-hour comparison before:
One-hour comparison after:
All-time overview before:
All-time overview after:
Period summaries, using invented Storybook data
Before:
After:
How did you test this code?
Period-summary regressions cover unequal experiment sizes, boolean and category rates, zero-success runs, and switching versions and date ranges.
Browser checks verify period scores and experiment counts at 1440px and 520px.
URL round-trip tests cover comparison periods ending now and explicit historical timestamps.
Shared-domain regressions cover sparse scorers, version changes, selection changes, refreshes, and stale responses.
Storybook checks cover short and multi-year axes, hidden legend series, and synchronized all-time hover at desktop and narrow widths in both themes.
The frontend TypeScript check, focused regression suites, formatter/linter, and CI preflight pass for these fixes.
Storybook checks verify existing controls at desktop/520px in light/dark themes. Partial-history fixtures verify the disclosure.
Browser checks cover history expansion, score selection/save, item inspection, payloads, and all scorer summaries.
The cleanup passes the full frontend TypeScript check, focused frontend suites, formatting/linting, and CI preflight.
A component regression verifies that incomplete history is disclosed and complete-history cards retain their existing presentation.
Backend regressions cover cell-query limits, experiment membership, scorer authorization, and serializers.
Frontend regressions cover stale responses, failed pagination, local preference isolation, redirects, version conflicts, and missing/zero scores.
Automated browser checks exercise score selection and order, 25 scorer columns, lazy payload reads, keyboard focus, and period comparison. Storybook renders cover narrow layouts, dark mode, and both flag states.
A real-view database roundtrip uploads numeric, boolean, and categorical results without a dataset, completes the experiment, and reads its history.
Feedback regressions cover automatic scorer selection, version navigation, and filter pagination. Browser checks cover collapsible payloads and compact tables at normal and narrow widths.
Hover regression stories verify timestamp alignment across sparse charts, switching the active chart, and clearing guides on mouse leave in light/dark and narrow layouts.
Browser checks confirm shared upload-state filtering, unchanged checkbox positions, onboarding at 520px, and readable light/dark tooltips.
Release status
Offline views, history links, and offline-specific copy use
ai-observability-offline-evaluations. Scorer relocation and manual-review management remain available without that flag.Automatic notifications
Docs update
Updated
docs/internal/ai-offline-evaluation-reporting.mdwith the UI behavior and bounded cell-read contract.🤖 Agent context
Autonomy: Human-driven (agent-assisted)
Agent: Codex,
gpt-6-astraTools included shell commands, Playwright, and collaborating agents. Fixtures and screenshots contain invented data. The local CodeRabbit review was skipped at the user's request.
Skills applied: UI components, frontend placement, Kea logic, generated API types, user-facing copy, charts, Storybook feature flags, test design, efficient validation, local development, CI preflight, PR descriptions, and GitHub operations. Feedback refinements add version navigation, compact tables, and collapsible payloads. The update keeps the existing screenshots at the author's request.
The cleanup following the code audit preserves existing UI behavior, adds partial-history disclosure, and verifies the rendered result against the previous PR head.