feat(aio): support OpenRouter decision models in evaluations - #110762
bernatixer wants to merge 4 commits into
Conversation
|
Merging to
After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here |
🤖 CI report
|
| Function | Location | Complexity | Limit |
|---|---|---|---|
AIObservabilityEvaluation |
products/ai_observability/frontend/evaluations/AIObservabilityEvaluation.tsx:84 |
119 | 10 |
setOutputType |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:821 |
21 | 10 |
saveEvaluation |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1273 |
21 | 10 |
<anonymous> |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1532 |
20 | 10 |
<anonymous> |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1594 |
17 | 10 |
<anonymous> |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1702 |
17 | 10 |
testHogOnSample |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:681 |
15 | 10 |
loadEvaluation |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1126 |
15 | 10 |
<anonymous> |
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1412 |
12 | 10 |
✅ Duplication (Python) — clean
New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
✅ Duplication (TypeScript) — clean
New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
⚠️ Bundle size — 🔺 +4.4 KiB (+0.0%)
Uncompressed size of every built .js bundle, compared against the base branch.
Total: 69.59 MiB · 🔺 +4.4 KiB (+0.0%)
| File | Size | Δ vs base |
|---|---|---|
render-query/src/render-query/render-query.js |
18.76 MiB | 🔺 +4.6 KiB (+0.0%) |
Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report
✅ Eager graph — within budget
How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.
| Root | Eager (shipped) | Δ vs base | Budget |
|---|---|---|---|
entry (logged-out pages, app bootstrap)src/index.tsx |
1.62 MiB · 22 files | no change | █████████░ 88.1% of 1.84 MiB |
logged-out boot: index + App + bootApp (preloaded by every page, including /login)src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts |
3.71 MiB · 660 files | no change | █████████░ 92.0% of 4.03 MiB |
authenticated shell (every logged-in page)src/scenes/AuthenticatedShell.tsx |
7.81 MiB · 2,507 files | no change | █████████░ 93.6% of 8.34 MiB |
dashboard scenesrc/scenes/dashboard/Dashboard.tsx |
9.65 MiB · 3,389 files | no change | ███████░░░ 71.6% of 13.48 MiB |
project home scenesrc/scenes/project-homepage/ProjectHomepage.tsx |
14.01 MiB · 5,020 files | no change | █████████░ 85.2% of 16.44 MiB |
events scenesrc/scenes/activity/explore/EventsScene.tsx |
9.27 MiB · 3,241 files | no change | ███████░░░ 73.3% of 12.64 MiB |
replay detail scenesrc/scenes/session-recordings/detail/SessionRecordingDetail.tsx |
12.08 MiB · 4,112 files | no change | ████████░░ 76.9% of 15.72 MiB |
🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
Largest files eagerly shipped from src/index.tsx
| Size | File |
|---|---|
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 24.6 KiB | ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js |
| 6.3 KiB | ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js |
| 4.5 KiB | ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js |
| 3.9 KiB | ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js |
| 1.4 KiB | ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js |
| 1.3 KiB | src/index.tsx |
| 1.3 KiB | src/RootErrorBoundary.tsx |
| 912 B | ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js |
| 854 B | src/scenes/ChunkLoadErrorBoundary.tsx |
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
| Size | File |
|---|---|
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 100.5 KiB | src/lib/api.ts |
| 92.0 KiB | src/products.tsx |
| 69.4 KiB | src/lib/lemon-ui/icons/icons.tsx |
| 40.1 KiB | src/lib/utils/eventUsageLogic.ts |
| 38.7 KiB | ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js |
| 33.9 KiB | ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js |
| 29.0 KiB | ../node_modules/.pnpm/zod@4.3.6/node_modules/zod/v4/core/schemas.js |
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
| Size | File |
|---|---|
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 279.9 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 110.0 KiB | ../packages/quill/packages/quill/dist/index.js |
| 100.5 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 92.0 KiB | src/products.tsx |
| 90.6 KiB | ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js |
Largest files eagerly shipped from src/scenes/dashboard/Dashboard.tsx
| Size | File |
|---|---|
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 279.9 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.8 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 110.0 KiB | ../packages/quill/packages/quill/dist/index.js |
| 100.5 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 92.0 KiB | src/products.tsx |
Largest files eagerly shipped from src/scenes/project-homepage/ProjectHomepage.tsx
| Size | File |
|---|---|
| 315.5 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js |
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 279.9 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.8 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 110.0 KiB | ../packages/quill/packages/quill/dist/index.js |
| 100.5 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
Largest files eagerly shipped from src/scenes/activity/explore/EventsScene.tsx
| Size | File |
|---|---|
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 279.9 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.8 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 110.0 KiB | ../packages/quill/packages/quill/dist/index.js |
| 100.5 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 92.0 KiB | src/products.tsx |
Largest files eagerly shipped from src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
| Size | File |
|---|---|
| 315.5 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js |
| 306.2 KiB | ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 279.9 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 219.9 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.8 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 110.0 KiB | ../packages/quill/packages/quill/dist/index.js |
| 100.5 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479
✅ Toolbar bundle — eager 2.20 MiB within budget
What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.
| Metric | Size | Δ vs base | Budget |
|---|---|---|---|
| Eager (shipped) entry + static imports |
2.20 MiB · 19 files | no change | ████░░░░░░ 38.4% of 5.72 MiB |
| Deferred (lazy) | 2.11 MiB · 44 files | no change | n/a — loads on demand |
Loader dist/toolbar.js |
1.2 KiB | no change | █░░░░░░░░░ 6.0% of 19.5 KiB |
Largest eagerly-shipped chunks
| Size | File |
|---|---|
| 833.8 KiB | dist/toolbar/toolbar-app-W2MCORNU.css |
| 657.3 KiB | dist/toolbar/chunk-chunk-G546HVXW.js |
| 259.4 KiB | dist/toolbar/chunk-chunk-DWA3PXCS.js |
| 138.2 KiB | dist/toolbar/chunk-chunk-6KOTDG7L.js |
| 131.8 KiB | dist/toolbar/chunk-chunk-FDH2IBXT.js |
| 75.2 KiB | dist/toolbar/toolbar-app-W3DICL63.js |
| 69.0 KiB | dist/toolbar/chunk-chunk-TSAL54PB.js |
| 35.6 KiB | dist/toolbar/chunk-chunk-RYZCVF6D.js |
| 21.0 KiB | dist/toolbar/chunk-chunk-53XUO32E.js |
| 6.8 KiB | dist/toolbar/chunk-chunk-DV7IWQNF.js |
Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile
✅ Dist folder size — 🔺 +70.9 KiB (+0.0%)
Total size of the built frontend/dist folder (all assets), compared against the base branch.
Total: 958.14 MiB · 🔺 +70.9 KiB (+0.0%)
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 📝 WalkthroughWalkthroughOpenRouter decision models can now use System One evaluations through an existing OpenRouter key. The backend identifies models from catalogue output modalities, exposes their capability through the models API, and applies System One validation and evaluation routing. The frontend uses that capability for numeric-bound validation, evaluation guidance, and model-picker filtering. Tests and internal instructions cover the updated behavior. Priority: ➖ Normal Merge Risk: 🟡 Moderate · up to Saved decision evaluations can become disabled if the feature flag is turned off. Fix that routing before rollout; malformed catalogue metadata can also affect which decision models are available. Security Architecture ReviewSecurity architecture risk: 🟡 Moderate · up to The new evaluation route retains project-scoped credentials and destination protections. However, turning off access can send saved decision evaluations through an incompatible fallback and leave them disabled, weakening rollback containment. Retained concerns
Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 1✅ Passed checks (1 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (2)
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts-1508-1524 (1)
1508-1524: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winTreat unresolved BYOK models as unknown.
byokModelsstarts as an empty array while its loader runs. During that state, an OpenRouter numeric evaluation with a System One-capable model does not satisfyusesSystemOne.numericBoundsRequiredis then false, so the frontend can treat the form as valid without bounds and hide the bounds guidance. IncludebyokModelsLoadingand require bounds for OpenRouter evaluations until the model list resolves.Suggested fix
- ['byokModels'], + ['byokModels', 'byokModelsLoading'], ... usesSystemOne: [ - (s) => [s.evaluation, s.byokModels], - (evaluation: EvaluationConfig | null, models: ModelOption[]): boolean => { + (s) => [s.evaluation, s.byokModels, s.byokModelsLoading], + ( + evaluation: EvaluationConfig | null, + models: ModelOption[], + byokModelsLoading: boolean + ): boolean => { const config = evaluation?.model_configuration return ( config?.provider === 'system_one' || (config?.provider === 'openrouter' && - models.some( - (model) => - model.id === config.model && - toLLMProvider(model.provider) === 'openrouter' && - (!config.provider_key_id || model.providerKeyId === config.provider_key_id) && - model.supportsSystemOne - )) + (byokModelsLoading || + models.some( + (model) => + model.id === config.model && + toLLMProvider(model.provider) === 'openrouter' && + (!config.provider_key_id || model.providerKeyId === config.provider_key_id) && + model.supportsSystemOne + ))) ) }, ],products/ai_observability/backend/llm/providers/openrouter.py-59-63 (1)
59-63: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winReject non-list modality values before caching them.
A truthy string such as
"decisions"passes the filter because"text" not in outputis true.decision_model_ids()then treats the string as containing the"decisions"modality and returns the model ID. The string does not make the dict comprehension raise;except Exceptiononly handles actual exceptions.Proposed fix
- modalities = { - model["id"]: model["architecture"]["output_modalities"] - for model in models - if "text" not in ((model.get("architecture") or {}).get("output_modalities") or ["text"]) - } + modalities: dict[str, list[str]] = {} + for model in models: + output = (model.get("architecture") or {}).get("output_modalities") or ["text"] + if isinstance(output, list) and "text" not in output: + modalities[model["id"]] = output
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: 6a0c06ba-c717-443f-b280-3964833ab40a
📒 Files selected for processing (20)
docs/internal/ai-observability-judge-inputs.mdposthog/temporal/ai_observability/evaluation_llm_judge.pyposthog/temporal/ai_observability/test_run_evaluation.pyproducts/ai_observability/backend/api/evaluations.pyproducts/ai_observability/backend/api/proxy.pyproducts/ai_observability/backend/api/test/test_evaluations.pyproducts/ai_observability/backend/api/test/test_proxy.pyproducts/ai_observability/backend/llm/__init__.pyproducts/ai_observability/backend/llm/providers/openrouter.pyproducts/ai_observability/backend/llm/providers/test/test_openrouter.pyproducts/ai_observability/backend/llm/system_one.pyproducts/ai_observability/frontend/evaluations/AIObservabilityEvaluation.tsxproducts/ai_observability/frontend/evaluations/components/EvaluationCodeEditor.test.tsxproducts/ai_observability/frontend/evaluations/components/EvaluationRunsTable.test.tsxproducts/ai_observability/frontend/evaluations/evaluationBackfillsLogic.test.tsproducts/ai_observability/frontend/evaluations/llmEvaluationLogic.test.tsproducts/ai_observability/frontend/evaluations/llmEvaluationLogic.tsproducts/ai_observability/frontend/evaluations/llmEvaluationsLogic.test.tsproducts/ai_observability/frontend/modelPickerLogic.test.tsproducts/ai_observability/frontend/modelPickerLogic.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.
|
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: aa451d25-55d7-43aa-a0a1-c45bd27e7abc
📒 Files selected for processing (8)
docs/internal/ai-observability-judge-inputs.mdposthog/temporal/ai_observability/evaluation_llm_judge.pyposthog/temporal/ai_observability/test_run_evaluation.pyproducts/ai_observability/backend/api/evaluations.pyproducts/ai_observability/backend/api/test/test_evaluations.pyproducts/ai_observability/backend/llm/system_one.pyproducts/ai_observability/backend/llm/test/test_system_one.pyproducts/ai_observability/frontend/playground/llmPlaygroundLogic.test.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.
| openrouter_enabled=provider == "openrouter" | ||
| and system_one_evaluations_enabled(team_id, base_url=OPENROUTER_BASE_URL), | ||
| ): |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🔎 Supported by static analysis
🏁 Script executed:
sed -n '530,665p' posthog/temporal/ai_observability/evaluation_llm_judge.py
sed -n '300,371p' posthog/temporal/ai_observability/test_run_evaluation.py
sed -n '50,64p' docs/internal/ai-observability-judge-inputs.mdRepository: PostHog/posthog
Length of output: 11063
🏁 Script executed:
#!/bin/bash
set -e
printf '%s\n' '--- changed hunk and surrounding control flow ---'
git diff --unified=35 b467e3fc4b136b65b0cc2c0a6ce3fc76bd80d1fb e137d2b587d7aa423bc91b6ba68c6a0c500acbc5 -- posthog/temporal/ai_observability/evaluation_llm_judge.py
printf '%s\n' '--- symbol definitions and references ---'
rg -n -S 'def is_system_one_model|is_system_one_model\(|system_one_evaluations_enabled|model_not_supported|UnsupportedModelError|disable.*evaluation|evaluation.*disable|skip.*run|saved evaluation|turning the flag off' posthog products docs --glob '*.py' --glob '*.md' | head -240
printf '%s\n' '--- exact judge implementation after the System One branch ---'
sed -n '640,790p' posthog/temporal/ai_observability/evaluation_llm_judge.py
printf '%s\n' '--- relevant test sections ---'
sed -n '240,390p' posthog/temporal/ai_observability/test_run_evaluation.py
sed -n '390,540p' posthog/temporal/ai_observability/test_run_evaluation.pyRepository: PostHog/posthog
Length of output: 42004
🏁 Script executed:
#!/bin/bash
set -e
rg -n -S 'def is_system_one_model|is_system_one_model\(|UnsupportedModelError|model_not_supported|disable.*evaluation|evaluation.*disable|system_one_evaluations_enabled' posthog products docs --glob '*.py' --glob '*.md' | head -220Repository: PostHog/posthog
Length of output: 32110
🏁 Script executed:
#!/bin/bash
set -e
printf '%s\n' '--- System One model identification ---'
sed -n '1,75p' products/ai_observability/backend/llm/system_one.py
printf '%s\n' '--- OpenRouter request and rejection mapping ---'
sed -n '1,135p' products/ai_observability/backend/llm/providers/openrouter.py
printf '%s\n' '--- evaluation error contract ---'
sed -n '1,135p' posthog/temporal/ai_observability/evaluation_errors.py
printf '%s\n' '--- judge UnsupportedModelError handling ---'
sed -n '825,860p' posthog/temporal/ai_observability/evaluation_llm_judge.py
printf '%s\n' '--- evaluation workflow terminal-error lifecycle ---'
sed -n '100,245p' posthog/temporal/ai_observability/run_evaluation.py
printf '%s\n' '--- API Jev/model-not-supported test ---'
sed -n '2145,2210p' products/ai_observability/backend/api/test/test_evaluations.pyRepository: PostHog/posthog
Length of output: 24132
Skip saved OpenRouter decision evaluations when the flag is disabled.
When the flag is off, is_system_one_model deliberately returns false, so call_llm_judge sends the saved model to Client.complete. If the OpenRouter catalogue identifies that model as a decision-only model, the request raises UnsupportedModelError. The workflow maps this to model_not_supported and disables the evaluation.
The disabled-flag test covers openai/gpt-4o, not a saved decision model. Keep chat fallback for ordinary OpenRouter chat models, but preserve the decision-model classification for saved evaluations and skip those runs when the flag is off. Do not guess the chat path when the catalogue is unavailable.
Problem
OpenRouter users must create a separate System One connection to use Jev as an evaluation judge. Selecting a decision model through an OpenRouter key is currently rejected.
Refs #106871. Extends the System One support from #103752, #108615 and #109215.
Changes
llm-analytics-system-one-evaluations.Catalogue lookups for decision routing run only for projects with the flag. Unflagged projects keep the existing chat path. Flagged projects require the catalogue for runs, model/output changes and enabling evaluations; renaming or disabling remains available during outages.
Out-of-credit responses disable the evaluation and mark its key as failing, matching the chat path.
Request path before and after
Before:
flowchart LR A[OpenRouter Jev selection] --> B[Rejected as non-chat] classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000; classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000; class A phYellow; class B phGray;After:
flowchart LR A[OpenRouter Jev selection] --> B[Existing System One client] --> C[OpenRouter System One API] classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000; classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff; classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff; class A phYellow; class B phBlue; class C phRed;How did you test this code?
Test rationale: Extended existing System One result cases to cover native OpenRouter routing and key attribution. API and picker cases cover rollout, numeric bounds and generative exclusions. Catalogue outage cases cover unflagged chat runs, renaming/disabling, and configuration validation. HTTP 402 cases cover terminal quota errors and key validation. Changed tests use mocked HTTP/event boundaries; no new ClickHouse query path.
Release status
Automatic notifications
Docs update
Updated
docs/internal/ai-observability-judge-inputs.md.🤖 Agent context
Autonomy: Human-driven (agent-assisted)
Agent: Codex, GPT-6
Tools: Git/GitHub CLI, Flox, Playwright and web lookup. Repo skills: routing-outbound-api-calls, improving-drf-endpoints, adopting-generated-api-types, writing-kea-logics, writing-tests, running-ci-preflight, reviewing-with-coderabbit and writing-pr-descriptions. Session link unavailable.
CodeRabbit CLI ran with
--deep: fixed null catalogue modalities and validation during catalogue outages. Rejected the endpoint finding: OpenRouter explicitly documents/api/v1/systemone; the unauthenticated probe returns 401. No change to/api/alpha/decisionsis needed.Duplicate search found #106131 for decision-model cost syncing, which remains separate. Fixtures are invented; no private session material is included. CI patch coverage remains to be checked.
PR review follow-up: addressed CodeRabbit’s unrelated-save catalogue lookup and the rollout/402 findings. Updated the playground test expectation for model capability metadata.