Skip to content

feat(aio): support OpenRouter decision models in evaluations - #110762

Draft
bernatixer wants to merge 4 commits into
masterfrom
feat/aio-openrouter-decision-judges
Draft

bernatixer wants to merge 4 commits into
masterfrom
feat/aio-openrouter-decision-judges

Conversation

@bernatixer

@bernatixer bernatixer commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Problem

OpenRouter users must create a separate System One connection to use Jev as an evaluation judge. Selecting a decision model through an OpenRouter key is currently rejected.

Refs #106871. Extends the System One support from #103752, #108615 and #109215.

Changes

  • Jev and other decision models appear under an existing OpenRouter key in the evaluation picker, behind llm-analytics-system-one-evaluations.
  • The existing System One client sends these evaluations to OpenRouter's documented endpoint. Boolean, categorical, numeric and N/A handling stay shared.
  • OpenRouter's catalogue identifies decision models, including aliases. Chat models such as Jev Router keep their chat path.
  • Decision models stay out of generative pickers and require numeric bounds. The UI explains their lack of written reasoning.
  • The probability cutoff stays at 0.5. Multiple category and applicability questions can share one request; separate evaluations still run separately.

Catalogue lookups for decision routing run only for projects with the flag. Unflagged projects keep the existing chat path. Flagged projects require the catalogue for runs, model/output changes and enabling evaluations; renaming or disabling remains available during outages.

Out-of-credit responses disable the evaluation and mark its key as failing, matching the chat path.

Before After
Before After
Request path before and after

Before:

flowchart LR
    A[OpenRouter Jev selection] --> B[Rejected as non-chat]
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
    class A phYellow;
    class B phGray;
Loading

After:

flowchart LR
    A[OpenRouter Jev selection] --> B[Existing System One client] --> C[OpenRouter System One API]
    classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
    classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
    classDef phRed fill:#f54e00,stroke:#f54e00,color:#fff;
    class A phYellow;
    class B phBlue;
    class C phRed;
Loading

How did you test this code?

  • Local provider, API and judge tests; evaluation and model-picker Jest suites; TypeScript, kea typegen, repo-wide mypy, formatting and preflight.
  • Playwright checked evaluation versus playground choices and a narrow picker in Storybook using invented fixtures.
  • Live catalogue metadata and endpoint authentication response checked. Authenticated OpenRouter inference has not been tested.

Test rationale: Extended existing System One result cases to cover native OpenRouter routing and key attribution. API and picker cases cover rollout, numeric bounds and generative exclusions. Catalogue outage cases cover unflagged chat runs, renaming/disabling, and configuration validation. HTTP 402 cases cover terminal quota errors and key validation. Changed tests use mocked HTTP/event boundaries; no new ClickHouse query path.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

Updated docs/internal/ai-observability-judge-inputs.md.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Codex, GPT-6

Tools: Git/GitHub CLI, Flox, Playwright and web lookup. Repo skills: routing-outbound-api-calls, improving-drf-endpoints, adopting-generated-api-types, writing-kea-logics, writing-tests, running-ci-preflight, reviewing-with-coderabbit and writing-pr-descriptions. Session link unavailable.

CodeRabbit CLI ran with --deep: fixed null catalogue modalities and validation during catalogue outages. Rejected the endpoint finding: OpenRouter explicitly documents /api/v1/systemone; the unauthenticated probe returns 401. No change to /api/alpha/decisions is needed.

Duplicate search found #106131 for decision-model cost syncing, which remains separate. Fixtures are invented; no private session material is included. CI patch coverage remains to be checked.

PR review follow-up: addressed CodeRabbit’s unrelated-save catalogue lookup and the rollout/402 findings. Updated the playground test expectation for model capability metadata.

@bernatixer bernatixer self-assigned this Oct 2, 2026
@trunk-io

trunk-io Bot commented Oct 2, 2026

Copy link
Copy Markdown

Merging to master in this repository is managed by Trunk.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 9 functions above the limit (max 119)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
AIObservabilityEvaluation products/ai_observability/frontend/evaluations/AIObservabilityEvaluation.tsx:84 119 10
setOutputType products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:821 21 10
saveEvaluation products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1273 21 10
<anonymous> products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1532 20 10
<anonymous> products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1594 17 10
<anonymous> products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1702 17 10
testHogOnSample products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:681 15 10
loadEvaluation products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1126 15 10
<anonymous> products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts:1412 12 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Bundle size — 🔺 +4.4 KiB (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 69.59 MiB · 🔺 +4.4 KiB (+0.0%)

File Size Δ vs base
render-query/src/render-query/render-query.js 18.76 MiB 🔺 +4.6 KiB (+0.0%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.62 MiB · 22 files no change █████████░ 88.1% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.71 MiB · 660 files no change █████████░ 92.0% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.81 MiB · 2,507 files no change █████████░ 93.6% of 8.34 MiB
dashboard scene
src/scenes/dashboard/Dashboard.tsx
9.65 MiB · 3,389 files no change ███████░░░ 71.6% of 13.48 MiB
project home scene
src/scenes/project-homepage/ProjectHomepage.tsx
14.01 MiB · 5,020 files no change █████████░ 85.2% of 16.44 MiB
events scene
src/scenes/activity/explore/EventsScene.tsx
9.27 MiB · 3,241 files no change ███████░░░ 73.3% of 12.64 MiB
replay detail scene
src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
12.08 MiB · 4,112 files no change ████████░░ 76.9% of 15.72 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 [object Object] stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx
🟢 [object Object] stays out of src/scenes/AuthenticatedShell.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
854 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
100.5 KiB src/lib/api.ts
92.0 KiB src/products.tsx
69.4 KiB src/lib/lemon-ui/icons/icons.tsx
40.1 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
29.0 KiB ../node_modules/.pnpm/zod@4.3.6/node_modules/zod/v4/core/schemas.js
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
110.0 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
92.0 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
Largest files eagerly shipped from src/scenes/dashboard/Dashboard.tsx
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.8 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
110.0 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
92.0 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/project-homepage/ProjectHomepage.tsx
Size File
315.5 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.8 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
110.0 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
Largest files eagerly shipped from src/scenes/activity/explore/EventsScene.tsx
Size File
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.8 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
110.0 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
92.0 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
Size File
315.5 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js
306.2 KiB ../node_modules/.pnpm/posthog-js@1.435.5_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
279.9 KiB src/taxonomy/core-filter-definitions-by-group.json
219.9 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.8 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
110.0 KiB ../packages/quill/packages/quill/dist/index.js
100.5 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.20 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.20 MiB · 19 files no change ████░░░░░░ 38.4% of 5.72 MiB
Deferred (lazy) 2.11 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
833.8 KiB dist/toolbar/toolbar-app-W2MCORNU.css
657.3 KiB dist/toolbar/chunk-chunk-G546HVXW.js
259.4 KiB dist/toolbar/chunk-chunk-DWA3PXCS.js
138.2 KiB dist/toolbar/chunk-chunk-6KOTDG7L.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-W3DICL63.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-RYZCVF6D.js
21.0 KiB dist/toolbar/chunk-chunk-53XUO32E.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +70.9 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 958.14 MiB · 🔺 +70.9 KiB (+0.0%)

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

📝 Walkthrough

Walkthrough

OpenRouter decision models can now use System One evaluations through an existing OpenRouter key. The backend identifies models from catalogue output modalities, exposes their capability through the models API, and applies System One validation and evaluation routing. The frontend uses that capability for numeric-bound validation, evaluation guidance, and model-picker filtering. Tests and internal instructions cover the updated behavior.

Priority: ➖ Normal

Merge Risk: 🟡 Moderate · up to e137d

Saved decision evaluations can become disabled if the feature flag is turned off. Fix that routing before rollout; malformed catalogue metadata can also affect which decision models are available.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to e137d

The new evaluation route retains project-scoped credentials and destination protections. However, turning off access can send saved decision evaluations through an incompatible fallback and leave them disabled, weakening rollback containment.

Retained concerns

  • Medium · reliability · observed: Rolling back the feature flag does not safely suspend saved OpenRouter decision evaluations. Classification instead selects chat completion, which attempts an outbound request before identifying unsupported non-chat models. A resulting model_not_supported terminal error disables the saved evaluation, contrary to the documented state-preserving rollback behavior. This newly reachable path weakens rollout containment and requires explicit recovery rather than merely restoring feature availability.
Security review details

Security Blast Radius

  • inferred — The rollback defect can affect saved decision evaluations across projects where the feature was enabled and later disabled. Runs reaching the live routing check after rollback can still send evaluation input to the same external provider through chat completion and can disable their associated evaluation. The inspected credential resolution does not expand this path to another project's keys.

Trust Boundaries and Controls

  • observed — User-selected provider-key IDs are resolved against the current project both during configuration persistence and execution. Unpinned configurations use that project's active key only when its provider matches; model persistence also validates the provider/key pairing.

Resilience and Maintainability Implications

  • observed — With the feature enabled, catalogue outages become retryable execution failures rather than guessed routing. Flag-off coverage exercises an ordinary chat model, not a saved decision model; it therefore does not validate rollback containment for the new route.

Hardening Proposals

  • proposed — Separate decision-model transport identity from live feature authorization so disabling access suspends saved decision evaluations without chat fallback or terminal disablement. Validate that invariant across queued execution, retries, concurrent runs and explicit recovery.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description covers the required Problem, Changes, testing, test rationale, release status, documentation, screenshots, and agent context sections. It clearly explains the user impact, feature-flag…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (2)
products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts-1508-1524 (1)

1508-1524: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Treat unresolved BYOK models as unknown.

byokModels starts as an empty array while its loader runs. During that state, an OpenRouter numeric evaluation with a System One-capable model does not satisfy usesSystemOne. numericBoundsRequired is then false, so the frontend can treat the form as valid without bounds and hide the bounds guidance. Include byokModelsLoading and require bounds for OpenRouter evaluations until the model list resolves.

Suggested fix
-            ['byokModels'],
+            ['byokModels', 'byokModelsLoading'],
...
         usesSystemOne: [
-            (s) => [s.evaluation, s.byokModels],
-            (evaluation: EvaluationConfig | null, models: ModelOption[]): boolean => {
+            (s) => [s.evaluation, s.byokModels, s.byokModelsLoading],
+            (
+                evaluation: EvaluationConfig | null,
+                models: ModelOption[],
+                byokModelsLoading: boolean
+            ): boolean => {
                 const config = evaluation?.model_configuration
                 return (
                     config?.provider === 'system_one' ||
                     (config?.provider === 'openrouter' &&
-                        models.some(
-                            (model) =>
-                                model.id === config.model &&
-                                toLLMProvider(model.provider) === 'openrouter' &&
-                                (!config.provider_key_id || model.providerKeyId === config.provider_key_id) &&
-                                model.supportsSystemOne
-                        ))
+                        (byokModelsLoading ||
+                            models.some(
+                                (model) =>
+                                    model.id === config.model &&
+                                    toLLMProvider(model.provider) === 'openrouter' &&
+                                    (!config.provider_key_id || model.providerKeyId === config.provider_key_id) &&
+                                    model.supportsSystemOne
+                            )))
                 )
             },
         ],
products/ai_observability/backend/llm/providers/openrouter.py-59-63 (1)

59-63: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Reject non-list modality values before caching them.

A truthy string such as "decisions" passes the filter because "text" not in output is true. decision_model_ids() then treats the string as containing the "decisions" modality and returns the model ID. The string does not make the dict comprehension raise; except Exception only handles actual exceptions.

Proposed fix
-        modalities = {
-            model["id"]: model["architecture"]["output_modalities"]
-            for model in models
-            if "text" not in ((model.get("architecture") or {}).get("output_modalities") or ["text"])
-        }
+        modalities: dict[str, list[str]] = {}
+        for model in models:
+            output = (model.get("architecture") or {}).get("output_modalities") or ["text"]
+            if isinstance(output, list) and "text" not in output:
+                modalities[model["id"]] = output

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 6a0c06ba-c717-443f-b280-3964833ab40a

📥 Commits

Reviewing files that changed from the base of the PR and between 3c747e5 and 4c1564b.

📒 Files selected for processing (20)
  • docs/internal/ai-observability-judge-inputs.md
  • posthog/temporal/ai_observability/evaluation_llm_judge.py
  • posthog/temporal/ai_observability/test_run_evaluation.py
  • products/ai_observability/backend/api/evaluations.py
  • products/ai_observability/backend/api/proxy.py
  • products/ai_observability/backend/api/test/test_evaluations.py
  • products/ai_observability/backend/api/test/test_proxy.py
  • products/ai_observability/backend/llm/__init__.py
  • products/ai_observability/backend/llm/providers/openrouter.py
  • products/ai_observability/backend/llm/providers/test/test_openrouter.py
  • products/ai_observability/backend/llm/system_one.py
  • products/ai_observability/frontend/evaluations/AIObservabilityEvaluation.tsx
  • products/ai_observability/frontend/evaluations/components/EvaluationCodeEditor.test.tsx
  • products/ai_observability/frontend/evaluations/components/EvaluationRunsTable.test.tsx
  • products/ai_observability/frontend/evaluations/evaluationBackfillsLogic.test.ts
  • products/ai_observability/frontend/evaluations/llmEvaluationLogic.test.ts
  • products/ai_observability/frontend/evaluations/llmEvaluationLogic.ts
  • products/ai_observability/frontend/evaluations/llmEvaluationsLogic.test.ts
  • products/ai_observability/frontend/modelPickerLogic.test.ts
  • products/ai_observability/frontend/modelPickerLogic.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.

Comment thread products/ai_observability/backend/api/evaluations.py
@trunk-io

trunk-io Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
compareTopLevelSections() reports a modifiers change when the current query overrides the team default A TypeError occurred because the code attempted to access the 'add' property of an undefined object. Logs ↗︎

View Full Report ↗︎ ⋅ Docs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: aa451d25-55d7-43aa-a0a1-c45bd27e7abc

📥 Commits

Reviewing files that changed from the base of the PR and between 4c1564b and e137d2b.

📒 Files selected for processing (8)
  • docs/internal/ai-observability-judge-inputs.md
  • posthog/temporal/ai_observability/evaluation_llm_judge.py
  • posthog/temporal/ai_observability/test_run_evaluation.py
  • products/ai_observability/backend/api/evaluations.py
  • products/ai_observability/backend/api/test/test_evaluations.py
  • products/ai_observability/backend/llm/system_one.py
  • products/ai_observability/backend/llm/test/test_system_one.py
  • products/ai_observability/frontend/playground/llmPlaygroundLogic.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.

Comment on lines +575 to +577
openrouter_enabled=provider == "openrouter"
and system_one_evaluations_enabled(team_id, base_url=OPENROUTER_BASE_URL),
):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🔎 Supported by static analysis

🏁 Script executed:

sed -n '530,665p' posthog/temporal/ai_observability/evaluation_llm_judge.py
sed -n '300,371p' posthog/temporal/ai_observability/test_run_evaluation.py
sed -n '50,64p' docs/internal/ai-observability-judge-inputs.md

Repository: PostHog/posthog

Length of output: 11063


🏁 Script executed:

#!/bin/bash
set -e
printf '%s\n' '--- changed hunk and surrounding control flow ---'
git diff --unified=35 b467e3fc4b136b65b0cc2c0a6ce3fc76bd80d1fb e137d2b587d7aa423bc91b6ba68c6a0c500acbc5 -- posthog/temporal/ai_observability/evaluation_llm_judge.py

printf '%s\n' '--- symbol definitions and references ---'
rg -n -S 'def is_system_one_model|is_system_one_model\(|system_one_evaluations_enabled|model_not_supported|UnsupportedModelError|disable.*evaluation|evaluation.*disable|skip.*run|saved evaluation|turning the flag off' posthog products docs --glob '*.py' --glob '*.md' | head -240

printf '%s\n' '--- exact judge implementation after the System One branch ---'
sed -n '640,790p' posthog/temporal/ai_observability/evaluation_llm_judge.py

printf '%s\n' '--- relevant test sections ---'
sed -n '240,390p' posthog/temporal/ai_observability/test_run_evaluation.py
sed -n '390,540p' posthog/temporal/ai_observability/test_run_evaluation.py

Repository: PostHog/posthog

Length of output: 42004


🏁 Script executed:

#!/bin/bash
set -e
rg -n -S 'def is_system_one_model|is_system_one_model\(|UnsupportedModelError|model_not_supported|disable.*evaluation|evaluation.*disable|system_one_evaluations_enabled' posthog products docs --glob '*.py' --glob '*.md' | head -220

Repository: PostHog/posthog

Length of output: 32110


🏁 Script executed:

#!/bin/bash
set -e
printf '%s\n' '--- System One model identification ---'
sed -n '1,75p' products/ai_observability/backend/llm/system_one.py

printf '%s\n' '--- OpenRouter request and rejection mapping ---'
sed -n '1,135p' products/ai_observability/backend/llm/providers/openrouter.py

printf '%s\n' '--- evaluation error contract ---'
sed -n '1,135p' posthog/temporal/ai_observability/evaluation_errors.py

printf '%s\n' '--- judge UnsupportedModelError handling ---'
sed -n '825,860p' posthog/temporal/ai_observability/evaluation_llm_judge.py

printf '%s\n' '--- evaluation workflow terminal-error lifecycle ---'
sed -n '100,245p' posthog/temporal/ai_observability/run_evaluation.py

printf '%s\n' '--- API Jev/model-not-supported test ---'
sed -n '2145,2210p' products/ai_observability/backend/api/test/test_evaluations.py

Repository: PostHog/posthog

Length of output: 24132


Skip saved OpenRouter decision evaluations when the flag is disabled.

When the flag is off, is_system_one_model deliberately returns false, so call_llm_judge sends the saved model to Client.complete. If the OpenRouter catalogue identifies that model as a decision-only model, the request raises UnsupportedModelError. The workflow maps this to model_not_supported and disables the evaluation.

The disabled-flag test covers openai/gpt-4o, not a saved decision model. Keep chat fallback for ordinary OpenRouter chat models, but preserve the decision-model classification for saved evaluations and skip those runs when the flag is off. Do not guess the chat path when the catalogue is unavailable.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant