Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
40 commits
Select commit Hold shift + click to select a range
cff9d1f
feat(aio): add Jev boolean evaluations with TypeSafe keys
bernatixer Sep 21, 2026
f6b6a0d
chore(aio): merge master into Jev evaluations
bernatixer Sep 21, 2026
98cf112
test(mcp): update unit test snapshots
tests-posthog[bot] Sep 21, 2026
6703c4d
feat(aio): support compatible System One judge endpoints
bernatixer Sep 25, 2026
ba7d354
chore(aio): merge master and preserve judge output contracts
bernatixer Sep 25, 2026
6f46380
fix(aio): distinguish rejected System One requests
bernatixer Sep 25, 2026
25623e9
test(mcp): update unit test snapshots
tests-posthog[bot] Sep 25, 2026
46d3045
feat(aio): support typed system one questions and numeric judges
bernatixer Sep 25, 2026
6a1a167
chore(aio): merge master and refresh generated contracts
bernatixer Sep 25, 2026
102cdcd
chore(aio): use fixed timestamps in system one judge tests
bernatixer Sep 25, 2026
02aeaff
chore(aio): preserve remote mcp schema snapshots
bernatixer Sep 25, 2026
b773d3b
chore(aio): regenerate numeric rubric mcp snapshots
bernatixer Sep 25, 2026
5457ae7
fix(aio): keep numeric mapping separate from system one client
bernatixer Sep 25, 2026
81ea2bb
chore(aio): sync master before system one validation
bernatixer Sep 25, 2026
d52c049
fix(aio): reuse typesafe egress for boolean evaluations
bernatixer Sep 25, 2026
7cbc470
chore(aio): sync master before egress validation
bernatixer Sep 25, 2026
2bcc0f8
fix(aio): gate system one evaluation experiments
bernatixer Sep 25, 2026
322035a
fix(aio): separate customer system one connections
bernatixer Sep 25, 2026
47b929d
chore(aio): sync master for connection access checks
bernatixer Sep 25, 2026
f452de6
fix(aio): leave custom system one costs unknown
bernatixer Sep 25, 2026
6ce1abc
chore(aio): sync master for pricing validation
bernatixer Sep 25, 2026
0a6ea3e
fix(aio): defer system one pricing without extra metadata
bernatixer Sep 25, 2026
91aa544
chore(aio): sync master before simplified pricing checks
bernatixer Sep 25, 2026
693576e
fix(aio): use existing model pricing for system one
bernatixer Sep 25, 2026
0099906
refactor(aio): reuse shared system one client
bernatixer Sep 25, 2026
8ccdf2a
feat(aio): support generic system one evaluation endpoints
bernatixer Sep 25, 2026
d22259e
test(mcp): update unit test snapshots
tests-posthog[bot] Sep 25, 2026
e10bc10
chore(aio): clarify compatible endpoint guidance
bernatixer Sep 25, 2026
bc8e42e
fix(aio): clarify system one base url field
bernatixer Sep 26, 2026
83cc3bc
fix(aio): address system one evaluation review findings
bernatixer Sep 26, 2026
badbc00
chore(aio): regenerate evaluation probability taxonomy
bernatixer Sep 26, 2026
09b0bba
fix(aio): restore evaluation badge reasoning fallback
bernatixer Sep 26, 2026
086a7ee
fix(aio): distinguish blocked endpoints from rejected inputs
bernatixer Sep 27, 2026
ff3bf97
fix(ci): update pytest-split fork pin
bernatixer Sep 27, 2026
12cd175
fix(aio): restore model picker story width
bernatixer Sep 27, 2026
0d2c220
fix(aio): show one add-key action in empty state
bernatixer Sep 27, 2026
da9f9a7
fix(aio): bound System One endpoint responses
bernatixer Sep 28, 2026
f8360e8
fix(aio): simplify bounded System One responses
bernatixer Sep 28, 2026
f59fec8
fix(aio): use streamed responses in System One tests
bernatixer Sep 28, 2026
177e3c3
chore(visual): update storybook baselines
posthog[bot] Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 52 additions & 0 deletions docs/internal/ai-observability-judge-inputs.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,58 @@ They sample the combined input, tool definitions, and output only when that text

Implementation: [trace judge](../../posthog/temporal/ai_observability/run_trace_evaluation.py), [session judge](../../posthog/temporal/ai_observability/run_session_evaluation.py), and [generation judge](../../posthog/temporal/ai_observability/evaluation_llm_judge.py).

## System One judges

System One-compatible models are available under the existing LLM judge option.
The `llm-analytics-system-one-evaluations` project-group feature flag controls access in the browser and background workers.
Deploy the ingestion and evaluation worker changes before enabling the flag.
Projects configure a System One-compatible deployment and its authentication.
Evaluation connections never fall back to an instance credential or gateway configuration.
PostHog's regional AI gateway endpoints additionally require an organization in `POSTHOG_INTERNAL_ORG_IDS`; customer projects cannot use those endpoints.
Connection validation and every evaluation check these gates; an absent flag or failed flag lookup blocks the call.
Turning the flag off stops subsequent runs, including queued work, without disabling the saved evaluation.
Keep the experimental flag limited to staff projects during rollout.
Add a connection under **System One** in provider key settings.
Enter the public HTTPS base URL and model ID of a compatible service; neither has a default.
TypeSafe's hosted endpoint is not supported by this integration.
The client appends `/systemone` to the base URL and sends the API key as a bearer token.
An empty key selects no authentication.
Changing the endpoint requires entering its credential again, or explicitly choosing no authentication, so an existing key is not forwarded to a new host.
Private network destinations and redirects are blocked by the shared DNS-pinned HTTP transport.
Saving a connection validates it with a short synthetic input and a Noul question, without sending evaluation data, using a 10-second request timeout.
Select the connection and configured model on each evaluation; these connections cannot become the shared active provider key used by other AI features.
Provider keys keep the provider they were created with; switching providers requires a new key.
The evaluation integration uses Noul for boolean outputs, with the same formatted text for generation, trace, and session targets.
The integration reuses the System One types and parser in `posthog/llm/system_one.py` and the explicit-connection client in `posthog/llm/system_one_client.py`.
Requests use the rate limiter and telemetry in `posthog/egress/typesafe`.
The selected connection supplies its own endpoint and credential; it never falls back to instance gateway settings.
Numeric and categorical support is separate from this integration.
Numeric evaluations retain their existing arbitrary ranges and completion-based judges.
API compatibility does not guarantee equivalent judgments or calibration across models.
Compare results on representative inputs when changing models.

For boolean evaluations, the prompt becomes a [Noul question](https://docs.typesafe.ai/primitives/noul).
A probability of at least 0.5 produces `true`; the evaluation's existing pass/fail polarity still applies.
The raw probability is stored in `$ai_evaluation_probability` when the criteria apply, with available token usage and the configured model ID.
Missing or invalid token counts remain unknown and do not discard a valid answer.
Ingestion estimates cost from the configured model and token usage when the model appears in the existing pricing catalog.
Models without a catalog match retain their usage with cost left unknown.
A custom deployment reporting a recognized model name can inherit that model's catalog estimate; this does not measure its hosting cost.

Evaluations that allow N/A send a separate Noul question about whether the criteria apply, using the 0.5 threshold.
Uncertainty alone does not produce N/A.
System One answers contain no written reasoning, so reports inspect the original source when explaining outcomes.

Each endpoint and credential pair has a separate, hashed rate-limit scope shared across workers.
Evaluations use the batch lane; connection validation uses the normal lane.
Local budget exhaustion, rate limits, and overload responses are retried through Temporal, honoring `Retry-After` up to one minute.
If retries fail, the run fails and the evaluation stays enabled.
Blocked endpoints and redirects disable the evaluation and mark the connection for revalidation, without recording model usage.
Requests rejected because of an individual input skip that run without changing the shared connection.
Invalid probabilities, missing answers, and mismatched answer types skip the item as an unparsable response.
Inputs rejected for exceeding the model's context window are skipped.
See TypeSafe's [API reference](https://docs.typesafe.ai/api) for the System One protocol.

## Model output limits

When the judge reply reaches the model's output limit, the evaluation skips that item with `output_limit_exceeded`.
Expand Down
28 changes: 28 additions & 0 deletions frontend/snapshots.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7456,6 +7456,10 @@ snapshots:
hash: v1.k794b7964.a260abcedfa1372c2c4caf36dbf8e51580ae6d6303b7d18617a6799fbdcb9841.pfjNZI_gu55WYAo4V3uxORMyk_GHRA7uQMkkBaMaYX8
scenes-app-ai-observability-byok-model-picker-notice--no-usable-provider-keys--light:
hash: v1.k794b7964.73105b9c91872982be5b2259a6c794520ac631d2e356babb1cdf235453a32475.L6eqzBeiEZSKXLMsMa3a2cfQvRDwxzia0bdVflixJQ0
scenes-app-ai-observability-byok-model-picker-notice--system-one--dark:
hash: v1.k794b7964.0a56e6f60237a190b55ca5299ce0ef5e4c115455515dc5702c1bf21b2692fdf3.nkYiu0RhIneQSqy4W0wvO0BXhgO1z1I4BzZmmNbTbC0
scenes-app-ai-observability-byok-model-picker-notice--system-one--light:
hash: v1.k794b7964.f15cdbf6cb327a49559c809539d095b7275dc4ef3709c0b2df17c5805d90c813.X6SPnX_tLcSqWR4bMLkftXKkrrDNLAG-4ni68Aratio
scenes-app-ai-observability-clustercard--default--dark:
hash: v1.k794b7964.745c8ac41b251f329b890d56d7b07fb31b4939a440a40eea7932f1081d3de702.ESnQVAOfySBfRJWfJsURYu-prgU3EUvQZCD62MyLlX4
scenes-app-ai-observability-clustercard--default--light:
Expand Down Expand Up @@ -7508,6 +7512,18 @@ snapshots:
hash: v1.k794b7964.c9cf0f699197f208ec18e3ad326c29d43b8034f37f5fc101a9e25773510895a5.f_qTgpzSlQlCRirQv0iD791ajpL-x8WnVi8T8qNYr1o
scenes-app-ai-observability-conversation-display--tools--light:
hash: v1.k794b7964.cc6f7770155f4b720ff0597c3dd5deff699a267cfbdd8fe76a131cb51740215d.lKCAOsQBSoBkK46EHSYdGmWxAi4CQ4wUrx2J1iRF_f0
scenes-app-ai-observability-evaluation-explanation--reasoning--dark:
hash: v1.k794b7964.df59798a73898fed2019aaa67ba440611c743d7bab0ff9c282db9eaeec4602dd.RDTleoCWmufVyJ6bte4BtB45JV51kbNvUbRZAa8lvUo
scenes-app-ai-observability-evaluation-explanation--reasoning--light:
hash: v1.k794b7964.d018ae75901811c91af25c5a8c68e3bf8b02a4b9ca0ffa2641faed46f572e385.LCSHPTx-1opaMtYqjT5TuO0ulPGtrNH9LK87BG5VdBY
scenes-app-ai-observability-evaluation-explanation--system-one--dark:
hash: v1.k794b7964.11fd15711bd135e7f8018f0a417835f17f0056f27d9262a211f8969316d7d1c8.a87FQdEsav9F-zw-IQFVAZWJ6i-9lqXPMCBCAOrisu4
scenes-app-ai-observability-evaluation-explanation--system-one--light:
hash: v1.k794b7964.5fcc731d0e4b1c3cd75c72bb4faeae96144dd4372085ee3db281b0c3c0f6ea11.S2KbBvKJr2ojmkP2R05o593GNVJKmcqv6tPcVWVcpwQ
scenes-app-ai-observability-evaluation-explanation--system-one-zero-probability--dark:
hash: v1.k794b7964.1cd199ebfc3ecf7f75c065b7f670778fed080997c6afb56bac02e49b49637c96.nFRmsF5CaYLyxKdaRS0UtS6ba8esGKqBv6MPu7HwOBo
scenes-app-ai-observability-evaluation-explanation--system-one-zero-probability--light:
hash: v1.k794b7964.f56325c23a185f032b8e7c331463345ea780e4e1b9a85f44a2f560869f06afa4.KcU1H2opmRuSoQcBHacqKya9WVvI5GATsFn_lgLD8rM
scenes-app-ai-observability-instrumentation-checklist-card--collecting--dark:
hash: v1.k794b7964.6f9f209d5daaf8573020fb50615421d426055c8090d0ee73a2629487d30cbf29.g4iTGpNCX68yiOgdFhRMYO2xEGf0WhTu4ji2Rqijroc
scenes-app-ai-observability-instrumentation-checklist-card--collecting--light:
Expand Down Expand Up @@ -7652,6 +7668,18 @@ snapshots:
hash: v1.k794b7964.ef922a0a06bf1db7141ffa5dda43e7106b3eec562f4dc4c0996b817b8e55b3be.Ij1-WjXoREtj3yIYkzbXuRg-ccNy-NPBUBW_31tlfRw
scenes-app-ai-observability-skills--skills-list--light:
hash: v1.k794b7964.d22fb946238114580130ed5d5978b187d6acd086868d1ebb8e326e02c65fc6f5.W7RfUzTvMqZEzuRPjF2h0XXTftd4X0bauTIRZAc48qk
scenes-app-ai-observability-system-one-connection--default--dark:
hash: v1.k794b7964.2729d2ff6ec5024b31cb584cfe8d15a07607c9c9723655feaa6cd5b401d4fd47.1MGXki6PWyOm9TClWTgzksbLuXkha1JfZN-z2sVQsmg
scenes-app-ai-observability-system-one-connection--default--light:
hash: v1.k794b7964.6a1956122a3c65bbc2e5f8a3c9f39ba580bba6e13c2dba3ebb9508ab981088ef.H_oIeIN9TydPQMe866PgGhV4nLvqkFdTwvG1VQPkI_k
scenes-app-ai-observability-system-one-connection--provider-settings--dark:
hash: v1.k794b7964.81ce7999f24877f6e4ac2d9ab89fb4e584e5bc0749e594742c15377c8da8a1a7.BlSbtKwSVk0n9eZKEpRMG7hADbYEDAXAedKVuCaBN_8
scenes-app-ai-observability-system-one-connection--provider-settings--light:
hash: v1.k794b7964.05b7de1c84c25e03badf722fa567182ff449e74a6299d6a543c554ecd20e3184.eccRHOq726ohqws6q_XlZPELbvyUZPfAZk1cwlog_QY
scenes-app-ai-observability-system-one-connection--provider-settings-flag-off--dark:
hash: v1.k794b7964.81ce7999f24877f6e4ac2d9ab89fb4e584e5bc0749e594742c15377c8da8a1a7.DaP59ecAHMR8eaDYuWmJyKEowz34lJ3qbuKNJ4DsN8U
scenes-app-ai-observability-system-one-connection--provider-settings-flag-off--light:
hash: v1.k794b7964.05b7de1c84c25e03badf722fa567182ff449e74a6299d6a543c554ecd20e3184.3SudNZ7cfTDqUdodzpr3A8w1rZBmSjpUHNJ-qyTQQC4
scenes-app-ai-observability-trace--full--dark:
hash: v1.k794b7964.142fc550717f2be0bc0c72f8b670fb5b70f1be8cf26ae2805e4a6ae6937c06a5.ms9_vuzQ5yFxSSCaQtMJLVHBMBR0fIKrrKJhivb5Zfg
scenes-app-ai-observability-trace--full--light:
Expand Down
1 change: 1 addition & 0 deletions frontend/src/lib/constants.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -363,6 +363,7 @@ export const FEATURE_FLAGS = {
LLM_ANALYTICS_EVAL_SETTLING_STRATEGY: 'llm-analytics-eval-settling-strategy', // owner: #team-ai-observability
LLM_ANALYTICS_NUMERIC_EVALS: 'llm-analytics-numeric-evaluations', // owner: #team-ai-observability
LLM_ANALYTICS_OFFLINE_EVALS: 'llm-analytics-offline-evals', // owner: #team-ai-observability
LLM_ANALYTICS_SYSTEM_ONE_EVALUATIONS: 'llm-analytics-system-one-evaluations', // owner: #team-ai-observability
LLM_ANALYTICS_TAGS: 'llm-analytics-tags', // owner: #team-ai-observability
LLM_ANALYTICS_TRACE_NAVIGATION: 'llm-analytics-trace-navigation', // owner: #team-ai-observability
LLM_OBSERVABILITY_TRACE_SEARCH: 'llm-observability-trace-search', // owner: #team-ai-observability
Expand Down
16 changes: 16 additions & 0 deletions frontend/src/taxonomy/core-filter-definitions-by-group.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading