Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 20 additions & 8 deletions docs/internal/ai-observability-judge-inputs.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,9 +39,9 @@ They sample the combined input, tool definitions, and output only when that text

Implementation: [trace judge](../../posthog/temporal/ai_observability/run_trace_evaluation.py), [session judge](../../posthog/temporal/ai_observability/run_session_evaluation.py), and [generation judge](../../posthog/temporal/ai_observability/evaluation_llm_judge.py).

## System One judges
## Decision model judges

System One-compatible models are available under the existing LLM judge option.
Decision models are available under the existing LLM judge option. They return typed answers without written reasoning.
The `llm-analytics-system-one-evaluations` project-group feature flag controls access in the browser and background workers.
Both use the project's UUID as its group key; the numeric project ID is a group property.
Deploy the ingestion and evaluation worker changes before enabling the flag.
Expand All @@ -51,17 +51,29 @@ The configured endpoint authorizes the supplied credential, including connection
Connection validation and every evaluation check these gates; an absent flag or failed flag lookup blocks the call.
Turning the flag off stops subsequent runs, including queued work, without disabling the saved evaluation.
Keep the experimental flag limited to staff projects during rollout.
Add a connection under **System One** in provider key settings.
OpenRouter decision models, including Jev, appear under an existing OpenRouter key in the evaluation model picker when this flag is enabled.
They use OpenRouter's alpha `/api/alpha/decisions` endpoint with that key; no custom endpoint or System One connection is needed.
The picker discovers models through the catalogue's `decisions` output modality, so new models and versions appear without a code change.
The picker does not exclude individual decision models. Listed models may not support every evaluation output type.
The catalogue does not expose supported question types. If a model rejects a request, the run is skipped without retrying, with a reason to check output-type and criteria compatibility.
Support added by a provider works on subsequent runs without a PostHog code change.
Chat models such as Jev Router keep using chat completions.
Decision models are excluded from the playground and tagger model pickers.
For projects with this flag enabled, an unavailable catalogue makes OpenRouter runs retry rather than guess which API to call.
These projects also need the catalogue when changing a model or output configuration, changing the evaluation type, or enabling an evaluation.
Renaming or disabling an evaluation does not require the catalogue. Projects without the flag keep the existing chat path.
An out-of-credits response disables the evaluation and marks its provider key as failing, as on the chat path.
For other compatible services, add a connection under **System One** in provider key settings.
Enter the public HTTPS base URL and model ID of a compatible service; neither has a default.
TypeSafe's hosted endpoint is not supported by this integration.
The client appends `/systemone` to the base URL and sends the API key as a bearer token.
An empty key selects no authentication.
Changing the endpoint requires entering its credential again, or explicitly choosing no authentication, so an existing key is not forwarded to a new host.
Private network destinations and redirects are blocked by the shared DNS-pinned HTTP transport.
Saving a connection validates it with a short synthetic input and a Noul question, without sending evaluation data, using a 10-second request timeout.
System One connections opt into the shared HTTPX client's bounded transport.
Decision model connections opt into the shared HTTPX client's bounded transport.
Each HTTP request has a total deadline covering connection setup, response headers, and the body; expiry cancels the network operation and closes the connection.
Validation uses 10 seconds and System One evaluations use 60 seconds.
Validation uses 10 seconds and decision model evaluations use 60 seconds.
The transport rejects response bodies above 1 MiB, including errors.
It requests uncompressed responses and rejects compressed responses to prevent decompression from bypassing the size limit.
Responses stream incrementally, with a separate connection per request; connections are not pooled across requests.
Expand Down Expand Up @@ -98,7 +110,7 @@ A custom deployment reporting a recognized model name can inherit that model's c

Evaluations that allow N/A send a separate Noul question about whether the criteria apply, using the 0.5 threshold.
Uncertainty alone does not produce N/A.
System One answers contain no written reasoning, so reports inspect the original source when explaining outcomes.
Decision model answers contain no written reasoning, so reports inspect the original source when explaining outcomes.

Endpoint rate limits and overload responses are retried through Temporal, honoring `Retry-After` up to one minute.
Evaluation events retain model, usage, latency, and error telemetry.
Expand All @@ -107,7 +119,7 @@ Blocked endpoints and redirects disable the evaluation and mark the connection f
Requests rejected because of an individual input skip that run without changing the shared connection.
Invalid probabilities, missing answers, and mismatched answer types skip the item as an unparsable response.
Inputs rejected for exceeding the model's context window are skipped.
See TypeSafe's [API reference](https://docs.typesafe.ai/api) for the System One protocol.
See OpenRouter's [Decisions API reference](https://openrouter.ai/docs/api/api-reference/alphadecisions/submit-a-decisions-request) and TypeSafe's [API reference](https://docs.typesafe.ai/api) for their compatible question and answer formats.

## Model output limits

Expand All @@ -125,7 +137,7 @@ Boolean online evaluations write their raw verdict to `$ai_evaluation_result`.
Numeric evaluations write their score to `$ai_evaluation_numeric_result`, with optional `$ai_evaluation_numeric_result_min` and `$ai_evaluation_numeric_result_max` bounds.
Result badges and mean scores display up to two decimal places, with two significant digits for values below one to keep small nonzero scores visible. Badges in the runs table expose the exact score on hover; storage, sorting, and passing rules use the original value.
If rounding would change whether the displayed score meets the passing rule, the badge shows the exact score instead.
For online LLM judges, including System One, `step` is a suggested score increment in the prompt; results are not rounded or restricted to its multiples.
For online LLM judges, including decision models, `step` is a suggested score increment in the prompt; results are not rounded or restricted to its multiples.
Hog evaluations use the numeric value returned by the code and do not apply `step`.
`$ai_evaluation_result_type` identifies the output type; events without it are legacy boolean results.
Categorical evaluations write a list of category keys to `$ai_evaluation_categorical_result`, including for single selection.
Expand Down
102 changes: 57 additions & 45 deletions posthog/temporal/ai_observability/evaluation_llm_judge.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,14 @@
from posthog.temporal.common.utils import close_db_connections

from products.ai_observability.backend.llm import DEFAULT_MODEL_BY_PROVIDER, Client, CompletionRequest, Usage
from products.ai_observability.backend.llm.decisions import (
DecisionClient,
DecisionEndpointBlockedError,
DecisionRateLimitError,
DecisionRequestRejectedError,
decision_evaluations_enabled,
is_decision_model,
)
from products.ai_observability.backend.llm.errors import (
AuthenticationError,
ContextWindowExceededError,
Expand All @@ -61,13 +69,7 @@
UnsupportedModelError,
provider_error_detail,
)
from products.ai_observability.backend.llm.system_one import (
SystemOneClient,
SystemOneEndpointBlockedError,
SystemOneRateLimitError,
SystemOneRequestRejectedError,
system_one_evaluations_enabled,
)
from products.ai_observability.backend.llm.providers.openrouter import OPENROUTER_DECISIONS_BASE_URL
from products.ai_observability.backend.llm.types import CompletionResponse
from products.ai_observability.backend.models.evaluation_configs import (
CategoricalOutputConfig,
Expand Down Expand Up @@ -506,7 +508,7 @@
)


def _system_one_numeric_score(minimum: float, maximum: float, index: float) -> float:
def _decision_numeric_score(minimum: float, maximum: float, index: float) -> float:
last_index = MAX_SCORE_LEVELS - 1
if index == 0:
return minimum
Expand All @@ -520,7 +522,7 @@
return minimum * (1 - weight) + maximum * weight


def call_llm_judge(

Check warning on line 525 in posthog/temporal/ai_observability/evaluation_llm_judge.py

View workflow job for this annotation

GitHub Actions / Python code quality (depot-ubuntu-24.04)

lint:complexity

`call_llm_judge` has cyclomatic complexity 61 (warn >10)

Check warning on line 525 in posthog/temporal/ai_observability/evaluation_llm_judge.py

View workflow job for this annotation

GitHub Actions / Python code quality (depot-ubuntu-24.04)

`call_llm_judge` has cyclomatic complexity 61 (warn >10)
*,
evaluation: dict[str, Any],
system_prompt: str,
Expand Down Expand Up @@ -553,23 +555,6 @@
is_byok = resolved.is_byok
key_id = str(provider_key.id) if provider_key else None

if provider == "system_one":
if output_type not in ("boolean", "categorical", "numeric"):
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning="System One supports boolean, categorical, and numeric evaluations.",
skip_reason="unsupported_output_type",
)
base_url = provider_key.encrypted_config.get("base_url", "") if provider_key else ""
if not system_one_evaluations_enabled(team_id, base_url=base_url):
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning="System One evaluations are not available for this project.",
skip_reason="system_one_unavailable",
)

type_config = get_output_type_config(allows_na, output_type=output_type, output_config=output_config)
response_format = type_config.response_format

Expand All @@ -582,9 +567,35 @@
)

probability: float | None = None
system_one_result = None
decision_result = None
try:
if provider == "system_one":
if is_decision_model(
provider,
model,
openrouter_enabled=provider == "openrouter"
and decision_evaluations_enabled(team_id, base_url=OPENROUTER_DECISIONS_BASE_URL),
):
Comment thread
coderabbitai[bot] marked this conversation as resolved.
if output_type not in ("boolean", "categorical", "numeric"):
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning="Decision models support boolean, categorical, and numeric evaluations.",
skip_reason="unsupported_output_type",
)
base_url = (
OPENROUTER_DECISIONS_BASE_URL
if provider == "openrouter"
else provider_key.encrypted_config.get("base_url", "")
if provider_key
else ""
)
if provider == "system_one" and not decision_evaluations_enabled(team_id, base_url=base_url):
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning="System One evaluations are not available for this project.",
skip_reason="system_one_unavailable",
)
prompt = evaluation["evaluation_config"]["prompt"]
categorical_config = (
CategoricalOutputConfig.model_validate(output_config) if output_type == "categorical" else None
Expand All @@ -597,11 +608,11 @@
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning="System One numeric evaluations require a minimum score below the maximum score.",
reasoning="Numeric evaluations with decision models require a minimum score below the maximum score.",
skip_reason="request_rejected",
)
numeric_levels = [
_system_one_numeric_score(numeric_config.min, numeric_config.max, index)
_decision_numeric_score(numeric_config.min, numeric_config.max, index)
for index in range(MAX_SCORE_LEVELS)
]
if numeric_config.step is not None:
Expand Down Expand Up @@ -642,16 +653,17 @@
+ prompt
)
)
system_one_result = SystemOneClient.evaluate(
decision_result = DecisionClient.evaluate(
api_key=provider_key.encrypted_config.get("api_key", "") if provider_key else "",
base_url=base_url,
path="decisions" if provider == "openrouter" else "systemone",
model=model,
state=user_prompt,
questions=questions,
)
applicable = True
if allows_na:
applicability_answer = system_one_result.answers["applicable"]
applicability_answer = decision_result.answers["applicable"]
if not isinstance(applicability_answer, NoulAnswer):
raise StructuredOutputParseError("The endpoint returned an invalid applicability answer.")
applicable = applicability_answer.probability >= 0.5
Expand All @@ -664,10 +676,10 @@
| NumericWithNAEvalResult
)
if numeric_levels is not None:
score_answer = system_one_result.answers["score"]
score_answer = decision_result.answers["score"]
if not isinstance(score_answer, ScoreAnswer):
raise StructuredOutputParseError("The endpoint returned an invalid score answer.")
score = _system_one_numeric_score(numeric_levels[0], numeric_levels[-1], score_answer.score)
score = _decision_numeric_score(numeric_levels[0], numeric_levels[-1], score_answer.score)
parsed = (
NumericWithNAEvalResult(reasoning="", score=score if applicable else None)
if allows_na
Expand All @@ -676,13 +688,13 @@
elif categorical_config is not None:
categories: list[str] = []
if categorical_config.selection_mode == "single":
category_answer = system_one_result.answers["category"]
category_answer = decision_result.answers["category"]
if not isinstance(category_answer, ChoiceAnswer):
raise StructuredOutputParseError("The endpoint returned an invalid category answer.")
categories = [category_answer.choice]
else:
for index, option in enumerate(categorical_config.options):
category_match = system_one_result.answers[f"category_{index}"]
category_match = decision_result.answers[f"category_{index}"]
if not isinstance(category_match, NoulAnswer):
raise StructuredOutputParseError("The endpoint returned an invalid category answer.")
if category_match.probability >= 0.5:
Expand All @@ -693,7 +705,7 @@
else CategoricalEvalResult(reasoning="", categories=categories)
)
else:
verdict_answer = system_one_result.answers["verdict"]
verdict_answer = decision_result.answers["verdict"]
if not isinstance(verdict_answer, NoulAnswer):
raise StructuredOutputParseError("The endpoint returned an invalid verdict answer.")
probability = verdict_answer.probability
Expand All @@ -710,9 +722,9 @@
model=model,
parsed=parsed,
usage=Usage(
input_tokens=(system_one_result.input_tokens or 0),
output_tokens=(system_one_result.output_tokens or 0),
total_tokens=(system_one_result.input_tokens or 0) + (system_one_result.output_tokens or 0),
input_tokens=(decision_result.input_tokens or 0),
output_tokens=(decision_result.output_tokens or 0),
total_tokens=(decision_result.input_tokens or 0) + (decision_result.output_tokens or 0),
),
)
else:
Expand All @@ -725,7 +737,7 @@
response_format=response_format,
)
)
except SystemOneEndpointBlockedError as e:
except DecisionEndpointBlockedError as e:
increment_user_errors("endpoint_blocked", provider=provider)
return terminal_user_error_result(
spec=require_user_error_spec("endpoint_blocked", is_byok=is_byok),
Expand All @@ -735,15 +747,15 @@
key_id=key_id,
is_byok=is_byok,
)
except SystemOneRequestRejectedError as e:
except DecisionRequestRejectedError as e:
increment_user_errors("request_rejected", provider=provider)
return build_skipped_evaluation_result(
output_type=output_type,
allows_na=allows_na,
reasoning=str(e),
skip_reason="request_rejected",
)
except SystemOneRateLimitError as e:
except DecisionRateLimitError as e:
increment_errors("rate_limit", provider=provider)
raise ApplicationError(
str(e),
Expand Down Expand Up @@ -981,9 +993,9 @@
"model": model,
"provider": provider,
}
if system_one_result is not None:
result_dict["input_tokens"] = system_one_result.input_tokens
result_dict["output_tokens"] = system_one_result.output_tokens
if decision_result is not None:
result_dict["input_tokens"] = decision_result.input_tokens
result_dict["output_tokens"] = decision_result.output_tokens

if isinstance(parsed_result, CategoricalEvalResult | CategoricalWithNAEvalResult):
result_dict["result_type"] = "categorical"
Expand Down
Loading
Loading