You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit bd5cff4
Browse filesBrowse the repository at this point in the historyBrowse files
fix(gemini): declare cache reporting as inclusive on generations (#860)
* fix(gemini): declare cache reporting as inclusive on generations
Gemini counts `cached_content_token_count` inside `prompt_token_count`, but the
SDK never said so, leaving ingestion to infer the accounting model from the token
counts alone.
That inference is unreliable here. Under explicit context caching the two counts
come from separate measurements, the cache object at creation time and the prompt
per request, so they can disagree by a few percent and the cache pool can land
just above the input total.
Set `cache_reporting_exclusive` to False on generations that report cache reads,
and carry it through the streaming merge and both capture paths so ingestion
prices cached tokens from the declared value.
Generated-By: PostHog Code
Task-Id: 06160e48-feb9-4d39-8b7b-3dcfd1d9ca24
* chore(ai): regenerate public API snapshot for TokenUsage field
Generated-By: PostHog Code
Task-Id: 06160e48-feb9-4d39-8b7b-3dcfd1d9ca24
fix: declare Gemini's cache accounting model on generations with cache reads, so ingestion prices cached tokens from `$ai_cache_reporting_exclusive` instead of inferring it from the token counts.
0 commit comments