Skip to content

fix: add prompt_cache_key and prompt_cache_retention fields to maintain OpenAI chat completion request interface - #1767

Draft
hustxiayang wants to merge 3 commits into
theagentrouter:mainfrom
hustxiayang:add-prompt-key
Draft

hustxiayang wants to merge 3 commits into
theagentrouter:mainfrom
hustxiayang:add-prompt-key

Conversation

@hustxiayang

@hustxiayang hustxiayang commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor

Description
add prompt_cache_key and prompt_cache_retention, which are used to allow users to influence the cache behaviour.
The api: https://platform.openai.com/docs/api-reference/chat/create
doc: https://platform.openai.com/docs/guides/prompt-caching

Signed-off-by: yxia216 <yxia216@bloomberg.net>
@hustxiayang
hustxiayang requested a review from a team as a code owner January 14, 2026 14:55
@dosubot dosubot Bot added the size:XS This PR changes 0-9 lines, ignoring generated files. label Jan 14, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 84.27%. Comparing base (342954a) to head (d013c17).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #1767   +/-   ##
=======================================
  Coverage   84.27%   84.27%           
=======================================
  Files         117      117           
  Lines       12862    12865    +3     
=======================================
+ Hits        10839    10842    +3     
+ Misses       1379     1378    -1     
- Partials      644      645    +1     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@nacx nacx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In PRs #1396 and #1681, support for configuring caching for Anthropic backends was already introduced.

Now that we're adding these fields, should we probably take them into account when translating to those backends? (cc @alexagriffith)

@mathetake

mathetake commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor

Without translation like @nacx mentioned, adding fields alone like the current change does nothing

@hustxiayang

Copy link
Copy Markdown
Contributor Author

@mathetake it's for openai, and does not need any translations

@mathetake

Copy link
Copy Markdown
Contributor

I am saying it's not translating meaning that this unmarshaling result is not used anywhere. could you read the source code and try to understand exactly where your change here affects the runtime behavior. It doesn't do anything

@mathetake
mathetake marked this pull request as draft January 14, 2026 18:45
@mathetake

mathetake commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor

Also please do not claim that you introduce the runtime behavior change without having any unit test like this. It's not how we develop a feature or any fix. otherwise it's only your word that we can trust regarding the change, which is completely unreliable and fragile

@hustxiayang

hustxiayang commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor Author

@mathetake I think it's just to maintain ChatCompletionRequest defined in this repo. I think some other fields defined in this struct are similar (LogitBias, PredictionContent , etc), they are not used in runtime, but just because ai-gateway maintained a ChatCompletionRequest

@mathetake

Copy link
Copy Markdown
Contributor

Then what you are saying now is different from what you described you were trying to do. Could you explain why is that

add prompt_cache_key and prompt_cache_retention to allow users to influence the cache behaviour.

@hustxiayang hustxiayang changed the title fix: add prompt_cache_key and prompt_cache_retention to allow users to influence the cache behaviour fix: add prompt_cache_key and prompt_cache_retention fields to maintain OpenAI chat completion request interface Jan 14, 2026
@hustxiayang

Copy link
Copy Markdown
Contributor Author

@mathetake ah, sorry that it might be a bit ambiguity in this sentence. The meaning I want to express is that add**, which is to allow**. I changed the title to make it clear.

@mathetake

Copy link
Copy Markdown
Contributor

No, I am saying this doesn't allow users to do anything. The description is completely wrong

@hustxiayang

Copy link
Copy Markdown
Contributor Author

@mathetake to allow ** was meant to describe the use of these fields. Probably the grammar was wrong, sorry for that.

@mathetake

Copy link
Copy Markdown
Contributor

so then this PR itself is useless and not necessary to merge unless you actually use these fields and introduce the runtime change as I suggested. Until then keeping this as draft

@hustxiayang

Copy link
Copy Markdown
Contributor Author

Hi, I was trying to encourage users to use this field as it allows you to influence routing and improve cache hit rates according to the OpenAI https://platform.openai.com/docs/guides/prompt-caching#how-it-works. I thought it's better to maintain it in ai-gateway's own definitions as other fields like PredictionContent because ai-gateway did not use openai-go, but never mind if you think it's unncessary.

@hustxiayang

hustxiayang commented Jan 14, 2026 •

Copy link
Copy Markdown
Contributor Author

with regards to @nacx's comments, I guess it's about making gcp anthropic cache control to be compatible/unified with OpenAI's prompt_cache_key and prompt_cache_retention. I feel it might be a bit difficult as they are bit different. For gcp anthropic cache control, it is used to allow users to control which contents to cache. This is not available for OpenAI models. prompt_cache_key is used as a prefix for routing. Let me know if it's wrong.

@nacx

nacx commented Jan 14, 2026

Copy link
Copy Markdown
Member

The contribution is valuable, but as @mathetake says, it is incomplete, and this PR can't be merged as-is.

The OpenAI model you're modifying provides a common interface that AI Gateway can forward directly to OpenAI, but also translates to other endpoints, such as Bedrock, GCP Vertex, etc.

The caching feature was already made available to Bedrock and GCP in the PRs I referenced, but those PRs enable the feature at a different place of the model, in some provider-specific fields. Now you're adding this feature in the common-API, and it is important that translation takes this into account to provide a proper experience to users. For this PR to progress we need:

  • Make sure it works for OpenAI -> OpenAI.
  • Works as well for OpenAI -> Others (especially the two I already mentioned).
  • There are comments/docs/discussion on when to use these new values or the ones that were introduced in the previous PRs.
    • Should those be deprecated in favour of these values?
    • Are they complementary? If not, which one takes precedence if users set both? Do we fail the request with some validation error?

@nacx

nacx commented Jan 14, 2026 •

Copy link
Copy Markdown
Member

prompt_cache_key may not be applicable there, but for example, the retention could be mapped to the anthropic's cache_control.ttl. This is just a small example that highlights the lack of detail in this PR. As said in my previous comment, there needs to be a thought-through reasoning about these fields, how they behave, and what their semantics are with regard to translation, as this belongs to a common abstraction API.

@missBerg missBerg added the area/translation Provider/endpoint coverage and schema translation (incl. fidelity bugs) label Jul 13, 2026
@netlify

netlify Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for theagentrouter canceled.

Name Link
🔨 Latest commit a1d50e2
🔍 Latest deploy log https://app.netlify.com/projects/theagentrouter/deploys/6abe7c1a8cbd3e00083387a5

@codecov

codecov Bot commented Sep 25, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/translation Provider/endpoint coverage and schema translation (incl. fidelity bugs) size:XS This PR changes 0-9 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants