You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Today /anthropic/v1/messages cannot target a GCPVertexAI backend directly: MessagesEndpointSpec.GetTranslator returns an error for that schema (endpointspec.go#L447-L458). The only way to send Anthropic-format clients (Claude Code, Anthropic SDKs) to Gemini is a two-hop translation, Anthropic → OpenAI → Gemini, which loses information that Gemini can natively consume.
I would like to contribute a direct anthropicToGCPVertexAITranslator (Anthropic Messages → Gemini generateContent / streamGenerateContent), in the same way anthropic_awsbedrock.go provides a direct Anthropic → Bedrock Converse path. Before opening a PR I'd like to confirm the direction is wanted and how it should relate to #2369 and #2040.
Motivation: what the two-hop path loses
Checked against internal/translator/openai_helper.go on main:
output_config.effort is dropped in the Anthropic → OpenAI hop; adaptive and disabled produce an empty thinkingConfig in getGenerationConfigThinkingConfig (openai_gcpvertexai.go#L493-L517), so the model falls back to default dynamic thinking
thinkingConfig.thinkingLevel (Gemini 3) / thinkingBudget (2.5), includeThoughts from display
These matter in practice for Claude Code: tool results that contain screenshots or PDFs (Read tool) and structured-output auxiliary requests both go through this path.
Proposal
Add internal/translator/anthropic_gcpvertexai.go (+ tests) and a filterapi.APISchemaGCPVertexAI case in MessagesEndpointSpec.GetTranslator. No changes to internal/apischema or shared helpers; the translator reuses jsonSchemaToGemini, responseJSONSchemaAvailable, mapReasoningEffortToThinkingLevel and the GCP path helpers.
Request side:
system / user / assistant / tool_result → systemInstruction / contents, including base64 and URL images, PDFs, search_result and document.source.content; tool_result media → functionResponse.parts.
Custom tools → functionDeclarations (parametersJsonSchema on Gemini 2.5/3, converted schema otherwise); tool_choice → functionCallingConfig. Built-in/server tools (web search, bash, text editor) are skipped with a debug log, and rejected with 400 if no custom tool remains.
thinking (enabled / adaptive / disabled) and output_config.effort → thinkingConfig (thinkingLevel on Gemini 3, thinkingBudget on 2.5); output_config.format (json_schema) → responseMimeType + responseJsonSchema / responseSchema, mirroring openAIReqToGeminiGenerationConfig.
Thought signatures round-trip inside Anthropic thinking blocks with a gemini: prefix (so Anthropic-issued signatures are never forwarded to Gemini); missing current-turn signatures get the documented skip_thought_signature_validator value.
Gemini constraints handled up front with Anthropic-style 400s: assistant prefill, consecutive same-role contents (merged), user contents mixing functionResponse and text (split), empty contents, empty media.
count_tokens: a separate CountTokensEndpointSpec with its own translator type; a natural follow-up using Gemini countTokens.
Server and client built-in tools (web_search_*, bash_*, text_editor_*, computer_*): skipped with a debug log, and rejected with 400 only when no custom tool remains. Gemini's googleSearch grounding returns groundingMetadata, not web_search_tool_result blocks, so mapping it would be a separate piece of work.
Assistant prefill: rejected with 400, since Vertex AI rejects requests ending with a model turn and emulating prefill would change the output contract.
metadata, service_tier, cache_control, container, mcp_servers, context_management, safeguards: ignored. They have no generateContent counterpart (explicit caching is a separate cachedContents resource; implicit caching still applies and is reported as cache_read_input_tokens) and the existing Anthropic translators ignore them as well. The list is documented on the translator type.
Status: implemented and tested on a fork — yoshimasak/agent-router@feat/anthropic-gcpvertexai (based on v1.1.0; will be rebased onto main before a PR). Unit tests cover request/response/streaming paths (~95% statement coverage of the new file), and the translator has been exercised end to end with Claude Code against Vertex AI gemini-3.8-flash (tool loops, images, PDFs, structured output, session resume).
fix: strip structured-output headers and body for gcp on message api path. #2040 (strip structured-output fields on GCPAnthropic): for GCPAnthropic stripping makes sense because Vertex AI's Claude endpoint rejects the fields. For Gemini the fields map to native features, so this translator translates output_config instead of stripping it. Please let me know if you'd rather keep both GCP paths consistent.
feat(translator): fix field gaps for multi-backend Claude Code #2332 (multi-backend Claude Code): overlapping goal, different approach (improving the two-hop path). This proposal doesn't replace it; the OpenAI hop remains the path for non-Gemini OpenAI-compatible backends.
Is a direct Anthropic → Gemini translator welcome, or would you prefer to keep investing in the two-hop path (e.g. multimodal tool results and output_config in openai_helper.go)?
PR shape: one PR (~1.6k lines + tests) or split into request-side / response-side PRs?
Should tests/data-plane (testupstream) cases and site/docs updates be part of the same PR?
Disclosure: the implementation was developed with an AI coding assistant; I have reviewed, tested and will maintain the code per the CONTRIBUTING AI policy.
Summary
Today
/anthropic/v1/messagescannot target aGCPVertexAIbackend directly:MessagesEndpointSpec.GetTranslatorreturns an error for that schema (endpointspec.go#L447-L458). The only way to send Anthropic-format clients (Claude Code, Anthropic SDKs) to Gemini is a two-hop translation, Anthropic → OpenAI → Gemini, which loses information that Gemini can natively consume.I would like to contribute a direct
anthropicToGCPVertexAITranslator(Anthropic Messages → GeminigenerateContent/streamGenerateContent), in the same wayanthropic_awsbedrock.goprovides a direct Anthropic → Bedrock Converse path. Before opening a PR I'd like to confirm the direction is wanted and how it should relate to #2369 and #2040.Motivation: what the two-hop path loses
Checked against
internal/translator/openai_helper.goonmain:tool_resultcontent withimage/documentblockstoolResultToTextkeeps text blocks only (openai_helper.go#L277-L291)functionResponse.parts(Gemini 3+)documentblocks (PDF, base64 / URL) in user messagestextandimageblocks are convertedinlineData/fileDatawithapplication/pdfoutput_config.format(json_schema, sent by Claude Code for session titles, branch names, hooks)responseMimeType+responseJsonSchemathinking_blocks→thoughtSignature(#2319, #2366), i.e. two conversions each waythoughtSignatureon partstool_choice: {type: tool},stop_sequences,top_k, mid-conversationsystemmessagestop_k; mid-conversationsystemrole becomes a developer message)thinking(adaptive/disabled) +output_config.effort(Claude Code's default)output_config.effortis dropped in the Anthropic → OpenAI hop;adaptiveanddisabledproduce an emptythinkingConfigingetGenerationConfigThinkingConfig(openai_gcpvertexai.go#L493-L517), so the model falls back to default dynamic thinkingthinkingConfig.thinkingLevel(Gemini 3) /thinkingBudget(2.5),includeThoughtsfromdisplayThese matter in practice for Claude Code: tool results that contain screenshots or PDFs (Read tool) and structured-output auxiliary requests both go through this path.
Proposal
Add
internal/translator/anthropic_gcpvertexai.go(+ tests) and afilterapi.APISchemaGCPVertexAIcase inMessagesEndpointSpec.GetTranslator. No changes tointernal/apischemaor shared helpers; the translator reusesjsonSchemaToGemini,responseJSONSchemaAvailable,mapReasoningEffortToThinkingLeveland the GCP path helpers.Request side:
tool_result→systemInstruction/contents, including base64 and URL images, PDFs,search_resultanddocument.source.content;tool_resultmedia →functionResponse.parts.functionDeclarations(parametersJsonSchemaon Gemini 2.5/3, converted schema otherwise);tool_choice→functionCallingConfig. Built-in/server tools (web search, bash, text editor) are skipped with a debug log, and rejected with 400 if no custom tool remains.thinking(enabled/adaptive/disabled) andoutput_config.effort→thinkingConfig(thinkingLevelon Gemini 3,thinkingBudgeton 2.5);output_config.format(json_schema) →responseMimeType+responseJsonSchema/responseSchema, mirroringopenAIReqToGeminiGenerationConfig.thinkingblocks with agemini:prefix (so Anthropic-issued signatures are never forwarded to Gemini); missing current-turn signatures get the documentedskip_thought_signature_validatorvalue.functionResponseand text (split), emptycontents, empty media.Response side:
message_start/content_block_*/message_delta/message_stop, usage (input_tokens,output_tokens, cached tokens),stop_reasonmapping (MAX_TOKENS→max_tokens, function calls →tool_use, prompt/safety blocks →refusal,MALFORMED_FUNCTION_CALLetc. →api_error), Gemini error → Anthropic error envelope.Out of scope / behavior notes:
count_tokens: a separateCountTokensEndpointSpecwith its own translator type; a natural follow-up using GeminicountTokens.web_search_*,bash_*,text_editor_*,computer_*): skipped with a debug log, and rejected with 400 only when no custom tool remains. Gemini'sgoogleSearchgrounding returnsgroundingMetadata, notweb_search_tool_resultblocks, so mapping it would be a separate piece of work.metadata,service_tier,cache_control,container,mcp_servers,context_management,safeguards: ignored. They have nogenerateContentcounterpart (explicit caching is a separatecachedContentsresource; implicit caching still applies and is reported ascache_read_input_tokens) and the existing Anthropic translators ignore them as well. The list is documented on the translator type.Status: implemented and tested on a fork — yoshimasak/agent-router@feat/anthropic-gcpvertexai (based on v1.1.0; will be rebased onto
mainbefore a PR). Unit tests cover request/response/streaming paths (~95% statement coverage of the new file), and the translator has been exercised end to end with Claude Code against Vertex AIgemini-3.8-flash(tool loops, images, PDFs, structured output, session resume).Relation to open work
parametersJsonSchema/responseJsonSchema/thinkingLevelon the existingresponseJSONSchemaAvailable/reasoningEffortAvailablehelpers. If feat: implement capability aware translation #2369 lands first I'll switch tomodelTranslationHints; if this lands first, the same three call sites would need to be covered by feat: implement capability aware translation #2369. Happy to coordinate either way.output_configinstead of stripping it. Please let me know if you'd rather keep both GCP paths consistent.effort: high|xhigh|maxtothinkingLevel: HIGHdirectly to avoid the current error on Pro; once translator: accept reasoning_effort high for Gemini 3 Pro models #2749 merges this special case can go away.Questions for maintainers
output_configinopenai_helper.go)?tests/data-plane(testupstream) cases andsite/docsupdates be part of the same PR?Disclosure: the implementation was developed with an AI coding assistant; I have reviewed, tested and will maintain the code per the CONTRIBUTING AI policy.