Skip to content

Proposal: direct /anthropic/v1/messages → GCPVertexAI (Gemini generateContent) translator #2773

Description

@yoshimasak

Summary

Today /anthropic/v1/messages cannot target a GCPVertexAI backend directly: MessagesEndpointSpec.GetTranslator returns an error for that schema (endpointspec.go#L447-L458). The only way to send Anthropic-format clients (Claude Code, Anthropic SDKs) to Gemini is a two-hop translation, Anthropic → OpenAI → Gemini, which loses information that Gemini can natively consume.

I would like to contribute a direct anthropicToGCPVertexAITranslator (Anthropic Messages → Gemini generateContent / streamGenerateContent), in the same way anthropic_awsbedrock.go provides a direct Anthropic → Bedrock Converse path. Before opening a PR I'd like to confirm the direction is wanted and how it should relate to #2369 and #2040.

Motivation: what the two-hop path loses

Checked against internal/translator/openai_helper.go on main:

Anthropic input Two-hop (Anthropic → OpenAI → Gemini) Gemini native capability
tool_result content with image / document blocks Dropped: toolResultToText keeps text blocks only (openai_helper.go#L277-L291) Multimodal functionResponse.parts (Gemini 3+)
document blocks (PDF, base64 / URL) in user messages Dropped: only text and image blocks are converted inlineData / fileData with application/pdf
output_config.format (json_schema, sent by Claude Code for session titles, branch names, hooks) Not mapped; the model returns plain text and the client fails to parse responseMimeType + responseJsonSchema
Thought signatures Carried through OpenAI thinking_blocks → thoughtSignature (#2319, #2366), i.e. two conversions each way Native thoughtSignature on parts
tool_choice: {type: tool}, stop_sequences, top_k, mid-conversation system messages Partially (OpenAI has no top_k; mid-conversation system role becomes a developer message) Direct mapping
thinking (adaptive / disabled) + output_config.effort (Claude Code's default) output_config.effort is dropped in the Anthropic → OpenAI hop; adaptive and disabled produce an empty thinkingConfig in getGenerationConfigThinkingConfig (openai_gcpvertexai.go#L493-L517), so the model falls back to default dynamic thinking thinkingConfig.thinkingLevel (Gemini 3) / thinkingBudget (2.5), includeThoughts from display

These matter in practice for Claude Code: tool results that contain screenshots or PDFs (Read tool) and structured-output auxiliary requests both go through this path.

Proposal

Add internal/translator/anthropic_gcpvertexai.go (+ tests) and a filterapi.APISchemaGCPVertexAI case in MessagesEndpointSpec.GetTranslator. No changes to internal/apischema or shared helpers; the translator reuses jsonSchemaToGemini, responseJSONSchemaAvailable, mapReasoningEffortToThinkingLevel and the GCP path helpers.

Request side:

  • system / user / assistant / tool_result → systemInstruction / contents, including base64 and URL images, PDFs, search_result and document.source.content; tool_result media → functionResponse.parts.
  • Custom tools → functionDeclarations (parametersJsonSchema on Gemini 2.5/3, converted schema otherwise); tool_choice → functionCallingConfig. Built-in/server tools (web search, bash, text editor) are skipped with a debug log, and rejected with 400 if no custom tool remains.
  • thinking (enabled / adaptive / disabled) and output_config.effort → thinkingConfig (thinkingLevel on Gemini 3, thinkingBudget on 2.5); output_config.format (json_schema) → responseMimeType + responseJsonSchema / responseSchema, mirroring openAIReqToGeminiGenerationConfig.
  • Thought signatures round-trip inside Anthropic thinking blocks with a gemini: prefix (so Anthropic-issued signatures are never forwarded to Gemini); missing current-turn signatures get the documented skip_thought_signature_validator value.
  • Gemini constraints handled up front with Anthropic-style 400s: assistant prefill, consecutive same-role contents (merged), user contents mixing functionResponse and text (split), empty contents, empty media.

Response side:

  • Non-streaming and SSE streaming conversion to Anthropic message_start / content_block_* / message_delta / message_stop, usage (input_tokens, output_tokens, cached tokens), stop_reason mapping (MAX_TOKENS → max_tokens, function calls → tool_use, prompt/safety blocks → refusal, MALFORMED_FUNCTION_CALL etc. → api_error), Gemini error → Anthropic error envelope.

Out of scope / behavior notes:

  • count_tokens: a separate CountTokensEndpointSpec with its own translator type; a natural follow-up using Gemini countTokens.
  • Server and client built-in tools (web_search_*, bash_*, text_editor_*, computer_*): skipped with a debug log, and rejected with 400 only when no custom tool remains. Gemini's googleSearch grounding returns groundingMetadata, not web_search_tool_result blocks, so mapping it would be a separate piece of work.
  • Assistant prefill: rejected with 400, since Vertex AI rejects requests ending with a model turn and emulating prefill would change the output contract.
  • metadata, service_tier, cache_control, container, mcp_servers, context_management, safeguards: ignored. They have no generateContent counterpart (explicit caching is a separate cachedContents resource; implicit caching still applies and is reported as cache_read_input_tokens) and the existing Anthropic translators ignore them as well. The list is documented on the translator type.

Status: implemented and tested on a fork — yoshimasak/agent-router@feat/anthropic-gcpvertexai (based on v1.1.0; will be rebased onto main before a PR). Unit tests cover request/response/streaming paths (~95% statement coverage of the new file), and the translator has been exercised end to end with Claude Code against Vertex AI gemini-3.8-flash (tool loops, images, PDFs, structured output, session resume).

Relation to open work

Questions for maintainers

  1. Is a direct Anthropic → Gemini translator welcome, or would you prefer to keep investing in the two-hop path (e.g. multimodal tool results and output_config in openai_helper.go)?
  2. Should the PR wait for feat: implement capability aware translation #2369, or land first with the existing model-name helpers?
  3. PR shape: one PR (~1.6k lines + tests) or split into request-side / response-side PRs?
  4. Should tests/data-plane (testupstream) cases and site/docs updates be part of the same PR?

Disclosure: the implementation was developed with an AI coding assistant; I have reviewed, tested and will maintain the code per the CONTRIBUTING AI policy.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions