Skip to content

feat: implement capability aware translation - #2369

Open
hustxiayang wants to merge 4 commits into
theagentrouter:mainfrom
hustxiayang:model-hints
Open

hustxiayang wants to merge 4 commits into
theagentrouter:mainfrom
hustxiayang:model-hints

Conversation

@hustxiayang

@hustxiayang hustxiayang commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Description

Fixes #2368

Motivation

Several translators make model-specific decisions by string-matching the model name — Anthropic max_tokens defaulting to 0, Gemini JSON-schema gating on "2.5"/"3" substrings, and hardcoded effortModels/outputConfigModels lists. These are silently wrong for new models (e.g. the substring gate breaks for a future gemini-4, and each Claude release needs a new list entry) and costly to maintain.

Change

Adds an optional modelTranslationHints field to AIServiceBackendSpec so operators can declare per-model capabilities that translators use instead of heuristics. It is a list keyed by modelNameOverride, with these hint fields per entry:

  • maxOutputTokens — default max_tokens for Anthropic when the client omits it
  • supportsResponseJsonSchema — Gemini JSON-schema gate
  • supportsReasoningEffort — Gemini thinking / Anthropic output_config.effort
  • supportsOutputConfig — Anthropic structured output
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
  name: gcp-anthropic-backend
spec:
  schema:
    name: GCPAnthropic
  modelTranslationHints:
  - modelNameOverride: "claude-opus-4-6@20250514"
    maxOutputTokens: 32000
    supportsReasoningEffort: true
    supportsOutputConfig: true

Why AIServiceBackendSpec

Based on the comments in Fixes #2368, translation hints are a property of a model on a specific backend, not of a routing decision. Placing them on AIServiceBackendSpec avoids duplication: an AIServiceBackend is a single object referenced by many routes/rules, so hints are declared once and every reference picks them up via the existing backend lookup — instead of repeating the same values on every backendRef.

Keying by the exact modelNameOverride (not a fuzzy match) keeps this provider-scoped and precise: because an AIServiceBackend has exactly one schema, the same logical model hosted on two providers (e.g. Claude on AWS Bedrock vs. GCP Vertex AI) is configured as separate entries on separate backends — which is correct, since providers can genuinely support different features for what is nominally the same model (e.g. output_config on Bedrock but not on Vertex).

Backward compatibility

This change is fully backward compatible and introduces no behavior change for existing users. The hardcoded lists and substring heuristics are retained as fallbacks; hints only take precedence when explicitly set:

  • No hints configured (existing users) → translators fall through to the existing list/heuristic — identical to today.
  • A hint field is set → that value is authoritative for that field.
  • Fallback is evaluated per field, so setting one hint (e.g. maxOutputTokens) never disables an unrelated feature that still relies on the heuristic.

This lets both mechanisms coexist during a transition window: operators migrate to explicit config at their own pace. A follow-up PR will remove the hardcoded lists/heuristics once config is the source of truth; the lists are marked deprecated in code to make that removal a mechanical change.

@hustxiayang
hustxiayang requested a review from a team as a code owner July 14, 2026 16:55
@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Jul 14, 2026
@hustxiayang
hustxiayang marked this pull request as draft July 14, 2026 16:56
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 84.85%. Comparing base (f91bb5b) to head (26f55a7).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #2369      +/-   ##
==========================================
+ Coverage   84.82%   84.85%   +0.03%     
==========================================
  Files         148      148              
  Lines       21838    21874      +36     
==========================================
+ Hits        18524    18562      +38     
+ Misses       2194     2193       -1     
+ Partials     1120     1119       -1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@missBerg missBerg added enhancement New feature or request area/translation Provider/endpoint coverage and schema translation (incl. fidelity bugs) area/api Control plane API (CRDs) labels Jul 15, 2026
@hustxiayang
hustxiayang marked this pull request as ready for review August 5, 2026 20:15
@dosubot dosubot Bot added size:XL This PR changes 500-999 lines, ignoring generated files. and removed size:L This PR changes 100-499 lines, ignoring generated files. labels Aug 5, 2026
@hustxiayang
hustxiayang requested a review from a team as a code owner September 10, 2026 17:29
@netlify

netlify Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for theagentrouter ready!

Name Link
🔨 Latest commit f6f946a
🔍 Latest deploy log https://app.netlify.com/projects/theagentrouter/deploys/6ac01479020dec00089853f4
😎 Deploy Preview https://deploy-preview-2369--theagentrouter.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

Signed-off-by: yxia216 <yxia216@bloomberg.net>
Signed-off-by: yxia216 <yxia216@bloomberg.net>
@codecov

codecov Bot commented Sep 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Signed-off-by: yxia216 <yxia216@bloomberg.net>
aabchoo pushed a commit that referenced this pull request Sep 24, 2026
…ed output and effort (#2673)

**Description**
The config setup is blocked by
#2369

---------

Signed-off-by: yxia216 <yxia216@bloomberg.net>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api Control plane API (CRDs) area/translation Provider/endpoint coverage and schema translation (incl. fidelity bugs) enhancement New feature or request size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

model capability based translation

4 participants