OpenAI-compatible inference for CodeOnTheGo's AI plugins. Registers itself as the
openai backend with ai-core's LlmInferenceService, which is
what ai-core's Agent chat, Code-Suggestions, Speech-to-Text
and Vector-Search actually talk to.
One backend, many servers. It speaks POST {baseUrl}/chat/completions, and
the base URL is a setting. Across OpenAI, Ollama, LM Studio, OpenRouter and
llama.cpp's llama-server the auth header, request JSON, SSE framing and error
shape are identical — only the host changes. So this is one backend with a URL
field rather than one plugin per provider:
| Base URL | What it is |
|---|---|
https://api.openai.com/v1 |
Default. OpenAI itself. |
http://localhost:11434/v1 |
Ollama on the device (e.g. in the bundled Termux). |
http://192.168.1.50:11434/v1 |
Ollama on the user's PC, over Wi-Fi. |
http://192.168.1.50:1234/v1 |
LM Studio's server. |
http://localhost:8080/v1 |
llama-server from llama.cpp. |
https://openrouter.ai/api/v1 |
OpenRouter — many models behind one key, some free. |
Calls the API directly over HttpURLConnection rather than an SDK: plugins run
in the host IDE's classloader, where okhttp3 resolves to the host's older
OkHttp, and an SDK bundling its own copy crashes generation with a
NoSuchMethodError.
Prerequisites: Android SDK (API 33+), JDK 17. Create local.properties with
sdk.dir=.... This plugin uses the shared wrapper at the repo root:
cd plugins/AI-Agent-OpenAI
../../gradlew assemblePlugin # release -> build/plugin/ai-agent-openai.cgp
../../gradlew assemblePluginDebug # debug variant
../../gradlew testDebugUnitTest # the JVM unit testsEverything is configured in AI Core → Agent settings, on the pane this plugin contributes: server URL (with presets), API key, model, and one Test Connection & List Models button. Nothing outside this plugin handles the key.
The pane adapts to the chosen server as it is picked, via
BaseUrlPolicy.keyRequirement(): REQUIRED for OpenAI's own host, EXPECTED for
another cloud provider, NOT_NEEDED for loopback or a private address — where the
key entry collapses to one muted line rather than showing an empty,
mandatory-looking field for a server that wants no credential. That happens whether
or not a key is already stored; a stored one leaves only Remove, so a key saved
for another server can still be cleared from here. Listing models and testing the connection
are the same GET {baseUrl}/models, so they are one control, and the model is a
single editable dropdown rather than a field beside a spinner.
Three rules that each break a real user if got wrong, and are covered by tests:
- The API key is optional.
isAvailable()requires a key only when the base URL is OpenAI's own host. For any other server a non-blank URL is enough — local Ollama and LM Studio need no credential, and demanding one would leave the backend permanently "not available" for exactly the users who wanted a custom server. - The model is a field, not a constant. Pointed at a local server the model
is whatever the user pulled (
qwen2.5-coder,llama3.2), so free-text entry always works andGET /v1/modelsis treated as optional — plenty of compatible servers do not implement it, which is why a 404 there reports "check the URL" rather than rejecting the key. - No auto-discovery. There is no probing of
localhost:11434; a background port scan is not something the user asked for. The URL field already reaches any server, on-device or on the LAN.
gpt-5.x and the o series reject max_tokens in favour of
max_completion_tokens, and several reject temperature. RequestTuning picks
the parameters from the model id and the server, and UnsupportedParameter reads
the offending name out of a 400 so the request is retried once without it —
compatible servers vary too much to hardcode a matrix.
https is required except for loopback and private ranges (RFC 1918, link-local,
IPv6 ULA, and bare LAN hostnames), where plain http is accepted and warned about
once on save. That is the "Ollama on my PC" case, and the host IDE's
network_security_config permits cleartext, so it works at runtime.
Stored encrypted (AES/GCM under a hardware-backed Android Keystore secret) and
sent as an Authorization: Bearer header, never in a URL query string. With
no key configured, no header is sent at all.
A key is bound to the server it was saved for. The base URL is recorded next to
the key (KEY_API_KEY_URL) and readApiKeyOrBlank() sends nothing when it does not
match the configured server's origin, so pointing the URL at a local or LAN address
after configuring OpenAI cannot put that bearer token on the network in the clear.
A key stored before the origin was recorded is still sent, since it cannot be shown
to belong elsewhere. The connection test applies the same rule.
security/SecureApiKeyStore.kt holds only this plugin's Keystore alias
(cotg_ai_openai_key_v1); the AES/GCM itself is the IDE's KeystoreSecretStore
(plugin-api, since 26.36 — hence this plugin's min_ide_version), so there
is one implementation in the process rather than a copy per plugin. The alias
is deliberately not shared: every plugin runs in the host app's process and UID
and therefore shares one Keystore, so a shared alias would let one plugin's
invalidated-key recovery (deleteEntry) destroy the other backend's stored key.
The plugins never read each other's ciphertext, so they have no reason to share
one.
Install ai-core as well — without the router this plugin has nothing to
register with. Order does not matter: this plugin re-registers when it sees
ai-core activate. Copy build/plugin/ai-agent-openai.cgp to the device, install
via CodeOnTheGo's Plugin Manager, then restart the IDE.
This backend declares ToolCallingBackend, so ai-core calls
generateStreamingWithTools and the agent's tools travel through the
Chat Completions function-calling API rather than a text envelope in the reply
(ADFA-5410). Declared tools arrive already structured, so a file whose contents
carry quotes or newlines cannot break the call.
- Request. Each tool is sent in
tools[]as{"type":"function","function":{name, description, parameters}}(OpenAiToolProtocol).tool_choiceis sent only when ai-core names a required tool (EXTRA_PARAM_REQUIRED_TOOL); a server that refuses it is retried once without it, under the same rule as the other 400 retries (RequestTuning). - Stream.
tool_callsdeltas are joined by index (SseChunk,OpenAiToolProtocol.CallAccumulator) and reported throughonToolCallonce the stream ends, so a call a retry replaced never reaches the caller. A call with no name or unparseable arguments is dropped. - Results.
ChatMessagegives an assistant turn no way to carrytool_calls, and atoolrole is only legal after one, so tool results go back asuserturns. - Prompt. ai-core sees the backend calls tools natively and passes no text call
syntax, so
tools.yml'snativeformat is sent. The text protocol, and itsno_native_channelline, is not used by this backend. - Servers without function calling. If a server refuses the
toolsdeclaration (400, 404 or 422 naming a tool field as unsupported), the turn is retried with no tools and a Toast says the agent cannot call tools on that server. The refusal is remembered per base URL, so only the first turn pays for it. The prompt for that run was built for native calling, so the model answers in prose.
The prompt this backend asks ai-core to send lives in src/main/assets/prompts/,
one YAML file per concern, apart from the code that sends it. Changing the tone,
adding a rule or translating the prompt is an edit to those files alone. ai-core
appends its own IDE CONTEXT block after the rendered prompt.
The files are loaded, validated and cached once, when the plugin is activated.
getSystemPrompt renders layout.yml from that cache for each request, since the
tool list, the protocol and the example path vary per run; it never waits. Until the
config has loaded, or if it cannot render, it returns null and ai-core sends its own
default prompt.
| File | Keys | What it is |
|---|---|---|
agent.yml |
schema_version, identity, include |
The entry point: the version (1; another is refused rather than misread), who the agent is, and the files below. |
scope.yml |
scope |
What the agent will answer: anything, with the project's tools only when the request is about the open project. |
rules.yml |
rules |
Rule groups, each a heading and its items; today one RULES group. Adding a rule is adding an item; a further group, e.g. by priority, renders as its own block. |
workflow.yml |
behavior, workflow |
How to go about building or changing something; the workflow's steps are numbered when rendered. |
tools.yml |
tools, tool_call_format |
What introduces the tool list, and how to call a tool: native under the function-calling API, text (with its examples) when calls travel in the reply. Exactly one is sent. |
layout.yml |
layout.system_prompt |
Where each text goes. |
The files, the names they are rendered under and the checks are AI-Agent-Gemini's
(see its README), and scope.yml and workflow.yml are identical to its copies;
edit the two plugins together. The one intended difference is
tool_call_format.text.no_native_channel, the line forbidding the provider's native
function-calling channel under the text protocol (TOOL_CALL_FORMAT_TEXT_NO_NATIVE_CHANNEL).
Rendering is strict: an unknown name throws, naming the text it was in, where the
file-per-section design this replaced dropped the file silently. Activation renders
the prompt for requests that open and close every section and logs any failure, and
OpenAiSystemPromptTest fails on one in the shipped files. A new key needs
OpenAiPromptConfig and its parser; a new name needs OpenAiPromptVariables.
The engine and the YAML plumbing (PromptTemplateEngine, PromptConfigLoader,
PromptConfigStore, PromptConfigObject, ...) are the IDE's, in plugin-api.jar's
com.itsaky.androidide.plugins.ai.prompt, shared with ai-core and the other backends.
Only OpenAiPromptConfig, its mapping in OpenAiPromptConfigParser, and sharedPromptConfig are this plugin's own.
Every source file sits in a package named for its layer; nothing is loose at the
root of com/itsaky/androidide/plugins/aiagentopenai/.
plugin/OpenAiPlugin.kt— plugin entry point; registers the backend with ai-corebackend/OpenAiBackend.kt— SSE streaming, native tool calling and the model catalogbackend/OpenAiHttpClient.kt— the HTTP transportbackend/OpenAiRequestBuilder.kt—messages[]mapping and request JSON, includingtools(pure)backend/OpenAiToolProtocol.kt—tools[],tool_choiceand the streamedtool_callsaccumulator (pure)backend/RequestTuning.kt— reasoning-model parameters, the 400-retry rule and tool-refusal detection (pure)backend/SseChunk.kt— one line of the token stream, text ortool_callsdeltas (pure)backend/ModelCatalogFilter.kt— splits one catalog into chat and embedding models (pure)backend/OpenAiEmbeddingProtocol.kt— the/v1/embeddingsbody, batching and index-ordered reply (pure)errors/OpenAiErrorFormatter.kt— turns a failure into one translated sentencesecurity/SecureApiKeyStore.kt— this plugin's Keystore alias, over the IDE'sKeystoreSecretStorepreferences/OpenAiPreferences.kt— this plugin's settings storeprompt/OpenAiSystemPrompt.kt— renderslayout.ymlfromOpenAiPromptVariables;prompt/config/mapsassets/prompts/onto this plugin's config type, which the IDE'sai.promptpackage loads, validates, caches and renderssettings/BaseUrlPolicy.kt— URL normalization and the cleartext rule (pure)settings/ServerPreset.kt— the one-tap server listsettings/ConnectionVerification.kt— what a live check established (pure)settings/— the pane this backend contributes to the selectorlogging/—LOG_PREFIX(AiAgentOpenAi), prefixing every logcat tag
The pure units carry the logic that would otherwise only fail on a device; they are covered by 245 JVM tests.
GPL-3.0 — same as AndroidIDE / CodeOnTheGo.