Provider Wire value Status Talyn Fleet selfhostedLive, and the DEFAULT. Firecracker microVMs on hardware we own, dispatched through the sandbox gateway. Runs on the workspace's OWN Claude or Codex subscription. PostHog Code posthog_codeLive. The fall-back: what a workspace not on the fleet allow-list runs on, and where a full fleet fails over to. Codex Cloud codex_cloudDeferred — OpenAI exposes no server-to-server cloud-task API. Note this is a different thing from running Codex on the fleet, which works today. Claude Codeclaude_codeRemoved (migration 0050). Anthropic Managed Agents billed metered API credits with no subscription option — the opposite of what the fleet offers — and the fleet runs Claude on the user's own subscription instead.One provider, two independently connectable credentials, and the model carries the vendor:
fleetProviderForModelreads the model id and the fleet builds the microVM's egress route table from it, so a run dispatched at a Codex model has no route toapi.anthropic.comat all. Picking an agent per task is therefore picking a model — there is no second field, because a second field is a second source of truth that can disagree with the first.
Claude —
claude setup-tokenyields a long-livedsk-ant-oat…, or a Console key for metered billing. Pasted.Codex — a ChatGPT-subscription token pair. The authorize leg cannot run on the backend: OpenAI publishes no third-party OAuth for subscription inference, and the only client the Codex backend accepts redirects to
http://localhost:1455/auth/callback. So the desktop runs the loopback PKCE flow in its main process (main/codexAuth.ts,originator=talyn, the same shape OpenCode uses) andapps/webpastes~/.codex/auth.json. The backend owns REFRESH only (services/selfHosted/codexOauth.ts, mirroringposthogCode/oauth.ts: in-process promise map + blocking advisory lock,invalid_grantterminal).Known risk, stated plainly: that flow reuses OpenAI's first-party Codex client id. It is established practice among third-party coding tools, but it is not a documented integration point — OpenAI can revoke the client or refuse unfamiliar originators, and that breaks every connected workspace at once. The fallback is to drop server-side refresh and prompt "Reconnect Codex" on the first
invalid_grant; measure the real access-token lifetime before choosing it.A dispatch always carries the workspace's own key for the vendor it is running, and suppresses the other with
policy.credentials. The sandbox gateway fills an absent or blankanthropicKey/openaiKeyfrom its own tenant's sealed custody, soopenaiKey: creds.openaiKey ?? ''was a silent route to spending somebody else's subscription. Nothing sits behind that door today (custody is only populated for GitHub-born tenants and ours is operator-minted) — which is a fact about one environment variable, not a property of the code. No credential for the model's vendor is a refusal, never a blank. Pinned byfleetCredentialCustody.test.ts.Three paths must agree on which vendor a run is spending, and
cloudTask.extra.llmis how:executor.tsrecords it,poller.tsrecredentialre-supplies it after a fleetd restart, andrunCredentials.tsserves it back to a host that asks. A row with nollmpredates the field and is an Anthropic run.
Status: Phases 1–2 shipped (June 2026, the cloud-only refactor). Goal: turn the one-off PostHog Code integration into a pluggable "cloud task provider" abstraction, then add OpenAI Codex Cloud and Claude Code Routines as two more delegators behind the same machinery.
What's done: the
CloudTaskProviderinterface + registry (services/cloudProviders/{types,registry,poller,environment}.ts), PostHog Code wrapped as the first provider (cloudProviders/posthog/provider.ts, delegating to the existingposthogCode/*executor/streamer/poller), the cloud-only task queue + generic poller, the neutralCloudTaskMetadata+readCloudTaskMetahelpers, and the generic/api/v1/cloud-providersroute. The whole local-execution layer (daemon, envs, agents, permissions, backlog) was removed in the same pass.Deviations from the plan below: (1)
dispatch(task, env)keeps itsenvparam — it's now the secret-free cloud-marker env (we chose to keep a thin env marker rather than moveprovideronto the task). (2) The deep streamer/poller generalisation intoTranscriptSource/TranscriptConverterwas deferred — with one provider, the PostHog poller/streamer are wrapped as-is; generalise them when Codex/Claude land. (3) The Settings/composer UI stays PostHog-specific until a 2nd provider exists. Phase 4 (Claude Code, Managed Agents) shipped (June 2026) —services/claudeCode/*
cloudProviders/claude/provider.ts, registered inindex.ts, with a genericCloudProviderCardSettings form. Phase 3 (Codex Cloud) is deferred — no server-to-server API (see the Phase 0 findings). Remaining: per-task provider picker in the composer;TranscriptSource/TranscriptConvertergeneralisation.
A posthog_code environment is a delegation marker: a task assigned to it
bypasses Owl's agent loop entirely. An executor kicks off a remote run, a
poller reconciles remote status → Owl task status, and a streamer +
converter ingest the remote transcript into task.transcript. Status/PR/
transcript flow back through Owl's normal task model.
Everything provider-specific lives under
packages/backend/src/services/posthogCode/ — client.ts, credentials.ts,
executor.ts, poller.ts, streamer.ts, acpConverter.ts. The frontend and
task model are already provider-agnostic (the transcript renderer keys off
task.transcript; the PR pill keys off task.metadata.pullRequest), so this work
is almost entirely backend plumbing.
Today the design is parallel-cloneable but not pluggable: dispatch is a hard
if (env.type === 'posthog_code') branch in taskQueue.ts, metadata is
posthog*-prefixed, and the converter reads PostHog-specific _meta. Going from
one provider to three justifies a small refactor first.
- Behaviour-preserving refactor first. Phase 1 moves PostHog Code under a provider interface with zero behaviour change and all existing tests green.
- One neutral target format. Providers differ only in (a) their API client,
(b) their wire→
AgentEventconverter, and (c) credential shape.AgentEventstays the universal transcript format. - Shared lifecycle, provider-specific edges. PR detection/linking, task
finalization, idle detection, and the streamer loop are generic. Each provider
only supplies
dispatch,reconcile(status mapping), and a converter. - De-risk the unknowns early. Each provider's "can we even create a cloud task + fetch its transcript over an API" is a spike (Phase 0) before we commit.
Feasibility hinges on each vendor exposing a programmatic create-cloud-task and fetch-transcript path. Confirm before building.
- Open question: the public
@openai/codex-sdk(codex.startThread().run()) appears to drive a local Codex process, not the hosted cloud sandbox. The cloud tasks (per-task sandbox, preloaded repo, proposes a PR) are surfaced via Codex Web and@codexGitHub mentions. Verify whether there is a REST/cloud task API to: create a cloud task against a repo + prompt, poll status, fetch the transcript, and read the resulting PR. - Fallbacks if no direct cloud API: (i) drive it via GitHub (
@codexmention on an issue/PR) and reconcile through our existing GitHub monitor; (ii) run the Codex SDK under Owl's local daemon instead (different model — see "Two models" below). Decide which is acceptable. - Deliverable: a throwaway script that creates a task and tails its output, or a written "not viable as cloud delegation yet → use GitHub/daemon path".
- Routines are a research preview (beta header
experimental-cc-routine-2026-04-01).POST …/routines/:id/firewith input text returns a session id + URL, runs on Anthropic's cloud, exposes results via webhook. - Open questions: (1) Can we create/run an ad-hoc prompt+repo per call, or must a routine be pre-created (prompt+repos+connectors saved up front)? If the latter, Owl's "arbitrary task" model maps awkwardly — we'd either create a routine per task (if the API allows) or require the user to bind an env to an existing routine. (2) How is the transcript retrieved — webhook payload, polling the session URL/API, or an SSE/stream endpoint? (3) Does a routine run open a PR, and is its URL in the result?
- Deliverable: a script that fires a routine and retrieves the transcript + result, plus a decision on the task→routine mapping.
Gate: only start Phase 3/4 for a provider once its spike passes. Phase 1–2 (the refactor) are worth doing regardless.
Ahead of running the spikes, web research against the vendors' official docs settled the two open questions and changed the Codex plan:
Codex Cloud → no server-to-server API; DEFERRED. OpenAI ships no public
REST/SDK endpoint to create a Codex Cloud task from a backend. The only
hosted-sandbox surfaces are the codex cloud CLI (exec/status/diff/
cancel, needs a self-hosted runner + pre-created opaque env IDs, unstable JSON
output — see openai/codex#24777) and @codex GitHub mentions. The
@openai/codex SDK and codex exec drive a local agent, not the cloud.
Decision: defer the Codex provider rather than re-introduce a runner or build
on the brittle CLI; revisit if OpenAI ships a scriptable cloud API. (The
GitHub-mention fallback remains a future option behind the same provider seam.)
Claude Code (web) → Managed Agents API; PROCEEDING. Anthropic exposes two
hosted surfaces: (a) the Routines API (POST /v1/claude_code/routines/:id/fire,
experimental beta, OAuth sk-ant-oat01-…) — fire-and-forget, returns only a
session id/URL, routines must be pre-created, no transcript/poll/cancel; and
(b) the Managed Agents API (POST /v1/agents + /v1/sessions, SSE transcript
via GET /v1/sessions/:id/stream, interrupt/archive to cancel, x-api-key,
beta header managed-agents-2026-04-01) — dynamic per-task session with a mounted
GitHub repo + prompt + model. We target Managed Agents for full PostHog-Code
parity; PR URL is parsed from the transcript (the agent prints it). Exact
payload/event shapes are unconfirmed — scripts/spikes/spike-claude.ts validates
them against a real account before the module is built.
scripts/spikes/spike-claude.ts (throwaway, git-ignored) exercised the full
lifecycle against a real account. Confirmed:
- Headers (all calls):
x-api-key,anthropic-version: 2023-06-01,anthropic-beta: managed-agents-2026-04-01,content-type: application/json. POST /v1/agents→ 200. Body:{ name, model, system, tools, mcp_servers }. Tooltypes areagent_toolset_20260401(prebuilt bash/read/write/edit/…, defaultspermission_policy.always_allow) andmcp_toolset({ type:'mcp_toolset', mcp_server_name:'github' }, defaultsalways_ask). GitHub MCP wired viamcp_servers:[{ type:'url', name:'github', url:'https://api.githubcopilot.com/mcp/' }]. Returnsagent_…id.POST /v1/environments→ 200 (env_…). The sandbox the session runs in (config.type:'cloud',networking:'unrestricted', has anenvironmentmap for env vars +init_script). Reusable per workspace.POST /v1/sessions→ 200 (sesn_…,status:'idle'). Body:{ agent:'agent_…', environment_id:'env_…', resources:[{ type:'github_repository', url, authorization_token, checkout? }] }. The repo mounts at/workspace/<repo>.checkoutis an OBJECT (string → 400 "value must be an object"); exact inner shape still TBD (omitting it defaults to the repo's default branch). The field isagent, notagent_id.- Prompt is NOT in the session create — send
POST /v1/sessions/{id}/eventswith{ events:[{ type:'user.message', content:[{ type:'text', text }] }] }. - Transcript = POLL
GET /v1/sessions/{id}/events?limit=100→{ data:[event…] }, oldest-first, dedup by eventid. (/events/streamonly replays the backlog then closes — not a long-lived tail, so the existing poll loop drives it.) Pagination for >100 events (anaftercursor) still to confirm. - Event taxonomy (semantic
event.type→ converter mapping toAgentEvent):user.message/agent.message{content:[{type:'text',text}]};agent.thinking;agent.tool_use{name, input, id, evaluated_permission};agent.tool_result{content:[{type:'text',text}], is_error, id};span.model_request_start|end(model_usagetoken counts → cost; otherwise ignorable);session.status_running|idle,session.thread_status_running|idle;session.error{error:{message,type,…}}(non-fatal unless terminal). - Terminal =
session.status_idlewithstop_reason.type:'end_turn'(sessions have no "completed" state — they sitidleafter finishing). Session GETstatusisrunning→idle. - Cancel =
POST …/events {events:[{type:'user.interrupt'}]}and/orDELETE /v1/sessions/{id}(→session_deleted). Both 200. - Cost available from
span.model_request_end.model_usagetoken counts.
PR-open path — CONFIRMED (write spike, real PR opened on owl). Plan B
(agent uses git/gh via bash) is dead: gh isn't installed in the sandbox,
the repo is served through a localhost git proxy, and the authorization_token
is not exposed to the agent's shell — the agent burned 19 tool calls hunting for
a credential and got 403s. The working path is the GitHub MCP + a vault:
POST /v1/vaults{ display_name, metadata }→{ id:'vlt_…', type:'vault' }.POST /v1/vaults/{id}/credentials{ display_name, auth:{ type:'static_bearer', mcp_server_url:'https://api.githubcopilot.com/mcp/', token:'<gh PAT>' } }→{ id:'vcrd_…', vault_id }. (Token is write-only; the storedmcp_server_urlis returned slash-normalized and still matches.)- Agent: set the github
mcp_toolsettodefault_config.permission_policy.always_allow(default isalways_ask, which would stall an unattended run). - Session: pass
vault_ids:[vaultId]— the credential is matched to the MCP by URL automatically. - The agent then opens the PR with the github MCP tools (
get_me,list_branches,create_branch,create_or_update_file,create_pull_request). The PR URL appears in theagent.mcp_tool_resultofcreate_pull_request(regexhttps://github\.com/[^/]+/[^/]+/pull/\d+over the event JSON catches it). - PAT scopes:
repo(classic) or fine-grainedcontents:rw+pull_requests:rw.
Remaining minor TBD: the checkout object shape (for mounting a PR's head
branch on pr_response/pr_review); code_writing works without it. For
PR-response tasks the agent can also operate on the existing branch via the MCP.
Verdict: Managed Agents gives full parity, including autonomous PRs. Provider
design — dispatch = ensure agent (toolset + github MCP always_allow) + ensure
environment + ensure vault (per workspace, reusable) → create session
(agent + environment_id + vault_ids + github_repository resource) → post
the prompt event; reconcile = poll GET /sessions/{id}/events, map → AgentEvent[],
terminal on session.status_idle/end_turn, detect the PR URL from the
create_pull_request agent.mcp_tool_result; cancel = user.interrupt + DELETE.
Introduce the interface + registry and move PostHog Code under it with no
behaviour change. New home: packages/backend/src/services/cloudProviders/.
export interface ReconcileOutcome {
status: 'in_progress' | 'completed' | 'failed' | 'cancelled';
prUrl?: string | null;
branch?: string | null;
error?: string | null;
logUrl?: string | null;
/** Provider-specific bits to merge into metadata.cloudTask.extra. */
extra?: Record<string, unknown>;
}
/** Normalises a provider's wire frames into Owl AgentEvents. */
export interface TranscriptConverter {
push(rawFrame: unknown): AgentEventInput[];
end(): AgentEventInput[];
}
/** Live tail + durable backfill for a provider's transcript. */
export interface TranscriptSource {
/** Async iterator of raw frames for a live run (SSE, websocket, poll loop). */
live?(signal: AbortSignal, cursor?: string): AsyncIterable<{ frame: unknown; cursor?: string }>;
/** One-shot durable fetch for terminal/finished runs. */
backfill?(): Promise<unknown[]>;
}
export interface CloudTaskProvider {
type: EnvironmentType; // 'posthog_code' | 'codex_cloud' | 'claude_code'
displayName: string;
defaultRenderer: EnvironmentRenderer; // 'structured'
/** Validate + persist credentials (Settings → Integrations). */
validateCredentials(workspaceId: string, input: unknown): Promise<{ ok: boolean; error?: string }>;
hasCredentials(workspaceId: string): Promise<boolean>;
/** Kick off a remote task; stamp metadata.cloudTask; flip task in_progress. */
dispatch(task: Task, env: Environment): Promise<{ ok: boolean; error?: string }>;
/** Read remote status for one in-progress task. */
reconcile(task: Task): Promise<ReconcileOutcome>;
/** Transcript ingestion. */
createConverter(): TranscriptConverter;
openTranscript(task: Task): Promise<TranscriptSource | null>;
}const providers = new Map<EnvironmentType, CloudTaskProvider>();
export function registerCloudProvider(p: CloudTaskProvider) { providers.set(p.type, p); }
export function getCloudProvider(type: EnvironmentType) { return providers.get(type) ?? null; }
export function isCloudEnv(type: EnvironmentType) { return providers.has(type); }
export function listCloudProviders() { return [...providers.values()]; }Register providers at boot in index.ts (alongside the existing poller init()).
- Relocate
services/posthogCode/→services/cloudProviders/posthog/(keep files). - Add
posthog/provider.tsimplementingCloudTaskProviderby delegating to the existingdispatchTaskToPostHogCode, the existing poller'sreconcilelogic (refactored to returnReconcileOutcome), and the existing ACP converter + SSE/session-logs streamer wrapped as aTranscriptSource.
Replace the special-case branch:
const provider = targetEnv && getCloudProvider(targetEnv.type);
if (provider) {
const r = await provider.dispatch(task, targetEnv);
if (!r.ok) { /* existing error handling */ }
continue;
}And in recoverStuckTasks (~L143–150), replace meta.posthogRunId with
meta.cloudTask?.remoteRunId (see Phase 2 metadata).
- Streamer (
cloudProviders/streamer.ts): generalise the currentposthogCode/streamer.tsto take aTranscriptSource+TranscriptConverterinstead of hard-coded SSE/session-logs. Keep the seq assignment, 25-event persist, truncation, reconnect-to-tail,flushNow, andseedTranscriptlogic. - Poller (
cloudProviders/poller.ts): generalise to: for each in-progress task withmetadata.cloudTask, look up the provider, callreconcile(), then run the shared finalize/PR-link/idle logic (liftlinkPr,findPullRequestUrl,maybeFinalizeIdle,finalizeout ofposthogCode/poller.tsinto shared helpers — they're already provider-neutral except PR-URL parsing, which is GitHub for all three).
Replace config.type === 'posthog_code' with isCloudEnv(config.type) →
{ success: true }.
Acceptance: existing prMonitorPoll, posthogCodePoller,
posthogCodeStreamer, posthogCodeAcpConverter, task-queue tests all pass
unchanged. No DB or wire change yet. PostHog Code behaves identically.
Add alongside (not replacing) PostHogCodeTaskMetadata:
export interface CloudTaskMetadata {
provider: EnvironmentType; // which provider owns this task
remoteTaskId: string;
remoteRunId?: string;
status?: string;
logUrl?: string;
prUrl?: string;
extra?: Record<string, unknown>; // provider-specific bag
}
// task.metadata.cloudTask?: CloudTaskMetadata- Back-compat: a read helper
readCloudTaskMeta(task)that prefersmetadata.cloudTaskand falls back to mapping legacyposthog*fields. All call sites (poller, streamer, routes, frontend) read through the helper. - Migration: a one-shot backfill (drizzle migration
00NN+ a small script, or lazy on first poll) writingcloudTaskfrom legacyposthog*for in-flight tasks. Completed tasks can stay legacy (read helper covers them).
integrationstable is already generic ((workspaceId, type)+ encryptedconfig). Add acloudProviders/credentials.tsbase:getIntegration(ws, type),storeIntegration(ws, type, config)with the existingtokenCryptoenvelope.- Each provider defines its own credential shape and
validateCredentials(PostHog: apiKey+projectId+host; Codex: apiKey/org + repo connection; Claude: OAuth/api key + beta header + routine binding).
GET /api/cloud-providers→ list registered providers + connected status.GET/PUT/DELETE/POST(test) /api/cloud-providers/:type/config→ delegate to the provider's credential methods; auto-provision the env on connect via a genericensureCloudEnvironment(workspaceId, type, defaults)(lift fromroutes/posthog.ts:ensurePostHogCodeEnvironment).- Keep
/api/posthog/*as thin aliases for one release, then migrate the frontend and delete.
EnvironmentType = 'local' | 'remote' | 'posthog_code' | 'codex_cloud' | 'claude_code'.- Add
CodexCloudEnvironmentConfig/ClaudeRoutineEnvironmentConfigto theEnvironmentConfigunion (model/adapter/routine-id defaults as needed).
- Settings → Integrations: render one panel per
GET /api/cloud-providersentry instead of a hard-coded PostHog panel. - Env picker / Add-Environment modal: list cloud providers generically; each connected provider exposes its auto-provisioned env.
- TaskComposer: drive model/adapter controls off provider capability metadata
(PostHog: model + reasoning effort; Codex: model; Claude: routine select) rather
than the
isCloudTaskspecial-case. - QueuePanel cloud banner: read provider
displayName+logUrlfrommetadata.cloudTaskinstead ofposthog*.
Acceptance: PostHog Code works end-to-end through the generic metadata/credentials/routes/UI. Legacy tasks still render. Tests updated to the neutral shapes.
Phase 0 found no server-to-server Codex Cloud API (only the
codex cloudCLI under a self-hosted runner, or@codexGitHub mentions). Deferred until OpenAI ships a scriptable cloud API, or we accept the GitHub-mention path behind this same provider seam. The sketch below is retained for that day.
services/cloudProviders/codex/:
client.ts— REST wrapper: create cloud task (repo + prompt + model), get status, fetch transcript (stream or poll), read PR. Auth via Codex API key/org.credentials.ts—type: 'codex'integration;validateCredentialspings.converter.ts— Codex event stream →AgentEvent[]. Codex emits a different shape than ACP (turn/message/tool-call events); map text → assistant/text, reasoning → thinking, tool calls → tool_use, results → tool_result, stderr/console → system. Emit livestream_eventdeltas like the ACP converter does.transcriptSource.ts—live()(SSE or poll loop) +backfill().provider.ts— implementsCloudTaskProvider;dispatchcreates the task- stamps
cloudTask;reconcilemaps Codex status →ReconcileOutcome+ PR URL.
- stamps
- Register in
index.ts; addcodex_cloudenv config; UI capability metadata (model select). - Tests: converter unit tests (fixture Codex frames → AgentEvents), a poller reconcile test (mocked client: in_progress → completed + PR), a dispatch test.
Implemented via Anthropic's Managed Agents API (not Routines — that path is subscription-billed but fire-and-forget / low-parity; see the Phase 0 findings and the billing analysis). Provider type
claude_code, displayName "Claude Code". Contract confirmed by the spike (see above).
services/claudeCode/ + services/cloudProviders/claude/provider.ts:
client.ts— Managed Agents wrapper: create agent / environment / vault (+ GitHub credential) / session, post the prompt event, poll the event list, get session, interrupt + delete.debugBusservice tagclaude_managed_agents.credentials.ts—type: 'claude_code'integration; encrypted Anthropic key + GitHub PAT; caches the reusable agent/environment/vault ids per workspace (cleared on credential rotation).converter.ts—managedAgentEventToAgentEvents(poll-based, complete events →AgentEvent[]; no chunk coalescing) +findPullRequestUrl/isTerminalEvent.executor.ts—dispatch: ensure resources → create session w/ repo resource + vault → post the prompt (starts the run) → stampcloudTask→in_progress.poller.ts—reconcile: poll events → transcript (persist + emit only what's new), detect+link the PR fromcreate_pull_request, finalize onsession.status_idle/end_turn;stopStreamingclears the in-memory cursor.provider.ts—CloudTaskProvider;cancel= interrupt + delete session.- Registered in
index.ts; DebugPanelSERVICE_INFO; genericCloudProviderCardSettings form (Anthropic key + GitHub PAT). - Tests:
claudeCodeConverter.test.ts(event mapping, PR detection, terminal),claudeCodeProvider.test.ts(conformance + registry). - Known follow-ups: per-task provider picker (PR-fix tasks currently prefer
PostHog, else Claude);
checkoutobject shape forpr_response/pr_reviewhead-branch mounting; executor/poller DB-mocked reconcile tests; reusing the workspace GitHub connection instead of a separate PAT (OAuth-token↔MCP compat unverified).
- Per-provider "Add Environment" cards with connect flows; status pills in Settings.
- Feature-flag each new provider (env var / settings) so they ship dark and enable per workspace.
- Docs: update
docs/ARCHITECTURE.md(provider abstraction),docs/ROADMAP.md(phase items),docs/SESSIONS.md(session note), and a user-facingdocs/CLOUD_PROVIDERS_USAGE.md.
PostHog Code originally authenticated with a personal API key the user pasted
into Settings, plus a project id they had to go and find. Both are still
supported, unchanged, forever-ish: a workspace connected that way is on
authMethod: 'personal_api_key' (or has no authMethod at all, which reads the
same — every install predating this change), and nothing migrates it or nags it.
New connections default to OAuth, because PostHog is a full OAuth2/OIDC
authorization server whose tasks API accepts pha_ bearer tokens with exactly
the same scope and per-project enforcement as a personal API key
(posthog/permissions.py treats the two identically). What that buys:
- No credential passes through Talyn's UI. The user authorizes on PostHog.
- The grant is narrow:
openid task:read task:write, on ONE project. - The project id stops being a form field.
required_access_level=projectmakes PostHog's consent screen render its single-project picker, and self-introspection (RFC 7662, allowed with nointrospectionscope when a token introspects itself) reports the chosen team back asscoped_teams. - Revocable from PostHog (Connected Apps), not just from here.
client_id is https://www.talyn.dev/oauth-client — a document we serve
(apps/marketing/app/oauth-client/route.ts) that PostHog fetches during the flow
(draft-ietf-oauth-client-id-metadata-document). No registration, no client secret;
Talyn is a public client and PKCE is the protection (PostHog requires PKCE of
every client, and prod advertises no private_key_jwt). Setup + the three ways to
misconfigure the pair are in docs/SETUP.md §6b.
POSTHOG_OAUTH_CLIENT_ID accepts an opaque id as well as a URL, because CIMD is
not the only way PostHog issues one: POST /oauth/register (RFC 7591 DCR) hands
back an id and hosts nothing. That is what lets a local backend run this flow —
register a dev client against http://localhost:4747/... (PostHog allows http for
loopback hosts only) rather than pointing dev at the production client, whose
document lists only the prod callback. Adding localhost to the published
document would widen the production client for every user; don't. Recipe in
docs/SETUP.md §6b.
PostHog rotates the refresh token on every use and enforces reuse protection with a 120-second grace — past which presenting a spent refresh token revokes the entire token family. Talyn calls this API from a poll loop, a streamer and the dispatcher, on two instances during every deploy, so an unguarded refresh doesn't fail: it logs the workspace out of a connection nobody touched.
Don't trust a number for the access-token lifetime. ACCESS_TOKEN_EXPIRE_SECONDS
in the PostHog repo says 1 hour, but us.posthog.com issued a 7-day access token
on 2026-08-06 (measured on the live grant). Nothing in the implementation assumes
either: expiresAt is computed from each response's expires_in. The practical
consequence is that on Cloud the refresh path runs roughly weekly, so a newly
connected workspace won't exercise it for days — which is why it's worth having the
unit tests rather than waiting to see it work in production.
services/posthogCode/oauth.ts single-flights refreshes twice over:
- an in-process promise map collapses the common case (many callers, one instance) with no round-trip, and
- a blocking Postgres advisory lock (
posthog-oauth-refresh:<workspaceId>) covers the cross-instance overlap. The winner's tokens are re-read inside the lock, so the loser uses the rotated pair rather than replaying a spent one.
Refreshing is lazy — there is no timer: whichever caller next asks for a token
and finds it within 5 minutes of expiry does the refresh. The skew doesn't take the
refresh off the request path, it hands it to an earlier caller with slack rather
than the unlucky one that would meet an expired token — and in practice that caller
is a background poll tick, since the pollers ask far more often than anything
user-facing. Nothing keeps a token warm, and it doesn't need to: the refresh window
is rolling, not absolute. Rotation pairs each new refresh token with a new
access token, validate_refresh_token enforces no expiry of its own, and PostHog's
daily sweep (products/growth/dags/oauth.py) only deletes an unrevoked refresh
token once its paired access token expired more than `REFRESH_TOKEN_EXPIRE_SECONDS
- OAUTH_EXPIRED_TOKEN_RETENTION_PERIOD
ago — 30d + 30d today. So every refresh buys roughly another two months, and an actively-used connection never lapses. We depend on none of those numbers:invalid_grantis terminal whenever it arrives, whatever the cause. A401on a token we believed valid forces exactly one refresh-and-retry (client.ts). Aninvalid_grantis terminal — it setsoauth.reauthRequiredAt` on the integration row so every surface agrees, the UI says "Reconnect needed", and nothing keeps hammering a grant that cannot come back. A 5xx or a network blip explicitly does NOT set that flag, and a preemptive refresh that fails transiently falls back to the access token still in hand (it has minutes left, by construction) rather than failing the caller — excluded on a forced refresh, where the API has just rejected that very token, and on a terminal error.
PostHog Desktop (PostHog/code) solves the same problem the same way, which is
a useful cross-check: AuthService.ensureValidSession dedupes onto one
refreshPromise assigned synchronously (their comment: "Resolving the stored
session must happen INSIDE refreshAndSync, else two callers both refresh and burn
the rotating token twice"), retries once on 401/403, and falls back to the current
token when a preemptive refresh fails. It differs where the deployment differs: a
1-minute skew instead of 5, no advisory lock (one process, no replicas), tokens
encrypted at rest under a machine-derived key (scrypt over
machineId|platform|arch, AES-256-GCM — deliberately not the OS keychain, so there
are no prompts and a cloud-synced backup is useless) instead of our
TALYN_TOKEN_KEY envelope, and an OAUTH_SCOPE_VERSION on the stored session that
forces re-consent when the app's scope set changes. That last one is worth copying
if we ever widen the scope ask; today a stored grant records its scope, but
nothing compares it.
| File | Role |
|---|---|
services/posthogCode/oauthConfig.ts |
The env pair + the feature gate (isPostHogOAuthEnabled) |
services/posthogCode/oauth.ts |
PKCE, authorize URL, code exchange, project resolution, refresh |
services/posthogCode/integrationRow.ts |
The posthog integration row's shape, read + upsert, shared by both auth paths |
services/posthogCode/credentials.ts |
One getToken() for both paths, so nothing downstream branches on auth method |
routes/posthog.ts |
POST /oauth/start (authenticated) + GET /oauth/callback (public, state-authenticated) |
db/migrations/0040_posthog_oauth_states.sql |
In-flight PKCE states — in Postgres, not a Map, because the callback can land on the other instance mid-deploy |
__tests__/posthogOauth.test.ts |
The lifecycle, incl. the concurrent-refresh case |
Deliberately not done: no project:read in the scope ask, so a grant that
covers a whole organization (or several projects) is refused at connect time with
a message rather than guessed at. If that turns out to be common, adding the scope
and a project picker is the fix — not silently filing tasks into the wrong project.
This plan covers cloud delegation (vendor hosts the sandbox + agent loop; Owl
kicks off + reconciles). The alternative is local execution via Owl's daemon
(Owl runs the claude/codex CLI itself). Owl already runs Claude locally; adding
the Codex CLI there is a runtime adapter on the agent service, not a cloud
provider, and is out of scope here. If 0a shows Codex has no cloud task API, the
local-daemon path is the pragmatic fallback for Codex.
- Codex cloud API existence (0a) — biggest unknown; may force GitHub-mention or local-daemon fallback.
- Routine ad-hoc vs pre-created (0b) — shapes the UX and whether "any task" is possible.
- Transcript retrieval variance — SSE (PostHog) vs webhook (Claude) vs poll
(Codex?) — handled by the
TranscriptSourceabstraction but each needs care (ordering, dedup, backfill). - PR creation — PostHog/Codex open PRs themselves; if Claude routines don't,
reuse Owl's
taskPullRequestto open from the produced branch. - Webhook reachability — fine on Railway; document a tunnel for local dev.
- Metadata migration — keep the read-through helper until all legacy tasks age
out; never hard-cut
posthog*.
| Phase | Scope | Est. |
|---|---|---|
| 0 | API spikes (Codex, Claude) | ~1–2 days |
| 1 | Provider interface + registry; move PostHog under it (no behaviour change) | ~1–2 days |
| 2 | Neutral metadata/credentials/routes/env-types + frontend plumbing | ~1–2 days |
| 3 | Codex Cloud provider | ~2–3 days |
| 4 | Claude Routines provider | ~2–3 days |
| 5 | UX polish, flags, docs | ~1 day |
After Phases 1–2, each new provider is a self-contained ~5–6 file module
(client + credentials + converter + transcriptSource + provider
- tests) with no changes to the core.