You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Tracking issue for self-optimizing workflows: PostHog's scout suggests a change to a workflow, a person approves it, and the change reaches people only through the normal publish. Design in the internal RFC requests-for-comments-internal#1108; this is its first phase.
Everything is behind the self-optimising-workflows flag, and off per workflow until the owner turns "Suggest improvements" on. Customer docs: Workflow suggestions, PostHog/posthog.com#19863, draft until rollout.
Goal
Close the loop a workflow owner never closes today: notice an email step performing badly, change it, and see whether the change helped. Suggestions only. No A/B testing, no auto-apply, and only PostHog's own scout files them; a person can only approve or reject.
The stack
Merge the whole stack with /trunk merge on the top PR. Each PR targets the branch below it so its diff shows one layer; nothing merges into a branch.
#92252 is closed. It carried the scout, the outcome rework and the evidence surfaces in one 77-file diff that rewrote components #91710 had just added, so it was split into #107879, #107795 and #107802. GitHub refuses to retarget a stacked PR's base, so the scout moved to a new PR rather than being retargeted in place.
created_via, created_by, source_type and resolution_note are gone from the model; source_id stays. hog_flow_proposal is an internal scope object, so only the scout can file.
Decisions that matter for review
One model holds a suggestion. It stores only the fields the change touches, who filed it, and its status through the queue. No second draft table. Fields exist only for what runs today; a new producer or surface adds columns when it arrives.
Approving stages a draft, never live config. Approve writes the suggestion into the workflow's draft, the same move as restoring a revision. Publish ships it and records the version. Discarding the draft, restoring a revision, or approving another suggestion returns the earlier one to the queue, so publish cannot mark as applied something that never shipped.
A suggestion carries only what it changes. Steps are keyed by id and merged field by field into the live step; an unrelated edit made while the suggestion waited survives approval. A suggestion is read against the version it names (base_version, required), so a step sent whole still merges as the fields that differ. Approve is refused (409) only when one of those fields moved to a value other than the proposed one; a draft edit hands an approved suggestion back only if it takes that change out of the draft. edges and variables replace the whole list, so for those any publish since counts. Create refuses a change publish would refuse.
Provenance is server-derived.created_via comes from the request (a Signals run records self_driving), so the "Suggested by PostHog" label cannot be set by a caller. source_id names the run, so a retry returns the row it already made.
Suggesting has its own scope, held only by the scout.hog_flow_proposal:write can file; publishing, editing and test-sending still need hog_flow:write, which no scout token carries. The scout declares the scope in its SKILL.md and the harness seeds it. It is not offered on personal API keys, but the server does not yet refuse it on one; making it server-minted only is listed under Later.
The numbers a person judges are PostHog's, not the scout's. When a suggestion is filed, the server reads the step's own metrics at base_version and stores that reading beside the scout's evidence; the card shows the reading, labeled "Measured by PostHog", and says so when the scout's number disagrees. The scout's evidence still has to carry a unit, a denominator and counter-metrics, or create refuses it. Open and click rates divide by tracked sends; under 20 observations the card says so instead of presenting a result.
Per-version metrics are how "did it help" is answered.metrics/totals?version=<n> reads one published version's series. A version is not a subset of the unversioned read (batch and broadcast runs key on the run), so versions are compared with each other. The scout waits two days of opens before judging a version.
The suggestion names the metric it aims at.evidence.metric picks the target from TARGET_METRICS (opens, clicks, bounces, complaints) and the outcome reads that one; anything else falls back to the open rate. Opens and clicks divide by tracked sends, the counter-metrics by all sends.
The opt-in is a row per workflow; turning it off keeps the row. Only a live workflow can be opted in or suggested against (a draft has no sends, an archived one is done), so the Suggestions tab stays on every saved workflow for discoverability, but its switch is disabled with "Suggestions need a live workflow. Enable it first." until the workflow is active. The server refuses a suggestion for a workflow that is off; the scout's work list is the workflows that are on. Its read scope still covers the whole project, so "reads only opted-in workflows" is a skill rule, not a server one. "Tried it and turned it off" is a rollout question. Cadence lives with the scout config, not on the row.
The scout files no inbox report and emits no signal. An actionable report can dispatch a code run and open a billable pull request; this change is configuration, not code. Its only output is the suggestion on the workflow.
Rollout
Before merging, add signals-scout-workflows to the signals-scout flag's withheld_skills, so the scout does not start on every Signals-enrolled team when the stack lands; release it per team as the workflows flag rolls out.
Billing: Signals bills per implemented report; this scout files none, so today it costs the customer nothing and PostHog inference. RFC open question 3. Until decided, no surface says it uses credits.
A run on a cloud project on the signals-scout allowlist, and a run the coordinator's own schedule fires. Everything so far ran in a local sandbox against the local stack.
After it ships, adoption reads from the hog_flow_optimization_enabled / _disabled and hog_flow_proposal_approved / _rejected events, and the suggestions themselves from the WorkflowProposal rows.
Later
The list's suggestion tag is read-only, so the one-click jump to the Self-driving tab is gone. An interactive element inside the row's anchor is unreachable by keyboard; bringing the shortcut back means rendering the tag outside LemonTableLink.
The pending and approved queues read 100 rows rather than following pagination, and say so when the count is higher. Reaching 100 needs roughly three months of a daily scout with nothing resolved, and only one suggestion can sit approved at a time, so paging waits until the queue can realistically fill.
Show the version's age on the card ("v5, live 3 hours") from its revision row, so a suggestion filed against a young version is visible as such rather than trusted to the scout's two-day rule.
Desktop delivery: a report kind the harness pins as non-implementable, carrying the proposal id with approve and reject actions. The model stays the record.
Say something when the scout cannot judge a workflow: tracking off, too few sends, metrics missing. Today that goes only to its scratchpad and the run record, so nobody hears that the thing blocking a suggestion is fixable.
Notify when a suggestion lands, rather than waiting for someone to open the workflow. In-app first, through the notifications facade, to everyone who can edit the workflow. Desktop delivery above is the step after that.
Unsubscribe rate as a counter-metric, once something emits email_unsubscribed.
The card charts a metric per version, not two arms, so a move is still not proof that the suggestion caused it. Deciding a winner is the RFC's A/B step.
Conversions as the headline metric, once conversion rates are stored per version and step. That is a metrics change rather than a proposal one.
One metric over time for the whole workflow, with a marker on every published version rather than only the ones around a suggestion, so a person can see what drove what.
An enable action on the workflow itself for workflows nobody opted in yet. The tab and the list already say when it is on; what is missing is the invitation when it is off.
Classify the secrets of a step a suggestion adds. strip_proposal_secrets reads the live step a patch patches, so a patch that adds a step carrying inputs but no type or template_id has nothing to classify against and a secret in it would be stored in plaintext. Narrow, since such a step cannot run, but it is the same class as the patch case.
A suggestion cannot remove a step. actions merges by id and never drops one, so "drop this email, it opens at 2%" is unexpressible: dropping only its edges leaves the step in the graph, and validate_graph treats an unreachable step as a warning rather than an error. It needs a removal marker the merge understands, plus the outcome card treating a removal as a change like any other.
More than one suggestion at a time. The model allows it, but approving a second one rebuilds the draft from live content and drops the first. Merging onto the staged draft instead would let two suggestions on different steps stack into one publish.
Feed the outcome back into the next scout run, so a change that made things worse can be suggested as a revert. Andy's report checks are the natural home for this.
Broadcasts: their per-version metrics already aggregate across batch runs, but a one-shot send has nothing to improve, so the work list should skip them or take only recurring ones.
Not in scope
The RFC's later phases: A/B testing between variants, auto-apply of winners, personalisation, structural changes to a workflow, and cross-workflow analysis.
Tracking issue for self-optimizing workflows: PostHog's scout suggests a change to a workflow, a person approves it, and the change reaches people only through the normal publish. Design in the internal RFC requests-for-comments-internal#1108; this is its first phase.
Everything is behind the
self-optimising-workflowsflag, and off per workflow until the owner turns "Suggest improvements" on. Customer docs: Workflow suggestions, PostHog/posthog.com#19863, draft until rollout.Goal
Close the loop a workflow owner never closes today: notice an email step performing badly, change it, and see whether the change helped. Suggestions only. No A/B testing, no auto-apply, and only PostHog's own scout files them; a person can only approve or reject.
The stack
Merge the whole stack with
/trunk mergeon the top PR. Each PR targets the branch below it so its diff shows one layer; nothing merges into a branch.WorkflowProposal: the record, tenant-scoped, with the migrationworkflows-suggest,workflows-list-proposals)signals-scout-workflows, the scout itself#92252 is closed. It carried the scout, the outcome rework and the evidence surfaces in one 77-file diff that rewrote components #91710 had just added, so it was split into #107879, #107795 and #107802. GitHub refuses to retarget a stacked PR's base, so the scout moved to a new PR rather than being retargeted in place.
created_via,created_by,source_typeandresolution_noteare gone from the model;source_idstays.hog_flow_proposalis an internal scope object, so only the scout can file.Decisions that matter for review
idand merged field by field into the live step; an unrelated edit made while the suggestion waited survives approval. A suggestion is read against the version it names (base_version, required), so a step sent whole still merges as the fields that differ. Approve is refused (409) only when one of those fields moved to a value other than the proposed one; a draft edit hands an approved suggestion back only if it takes that change out of the draft.edgesandvariablesreplace the whole list, so for those any publish since counts. Create refuses a change publish would refuse.created_viacomes from the request (a Signals run recordsself_driving), so the "Suggested by PostHog" label cannot be set by a caller.source_idnames the run, so a retry returns the row it already made.hog_flow_proposal:writecan file; publishing, editing and test-sending still needhog_flow:write, which no scout token carries. The scout declares the scope in itsSKILL.mdand the harness seeds it. It is not offered on personal API keys, but the server does not yet refuse it on one; making it server-minted only is listed under Later.base_versionand stores that reading beside the scout's evidence; the card shows the reading, labeled "Measured by PostHog", and says so when the scout's number disagrees. The scout's evidence still has to carry a unit, a denominator and counter-metrics, or create refuses it. Open and click rates divide by tracked sends; under 20 observations the card says so instead of presenting a result.metrics/totals?version=<n>reads one published version's series. A version is not a subset of the unversioned read (batch and broadcast runs key on the run), so versions are compared with each other. The scout waits two days of opens before judging a version.evidence.metricpicks the target fromTARGET_METRICS(opens, clicks, bounces, complaints) and the outcome reads that one; anything else falls back to the open rate. Opens and clicks divide by tracked sends, the counter-metrics by all sends.Rollout
signals-scout-workflowsto thesignals-scoutflag'swithheld_skills, so the scout does not start on every Signals-enrolled team when the stack lands; release it per team as the workflows flag rolls out.signals-scoutallowlist, and a run the coordinator's own schedule fires. Everything so far ran in a local sandbox against the local stack.After it ships, adoption reads from the
hog_flow_optimization_enabled/_disabledandhog_flow_proposal_approved/_rejectedevents, and the suggestions themselves from theWorkflowProposalrows.Later
The list's suggestion tag is read-only, so the one-click jump to the Self-driving tab is gone. An interactive element inside the row's anchor is unreachable by keyboard; bringing the shortcut back means rendering the tag outside
LemonTableLink.The pending and approved queues read 100 rows rather than following pagination, and say so when the count is higher. Reaching 100 needs roughly three months of a daily scout with nothing resolved, and only one suggestion can sit approved at a time, so paging waits until the queue can realistically fill.
Show the version's age on the card ("v5, live 3 hours") from its revision row, so a suggestion filed against a young version is visible as such rather than trusted to the scout's two-day rule.
Desktop delivery: a report kind the harness pins as non-implementable, carrying the proposal id with approve and reject actions. The model stays the record.
Say something when the scout cannot judge a workflow: tracking off, too few sends, metrics missing. Today that goes only to its scratchpad and the run record, so nobody hears that the thing blocking a suggestion is fixable.
Notify when a suggestion lands, rather than waiting for someone to open the workflow. In-app first, through the notifications facade, to everyone who can edit the workflow. Desktop delivery above is the step after that.
Unsubscribe rate as a counter-metric, once something emits
email_unsubscribed.Engagement splits by version only for sends made after feat(workflows): emit the versioned email tracking code #91487; older versions read
n=0, which the sample floor labels.The card charts a metric per version, not two arms, so a move is still not proof that the suggestion caused it. Deciding a winner is the RFC's A/B step.
Conversions as the headline metric, once conversion rates are stored per version and step. That is a metrics change rather than a proposal one.
One metric over time for the whole workflow, with a marker on every published version rather than only the ones around a suggestion, so a person can see what drove what.
An enable action on the workflow itself for workflows nobody opted in yet. The tab and the list already say when it is on; what is missing is the invitation when it is off.
Classify the secrets of a step a suggestion adds.
strip_proposal_secretsreads the live step a patch patches, so a patch that adds a step carrying inputs but notypeortemplate_idhas nothing to classify against and a secret in it would be stored in plaintext. Narrow, since such a step cannot run, but it is the same class as the patch case.A suggestion cannot remove a step.
actionsmerges by id and never drops one, so "drop this email, it opens at 2%" is unexpressible: dropping only its edges leaves the step in the graph, andvalidate_graphtreats an unreachable step as a warning rather than an error. It needs a removal marker the merge understands, plus the outcome card treating a removal as a change like any other.More than one suggestion at a time. The model allows it, but approving a second one rebuilds the draft from live content and drops the first. Merging onto the staged draft instead would let two suggestions on different steps stack into one publish.
Feed the outcome back into the next scout run, so a change that made things worse can be suggested as a revert. Andy's report checks are the natural home for this.
Broadcasts: their per-version metrics already aggregate across batch runs, but a one-shot send has nothing to improve, so the work list should skip them or take only recurring ones.
Not in scope
The RFC's later phases: A/B testing between variants, auto-apply of winners, personalisation, structural changes to a workflow, and cross-workflow analysis.