Skip to content

fix(signals): give operational scouts their own dispatch budget - #107613

Open
posthog[bot] wants to merge 4 commits into
masterfrom
posthog-self-driving/fixsignals-stop-resumed-inbox-eafc82
Open

posthog[bot] wants to merge 4 commits into
masterfrom
posthog-self-driving/fixsignals-stop-resumed-inbox-eafc82

Conversation

@posthog

@posthog posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • Product scouts (general, product analytics, web analytics) ran at about half their usual hourly volume after #106933 shipped. As a result, projects get fewer scout findings.
  • fix(signals): resume operational scouts on wildcard-enrolled teams #106933 resumes signals-scout-inbox-validation on every wildcard-enrolled project that has a seed-disabled row. Every resumed row has never run, so _overdue_seconds returns inf.
  • _allocate_tick_budget fills one shared MAX_RUNS_PER_TICK (1000 per 30-minute tick) most-overdue-team first. Each resumed team's first pick is inbox validation, so these runs take the tick before product scouts get a slot.
  • The spike is mostly a backlog. Operational scouts use the model default interval (24 hours), so volume drops after each row runs once. During the drain, product scouts are deferred.

Origin

  • Product analytics
  • First signal: 2026-09-28
  • Inbox report: open
  • Likely cause: 67b63fe
  • Task started by: auto-start, after the report was rated P2 and ready to fix

Changes

  • Separate budget. Planning splits due runs into product and operational pools (is_operational_scout). Each pool goes through _allocate_tick_budget with its own global cap.
    • Product scouts keep all of MAX_RUNS_PER_TICK, so operational runs never defer them.
    • Operational scouts get MAX_OPERATIONAL_RUNS_PER_TICK = 200. A wave such as this one drains at a bounded rate instead of in a burst.
  • Flag-tunable. max_operational_runs_per_tick_global in the signals-scout payload overrides the operational cap without a deploy. Invalid values fall back to the default, as max_runs_per_tick_global does.
  • Per-team tick caps and the daily budget count both pools together, so a team's bounds still hold.
  • Only the harness-seeded canonical operational skill takes the operational budget. A team's own scout with the same name stays a product scout.
  • scout_harness/AGENTS.md documents the new key.
Before After
Product scout slots per tick 1000 minus operational due 1000
Operational slots per tick up to 1000 200 (flag-tunable)

Note

200 per tick is a guess at a safe ceiling. The intended steady-state volume of inbox validation is not confirmed yet. If report checks lag, raise the flag key. If compute is the concern, lower it.

Out of scope: the edit-rate alert on insight 8KWJNZMK still counts operational runs in series C. That needs an insight edit in PostHog, not a code change.

How did you test this code?

  • Added test_operational_scouts_have_their_own_tick_budget. It covers this regression: never-run operational lanes on two teams take the whole global cap from a due product scout. The test fails with the split reverted and passes with it.
  • Added test_per_team_caps_count_both_budgets. It checks that a team's per-tick and daily caps still hold when the team has runs in both pools.
  • Added test_resolve_global_max_operational_runs_per_tick. It checks that the key parses, that invalid values fall back, and that the product key does not widen the operational cap.
  • Ran hogli test products/signals/backend/test/test_scout_coordinator.py locally. All tests pass.
  • Not checked: production volume after deploy. Watch hourly scout runs by skill: general scout runs should go back toward their earlier baseline, and inbox validation should stay at or below about 400 an hour.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

None. Internal dispatch limits only.

🤖 Agent context

Autonomy: Fully autonomous

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5)

  • Started from a Self-driving inbox report. Skills used: /writing-pr-descriptions.
  • Duplicate search: no open PR changes the scout dispatch budget.
  • Rejected option: allocate operational runs only from the budget that product scouts leave. When product scouts fill the cap, operational runs would starve, and the report checks on them would stop.
  • Public artifact: the diff and this body contain no session material.

Created with PostHog Desktop from this inbox report.

🤖 Generated with Claude Code

Operational scouts now draw from a separate per-tick ceiling (MAX_OPERATIONAL_RUNS_PER_TICK, flag key max_operational_runs_per_tick_global). A wave of never-run operational lanes can no longer take the global slots of product scouts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
@posthog posthog Bot added the self-driving label Sep 28, 2026
@posthog
posthog Bot marked this pull request as ready for review September 28, 2026 11:21
@trunk-io

trunk-io Bot commented Sep 28, 2026

Copy link
Copy Markdown

Merging to master in this repository is managed by Trunk.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

@posthog

posthog Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

🦔 PostHog Review reviewed this pull request

Found 0 must fix, 0 should fix, 1 consider.

Published 1 finding (view the review).

Resolved comments: 2 already settled

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

🚨 Comment density — 11% of added code lines are comments (22 of 197)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/signals/backend/temporal/agentic/scout_coordinator.py 12 81
products/signals/backend/test/test_scout_coordinator.py 6 89
products/signals/backend/scout_harness/team_limits.py 4 13

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Backend coverage — 98.0% of changed backend lines covered — 1 uncovered

🧪 Backend test coverage

Patch coverage — changed backend lines (products + core): ████████████████████ 98.0% (79 / 80)

File Patch Uncovered changed lines
products/signals/backend/scout_harness/config_registry.py 80.0% 476

🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 117880218227538 -n patch-coverage), or the coverage-data block at the end of this comment.

Per-product line coverage (touched products)
Product Coverage Lines
platform_features ██░░░░░░░░░░░░░░░░░░ 12.1% 7 / 58
warehouse_sources_queue ██████░░░░░░░░░░░░░░ 29.1% 92 / 316
demo ████████████░░░░░░░░ 57.8% 1,545 / 2,673
data_tools ████████████░░░░░░░░ 61.2% 90 / 147
aeo ██████████████░░░░░░ 70.5% 467 / 662
ai_gateway ███████████████░░░░░ 75.0% 9 / 12
batch_exports ████████████████░░░░ 81.2% 21,459 / 26,431
apm █████████████████░░░ 84.1% 1,306 / 1,553
ml_inference █████████████████░░░ 87.2% 482 / 553
cdp ██████████████████░░ 88.2% 4,548 / 5,155
mcp_analytics ██████████████████░░ 88.9% 4,910 / 5,523
product_tours ██████████████████░░ 89.3% 1,331 / 1,491
dashboards ██████████████████░░ 89.5% 6,839 / 7,641
signals ██████████████████░░ 89.9% 54,713 / 60,868
data_warehouse ██████████████████░░ 89.9% 13,912 / 15,470
notebooks ██████████████████░░ 90.2% 15,287 / 16,945
cohorts ██████████████████░░ 90.4% 8,420 / 9,316
streamlit_apps ██████████████████░░ 90.7% 2,625 / 2,895
managed_warehouse ██████████████████░░ 90.9% 10,215 / 11,234
tasks ██████████████████░░ 91.1% 73,993 / 81,212
data_modeling ██████████████████░░ 91.5% 10,525 / 11,498
business_knowledge ██████████████████░░ 91.6% 6,899 / 7,528
engineering_analytics ██████████████████░░ 91.7% 11,017 / 12,014
exports ██████████████████░░ 91.8% 9,685 / 10,555
ai_training ██████████████████░░ 92.2% 356 / 386
conversations ███████████████████░ 92.5% 28,726 / 31,047
early_access_features ███████████████████░ 92.6% 1,341 / 1,448
managed_migrations ███████████████████░ 92.7% 1,581 / 1,705
visual_review ███████████████████░ 92.8% 9,244 / 9,966
canvas ███████████████████░ 92.8% 6,877 / 7,409
approvals ███████████████████░ 93.0% 3,919 / 4,214
mcp_registry ███████████████████░ 93.1% 1,670 / 1,794
error_tracking ███████████████████░ 93.1% 15,843 / 17,010
notifications ███████████████████░ 93.2% 1,145 / 1,229
slack_app ███████████████████░ 93.2% 13,677 / 14,674
stamphog ███████████████████░ 93.2% 7,885 / 8,456
surveys ███████████████████░ 93.3% 6,571 / 7,040
context_layer ███████████████████░ 93.8% 3,373 / 3,595
web_analytics ███████████████████░ 93.9% 21,653 / 23,051
alerts ███████████████████░ 94.0% 8,541 / 9,082
billing_alerts ███████████████████░ 94.1% 2,094 / 2,226
mcp_store ███████████████████░ 94.4% 8,940 / 9,472
ai_observability ███████████████████░ 94.4% 22,534 / 23,870
wizard ███████████████████░ 94.7% 6,151 / 6,496
reminders ███████████████████░ 94.8% 760 / 802
workflows ███████████████████░ 94.9% 14,335 / 15,113
review_hog ███████████████████░ 94.9% 11,490 / 12,109
annotations ███████████████████░ 95.1% 817 / 859
endpoints ███████████████████░ 95.1% 9,211 / 9,681
customer_analytics ███████████████████░ 95.2% 24,899 / 26,167
legal_documents ███████████████████░ 95.2% 2,311 / 2,427
marketing_analytics ███████████████████░ 95.3% 19,216 / 20,161
posthog_ai ███████████████████░ 95.4% 2,489 / 2,610
experiments ███████████████████░ 95.4% 32,645 / 34,211
growth ███████████████████░ 95.4% 9,812 / 10,282
logs ███████████████████░ 95.4% 15,290 / 16,022
actions ███████████████████░ 95.5% 756 / 792
data_catalog ███████████████████░ 95.5% 4,401 / 4,606
tracing ███████████████████░ 95.6% 3,518 / 3,680
autoresearch ███████████████████░ 95.7% 8,481 / 8,865
messaging ███████████████████░ 95.8% 3,798 / 3,963
skills ███████████████████░ 95.8% 6,972 / 7,274
replay_vision ███████████████████░ 95.9% 27,154 / 28,310
product_analytics ███████████████████░ 96.2% 28,495 / 29,617
revenue_analytics ███████████████████░ 96.4% 1,876 / 1,946
access_control ███████████████████░ 96.4% 7,122 / 7,386
user_interviews ███████████████████░ 96.5% 2,859 / 2,963
feature_flags ███████████████████░ 96.5% 25,499 / 26,416
warehouse_sources ███████████████████░ 97.2% 452,814 / 465,679
data_quality ████████████████████ 97.7% 7,592 / 7,774
links ████████████████████ 97.9% 234 / 239
security ████████████████████ 98.0% 1,203 / 1,228
metrics ████████████████████ 98.1% 4,085 / 4,166
analytics_platform ████████████████████ 98.3% 2,783 / 2,832
pulse ████████████████████ 98.5% 2,043 / 2,075
live_debugger ████████████████████ 99.2% 626 / 631
field_notes ████████████████████ 99.4% 172 / 173

Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.

@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team September 28, 2026 11:21
stamphog[bot]

This comment was marked as outdated.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 94b5ce77-2a9b-43a8-aff0-c91fd742b529

📥 Commits

Reviewing files that changed from the base of the PR and between 2c7d1c7 and 3ffb207.

📒 Files selected for processing (2)
  • products/signals/backend/temporal/agentic/scout_coordinator.py
  • products/signals/backend/test/test_scout_coordinator.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The coordinator resolves a separate, flag-tunable global cap for operational scouts, with a default of 200. The planner allocates operational and product scouts under separate global caps. Per-team tick and daily caps still apply to both groups. Tests cover cap resolution and allocation across both scout groups.

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 3ffb2

The separate scout budgets are mergeable after normal checks; production dispatch volume can be monitored after deployment.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 3ffb2

The scheduling change affects how many scouts can run across enrolled teams. Existing team limits remain in place, and no new public access path was identified. Overlapping scheduled runs and deployment behavior remain unverified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The changed dispatch capacity applies across enrolled teams, but each selected run remains subject to its team's shared tick and daily limits.

Trust Boundaries and Controls

  • observed — Planning retains enrollment, withheld-skill, enabled-configuration, and skill-provenance checks before budget allocation. The inspected public-API-marked change is a test.

Resilience and Maintainability Implications

  • observed — Post-dispatch stamping avoids advancing a run that was not started. Duplicate rejection is keyed to a tick-specific child ID and does not itself establish protection against overlap between different ticks.

Hardening Proposals

  • proposed — Before raising the operational override, verify the scheduler's overlap policy and monitor combined dispatch volume; document that setting the override to zero restores the default rather than stopping operational dispatch.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and explains the problem, impact, changes, testing, release status, operational limits, and agent context. It identifies that production validation remains outstanding. The…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: e2d3ec6b-ee4e-4309-a574-ce14bcf5ed12

📥 Commits

Reviewing files that changed from the base of the PR and between 2fb7cf7 and 1e6e3ef.

📒 Files selected for processing (4)
  • products/signals/backend/scout_harness/AGENTS.md
  • products/signals/backend/scout_harness/team_limits.py
  • products/signals/backend/temporal/agentic/scout_coordinator.py
  • products/signals/backend/test/test_scout_coordinator.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 5 remain after this review.

Comment thread products/signals/backend/scout_harness/team_limits.py Outdated
Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
Comment thread products/signals/backend/test/test_scout_coordinator.py Outdated
@posthog

posthog Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏

@posthog posthog Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PostHog Review

Found 2 consider.

Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
Per-team tick and daily caps now trim a team's product and operational runs together, before the two global budgets are filled. Only a harness-seeded canonical operational skill takes the operational budget. A team's own scout with the same name stays in the product pool.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
@stamphog
stamphog Bot dismissed their stale review September 28, 2026 11:39

A new stamphog review started for this PR — the fresh verdict replaces this approval.

stamphog[bot]

This comment was marked as outdated.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: eaeb1146-bc81-4533-88a3-e54543fcfb39

📥 Commits

Reviewing files that changed from the base of the PR and between 1e6e3ef and 2e6e333.

📒 Files selected for processing (4)
  • products/signals/backend/scout_harness/config_registry.py
  • products/signals/backend/scout_harness/team_limits.py
  • products/signals/backend/temporal/agentic/scout_coordinator.py
  • products/signals/backend/test/test_scout_coordinator.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 4 remain after this review.

Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
@trunk-io

trunk-io Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

When an operational run survives the per-team trim but loses its global slot, the team's per-team slot now goes back to its most overdue dropped product run. `_DueRun` declares frozen=False for the prefer-frozen-dataclasses ratchet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
@stamphog
stamphog Bot dismissed their stale review September 28, 2026 11:57

A new stamphog review started for this PR — the fresh verdict replaces this approval.

stamphog[bot]

This comment was marked as outdated.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: c57426f0-e703-41c2-be69-f4e27328333a

📥 Commits

Reviewing files that changed from the base of the PR and between 2e6e333 and 2c7d1c7.

📒 Files selected for processing (2)
  • products/signals/backend/temporal/agentic/scout_coordinator.py
  • products/signals/backend/test/test_scout_coordinator.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 3 remain after this review.

Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated

@posthog posthog Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PostHog Review

Found 1 consider.

Comment thread products/signals/backend/temporal/agentic/scout_coordinator.py Outdated
Each round, a team takes its most overdue run from a budget that still has room, up to its per-team cap. When one budget is full, the team's slot goes to its next run in the other budget, in both directions. This replaces the trim-then-refill pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
@stamphog
stamphog Bot dismissed their stale review September 28, 2026 12:14

A new stamphog review started for this PR — the fresh verdict replaces this approval.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved.

Internal scheduling/dispatch logic for the self-driving scout coordinator — not risky territory (no schema, auth, billing, public API, dependency, or CI/infra changes). The major bidirectional-refill bug CodeRabbit/PostHog bot flagged on earlier commits appears fixed in the current diff: I manually traced the new round-robin allocation algorithm against both directions of the reported scenario and it produces the correct result, matching the author's inline reply describing the fix, and a dedicated regression test for exactly this case is included.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 150L, 3F substantive, 255L/5F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1c-medium (255L, 5F, single-area, fix)
stamphog 2.2.0 .stamphog/policy.yml @ 3ffb207 · reviewed head 3ffb207

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant