fix(signals): give operational scouts their own dispatch budget - #107613
posthog[bot] wants to merge 4 commits into
Conversation
Operational scouts now draw from a separate per-tick ceiling (MAX_OPERATIONAL_RUNS_PER_TICK, flag key max_operational_runs_per_tick_global). A wave of never-run operational lanes can no longer take the global slots of product scouts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
|
Merging to
After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here |
🦔 PostHog Review reviewed this pull requestFound 0 must fix, 0 should fix, 1 consider. Published 1 finding (view the review). Resolved comments: 2 already settled |
🤖 CI report
|
| File | Comment lines | Added lines |
|---|---|---|
products/signals/backend/temporal/agentic/scout_coordinator.py |
12 | 81 |
products/signals/backend/test/test_scout_coordinator.py |
6 | 89 |
products/signals/backend/scout_harness/team_limits.py |
4 | 13 |
This check does not block merging. It updates on every push and clears when the share drops.
⚠️ Backend coverage — 98.0% of changed backend lines covered — 1 uncovered
🧪 Backend test coverage
Patch coverage — changed backend lines (products + core): ████████████████████ 98.0% (79 / 80)
| File | Patch | Uncovered changed lines |
|---|---|---|
products/signals/backend/scout_harness/config_registry.py |
80.0% | 476 |
🤖 Agents: add a test covering the lines above, or note why under "How did you test this code?". Machine-readable gap list: the patch-coverage artifact on this run (gh run download 117880218227538 -n patch-coverage), or the coverage-data block at the end of this comment.
Per-product line coverage (touched products)
| Product | Coverage | Lines |
|---|---|---|
platform_features |
██░░░░░░░░░░░░░░░░░░ 12.1% |
7 / 58 |
warehouse_sources_queue |
██████░░░░░░░░░░░░░░ 29.1% |
92 / 316 |
demo |
████████████░░░░░░░░ 57.8% |
1,545 / 2,673 |
data_tools |
████████████░░░░░░░░ 61.2% |
90 / 147 |
aeo |
██████████████░░░░░░ 70.5% |
467 / 662 |
ai_gateway |
███████████████░░░░░ 75.0% |
9 / 12 |
batch_exports |
████████████████░░░░ 81.2% |
21,459 / 26,431 |
apm |
█████████████████░░░ 84.1% |
1,306 / 1,553 |
ml_inference |
█████████████████░░░ 87.2% |
482 / 553 |
cdp |
██████████████████░░ 88.2% |
4,548 / 5,155 |
mcp_analytics |
██████████████████░░ 88.9% |
4,910 / 5,523 |
product_tours |
██████████████████░░ 89.3% |
1,331 / 1,491 |
dashboards |
██████████████████░░ 89.5% |
6,839 / 7,641 |
signals |
██████████████████░░ 89.9% |
54,713 / 60,868 |
data_warehouse |
██████████████████░░ 89.9% |
13,912 / 15,470 |
notebooks |
██████████████████░░ 90.2% |
15,287 / 16,945 |
cohorts |
██████████████████░░ 90.4% |
8,420 / 9,316 |
streamlit_apps |
██████████████████░░ 90.7% |
2,625 / 2,895 |
managed_warehouse |
██████████████████░░ 90.9% |
10,215 / 11,234 |
tasks |
██████████████████░░ 91.1% |
73,993 / 81,212 |
data_modeling |
██████████████████░░ 91.5% |
10,525 / 11,498 |
business_knowledge |
██████████████████░░ 91.6% |
6,899 / 7,528 |
engineering_analytics |
██████████████████░░ 91.7% |
11,017 / 12,014 |
exports |
██████████████████░░ 91.8% |
9,685 / 10,555 |
ai_training |
██████████████████░░ 92.2% |
356 / 386 |
conversations |
███████████████████░ 92.5% |
28,726 / 31,047 |
early_access_features |
███████████████████░ 92.6% |
1,341 / 1,448 |
managed_migrations |
███████████████████░ 92.7% |
1,581 / 1,705 |
visual_review |
███████████████████░ 92.8% |
9,244 / 9,966 |
canvas |
███████████████████░ 92.8% |
6,877 / 7,409 |
approvals |
███████████████████░ 93.0% |
3,919 / 4,214 |
mcp_registry |
███████████████████░ 93.1% |
1,670 / 1,794 |
error_tracking |
███████████████████░ 93.1% |
15,843 / 17,010 |
notifications |
███████████████████░ 93.2% |
1,145 / 1,229 |
slack_app |
███████████████████░ 93.2% |
13,677 / 14,674 |
stamphog |
███████████████████░ 93.2% |
7,885 / 8,456 |
surveys |
███████████████████░ 93.3% |
6,571 / 7,040 |
context_layer |
███████████████████░ 93.8% |
3,373 / 3,595 |
web_analytics |
███████████████████░ 93.9% |
21,653 / 23,051 |
alerts |
███████████████████░ 94.0% |
8,541 / 9,082 |
billing_alerts |
███████████████████░ 94.1% |
2,094 / 2,226 |
mcp_store |
███████████████████░ 94.4% |
8,940 / 9,472 |
ai_observability |
███████████████████░ 94.4% |
22,534 / 23,870 |
wizard |
███████████████████░ 94.7% |
6,151 / 6,496 |
reminders |
███████████████████░ 94.8% |
760 / 802 |
workflows |
███████████████████░ 94.9% |
14,335 / 15,113 |
review_hog |
███████████████████░ 94.9% |
11,490 / 12,109 |
annotations |
███████████████████░ 95.1% |
817 / 859 |
endpoints |
███████████████████░ 95.1% |
9,211 / 9,681 |
customer_analytics |
███████████████████░ 95.2% |
24,899 / 26,167 |
legal_documents |
███████████████████░ 95.2% |
2,311 / 2,427 |
marketing_analytics |
███████████████████░ 95.3% |
19,216 / 20,161 |
posthog_ai |
███████████████████░ 95.4% |
2,489 / 2,610 |
experiments |
███████████████████░ 95.4% |
32,645 / 34,211 |
growth |
███████████████████░ 95.4% |
9,812 / 10,282 |
logs |
███████████████████░ 95.4% |
15,290 / 16,022 |
actions |
███████████████████░ 95.5% |
756 / 792 |
data_catalog |
███████████████████░ 95.5% |
4,401 / 4,606 |
tracing |
███████████████████░ 95.6% |
3,518 / 3,680 |
autoresearch |
███████████████████░ 95.7% |
8,481 / 8,865 |
messaging |
███████████████████░ 95.8% |
3,798 / 3,963 |
skills |
███████████████████░ 95.8% |
6,972 / 7,274 |
replay_vision |
███████████████████░ 95.9% |
27,154 / 28,310 |
product_analytics |
███████████████████░ 96.2% |
28,495 / 29,617 |
revenue_analytics |
███████████████████░ 96.4% |
1,876 / 1,946 |
access_control |
███████████████████░ 96.4% |
7,122 / 7,386 |
user_interviews |
███████████████████░ 96.5% |
2,859 / 2,963 |
feature_flags |
███████████████████░ 96.5% |
25,499 / 26,416 |
warehouse_sources |
███████████████████░ 97.2% |
452,814 / 465,679 |
data_quality |
████████████████████ 97.7% |
7,592 / 7,774 |
links |
████████████████████ 97.9% |
234 / 239 |
security |
████████████████████ 98.0% |
1,203 / 1,228 |
metrics |
████████████████████ 98.1% |
4,085 / 4,166 |
analytics_platform |
████████████████████ 98.3% |
2,783 / 2,832 |
pulse |
████████████████████ 98.5% |
2,043 / 2,075 |
live_debugger |
████████████████████ 99.2% |
626 / 631 |
field_notes |
████████████████████ 99.4% |
172 / 173 |
Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: PostHog/posthog/.coderabbit.yaml Review profile: QUIET Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe coordinator resolves a separate, flag-tunable global cap for operational scouts, with a default of 200. The planner allocates operational and product scouts under separate global caps. Per-team tick and daily caps still apply to both groups. Tests cover cap resolution and allocation across both scout groups. Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to The separate scout budgets are mergeable after normal checks; production dispatch volume can be monitored after deployment. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The scheduling change affects how many scouts can run across enrolled teams. Existing team limits remain in place, and no new public access path was identified. Overlapping scheduled runs and deployment behavior remain unverified. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 1✅ Passed checks (1 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 4
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: e2d3ec6b-ee4e-4309-a574-ce14bcf5ed12
📒 Files selected for processing (4)
products/signals/backend/scout_harness/AGENTS.mdproducts/signals/backend/scout_harness/team_limits.pyproducts/signals/backend/temporal/agentic/scout_coordinator.pyproducts/signals/backend/test/test_scout_coordinator.py
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 5 remain after this review.
|
PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏 |
Per-team tick and daily caps now trim a team's product and operational runs together, before the two global budgets are filled. Only a harness-seeded canonical operational skill takes the operational budget. A team's own scout with the same name stays in the product pool. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
A new stamphog review started for this PR — the fresh verdict replaces this approval.
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: eaeb1146-bc81-4533-88a3-e54543fcfb39
📒 Files selected for processing (4)
products/signals/backend/scout_harness/config_registry.pyproducts/signals/backend/scout_harness/team_limits.pyproducts/signals/backend/temporal/agentic/scout_coordinator.pyproducts/signals/backend/test/test_scout_coordinator.py
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 4 remain after this review.
When an operational run survives the per-team trim but loses its global slot, the team's per-team slot now goes back to its most overdue dropped product run. `_DueRun` declares frozen=False for the prefer-frozen-dataclasses ratchet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
A new stamphog review started for this PR — the fresh verdict replaces this approval.
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: c57426f0-e703-41c2-be69-f4e27328333a
📒 Files selected for processing (2)
products/signals/backend/temporal/agentic/scout_coordinator.pyproducts/signals/backend/test/test_scout_coordinator.py
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 3 remain after this review.
Each round, a team takes its most overdue run from a budget that still has room, up to its per-team cap. When one budget is full, the team's slot goes to its next run in the other budget, in both directions. This replaces the trim-then-refill pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: f95f6814-a23d-41d7-9834-df85ee6fe7c6
A new stamphog review started for this PR — the fresh verdict replaces this approval.
There was a problem hiding this comment.
Approved.
Internal scheduling/dispatch logic for the self-driving scout coordinator — not risky territory (no schema, auth, billing, public API, dependency, or CI/infra changes). The major bidirectional-refill bug CodeRabbit/PostHog bot flagged on earlier commits appears fixed in the current diff: I manually traced the new round-robin allocation algorithm against both directions of the reported scenario and it produces the correct result, matching the author's inline reply describing the fix, and a dedicated regression test for exactly this case is included.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✓ | 150L, 3F substantive, 255L/5F incl. docs/generated/snapshots — within ceiling |
| tier | ✓ | T1-agent / T1c-medium (255L, 5F, single-area, fix) |
| stamphog 2.2.0 | .stamphog/policy.yml @ 3ffb207 · reviewed head 3ffb207 |
Problem
signals-scout-inbox-validationon every wildcard-enrolled project that has a seed-disabled row. Every resumed row has never run, so_overdue_secondsreturnsinf._allocate_tick_budgetfills one sharedMAX_RUNS_PER_TICK(1000 per 30-minute tick) most-overdue-team first. Each resumed team's first pick is inbox validation, so these runs take the tick before product scouts get a slot.Origin
67b63feChanges
is_operational_scout). Each pool goes through_allocate_tick_budgetwith its own global cap.MAX_RUNS_PER_TICK, so operational runs never defer them.MAX_OPERATIONAL_RUNS_PER_TICK = 200. A wave such as this one drains at a bounded rate instead of in a burst.max_operational_runs_per_tick_globalin thesignals-scoutpayload overrides the operational cap without a deploy. Invalid values fall back to the default, asmax_runs_per_tick_globaldoes.scout_harness/AGENTS.mddocuments the new key.Note
200 per tick is a guess at a safe ceiling. The intended steady-state volume of inbox validation is not confirmed yet. If report checks lag, raise the flag key. If compute is the concern, lower it.
Out of scope: the edit-rate alert on insight 8KWJNZMK still counts operational runs in series C. That needs an insight edit in PostHog, not a code change.
How did you test this code?
test_operational_scouts_have_their_own_tick_budget. It covers this regression: never-run operational lanes on two teams take the whole global cap from a due product scout. The test fails with the split reverted and passes with it.test_per_team_caps_count_both_budgets. It checks that a team's per-tick and daily caps still hold when the team has runs in both pools.test_resolve_global_max_operational_runs_per_tick. It checks that the key parses, that invalid values fall back, and that the product key does not widen the operational cap.hogli test products/signals/backend/test/test_scout_coordinator.pylocally. All tests pass.Release status
Automatic notifications
Docs update
None. Internal dispatch limits only.
🤖 Agent context
Autonomy: Fully autonomous
Agent: Claude Code, Claude Opus 5.5 (
claude-opus-5-5)/writing-pr-descriptions.Created with PostHog Desktop from this inbox report.
🤖 Generated with Claude Code