Skip to content

perf(metrics): read the hourly projection in the overview services query - #110789

Open
frankh wants to merge 3 commits into
masterfrom
posthog/metrics-overview-read-projection
Open

frankh wants to merge 3 commits into
masterfrom
posthog/metrics-overview-read-projection

Conversation

@frankh

@frankh frankh commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Problem

The metrics overview page is slow for teams with many active series.

  • The services query reads every series-hour row of the last day in metrics4_series.
  • #110666 adds the services_by_hour projection, which holds the same rollup per team, hour and service.
  • The query in its current shape cannot use that projection, so it still scans the raw rows.

Changes

  • The overview loads faster. The services query reads a few projection rows per service and hour, not every series-hour row.
  • The window covers whole UTC hours, so it is 24 to 25 hours long, not exactly 24.
  • The series count uses uniq, so for large teams the number can differ slightly from an exact count.
  • metric_series in HogQL gets a time_bucket column, so the query can filter on it.
  • The query runs with convertToProjectTimezone off and converts to UTC outside max().
  • HogQL settings accept force_optimize_projection. Only the test sets it.

ClickHouse used the projection only for one shape of the query (EXPLAIN on the dev metrics cluster):

Filter Series count Last seen Uses projection
timestamp uniqExact max(toTimeZone(timestamp)) No (shape on master)
time_bucket uniq max(toTimeZone(timestamp)) No
timestamp uniq max(timestamp) No
time_bucket uniq toTimeZone(max(timestamp)) Yes (this PR)

HogQL wraps every DateTime column in toTimeZone, so max(last_seen) prints as max(toTimeZone(timestamp, ...)). That expression does not match the projection's max(timestamp), which is why the modifier is off for this query.

Warning

Merge this only after #110666 is deployed and every part in the last day has the projection. In dev, a query that used the projection while only some parts had it returned incomplete results.

How did you test this code?

  • pytest products/metrics/backend/tests/test_metrics_overview_query_runner.py passes locally.
  • EXPLAIN on the dev metrics cluster, for the SQL this query prints, shows AggregatingProjection and a read from services_by_hour.
  • Not checked: the page load time in dev with this code, because the code is not deployed.

Test rationale: test_services_query_reads_the_hourly_projection runs the services query with force_optimize_projection, so ClickHouse rejects the query when it does not use a projection. It fails if the query shape drifts, for example a return to uniqExact or a removed modifier. That drift would only show as a slow page. With the modifier removed, this test failed with PROJECTION_NOT_USED and every other overview test passed.

👉 Stay up-to-date with PostHog coding conventions for a smoother review.

Release status

  • No feature flag controls this change
  • This change is behind a feature flag and is not available to users
  • This change makes a previously flagged feature available to everyone

Automatic notifications

  • Publish to changelog?

Docs update

None.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Opus 5.5

  • This PR is the top layer of a two-layer stack on feat(metrics): add hourly per-service projection to metrics4_series #110666. It stays separate because a ClickHouse migration PR must be migration-only.
  • The test forces the projection only on the services query. The name count reads metric_names, which has no projection, so a module-wide setting would fail that query.
  • Duplicate search: no other open PR changes the overview services query.
  • Repo skills read: /writing-tests, /writing-code-comments, /writing-pr-descriptions.

🤖 Generated with Claude Code


Created with PostHog Desktop

@frankh frankh self-assigned this Oct 2, 2026
@frankh
frankh added this pull request to stack #110790 October 2, 2026 13:07
@frankh
frankh marked this pull request as ready for review October 2, 2026 13:08
@frankh frankh added the stamphog Request AI approval (no full review) label Oct 2, 2026
@parameterai

parameterai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Risk: No findings

This increment adds force_optimize_projection to HogQLQuerySettings and rewrites the projection test to force that setting via a patch of execute_hogql_query instead of reading system.query_log. I traced the new setting into the settings pipeline: it is None by default (never emitted), the printer validates keys with ^[a-zA-Z0-9_]+$ and escapes values (posthog/hogql/printer/base.py:1641), and the team_id WHERE guard is applied at print time independently of any SETTINGS clause, so the new field does not weaken tenant isolation or add an injection path. The one new exposure is that users can now write SETTINGS force_optimize_projection=1 in their own HogQL (the grammar allows a settings clause), but the worst outcome is their own query erroring or returning incomplete same-team results while the projection is partially deployed — a rollout-ordering concern, not a security issue.

Sentinel reviewed e683ee6 · Review settings

@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team October 2, 2026 13:09
@pr-assigner-resolver-posthog

Copy link
Copy Markdown

👀 Auto-assigned reviewers

These soft owners were skipped because they only have minor changes here. Nothing blocks merge, so self-assign if you'd like a look:

  • @PostHog/team-data-tools (posthog/owners.yaml)

Soft owners come from each directory's owners.yaml and each product's product.yaml (resolved nearest-file-wins). For a skipped owner, the locator is the file that decided it. Generated files and lockfiles are ignored when deciding ownership.

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Comment density — 11% of added code lines are comments (3 of 28)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/metrics/backend/metrics_overview_query_runner.py 3 11

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Playwright — 1 flaky

🎭 Playwright report · View test results →

⚠️ 1 flaky test:

  • create experiment via wizard, add metrics, and launch (chromium)

These issues are not necessarily caused by your changes.
Annoyed by this section? Help fix flakies and failures and it will go green!

@greptile-apps

greptile-apps Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[Medium risk] Adjusts metrics query to use hourly projection.

The PR is not ready to merge because its active-series window disagrees with the window reported to users, and the new test must meet the repository’s testing requirement.

Reviews (1) · Last reviewed commit: "perf(metrics): read the hourly projectio..."

toTimeZone(max(max(last_seen)) OVER (), 'UTC') AS total_last_seen_at
FROM posthog.metric_series
WHERE last_seen > now() - {lookback}
WHERE time_bucket >= toStartOfHour(toTimeZone(now() - {lookback}, 'UTC'))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Lookback includes stale series. If a series was last seen before the lookback cutoff but within the same UTC hour, this bucket-only filter still includes it. The overview can therefore show nearly an extra hour of stale services and active series while labeling the counts “Last 24 hours.” Keep an exact last_seen cutoff alongside the bucket bound, or report the window the query actually uses.

Prompt To Fix With AI
This is a comment left during a code review.
Path: products/metrics/backend/metrics_overview_query_runner.py
Line: 128

Comment:
**Lookback includes stale series.** If a series was last seen before the lookback cutoff but within the same UTC hour, this bucket-only filter still includes it. The overview can therefore show nearly an extra hour of stale services and active series while labeling the counts “Last 24 hours.” Keep an exact `last_seen` cutoff alongside the bucket bound, or report the window the query actually uses.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it really doesn't matter if it's 25hrs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — the 24–25 hour whole-UTC-hour window is explicitly intentional, so including rows from the cutoff hour is part of the desired behavior. My comment treated the lookback as an exact 24-hour cutoff; it isn't a valid finding for this change, and no adjustment is needed.

Comment thread products/metrics/backend/tests/test_metrics_overview_query_runner.py Outdated
Comment thread products/metrics/backend/tests/test_metrics_overview_query_runner.py Outdated

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

@greptile-apps[bot] still has a review in flight (👀) — not approving over an unfinished review. The review re-runs on the next push, or re-request one once the reviewer finishes.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 22L, 2F substantive, 47L/3F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1b-small (47L, 3F, two-areas, perf)
stamphog 2.3.1 .stamphog/policy.yml @ 06a9481 · reviewed head 06a9481

@frankh
frankh requested a review from a team as a code owner October 2, 2026 13:17
stamphog[bot]

This comment was marked as outdated.

@frankh

frankh commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/trunk merge

@trunk-io

trunk-io Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

⏳ Stack waiting to start tests on this stack - details.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 09a9e1c3-58fb-455f-8d09-3409a8ea1f1d

📥 Commits

Reviewing files that changed from the base of the PR and between e683ee6 and de574c6.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The metrics overview query now filters metric series by the UTC hour bucket, uses uniq for series counts, and converts aggregated latest timestamps to UTC. It disables HogQL project-timezone conversion. HogQL query settings add an optional forced projection optimization setting, and a test checks the service rollup when that setting is enabled.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to e683e

The overview query’s intended UTC-hour window is supported, but its new regression test leaves the boundary unprotected. Add that coverage before merging if practical, and deploy only after the projection is ready as the PR specifies.

Security Architecture Review

Security architecture risk: 🔵 Low · up to e683e

The inspected query remains team-scoped and read-only, with no demonstrated increase in access or privileges. The main uncertainty is whether projection coverage is complete before rollout and whether mixed coverage preserves complete overview results.

Retained concerns

  • Low · reliability · inferred: Deploying the projection definition alone does not establish the documented readiness prerequisite: the migration does not materialize existing parts. Activating the new read shape during mixed coverage therefore has unresolved result-completeness and failure-containment behavior. Production does not force projection use, so unsafe fallback behavior is not established.
Security review details

Security Blast Radius

  • inferred — The directly supported exposure is a team’s metrics overview and its reads on the shared metrics store. The inspected change alters aggregation and planner eligibility, not execution credentials or write authority; broader exposure was not established.

Trust Boundaries and Controls

  • observed — The new planner setting defaults to unset. The production runner does not enable it, while the test enables it locally for the services query. The existing team-isolation fixture asserts that another team’s metric produces no service result.

Resilience and Maintainability Implications

  • inferred — The changed services operation is read-only and does not commit metric state. Interruption or repetition therefore does not introduce a partial-write or cleanup lifecycle in this path, although reads during changing projection coverage still have unresolved completeness behavior.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description is complete and matches the repository template. It explains the problem, user-visible changes, testing, test rationale, release status, deployment dependency, and agent context. The a…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
products/metrics/backend/tests/test_metrics_overview_query_runner.py (1)

73-85: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Cover the hourly boundary in the forced-projection test.

The query filters from toStartOfHour(now() - lookback), so the test should include one earlier point from the same series in another included UTC hour. Add a separate series in the bucket immediately before that boundary and assert that it is excluded. Use a distinct label set for the excluded series so uniq(series_fingerprint) makes the assertion observable.

Existing tests cover same-window aggregation and data several days outside the window. They do not reliably detect cross-hour aggregation or the exact whole-hour boundary.


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: PostHog/posthog/.coderabbit.yaml

Review profile: QUIET

Plan: Enterprise

Run ID: 8291365d-8be2-4d5b-8f6c-1acebfdd84f7

📥 Commits

Reviewing files that changed from the base of the PR and between 0981fd4 and e683ee6.

📒 Files selected for processing (4)
  • posthog/hogql/constants.py
  • posthog/hogql/database/schema/metrics.py
  • products/metrics/backend/metrics_overview_query_runner.py
  • products/metrics/backend/tests/test_metrics_overview_query_runner.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@trunk-io

trunk-io Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

@frankh
frankh requested a review from Gilbert09 October 2, 2026 13:53
Base automatically changed from posthog/metrics4-series-services-projection to master October 2, 2026 14:09
@trunk-io
trunk-io Bot requested a review from a team as a code owner October 2, 2026 14:09
@stamphog
stamphog Bot dismissed their stale review October 2, 2026 14:09

The PR was retargeted to a different base branch, so the approved diff is no longer what was reviewed. Stamphog re-reviews automatically.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

Two gates refused this pull request, so stamphog can't review it. The deny-list gate matched the migrations path: posthog/clickhouse/migrations/0345_metrics4_series_services_projection.py and max_migration.txt are changed, and migrations always need a human reviewer. The tier gate also classified the change as T2-never, because it spans two areas (ClickHouse schema and HogQL/query runner code) and is a performance change, across 19 files and 311 lines.

To move forward, ask a human reviewer, such as someone on the owning team, to review it. The description says this is already stacked on #110666, so you could also move the schema and migration pieces into that PR and leave only the query runner and HogQL changes here. That would give a smaller change in one area for a later pass.

  • coderabbitai[bot] reviewed the current head.
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: migrations
size ✓ 258L, 16F substantive, 311L/19F incl. docs/generated/snapshots — within ceiling
tier ✗ classified as T2-never: T2-never (311L, 19F, two-areas, perf)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head e683ee6

@stamphog stamphog Bot removed the stamphog Request AI approval (no full review) label Oct 2, 2026
robbie-c and others added 3 commits October 2, 2026 15:13
…er (#110771)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The services query scanned every series-hour row of the last day. The
services_by_hour projection on metrics4_series holds the same rollup per
team, hour and service, but the query shape kept ClickHouse from using it.

The query now filters on time_bucket, counts series with uniq, and
aggregates the bare last_seen column with convertToProjectTimezone off,
converting to UTC outside max(). metric_series exposes time_bucket for
the filter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 3538a544-5668-418a-aebd-c3430d18864a
The overview test now runs the services query with
force_optimize_projection, so ClickHouse rejects the query when the
projection is not used. HogQL settings accept the new field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: 3538a544-5668-418a-aebd-c3430d18864a
@frankh
frankh force-pushed the posthog/metrics-overview-read-projection branch from e683ee6 to de574c6 Compare October 2, 2026 14:13
@frankh frankh added the stamphog Request AI approval (no full review) label Oct 2, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved.

Contained, reversible perf change to the metrics overview query, with a test that forces projection use. The author has strong familiarity with the code. The one reviewer concern (24 to 25 hour window) was answered as intentional, and the rollout-ordering warning is disclosed in the description.

  • Author wrote 100% of the modified lines and has 38 merged PRs in these paths (familiarity STRONG).
  • 👍 on the PR from greptile-apps[bot].
  • Rollout ordering: the services_by_hour projection must be deployed and cover all parts in the last day before this ships, or the overview can show incomplete results. The author disclosed this.
Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✓ 23L, 3F substantive, 42L/4F incl. docs/generated/snapshots — within ceiling
tier ✓ T1-agent / T1b-small (42L, 4F, two-areas, perf)
stamphog 2.3.1 .stamphog/policy.yml @ de574c6 · reviewed head de574c6

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stamphog Request AI approval (no full review)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants