feat(data-warehouse): add guru analytics, folder, tag and group tables - #110788
Conversation
Adds analytics_events (incremental on eventDate via fromDate), folders, folder_items, tag_categories, tags and group_members to the Guru source, and ticks the matching coverage-gap entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: e85f1d6b-ee21-4d05-a121-f03e27b121f9
|
😎 Merged successfully - details. |
|
Risk: No findings This increment adds a stable content-hash fallback ( Sentinel reviewed |
🤖 CI report
|
| First copy | Second copy | Lines | Tokens |
|---|---|---|---|
products/warehouse_sources/backend/temporal/data_imports/sources/better_stack/better_stack.py:22 |
products/warehouse_sources/backend/temporal/data_imports/sources/guru/guru.py:22 |
12 | 109 |
✅ Duplication (TypeScript) — clean
New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
🚨 Comment density — 7% of added code lines are comments (25 of 376)
This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.
Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.
Files with the most added comment lines:
| File | Comment lines | Added lines |
|---|---|---|
products/warehouse_sources/backend/temporal/data_imports/sources/guru/guru.py |
12 | 101 |
products/warehouse_sources/backend/temporal/data_imports/sources/guru/settings.py |
11 | 74 |
products/warehouse_sources/backend/temporal/data_imports/sources/guru/tests/test_guru.py |
1 | 126 |
products/warehouse_sources/backend/temporal/data_imports/sources/guru/tests/test_guru_source.py |
1 | 2 |
This check does not block merging. It updates on every push and clears when the share drops.
⚠️ Playwright — 1 failed
🎭 Playwright report · View test results →
❌ 1 failed test:
- create experiment via wizard, add metrics, and launch (chromium)
These issues are not necessarily caused by your changes.
Annoyed by this section? Help fix flakies and failures and it will go green!
⚠️ Backend coverage — 97.0% of changed backend lines covered — 3 uncovered
🧪 Backend test coverage
Patch coverage — changed backend lines (products + core): ███████████████████░ 97.0% (125 / 128)
| File | Patch | Uncovered changed lines |
|---|---|---|
products/warehouse_sources/backend/temporal/data_imports/sources/guru/settings.py |
81.8% | 28, 33 |
products/warehouse_sources/backend/temporal/data_imports/sources/guru/guru.py |
97.9% | 63 |
🤖 Agents: add a test only if an uncovered line exposes a realistic regression that existing tests miss. Otherwise explain why no new test is needed under "How did you test this code?". Gap list: the patch-coverage artifact on this run (gh run download 25229674782120 -n patch-coverage), or the coverage-data block at the end of this comment.
Per-product line coverage (touched products)
| Product | Coverage | Lines |
|---|---|---|
platform_features |
██░░░░░░░░░░░░░░░░░░ 12.1% |
7 / 58 |
demo |
███████████░░░░░░░░░ 53.4% |
1,445 / 2,707 |
data_tools |
████████████░░░░░░░░ 61.2% |
90 / 147 |
warehouse_sources_queue |
█████████████░░░░░░░ 65.8% |
1,646 / 2,500 |
ai_gateway |
███████████████░░░░░ 75.0% |
9 / 12 |
aeo |
███████████████░░░░░ 76.3% |
617 / 809 |
batch_exports |
████████████████░░░░ 81.3% |
21,574 / 26,552 |
apm |
█████████████████░░░ 84.1% |
1,306 / 1,553 |
cdp |
██████████████████░░ 88.3% |
4,559 / 5,164 |
ml_inference |
██████████████████░░ 88.8% |
539 / 607 |
mcp_analytics |
██████████████████░░ 89.2% |
5,038 / 5,651 |
product_tours |
██████████████████░░ 89.3% |
1,340 / 1,500 |
dashboards |
██████████████████░░ 89.6% |
6,924 / 7,727 |
notebooks |
██████████████████░░ 90.2% |
15,304 / 16,971 |
signals |
██████████████████░░ 90.4% |
60,150 / 66,523 |
cohorts |
██████████████████░░ 90.5% |
8,534 / 9,434 |
data_warehouse |
██████████████████░░ 90.6% |
14,375 / 15,860 |
streamlit_apps |
██████████████████░░ 90.8% |
2,684 / 2,956 |
managed_warehouse |
██████████████████░░ 91.0% |
10,252 / 11,263 |
data_modeling |
██████████████████░░ 91.2% |
10,562 / 11,584 |
tasks |
██████████████████░░ 91.3% |
79,533 / 87,103 |
exports |
██████████████████░░ 91.7% |
9,684 / 10,566 |
business_knowledge |
██████████████████░░ 92.0% |
8,472 / 9,208 |
engineering_analytics |
██████████████████░░ 92.2% |
11,497 / 12,475 |
ai_training |
██████████████████░░ 92.2% |
356 / 386 |
today |
██████████████████░░ 92.3% |
999 / 1,082 |
webmcp |
███████████████████░ 92.5% |
248 / 268 |
early_access_features |
███████████████████░ 92.6% |
1,339 / 1,446 |
conversations |
███████████████████░ 92.6% |
29,226 / 31,559 |
visual_review |
███████████████████░ 92.7% |
10,000 / 10,785 |
managed_migrations |
███████████████████░ 92.7% |
1,581 / 1,705 |
stamphog |
███████████████████░ 92.8% |
8,109 / 8,742 |
canvas |
███████████████████░ 92.9% |
7,155 / 7,703 |
approvals |
███████████████████░ 93.0% |
3,974 / 4,271 |
mcp_registry |
███████████████████░ 93.1% |
1,670 / 1,794 |
notifications |
███████████████████░ 93.2% |
1,144 / 1,228 |
error_tracking |
███████████████████░ 93.2% |
16,370 / 17,557 |
surveys |
███████████████████░ 93.4% |
6,644 / 7,113 |
autoresearch |
███████████████████░ 93.6% |
8,837 / 9,442 |
slack_app |
███████████████████░ 93.7% |
14,611 / 15,600 |
context_layer |
███████████████████░ 93.8% |
3,414 / 3,639 |
web_analytics |
███████████████████░ 93.9% |
23,680 / 25,229 |
billing_alerts |
███████████████████░ 94.1% |
2,094 / 2,226 |
mcp_store |
███████████████████░ 94.3% |
8,959 / 9,501 |
alerts |
███████████████████░ 94.7% |
9,319 / 9,844 |
wizard |
███████████████████░ 94.7% |
6,150 / 6,496 |
ai_observability |
███████████████████░ 94.7% |
25,860 / 27,307 |
workflows |
███████████████████░ 94.7% |
15,187 / 16,034 |
reminders |
███████████████████░ 94.8% |
760 / 802 |
review_hog |
███████████████████░ 95.0% |
11,750 / 12,362 |
annotations |
███████████████████░ 95.1% |
817 / 859 |
endpoints |
███████████████████░ 95.1% |
9,231 / 9,703 |
customer_analytics |
███████████████████░ 95.2% |
26,116 / 27,428 |
legal_documents |
███████████████████░ 95.2% |
2,311 / 2,427 |
marketing_analytics |
███████████████████░ 95.3% |
19,450 / 20,413 |
posthog_ai |
███████████████████░ 95.3% |
2,491 / 2,614 |
actions |
███████████████████░ 95.5% |
756 / 792 |
logs |
███████████████████░ 95.5% |
15,468 / 16,200 |
experiments |
███████████████████░ 95.5% |
33,000 / 34,548 |
data_catalog |
███████████████████░ 95.5% |
4,401 / 4,606 |
tracing |
███████████████████░ 95.6% |
3,536 / 3,699 |
replay_vision |
███████████████████░ 95.6% |
29,371 / 30,708 |
growth |
███████████████████░ 95.7% |
11,381 / 11,888 |
skills |
███████████████████░ 95.8% |
6,972 / 7,274 |
messaging |
███████████████████░ 95.9% |
3,834 / 3,999 |
product_analytics |
███████████████████░ 96.0% |
28,521 / 29,696 |
revenue_analytics |
███████████████████░ 96.4% |
1,889 / 1,959 |
user_interviews |
███████████████████░ 96.5% |
2,870 / 2,974 |
feature_flags |
███████████████████░ 96.6% |
27,119 / 28,060 |
access_control |
███████████████████░ 96.7% |
7,739 / 8,007 |
warehouse_sources |
███████████████████░ 97.3% |
469,786 / 482,717 |
data_quality |
████████████████████ 97.5% |
7,701 / 7,895 |
links |
████████████████████ 97.9% |
234 / 239 |
metrics |
████████████████████ 98.0% |
4,252 / 4,338 |
security |
████████████████████ 98.0% |
1,304 / 1,330 |
analytics_platform |
████████████████████ 98.3% |
2,784 / 2,833 |
pulse |
████████████████████ 98.5% |
2,046 / 2,078 |
live_debugger |
████████████████████ 99.2% |
626 / 631 |
field_notes |
████████████████████ 99.4% |
172 / 173 |
Report-only. Patch coverage = changed backend lines covered vs origin/master. Sorted lowest first.
Known gaps: lines covered only by Temporal tests show as uncovered; core line numbers may drift if master changed the same file.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: PostHog/posthog/.coderabbit.yaml Review profile: QUIET Plan: Enterprise Run ID: 📒 Files selected for processing (4)
💤 Files with no reviewable changes (1)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughGuru source configuration and handling add support for group members, folders, folder items, tag categories, tags, and analytics events. Analytics requests use a configured date filter for incremental processing. Team-scoped endpoints resolve the team ID through Priority: ⬇️ Low Merge Risk: ⚪ Minimal · up to Analytics rows now receive the required merge key. No confirmed merge-blocking issue remains; behavior with a live Guru account remains unverified. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The additional data is fetched using existing account credentials, with the provider origin and redirect restrictions preserved. No introduced security vulnerability was established, but the expanded employee-data scope and incomplete verification of downstream access and recovery behavior prevent a minimal-risk assessment. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 1✅ Passed checks (1 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
Review profile: QUIET
Plan: Enterprise
Run ID: c6e7fb01-78db-4c60-9c32-43be24f67e49
📒 Files selected for processing (6)
products/warehouse_sources/backend/temporal/data_imports/sources/COVERAGE_GAPS_APPENDIX.mdproducts/warehouse_sources/backend/temporal/data_imports/sources/guru/canonical_descriptions.pyproducts/warehouse_sources/backend/temporal/data_imports/sources/guru/guru.pyproducts/warehouse_sources/backend/temporal/data_imports/sources/guru/settings.pyproducts/warehouse_sources/backend/temporal/data_imports/sources/guru/tests/test_guru.pyproducts/warehouse_sources/backend/temporal/data_imports/sources/guru/tests/test_guru_source.py
Limit details: You’ve used all 12 included reviews currently available.
Falls back to a content hash when an analytics event has no id, and declares the Guru endpoint config dataclass frozen. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Generated-By: PostHog Desktop Task-Id: e85f1d6b-ee21-4d05-a121-f03e27b121f9
A new stamphog review started for this PR — the fresh verdict replaces this approval.
There was a problem hiding this comment.
Approved.
Additive warehouse source tables by an owning-team author with STRONG familiarity, with tests. The one substantive reviewer concern (a missing analytics event key) was addressed with a fallback key and a test.
- Author wrote 97% of the modified lines and has 51 merged PRs in these paths (familiarity STRONG).
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✓ | 288L, 3F substantive, 458L/7F incl. docs/generated/snapshots — within ceiling |
| tier | ✓ | T1-agent / T1d-complex (458L, 7F, two-areas, feat) |
| stamphog 2.3.1 | .stamphog/policy.yml @ 2cefa4f · reviewed head 2cefa4f |
|
/trunk merge |
Problem
COVERAGE_GAPS_APPENDIX.md.Changes
Six new tables, each checked against the Guru v1 API reference before it was built:
analytics_eventsGET /v1/teams/{teamId}/analyticseventDate(server-sidefromDate)idfoldersGET /v1/foldersidfolder_itemsGET /v1/folders/{id}/items, fanned out over foldersfolder_id,idtag_categoriesGET /v1/teams/{teamId}/tagcategoriesidtagsidgroup_membersGET /v1/groups/{id}/members, fanned out over groupsgroup_id,email/whoamibefore the sync starts.tagstable flattens each category'stagsarray and addscategoryIdandcategoryName.analytics_eventsdeclaressort_mode="desc", so the watermark commits only when the run ends.id, but the guide sample omits it. An event withoutidgets a SHA-256 content hash as its key.fromDatetakes at most millisecond precision, so the cursor goes out as...T00:00:00.000+00:00./tagcategories/tagslist call, but Guru only has get-by-id for tags. The tag rows come from the categories response.How did you test this code?
Ran the Guru test module and
sources/testslocally (the source-wide invariants: versions, categories, catalog). Ran mypy over the Guru package. Not checked: a live sync against a real Guru account.Test rationale:
TestBuildParams: analytics sendsfromDatewith millisecond precision and never the card GQLq/sort params. Full refresh sends nofromDate. Extended the existing parameterized class.TestTeamScopedEndpoints: the request path uses the/whoamiteam ID, tags are flattened with their category, and a/whoami401 surfaces as the error stringget_non_retryable_errorsmatches.TestFanoutEndpoints: child rows carrygroup_id/folder_id, and the primary key columns are never null. A null key would seed duplicate rows on merge.TestGuruSourceResponsenow mocks/whoamiand skips fan-out endpoints.test_get_schemasnow expectsanalytics_eventsto be incremental.Release status
Automatic notifications
Docs update
None. The Guru docs page renders its table list from
public_source_configs, so the new tables appear there without a docs change.🤖 Agent context
Autonomy: Fully autonomous
Agent: Claude Code, Claude Opus 5.5 (
claude-opus-5-5)/implementing-warehouse-sources,/writing-tests,/writing-pr-descriptions.Created with PostHog Desktop
🤖 Generated with Claude Code