Skip to content

Commit b6d44df

Browse files
evansenterclaude
andauthored
docs: Add schema design documentation (#66)
- Create docs/SCHEMA.md with table definitions, indexes, and migration history - Link from CLAUDE.md and README.md - Documents design principles: don't over-distill, aggregate→drill-down - Includes all indexes across all tables 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
1 parent 914ef1c commit b6d44df

3 files changed

Lines changed: 269 additions & 0 deletions

File tree

‎CLAUDE.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -4,6 +4,8 @@ Queryable analytics for Claude Code session logs, exposed as an MCP server and C
44

55
**API Reference**: `session-analytics-cli --help` or `src/session_analytics/guide.md` (MCP resource: `session-analytics://guide`).
66

7+
**Schema Design**: See [docs/SCHEMA.md](docs/SCHEMA.md) for database tables, indexes, and migration history.
8+
79
---
810

911
## ⚠️ DATABASE PROTECTION

‎README.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -138,6 +138,8 @@ make check
138138
3. **Auto-refresh**: Queries detect stale data (>5 min) and trigger re-ingestion
139139
4. **Patterns**: Pre-computes tool sequences and permission gaps for fast queries
140140

141+
See [docs/SCHEMA.md](docs/SCHEMA.md) for detailed database schema documentation.
142+
141143
## Architecture
142144

143145
Key patterns used in the codebase:

‎docs/SCHEMA.md‎

Lines changed: 265 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,265 @@
1+
# Database Schema Design
2+
3+
This document describes the SQLite database schema for session-analytics.
4+
5+
**Location**: `~/.claude/contrib/analytics/data.db`
6+
7+
---
8+
9+
## Design Principles
10+
11+
1. **Don't over-distill** - Store raw signals (error counts, timestamps, parameters) rather than pre-computed interpretations. The consuming LLM handles context.
12+
13+
2. **Aggregate → drill-down** - Every aggregate must be traceable to specifics. If "821 Bash errors" appears, the schema must support finding which commands failed.
14+
15+
3. **Denormalize for common queries** - Extract frequently-filtered fields (command, file_path, skill_name) into columns rather than requiring JSON parsing.
16+
17+
---
18+
19+
## Tables Overview
20+
21+
| Table | Purpose | Rows (typical) |
22+
|-------|---------|----------------|
23+
| `events` | All tool calls, messages, and summaries from JSONL logs | 100K+ |
24+
| `sessions` | Aggregated session metadata | 1K+ |
25+
| `ingestion_state` | Tracks which JSONL files have been processed | ~100 |
26+
| `patterns` | Pre-computed patterns (re-computable, safe to drop) | ~1K |
27+
| `git_commits` | Git history for correlation | ~5K |
28+
| `session_commits` | Junction table linking sessions to commits | ~3K |
29+
| `bus_events` | Cross-session events from event-bus | ~2K |
30+
| `events_fts` | FTS5 virtual table for user message search | N/A |
31+
32+
---
33+
34+
## Core Tables
35+
36+
### events
37+
38+
The primary table storing all parsed JSONL entries.
39+
40+
```sql
41+
CREATE TABLE events (
42+
id INTEGER PRIMARY KEY,
43+
uuid TEXT NOT NULL, -- Unique within session (see UNIQUE constraint)
44+
timestamp TIMESTAMP NOT NULL,
45+
session_id TEXT NOT NULL,
46+
project_path TEXT,
47+
entry_type TEXT, -- 'user', 'assistant', 'summary', 'tool_use', 'tool_result'
48+
49+
-- Tool-specific (null if not a tool call)
50+
tool_name TEXT,
51+
tool_input_json TEXT, -- Full JSON for drill-down
52+
tool_id TEXT, -- Correlates tool_use with tool_result
53+
is_error INTEGER DEFAULT 0,
54+
55+
-- Denormalized for common filters
56+
command TEXT, -- Bash: first word (e.g., "git")
57+
command_args TEXT, -- Bash: remaining args
58+
file_path TEXT, -- Read/Edit/Write target
59+
skill_name TEXT, -- Skill invocation name
60+
61+
-- Token tracking (only on assistant events to avoid duplication)
62+
input_tokens INTEGER,
63+
output_tokens INTEGER,
64+
cache_read_tokens INTEGER,
65+
cache_creation_tokens INTEGER,
66+
model TEXT,
67+
68+
-- Context
69+
git_branch TEXT,
70+
cwd TEXT,
71+
72+
-- User journey (RFC #17)
73+
user_message_text TEXT, -- For FTS search
74+
exit_code INTEGER, -- Reserved for future extraction
75+
76+
-- Agent tracking (RFC #41)
77+
parent_uuid TEXT, -- Links tool_use to parent assistant event
78+
agent_id TEXT, -- Task subagent ID from agent-*.jsonl
79+
is_sidechain INTEGER DEFAULT 0,
80+
version TEXT, -- Claude Code version
81+
82+
UNIQUE(session_id, uuid) -- UUID unique within each session
83+
)
84+
```
85+
86+
**Key patterns**:
87+
- `entry_type='tool_use'` + `entry_type='tool_result'` are correlated by `tool_id`
88+
- Token columns only populated on `entry_type='assistant'` to avoid double-counting
89+
- `user_message_text` enables FTS via `events_fts` virtual table
90+
- `tool_input_json` preserves full parameters for drill-down queries
91+
92+
### sessions
93+
94+
Aggregated metadata per session.
95+
96+
```sql
97+
CREATE TABLE sessions (
98+
id TEXT PRIMARY KEY, -- UUID from session file
99+
project_path TEXT,
100+
first_seen TIMESTAMP,
101+
last_seen TIMESTAMP,
102+
entry_count INTEGER DEFAULT 0,
103+
tool_use_count INTEGER DEFAULT 0,
104+
total_input_tokens INTEGER DEFAULT 0,
105+
total_output_tokens INTEGER DEFAULT 0,
106+
primary_branch TEXT,
107+
slug TEXT, -- Human-readable session name
108+
context_switch_count INTEGER DEFAULT 0 -- RFC #26
109+
)
110+
```
111+
112+
### git_commits
113+
114+
Git history for session correlation.
115+
116+
```sql
117+
CREATE TABLE git_commits (
118+
sha TEXT PRIMARY KEY,
119+
timestamp TIMESTAMP,
120+
message TEXT,
121+
session_id TEXT, -- Inferred from timestamp proximity
122+
project_path TEXT
123+
)
124+
```
125+
126+
### session_commits
127+
128+
Junction table for time-to-commit analysis.
129+
130+
```sql
131+
CREATE TABLE session_commits (
132+
session_id TEXT NOT NULL,
133+
commit_sha TEXT NOT NULL,
134+
time_to_commit_seconds INTEGER,
135+
is_first_commit INTEGER DEFAULT 0,
136+
PRIMARY KEY (session_id, commit_sha)
137+
)
138+
```
139+
140+
### bus_events
141+
142+
Events from the event-bus for cross-session insights.
143+
144+
```sql
145+
CREATE TABLE bus_events (
146+
id INTEGER PRIMARY KEY,
147+
event_id INTEGER UNIQUE NOT NULL, -- Original ID from event-bus
148+
timestamp TIMESTAMP NOT NULL,
149+
event_type TEXT NOT NULL, -- 'gotcha_discovered', 'pattern_found', etc.
150+
channel TEXT,
151+
session_id TEXT,
152+
repo TEXT, -- Extracted from channel
153+
payload TEXT,
154+
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
155+
)
156+
```
157+
158+
### ingestion_state
159+
160+
Tracks which JSONL files have been processed for incremental ingestion.
161+
162+
```sql
163+
CREATE TABLE ingestion_state (
164+
file_path TEXT PRIMARY KEY,
165+
file_size INTEGER,
166+
last_modified TIMESTAMP,
167+
entries_processed INTEGER,
168+
last_processed TIMESTAMP
169+
)
170+
```
171+
172+
### patterns
173+
174+
Pre-computed patterns for fast querying (re-computable, safe to delete).
175+
176+
```sql
177+
CREATE TABLE patterns (
178+
id INTEGER PRIMARY KEY,
179+
pattern_type TEXT NOT NULL, -- 'tool_frequency', 'sequence', etc.
180+
pattern_key TEXT NOT NULL, -- e.g., "Bash" or "Read → Edit"
181+
count INTEGER DEFAULT 0,
182+
last_seen TIMESTAMP,
183+
metadata_json TEXT,
184+
computed_at TIMESTAMP,
185+
UNIQUE(pattern_type, pattern_key)
186+
)
187+
```
188+
189+
---
190+
191+
## Indexes
192+
193+
Performance-critical indexes on the `events` table:
194+
195+
| Index | Columns | Purpose |
196+
|-------|---------|---------|
197+
| `idx_events_timestamp` | `timestamp` | Time-range queries (days parameter) |
198+
| `idx_events_session` | `session_id` | Session-specific event lookup |
199+
| `idx_events_tool` | `tool_name` | Tool frequency analysis |
200+
| `idx_events_project` | `project_path` | Project filtering |
201+
| `idx_events_tool_id` | `tool_id` | Self-join for tool_use ↔ tool_result correlation |
202+
| `idx_events_parent_uuid` | `parent_uuid` | Token deduplication queries |
203+
| `idx_events_agent_id` | `agent_id` | Agent activity breakdown |
204+
| `idx_events_has_user_message` | Partial on `id` | FTS join optimization |
205+
206+
**Performance note**: The `idx_events_tool_id` index is critical for `query_error_details()` which self-joins events to correlate errors with their input parameters. Without it, queries take ~25s on 160K rows; with it, ~0.3s.
207+
208+
### Other Table Indexes
209+
210+
| Table | Index | Columns |
211+
|-------|-------|---------|
212+
| `git_commits` | `idx_git_commits_timestamp` | `timestamp` |
213+
| `git_commits` | `idx_git_commits_session` | `session_id` |
214+
| `git_commits` | `idx_git_commits_project` | `project_path` |
215+
| `session_commits` | `idx_session_commits_session` | `session_id` |
216+
| `session_commits` | `idx_session_commits_commit` | `commit_sha` |
217+
| `bus_events` | `idx_bus_events_timestamp` | `timestamp` |
218+
| `bus_events` | `idx_bus_events_type` | `event_type` |
219+
| `bus_events` | `idx_bus_events_session` | `session_id` |
220+
| `bus_events` | `idx_bus_events_repo` | `repo` |
221+
222+
---
223+
224+
## Full-Text Search
225+
226+
User messages are indexed via FTS5:
227+
228+
```sql
229+
CREATE VIRTUAL TABLE events_fts USING fts5(
230+
user_message_text,
231+
content='events',
232+
content_rowid='id'
233+
)
234+
```
235+
236+
Sync triggers maintain index consistency:
237+
- `events_fts_insert`: Populates FTS on new events
238+
- `events_fts_delete`: Removes from FTS on delete
239+
- `events_fts_update`: Handles message text changes
240+
241+
---
242+
243+
## Migration History
244+
245+
| Version | Name | Changes |
246+
|---------|------|---------|
247+
| 1 | Initial | Core tables: events, sessions, ingestion_state, patterns |
248+
| 2 | add_rfc17_phase1_columns | user_message_text, exit_code, git_commits table |
249+
| 3 | add_user_message_fts | FTS5 virtual table and sync triggers |
250+
| 4 | add_session_enrichment | session_commits junction, context_switch_count |
251+
| 5 | add_agent_tracking | parent_uuid, agent_id, is_sidechain, version |
252+
| 6 | add_event_bus_integration | bus_events table |
253+
| 7 | add_tool_id_index | Performance index for self-joins |
254+
255+
---
256+
257+
## Schema Evolution
258+
259+
When adding schema changes:
260+
261+
1. Add migration function with `@migration(N, "name")` decorator
262+
2. Update `SCHEMA_VERSION = N` constant
263+
3. Add to `_init_db()` for fresh installs
264+
4. Use `IF NOT EXISTS` for idempotency
265+
5. Test with both fresh DB and migration path

0 commit comments

Comments
 (0)