Skip to content

fix(last9-cloudwatch): cite stream timestamp and statistic mappings - #17

Merged
prathamesh-sonpatki merged 2 commits into
masterfrom
prathamesh/eng-1891-cloudwatch-prose-followup
Sep 10, 2026
Merged

prathamesh-sonpatki merged 2 commits into
masterfrom
prathamesh/eng-1891-cloudwatch-prose-followup

Conversation

@prathamesh-sonpatki

@prathamesh-sonpatki prathamesh-sonpatki commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

Why

The paired eval in last9/last9-mcp-evals#60 passed 12/12 structured cases for the last9-cloudwatch skill, but manual review of the candidate reports recorded five residual prose defects. Each traces to a wording gap in the skill, one of them caused by the skill's own text.

What changed

Defect seen in eval report Fix
Period-end timestamps asserted as a "CloudWatch Metric Streams convention" with no source (DynamoDB, EC2, mixed RDS) SKILL.md step 3 now cites the AWS OpenTelemetry translation docs (0.7.0 and 1.0.0): time_unix_nano = period endTime. Requires lineage evidence; exporters stay unverified. dynamodb.md aligned.
300-second spacing read as "basic monitoring confirmed" (EC2) ec2.md: spacing establishes cadence only, not monitoring tier or period boundaries.
Sample time called "source publication Unix epoch" (S3) s3.md said "raw publication timestamp". Now "raw observation timestamp", explicitly not AWS publish time nor Last9 ingest time.
Empty result described as "no data was ingested" (zero/empty case) SKILL.md query outcomes: describe as no matching samples for that selector/window; forbid not-ingested/not-delivered/data-gap wording without separate evidence.
quantile="1" called "maximum/latest" with inferred mapping (Billing) SKILL.md statistics table cites the documented mapping (quantile 0 = period Min, 1 = period Max) and states a period Maximum is not the latest reading. billing.md adds the same.

Net: 7 lines changed across 5 files. No new files, no structural changes.

Verification

  • scripts/check-skill-pack.sh and check-skill-pack-selftest.sh pass.
  • AWS docs fetched and confirmed for both stream formats: Summary datapoint, time_unix_nano = endTime, quantile 0.0 = Min, 1.0 = Max. Only attribute encoding differs between 0.7.0 and 1.0.0.
  • Paired eval dispatched against this branch's SHA b480e76 via CloudWatch skill pilot workflow_dispatch; run link and before/after prose review will be posted in a comment once complete.
  • Loading smoke check (Claude Code, Codex, Cursor) re-run on the revised install; result posted in a comment.

Refs ENG-1891.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Follow-up to the paired eval in last9-mcp-evals#60. Structured grades passed
12/12 but manual review of candidate reports found five prose defects that
trace to wording gaps in the skill:

- Period-end timestamps were asserted as a universal convention without a
  source. Cite the documented Metric Streams OpenTelemetry translation
  (time_unix_nano = period endTime) and require lineage evidence; exporters
  stay unverified.
- Sample spacing was read as proof of the monitoring tier. State that spacing
  establishes cadence only.
- The S3 reference called the sample time a "publication timestamp". Rename
  to observation timestamp and say what it is not.
- Empty results were described as "no data ingested". Forbid causal wording
  without separate evidence.
- quantile="1" was called "maximum/latest" with an inferred mapping. Document
  quantile 0/1 = period Min/Max from the AWS translation and state that a
  period Maximum is not the latest reading.

Refs ENG-1891.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@prathamesh-sonpatki

Copy link
Copy Markdown
Member Author

Loading smoke check on this branch (b480e76)

Install path, fresh empty project:

npx skills add last9/ai-toolkit#prathamesh/eng-1891-cloudwatch-prose-followup --skill last9-cloudwatch -a claude-code -a codex -a cursor -y

diff -r of the installed tree against this branch's skills/last9-cloudwatch: identical.

Canary prompt (no MCP/network): four lines, two of which can only be answered from text introduced in this PR (the time_unix_nano = period endTime rule in SKILL.md, and the S3 reference's "neither AWS publish time nor Last9 ingest time" sentence), plus the absolute path of the SKILL.md read.

Host Version New SKILL.md text New S3 reference text Path read
Claude Code claude -p 2.1.267 ✅ "CloudWatch period endTime" ✅ .claude/skills/last9-cloudwatch/SKILL.md
Codex codex exec 0.153.4 ✅ ✅ .agents/skills/last9-cloudwatch/SKILL.md
Cursor cursor-agent -p 2026.09.08 ✅ ✅ .agents/skills/last9-cloudwatch/SKILL.md

Cursor's first attempt exited 1 with no output after two "Connection lost" retries to its backend; the immediate retry passed. Host-side transient, not a skill issue.

Paired eval against this SHA: https://github.com/last9/last9-mcp-evals/actions/runs/34437702939 (in progress; result to follow).

@prathamesh-sonpatki

Copy link
Copy Markdown
Member Author

Paired eval on this branch (b480e76)

Three workflow_dispatch runs of CloudWatch skill pilot in last9-mcp-evals, same fixtures, model claude-sonnet-4-6, both arms identical except the skill bundle.

Run Repeats Candidate Baseline Paired regressions Notes
34437702939 1 11/12 1/12 0 daily-s3-storage failed tool_protocol: model double-quoted an ISO string in two convert_timestamp args (eval-only helper the skill never mentions), got invalid_arguments, then self-corrected and reported correct values.
34438942737 1 11/12 4/12 0 mixed-cloudwatch-exporter hit the 16-call cap one call before report_findings; correct values (0.02 s, 7 connections) were already computed. Same case passed in the other runs and in #60.
34440634322 2 24/24 0/24 0 Clean. Avg 10.4 calls per candidate execution (identical to #60's 10.4), max 14.

Aggregate: candidate 46/48, zero grading-rule failures, zero paired regressions. Both misses are protocol slips of different kinds; neither traces to the changed text.

Prose review of the five defects this PR targets

Checked the candidate report_findings method/limitation text in runs 1 and 2 for the exact phrases flagged in #60.

#60 defect Now
Period-end stamps asserted as "convention" with no source (DynamoDB, EC2, mixed RDS) Cites "AWS CloudWatch Metric Streams OpenTelemetry 0.7.0/1.0.0 translation docs" and adds "lineage is assumed… not separately verified".
300 s spacing = "basic monitoring confirmed" (EC2) "Five-minute spacing is consistent with basic monitoring but does not confirm it."
"source publication Unix epoch" (S3) "The raw sample timestamp is the observation timestamp as returned by the series."
Empty result = "no data was ingested" (zero/empty) "Empty result proves no matching samples in this window and scope; it does not establish whether the metric was never ingested, delivery was delayed, or…"
quantile="1" = "maximum/latest", mapping inferred (Billing) Mapping now cited: "in the documented OpenTelemetry Metric Streams translation maps to the period Maximum". Residual: still reasons that for a monotonic month-to-date gauge the period Maximum is the latest accumulated estimate. Defensible; not tightened further.

Docs verified directly: both AWS translation pages state Summary datapoint, time_unix_nano = period endTime, quantile 0.0 = Min, quantile 1.0 = Max. Only attribute encoding differs between 0.7.0 and 1.0.0.

Loading smoke check on this branch is in the comment above (3/3 hosts). Ready for review.

…rding

Review follow-up on the prose revision:

- The 1.0.0 link pointed at the wire-format page, which only implies the
  timestamp mapping through an example. Point at the translation page,
  which states that timeUnixNano contains the CloudWatch endTime, matching
  the 0.7.0 link already used.
- Scope the period-end conclusion to OpenTelemetry-format streams so the
  rule cannot be read as covering JSON-format streams.
- DynamoDB: say why a sample stamped exactly at the interval start is not
  needed instead of implying its absence provides coverage.

Refs ENG-1891.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@prathamesh-sonpatki

Copy link
Copy Markdown
Member Author

Code review response (ca085f7)

Independent review of 440d93d..b480e76: no Critical, one Important, four Minor. Applied:

  • Important, fixed: the 1.0.0 link pointed at the wire-format page (…-opentelemetry-100.html), which only implies the timestamp mapping through an example. Now points at …-opentelemetry-translation-100.html, which states "timeUnixNano contains the CloudWatch endTime" (fetched and confirmed). Symmetric with the 0.7.0 link.
  • Minor, fixed: "a verified stream sample is stamped at period end" → "a verified OpenTelemetry-format stream sample…", so the rule cannot be read as covering JSON-format streams.
  • Minor, fixed: dynamodb.md no longer implies that a missing sample at start provides coverage; it now says a sample at start is not needed because it belongs to the preceding period.

Not applied, by judgment:

  • Adding "use the documented mapping in SKILL.md, not spacing" to ec2.md: the ec2 sentence is about inference from spacing; SKILL.md step 3 already owns the mapping. Cross-reference would be redundant.
  • Splitting step 3 into its own numbered step: still one idea per sentence; revisit if it grows again.

Reviewer independently verified the AWS claims and confirmed no packaging or hygiene impact. Eval re-dispatch on ca085f7 in progress; link to follow.

@prathamesh-sonpatki

Copy link
Copy Markdown
Member Author

Eval on ca085f7 (current head)

Run 34443835196, repeats=1: candidate 12/12, baseline 2/12, paired regressions 0. One pair inconclusive (msk-broker-ingress): the baseline arm issued a query the finite replay does not support (fixture_miss), excluded from uplift by design. Average 10.8 calls per candidate execution; mixed-cloudwatch-exporter used exactly 16 with the report as the final call, so that case sits at the budget edge because of its generic-label trap.

Totals across all four dispatches on this PR's two commits: candidate 58/60, zero paired regressions, zero grading-rule failures.

@prathamesh-sonpatki
prathamesh-sonpatki merged commit 3b39308 into master Sep 10, 2026
5 checks passed
@prathamesh-sonpatki
prathamesh-sonpatki deleted the prathamesh/eng-1891-cloudwatch-prose-followup branch September 10, 2026 08:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant