From 3332f9ce056c073bfd7d41390b8cd4a453e9b536 Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Thu, 10 Sep 2026 12:21:54 +0200 Subject: [PATCH 1/2] fix(last9-cloudwatch): name the averaging window for MSK throughput Confirming eval run 34454314035 caught a candidate reporting 200 instead of 250 bytes/second for a requested interval mean: it issued bare instant queries, saw only the latest period, and divided that pair. The grader failed it on value, evidence, window coverage, and raw observation. SKILL.md carries the weighted-window rule and the RDS, DMS, EC2, and ElastiCache references each name the latest-period versus whole-window distinction. The MSK throughput bullet said only "use the requested level or average", so it was the one family reference that left the choice open. Name both cases and say to read the raw companion series over the window rather than a single instant query. Refs ENG-1891. Co-Authored-By: Claude Opus 5 (1M context) --- skills/last9-cloudwatch/references/msk.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/last9-cloudwatch/references/msk.md b/skills/last9-cloudwatch/references/msk.md index 8f25e37..15d068c 100644 --- a/skills/last9-cloudwatch/references/msk.md +++ b/skills/last9-cloudwatch/references/msk.md @@ -10,7 +10,7 @@ Broker metrics, cluster aggregates, and topic breakdowns are not interchangeable ## Throughput and lag -- `BytesInPerSec`, `BytesOutPerSec`, and `MessagesInPerSec` already describe rates. Use the requested level or average; applying `rate()` again changes the quantity. Combining broker/topic rows needs a verified non-overlapping population. +- `BytesInPerSec`, `BytesOutPerSec`, and `MessagesInPerSec` already describe rates, so applying `rate()` again changes the quantity. Match the average to the requested window: for the latest period use that matched Sum/SampleCount pair; for a mean over the requested interval use total Sum / total SampleCount across the disjoint periods, read from the raw companion series over that window rather than a single instant query returning only the last period. Combining broker/topic rows needs a verified non-overlapping population. - `MaxOffsetLag` is the maximum offset lag across the applicable partitions; `SumOffsetLag` is a different aggregate. Neither is a count of unique messages processed during the observation window. - `EstimatedMaxTimeLag` is in seconds. An offset count cannot be converted to time without additional evidence. From 6a5ec420ba7c35382d563d067f2ba0ff333387af Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Thu, 10 Sep 2026 12:28:40 +0200 Subject: [PATCH 2/2] fix(last9-cloudwatch): state what a selector without a range returns The MSK failure was not only an MSK wording gap. SKILL.md gives the weighted-window formula in its statistics table, and gives the range selector separately, in a paragraph about inspecting raw timestamps. Nothing said that a selector carrying no range returns only the latest period, so a model could follow the formula correctly and still feed it a single period, which is what the failing run did. State it once in the tool reference, where query shape is already discussed, and give a self-check: the returned sample count must match the periods the interval should contain. This covers every family rather than the one whose reference text happened to be thinnest. Refs ENG-1891. Co-Authored-By: Claude Opus 5 (1M context) --- skills/last9-cloudwatch/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index f34a6c4..8a8933a 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -47,7 +47,7 @@ Installed schemas take precedence over this reference: | `prometheus_instant_query` | `query`, `datasource`, `time_iso` | | `prometheus_range_query` | `query`, `datasource`, `start_time_iso`, `end_time_iso` | -Discover names with `label: "__name__"`. Do not invent a `step` parameter. Distinguish raw range selectors evaluated once from repeatedly evaluated charts. +Discover names with `label: "__name__"`. Do not invent a `step` parameter. Distinguish raw range selectors evaluated once from repeatedly evaluated charts. A selector carrying no range returns only the latest period, so any total, average, or coverage claim about a whole interval needs a range selector spanning that interval; before claiming one, check that the returned sample count matches the periods the interval should contain. ## Discover and execute efficiently