From 8664be1256bdbe39ea08935cb2ceaace8f402aad Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Mon, 7 Sep 2026 17:54:14 +0200 Subject: [PATCH 1/8] feat(skills): add CloudWatch investigation workflow --- README.md | 1 + skills/last9-cloudwatch/SKILL.md | 121 +++++++++++++++++++++++++++++++ 2 files changed, 122 insertions(+) create mode 100644 skills/last9-cloudwatch/SKILL.md diff --git a/README.md b/README.md index 2180c89..ccfe66f 100644 --- a/README.md +++ b/README.md @@ -16,6 +16,7 @@ MCP gives your agent access. Skills give it judgment. | [`go-agent-install`](skills/go-agent-install/SKILL.md) | Instrument a Go service with Last9 go-agent: detect the stack, wire chi + `database/sql` tracing, promote opt-in body capture, and verify spans land — without double-instrumenting | | [`last9-logs`](skills/last9-logs/SKILL.md) | Log investigation: scope to a service first, attribute filters over body search, aggregate before drilling into raw lines | | [`last9-traces`](skills/last9-traces/SKILL.md) | Trace investigation: a five-question interview that lands on the right tool call, plus a `tracejson` syntax reference card | +| [`last9-cloudwatch`](skills/last9-cloudwatch/SKILL.md) | CloudWatch investigation: discover AWS resource dimensions, interpret period statistics and units, and distinguish sparse data from missing ingestion | ## Installation diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md new file mode 100644 index 0000000..33deb4d --- /dev/null +++ b/skills/last9-cloudwatch/SKILL.md @@ -0,0 +1,121 @@ +--- +name: last9-cloudwatch +description: Investigate AWS CloudWatch metrics in Last9 through read-only MCP queries. Use for CloudWatch metric discovery, RDS or Aurora resource metrics, CloudWatch statistics and units, or sparse S3 daily metrics ("CloudWatch in Last9", "RDS CPU", "Aurora latency", "S3 metric is empty"). +compatibility: Requires the Last9 MCP server connected to the session +metadata: + author: last9 +--- + +# last9-cloudwatch — investigate CloudWatch metrics in Last9 + +**Operating principle: establish the resource and the published statistic before choosing the calculation.** CloudWatch metrics can describe AWS resources without an application service, environment label, or APM instrumentation. Start with the requested AWS population; do not require traces to query its infrastructure metrics. + +## Prerequisites and scope + +Use the authenticated [Last9 MCP server](https://github.com/last9/last9-mcp-server) and its advertised tool schemas. This workflow reads existing telemetry. Creating streams, changing AWS permissions, configuring collectors, generating Terraform, and building dashboards are separate tasks. + +Confirm the connected organization and selected datasource from the conversation and available connection or datasource information. Call `list_datasources` when selection is not already established. Carry the selected datasource into every discovery and query call; a different default is not a reason to change the user's selection. If the connection cannot access the intended organization, report that boundary before querying another one. + +Pin the AWS account, region, resource, and UTC start/end times. Resolve relative times against an available clock once and reuse those bounds for comparisons. Discover the actual dimension keys before inserting filters. Account or region may be established by a dedicated datasource or integration configuration instead of a series label; state that evidence. If scope remains ambiguous, surface the available choices and resolve it before combining measurements. + +If the needed tools are unavailable, explain the missing connection or capability. Analyze supplied observations when sufficient, but identify proposed queries as unexecuted. Never present invented calls or results as evidence. + +## Tool reference + +Read the live descriptions first; these are the relevant parameter shapes, not a replacement for the installed schemas. + +| Tool | Use | Parameters | +|---|---|---| +| `list_datasources` | Resolve the datasource | Use its advertised schema | +| `prometheus_label_values` | Discover metric names or dimension values | `label`, `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | +| `prometheus_labels` | Discover label names for a selector | `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | +| `prometheus_instant_query` | Evaluate an expression at a fixed time, or read a raw range selector | `query`, `datasource`, `time_iso` | +| `prometheus_range_query` | Inspect an expression over the requested interval | `query`, `datasource`, `start_time_iso`, `end_time_iso` | + +Use `label: "__name__"` for metric-name discovery. These query tools do not expose a `step` parameter in this interface; do not invent one. Preserve the distinction between a raw range selector evaluated once and a chart expression evaluated repeatedly. + +## Discover, verify, then calculate + +1. **Find candidate names.** Search the selected datasource and time window for the requested AWS namespace or metric. The [CloudWatch integration guide](https://last9.io/docs/integrations/observability/aws-cloudwatch-metrics/) documents the `amazonaws_com_AWS` prefix for its stream path; use it as a discovery hint, not a universal name contract. Exporters and other ingestion paths can use different names. A missing result under one prefix is not proof that the resource has no metrics. +2. **Inspect dimensions and actual samples.** Discover label names and relevant values, then execute a narrow selector for the chosen account, region, and resource. A catalog entry proves discoverability, not current delivery. `prometheus_labels` may return a generic catalog even with `match_query`; if it does, state that limitation and inspect a bounded candidate metric query's actual returned label keys. Inspect returned series labels and raw timestamps before trusting a filter or interpreting emptiness. Do not add `service_name`, `env`, `namespace`, or a guessed AWS dimension simply because it appears on another metric. +3. **Identify lineage and one dimension level.** Distinguish CloudWatch Metric Streams, exporters, and trace-derived metrics using integration information and observed series. Preserve valid mixed configurations. Keep sources separate until their populations, periods, and definitions justify combining them; overlapping copies cannot be added. Similarly, choose instance, cluster, or role aggregation deliberately. Removing labels with `sum by (...)` after selecting overlapping rows does not remove double counting. +4. **Establish the measurement contract.** Record the AWS metric, unit, statistic, publication period, timestamp meaning, and coverage. Check whether a sample is a period summary, instantaneous gauge, or cumulative counter. Establish units from the AWS metric definition and the ingestion mapping; a unit label is not guaranteed, and magnitude alone is not evidence. Follow the current integration guide and verify additional statistics and the actual ingestion format when relevant; do not infer a permanent format-version requirement from an old example. If dimensions appear only as an opaque encoded value, report the observed shape and the missing resource-level selection capability; do not invent direct label keys or recreate the stream. +5. **Execute the required calculation.** Build queries from discovered names and verified selectors. Execute them with the pinned datasource and times, inspect results, and cross-check units and arithmetic. If the inputs cannot support the requested statistic, return it as unavailable and name the evidence needed next. + +Keep query outcomes distinct: a successful empty result contains no matching observations, while an explicit numeric zero is a measurement. An invalid query, authentication failure, or timeout establishes neither. Correct invalid arguments or expressions using the actual schema and discovered names, then retry the scoped read; otherwise report the blocking error rather than treating it as absence. + +Discovery example for an RDS investigation, after substituting the selected datasource and UTC bounds: + +```json +{ + "label": "__name__", + "match_query": "{__name__=~\"amazonaws_com_AWS_RDS_.+\"}", + "datasource": "", + "start_time_iso": "", + "end_time_iso": "" +} +``` + +This is input to `prometheus_label_values`, not a query result. Narrow discovery further with already-verified resource filters. Use returned metric names in the subsequent label and sample reads. + +## Statistics and calculation guardrails + +[CloudWatch Metric Streams](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-Metric-Streams.html) carry period statistics including Sum and SampleCount; additional statistics can be configured. The [OpenTelemetry translation](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-metric-streams-formats-opentelemetry-translation.html) explains the summary mapping. Verify that mapping for the series being queried before applying these recipes. + +| Observed measurement | Interpretation and calculation | +|---|---| +| Confirmed stream-summary `_sum` and `_count` companions | Period Sum and SampleCount. `_sum / _count` gives the period's sample average when labels, period, and population match and count is positive. `_count` counts observations; it is not automatically a request count. | +| Several confirmed, disjoint period summaries | A sample-weighted window average is total Sum divided by total SampleCount. Averaging the period averages is wrong when counts differ. Verify coverage and boundaries before claiming a whole-window result. | +| An AWS metric whose Sum counts events | Add disjoint period Sums for an event total; divide by covered elapsed seconds for average events/sec only with the required complete coverage. Establish this from the metric's definition, not its suffix. | +| A gauge such as storage size or resource utilization | Report the requested level or a clearly defined aggregate. Summing observations over time does not give total storage or total utilization. | +| A verified cumulative exporter counter | `rate()` or `increase()` may apply with suitable history and reset handling. Do not transfer that treatment to CloudWatch period-summary companions because their names end in `_sum` or `_count`. | +| Observed minimum, maximum, or percentile series | Select the actual statistic label, including `quantile` where present. A maximum of period p99s is the peak period p99; it cannot establish a whole-window p99. Do not run `histogram_quantile()` on already-quantiled values. | +| Instance rows plus cluster, role, or other rollup rows | Select a non-overlapping population at one intended dimension level. Never sum both an aggregate and its constituent resources. | + +For a confirmed summary family, the following are **templates**. Replace metric names, selectors, and window with observed values before execution: + +```promql +{} / {} +``` + +```promql +sum(sum_over_time({}[])) +/ +sum(sum_over_time({}[])) +``` + +The first expression needs one-to-one matching companion labels. The second additionally needs matching populations and periods, positive total SampleCount, and raw samples representing disjoint reports. Do not add arbitrary label-dropping modifiers to make an unexplained mismatch disappear. Do not silently turn a zero or missing denominator into zero utilization or latency. + +Inspect raw sample times, for example with `{}[]` through `prometheus_instant_query` at the fixed end time. Chart points may repeat or resample earlier observations; they are not automatically independent stream publications. Missing periods, duplicate delivery, late data, boundary misalignment, or unknown timestamp semantics limit a whole-window total or average. State the supported coverage rather than filling gaps with zero. + +## Worked example: RDS or Aurora CPU and read latency + +For “show CPU and read latency for this database over this interval”: + +1. Resolve the chosen account, region, and database using actual dimensions. For Aurora, determine whether the request concerns an instance, the cluster, or a role. AWS publishes [different dimension combinations](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/dimensions.html); inspect the selected series instead of combining every row with the cluster's name. +2. Discover the CPUUtilization and ReadLatency families in that scope. Verify whether the returned metrics are stream-summary companions, exporter gauges, or a different source. Read a small raw window and inspect labels, timestamps, and units before building the final expression. +3. If the observed family uses period Sum/SampleCount, execute the companion ratio for the period averages. For a requested full-window average, execute the weighted template only after verifying its coverage conditions. Keep CPU and latency calculations separate, and report their actual aggregation level. +4. [RDS CloudWatch metrics](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) define CPUUtilization as Percent and ReadLatency as seconds. Confirm the ingested unit mapping. Convert seconds to milliseconds with a factor of 1,000 only when that is the observed input unit. Enhanced Monitoring or exporter metrics with similar names can have different definitions or units. +5. Report the measured values with the exact executed selectors and interval, or say which result is unavailable. Without distributions or equivalent raw observations, period latency averages and percentiles do not establish a whole-window latency percentile. + +## Sparse metrics: S3 daily storage + +An empty recent lookup for BucketSizeBytes or NumberOfObjects does not establish zero, deletion, or a stopped stream. These are [daily S3 storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics. AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html) for viewing them. + +Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. + +If expected history is also empty, verify the datasource, filters, source configuration, statistic, and publication cadence before attributing the gap to ingestion. Use available read-only delivery evidence when present. If it is absent, report that whether delivery stopped remains unconfirmed. Do not infer an exact next publication time from “daily.” + +## Metric availability and Discover visibility + +A resource can have queryable CloudWatch metrics without appearing in a particular Discover view. Verify ingestion by executing a scoped metric query in the selected datasource. Investigate Discover identity and supported resource discovery separately, using the actual identity dimensions and current integration guidance. Missing application traces or a missing Discover row is not evidence that CloudWatch ingestion failed. + +## Report the evidence + +For each requested measurement, give the value and unit or **unavailable**, the AWS resource and dimension level, datasource and UTC interval, selected source/statistic, and the exact executed query with its returned evidence or tool-call reference. Separate observations from explanations and unresolved causes. Report partial coverage, last-known sample times, and any unsupported requested statistic explicitly; check that the prose arithmetic agrees with the samples and the stated result. + +For an ingestion-versus-UI question, state what the metric query established and what remains unknown about discovery. Link the [Last9 metric explorer](https://app.last9.io/metrics) when useful. + +## Related skills + +Use `last9-logs` for a separate log investigation or `last9-traces` for application span analysis when relevant and available. Neither is a prerequisite for this CloudWatch workflow. From e4e304eddcae5cad0da05f816ae53ad6bde0c00d Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Mon, 7 Sep 2026 18:18:13 +0200 Subject: [PATCH 2/8] fix(skills): preserve CloudWatch timestamps and evidence --- skills/last9-cloudwatch/SKILL.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index 33deb4d..795b18e 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -102,9 +102,9 @@ For “show CPU and read latency for this database over this interval”: An empty recent lookup for BucketSizeBytes or NumberOfObjects does not establish zero, deletion, or a stopped stream. These are [daily S3 storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics. AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html) for viewing them. -Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. +Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. Preserve the exact returned timestamp: copy raw Unix seconds when needed, or verify the UTC conversion with an available tool. Do not round it to a date or midnight. -If expected history is also empty, verify the datasource, filters, source configuration, statistic, and publication cadence before attributing the gap to ingestion. Use available read-only delivery evidence when present. If it is absent, report that whether delivery stopped remains unconfirmed. Do not infer an exact next publication time from “daily.” +If expected history is also empty, verify the datasource, filters, source configuration, statistic, and publication cadence before attributing the gap to ingestion. Use available read-only delivery evidence when present. If it is absent, report that whether delivery stopped remains unconfirmed. Do not infer the next publication date or time from “daily” or from the previous sample. ## Metric availability and Discover visibility @@ -112,7 +112,7 @@ A resource can have queryable CloudWatch metrics without appearing in a particul ## Report the evidence -For each requested measurement, give the value and unit or **unavailable**, the AWS resource and dimension level, datasource and UTC interval, selected source/statistic, and the exact executed query with its returned evidence or tool-call reference. Separate observations from explanations and unresolved causes. Report partial coverage, last-known sample times, and any unsupported requested statistic explicitly; check that the prose arithmetic agrees with the samples and the stated result. +For each requested measurement, preserve the user's requested output name and give the value and unit or **unavailable**, the AWS resource and dimension level, datasource and UTC interval, selected source/statistic, and the exact executed query with its returned evidence or tool-call reference. Cite all operands actually used, including both Sum and SampleCount when computing an average. Separate observations from explanations and unresolved causes. Report partial coverage, last-known sample times, and any unsupported requested statistic explicitly; check that the prose arithmetic agrees with the samples and the stated result. For an ingestion-versus-UI question, state what the metric query established and what remains unknown about discovery. Link the [Last9 metric explorer](https://app.last9.io/metrics) when useful. From fe82ba8f3b32307b5cb50630ac401626773084da Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Mon, 7 Sep 2026 18:41:48 +0200 Subject: [PATCH 3/8] fix(skills): verify CloudWatch timestamp interpretation --- skills/last9-cloudwatch/SKILL.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index 795b18e..e464e41 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -86,7 +86,7 @@ sum(sum_over_time({}[])) The first expression needs one-to-one matching companion labels. The second additionally needs matching populations and periods, positive total SampleCount, and raw samples representing disjoint reports. Do not add arbitrary label-dropping modifiers to make an unexplained mismatch disappear. Do not silently turn a zero or missing denominator into zero utilization or latency. -Inspect raw sample times, for example with `{}[]` through `prometheus_instant_query` at the fixed end time. Chart points may repeat or resample earlier observations; they are not automatically independent stream publications. Missing periods, duplicate delivery, late data, boundary misalignment, or unknown timestamp semantics limit a whole-window total or average. State the supported coverage rather than filling gaps with zero. +Inspect raw sample times, for example with `{}[]` through `prometheus_instant_query` at the fixed end time. Check the returned result type and timestamp meaning: use a matrix's timestamps as publication evidence only when they represent original samples. An evaluated vector's tuple timestamps do not establish raw publication times. Chart points may repeat or resample earlier observations; they are not automatically independent stream publications. Missing periods, duplicate delivery, late data, boundary misalignment, or unknown timestamp semantics limit a whole-window total or average. State the supported coverage rather than filling gaps with zero. ## Worked example: RDS or Aurora CPU and read latency @@ -102,7 +102,7 @@ For “show CPU and read latency for this database over this interval”: An empty recent lookup for BucketSizeBytes or NumberOfObjects does not establish zero, deletion, or a stopped stream. These are [daily S3 storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics. AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html) for viewing them. -Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. Preserve the exact returned timestamp: copy raw Unix seconds when needed, or verify the UTC conversion with an available tool. Do not round it to a date or midnight. +Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. Preserve the exact returned timestamp: copy raw Unix seconds when needed, or verify the UTC conversion with an available tool. If no reliable conversion tool is available, report the exact epoch only and omit converted UTC dates, times, and computed ages. Do not round it to a date or midnight. If expected history is also empty, verify the datasource, filters, source configuration, statistic, and publication cadence before attributing the gap to ingestion. Use available read-only delivery evidence when present. If it is absent, report that whether delivery stopped remains unconfirmed. Do not infer the next publication date or time from “daily” or from the previous sample. From ede69406d1e567ba886d72d1efe6d44538f6b57e Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Mon, 7 Sep 2026 19:26:10 +0200 Subject: [PATCH 4/8] fix(skills): verify raw metrics before inferring absence --- skills/last9-cloudwatch/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index e464e41..ddf3470 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -42,7 +42,7 @@ Use `label: "__name__"` for metric-name discovery. These query tools do not expo 4. **Establish the measurement contract.** Record the AWS metric, unit, statistic, publication period, timestamp meaning, and coverage. Check whether a sample is a period summary, instantaneous gauge, or cumulative counter. Establish units from the AWS metric definition and the ingestion mapping; a unit label is not guaranteed, and magnitude alone is not evidence. Follow the current integration guide and verify additional statistics and the actual ingestion format when relevant; do not infer a permanent format-version requirement from an old example. If dimensions appear only as an opaque encoded value, report the observed shape and the missing resource-level selection capability; do not invent direct label keys or recreate the stream. 5. **Execute the required calculation.** Build queries from discovered names and verified selectors. Execute them with the pinned datasource and times, inspect results, and cross-check units and arithmetic. If the inputs cannot support the requested statistic, return it as unavailable and name the evidence needed next. -Keep query outcomes distinct: a successful empty result contains no matching observations, while an explicit numeric zero is a measurement. An invalid query, authentication failure, or timeout establishes neither. Correct invalid arguments or expressions using the actual schema and discovered names, then retry the scoped read; otherwise report the blocking error rather than treating it as absence. +Keep query outcomes distinct: a successful empty result means the executed expression returned no values, while an explicit numeric zero is a measurement. An empty derived expression, such as `_sum / _count`, does not prove that raw observations are missing. Query each raw operand with the same verified scope and window before claiming no recent data; check for missing samples, zero or invalid denominators, and vector-matching label differences. An invalid query, authentication failure, or timeout establishes neither zero nor absence. Correct invalid arguments or expressions using the actual schema and discovered names, then retry the scoped read; otherwise report the blocking error rather than treating it as absence. Discovery example for an RDS investigation, after substituting the selected datasource and UTC bounds: From 78eb1738467a347001409e5b39e17e7557c6e8be Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Tue, 8 Sep 2026 07:23:51 +0200 Subject: [PATCH 5/8] fix(skills): focus CloudWatch evidence and freshness --- skills/last9-cloudwatch/SKILL.md | 103 ++++++++++++++----------------- 1 file changed, 48 insertions(+), 55 deletions(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index ddf3470..d673efc 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -6,45 +6,40 @@ metadata: author: last9 --- -# last9-cloudwatch — investigate CloudWatch metrics in Last9 +# last9-cloudwatch -**Operating principle: establish the resource and the published statistic before choosing the calculation.** CloudWatch metrics can describe AWS resources without an application service, environment label, or APM instrumentation. Start with the requested AWS population; do not require traces to query its infrastructure metrics. +**Establish resource and statistic before calculating.** CloudWatch needs no application service, environment label, or APM instrumentation. ## Prerequisites and scope -Use the authenticated [Last9 MCP server](https://github.com/last9/last9-mcp-server) and its advertised tool schemas. This workflow reads existing telemetry. Creating streams, changing AWS permissions, configuring collectors, generating Terraform, and building dashboards are separate tasks. +Use the authenticated [Last9 MCP server](https://github.com/last9/last9-mcp-server) for read-only queries. AWS/collector changes, Terraform, and dashboards are separate tasks. -Confirm the connected organization and selected datasource from the conversation and available connection or datasource information. Call `list_datasources` when selection is not already established. Carry the selected datasource into every discovery and query call; a different default is not a reason to change the user's selection. If the connection cannot access the intended organization, report that boundary before querying another one. +Establish the connected organization, selected datasource, AWS account, region, resource, and UTC bounds. Use `list_datasources` if unresolved; carry the selection into every call regardless of defaults. Resolve relative times once with an available clock. Datasource/integration configuration can establish account/region without labels; cite that evidence. Resolve ambiguous scope before combining data; never substitute an accessible organization for the intended one. -Pin the AWS account, region, resource, and UTC start/end times. Resolve relative times against an available clock once and reuse those bounds for comparisons. Discover the actual dimension keys before inserting filters. Account or region may be established by a dedicated datasource or integration configuration instead of a series label; state that evidence. If scope remains ambiguous, surface the available choices and resolve it before combining measurements. - -If the needed tools are unavailable, explain the missing connection or capability. Analyze supplied observations when sufficient, but identify proposed queries as unexecuted. Never present invented calls or results as evidence. +If tools are missing, use sufficient supplied observations or report the capability gap. Mark unexecuted queries; never invent results. ## Tool reference -Read the live descriptions first; these are the relevant parameter shapes, not a replacement for the installed schemas. - -| Tool | Use | Parameters | -|---|---|---| -| `list_datasources` | Resolve the datasource | Use its advertised schema | -| `prometheus_label_values` | Discover metric names or dimension values | `label`, `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | -| `prometheus_labels` | Discover label names for a selector | `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | -| `prometheus_instant_query` | Evaluate an expression at a fixed time, or read a raw range selector | `query`, `datasource`, `time_iso` | -| `prometheus_range_query` | Inspect an expression over the requested interval | `query`, `datasource`, `start_time_iso`, `end_time_iso` | +Installed schemas take precedence over this reference: -Use `label: "__name__"` for metric-name discovery. These query tools do not expose a `step` parameter in this interface; do not invent one. Preserve the distinction between a raw range selector evaluated once and a chart expression evaluated repeatedly. +| Tool | Relevant parameters | +|---|---| +| `list_datasources` | Advertised schema | +| `prometheus_label_values` | `label`, `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | +| `prometheus_labels` | `match_query`, `datasource`, `start_time_iso`, `end_time_iso` | +| `prometheus_instant_query` | `query`, `datasource`, `time_iso` | +| `prometheus_range_query` | `query`, `datasource`, `start_time_iso`, `end_time_iso` | -## Discover, verify, then calculate +Discover names with `label: "__name__"`. Do not invent a `step` parameter. Distinguish raw range selectors evaluated once from repeatedly evaluated charts. -1. **Find candidate names.** Search the selected datasource and time window for the requested AWS namespace or metric. The [CloudWatch integration guide](https://last9.io/docs/integrations/observability/aws-cloudwatch-metrics/) documents the `amazonaws_com_AWS` prefix for its stream path; use it as a discovery hint, not a universal name contract. Exporters and other ingestion paths can use different names. A missing result under one prefix is not proof that the resource has no metrics. -2. **Inspect dimensions and actual samples.** Discover label names and relevant values, then execute a narrow selector for the chosen account, region, and resource. A catalog entry proves discoverability, not current delivery. `prometheus_labels` may return a generic catalog even with `match_query`; if it does, state that limitation and inspect a bounded candidate metric query's actual returned label keys. Inspect returned series labels and raw timestamps before trusting a filter or interpreting emptiness. Do not add `service_name`, `env`, `namespace`, or a guessed AWS dimension simply because it appears on another metric. -3. **Identify lineage and one dimension level.** Distinguish CloudWatch Metric Streams, exporters, and trace-derived metrics using integration information and observed series. Preserve valid mixed configurations. Keep sources separate until their populations, periods, and definitions justify combining them; overlapping copies cannot be added. Similarly, choose instance, cluster, or role aggregation deliberately. Removing labels with `sum by (...)` after selecting overlapping rows does not remove double counting. -4. **Establish the measurement contract.** Record the AWS metric, unit, statistic, publication period, timestamp meaning, and coverage. Check whether a sample is a period summary, instantaneous gauge, or cumulative counter. Establish units from the AWS metric definition and the ingestion mapping; a unit label is not guaranteed, and magnitude alone is not evidence. Follow the current integration guide and verify additional statistics and the actual ingestion format when relevant; do not infer a permanent format-version requirement from an old example. If dimensions appear only as an opaque encoded value, report the observed shape and the missing resource-level selection capability; do not invent direct label keys or recreate the stream. -5. **Execute the required calculation.** Build queries from discovered names and verified selectors. Execute them with the pinned datasource and times, inspect results, and cross-check units and arithmetic. If the inputs cannot support the requested statistic, return it as unavailable and name the evidence needed next. +## Discover and execute efficiently -Keep query outcomes distinct: a successful empty result means the executed expression returned no values, while an explicit numeric zero is a measurement. An empty derived expression, such as `_sum / _count`, does not prove that raw observations are missing. Query each raw operand with the same verified scope and window before claiming no recent data; check for missing samples, zero or invalid denominators, and vector-matching label differences. An invalid query, authentication failure, or timeout establishes neither zero nor absence. Correct invalid arguments or expressions using the actual schema and discovered names, then retry the scoped read; otherwise report the blocking error rather than treating it as absence. +1. **Find names and inspect samples.** The [integration guide](https://last9.io/docs/integrations/observability/aws-cloudwatch-metrics/) documents `amazonaws_com_AWS` for its stream path. Treat it as a discovery hint: exporters can use other names. Missing one prefix does not prove missing resource metrics. Discover dimensions, then execute a bounded selector. Catalog membership proves discoverability, not delivery. `prometheus_labels` may return generic keys even with `match_query`; state that limitation and use actual returned series labels. Do not invent `service_name`, `env`, `namespace`, or AWS dimensions from another metric's catalog. +2. **Verify lineage and population.** Distinguish Metric Streams, exporters, and traces using integration information and observed series. Preserve valid mixed sources; combine only when definitions, periods, and populations justify it. Choose instance, cluster, or role scope deliberately. Removing labels after selecting overlapping copies or rollups does not remove double counting. +3. **Establish semantics.** Identify the metric, unit, statistic, period, timestamp meaning, and coverage; distinguish period summaries, gauges, and cumulative counters. Use AWS definitions and ingestion mapping for units, not magnitude or an assumed unit label. Check the current guide and observed format/additional statistics rather than imposing a historical format version. If dimensions are opaque encoded values, report the resource-selection limitation; do not invent direct keys or recreate the stream. +4. **Batch independent reads and reuse evidence.** Once scopes are known, batch companion Sum/Count reads and independent discovery across requested metrics. Reuse verified names, labels, raw samples, and calculations at the same scope/time. Skip redundant catalog lookups and extra calculations: a latest-period request does not need a separate weighted window average. Execute the required expression or derive the requested result from returned raw operands, check arithmetic and units, then report as soon as the requested measurements are supported. If support is unavailable, report the gap and needed evidence instead of expanding the investigation. -Discovery example for an RDS investigation, after substituting the selected datasource and UTC bounds: +`prometheus_label_values` example: substitute datasource/UTC bounds and add verified resource filters. ```json { @@ -56,23 +51,23 @@ Discovery example for an RDS investigation, after substituting the selected data } ``` -This is input to `prometheus_label_values`, not a query result. Narrow discovery further with already-verified resource filters. Use returned metric names in the subsequent label and sample reads. +## Query outcomes and statistics -## Statistics and calculation guardrails +A successful empty expression returned no values; an explicit numeric zero is a measurement. An empty ratio does **not** prove missing raw observations. Inspect each raw operand at the same verified scope/window for absent samples, zero or invalid denominators, and vector-matching differences. Invalid queries, authentication failures, and timeouts establish neither zero nor absence. Repair invalid arguments/expressions using the actual schema and discovered names, retry the scoped read, or report the blocking error. -[CloudWatch Metric Streams](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-Metric-Streams.html) carry period statistics including Sum and SampleCount; additional statistics can be configured. The [OpenTelemetry translation](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-metric-streams-formats-opentelemetry-translation.html) explains the summary mapping. Verify that mapping for the series being queried before applying these recipes. +[Metric Streams](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-Metric-Streams.html) carry period Sum and SampleCount plus configurable statistics. Verify the series' [summary mapping](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/CloudWatch-metric-streams-formats-opentelemetry-translation.html) before using these recipes: -| Observed measurement | Interpretation and calculation | +| Observed measurement | Calculation and guardrail | |---|---| -| Confirmed stream-summary `_sum` and `_count` companions | Period Sum and SampleCount. `_sum / _count` gives the period's sample average when labels, period, and population match and count is positive. `_count` counts observations; it is not automatically a request count. | -| Several confirmed, disjoint period summaries | A sample-weighted window average is total Sum divided by total SampleCount. Averaging the period averages is wrong when counts differ. Verify coverage and boundaries before claiming a whole-window result. | -| An AWS metric whose Sum counts events | Add disjoint period Sums for an event total; divide by covered elapsed seconds for average events/sec only with the required complete coverage. Establish this from the metric's definition, not its suffix. | -| A gauge such as storage size or resource utilization | Report the requested level or a clearly defined aggregate. Summing observations over time does not give total storage or total utilization. | -| A verified cumulative exporter counter | `rate()` or `increase()` may apply with suitable history and reset handling. Do not transfer that treatment to CloudWatch period-summary companions because their names end in `_sum` or `_count`. | -| Observed minimum, maximum, or percentile series | Select the actual statistic label, including `quantile` where present. A maximum of period p99s is the peak period p99; it cannot establish a whole-window p99. Do not run `histogram_quantile()` on already-quantiled values. | -| Instance rows plus cluster, role, or other rollup rows | Select a non-overlapping population at one intended dimension level. Never sum both an aggregate and its constituent resources. | +| Confirmed `_sum` / `_count` companions | Period Sum / SampleCount is the sample average when labels, period, and population match and count is positive. SampleCount counts observations, not automatically requests. | +| Disjoint period summaries | Window average = total Sum / total SampleCount. Averaging period averages is wrong when counts differ. Whole-window claims require matching coverage and boundaries. | +| Metric whose Sum counts events | Add disjoint period Sums; divide by elapsed seconds for average events/sec only with complete coverage. Establish event semantics from the metric definition, not its suffix. | +| Gauge: storage or utilization | Report the requested level or defined aggregate. Adding observations over time is not total storage or utilization. | +| Verified cumulative exporter counter | `rate()` / `increase()` may apply with sufficient history and reset handling. Never apply them to period summaries merely because of `_sum` / `_count` suffixes. | +| Minimum, maximum, percentile | Select the actual statistic label (`quantile` if present). Maximum period p99 is peak period p99, not whole-window p99. Do not apply `histogram_quantile()` to already-quantiled values. | +| Instance, cluster, role, or other rollups | Select one non-overlapping population; never add an aggregate and its constituents. | -For a confirmed summary family, the following are **templates**. Replace metric names, selectors, and window with observed values before execution: +For verified summaries, substitute observed names, selectors, and window: ```promql {} / {} @@ -84,38 +79,36 @@ sum(sum_over_time({}[])) sum(sum_over_time({}[])) ``` -The first expression needs one-to-one matching companion labels. The second additionally needs matching populations and periods, positive total SampleCount, and raw samples representing disjoint reports. Do not add arbitrary label-dropping modifiers to make an unexplained mismatch disappear. Do not silently turn a zero or missing denominator into zero utilization or latency. +The period ratio needs one-to-one companion matching. The window ratio additionally needs disjoint reports, matching populations/periods, and positive total SampleCount. Do not hide unexplained mismatches by dropping labels or convert a missing/zero denominator into zero latency or utilization. -Inspect raw sample times, for example with `{}[]` through `prometheus_instant_query` at the fixed end time. Check the returned result type and timestamp meaning: use a matrix's timestamps as publication evidence only when they represent original samples. An evaluated vector's tuple timestamps do not establish raw publication times. Chart points may repeat or resample earlier observations; they are not automatically independent stream publications. Missing periods, duplicate delivery, late data, boundary misalignment, or unknown timestamp semantics limit a whole-window total or average. State the supported coverage rather than filling gaps with zero. +Inspect raw times using `{}[]` through `prometheus_instant_query` at the fixed end. Check result type and timestamp meaning: source observation times need original samples. Vector tuples and chart points can carry evaluation times or repeat earlier observations. Missing periods, duplicates, late data, misaligned boundaries, or unknown timestamp semantics limit totals/averages; report supported coverage instead of filling gaps with zero. ## Worked example: RDS or Aurora CPU and read latency -For “show CPU and read latency for this database over this interval”: +Resolve account, region, and database dimensions. For Aurora, choose instance, cluster, or role: AWS publishes [different dimension combinations](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/dimensions.html), so selecting every row with a cluster name can overlap resources. + +Discover CPUUtilization and ReadLatency families, batch their companion reads, and verify source, labels, timestamps, and units. For stream summaries, use the latest companion pair for requested period averages; use the weighted template only for a requested window average with verified coverage. Report each metric's aggregation level separately. -1. Resolve the chosen account, region, and database using actual dimensions. For Aurora, determine whether the request concerns an instance, the cluster, or a role. AWS publishes [different dimension combinations](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/dimensions.html); inspect the selected series instead of combining every row with the cluster's name. -2. Discover the CPUUtilization and ReadLatency families in that scope. Verify whether the returned metrics are stream-summary companions, exporter gauges, or a different source. Read a small raw window and inspect labels, timestamps, and units before building the final expression. -3. If the observed family uses period Sum/SampleCount, execute the companion ratio for the period averages. For a requested full-window average, execute the weighted template only after verifying its coverage conditions. Keep CPU and latency calculations separate, and report their actual aggregation level. -4. [RDS CloudWatch metrics](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) define CPUUtilization as Percent and ReadLatency as seconds. Confirm the ingested unit mapping. Convert seconds to milliseconds with a factor of 1,000 only when that is the observed input unit. Enhanced Monitoring or exporter metrics with similar names can have different definitions or units. -5. Report the measured values with the exact executed selectors and interval, or say which result is unavailable. Without distributions or equivalent raw observations, period latency averages and percentiles do not establish a whole-window latency percentile. +[RDS metrics](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) define CPUUtilization in Percent and ReadLatency in seconds. Confirm ingestion mapping before converting latency to milliseconds by multiplying by 1,000. Enhanced Monitoring and similarly named exporters can differ in units/definitions. Period latency averages or percentiles cannot establish whole-window percentiles without distributions or equivalent raw observations. -## Sparse metrics: S3 daily storage +## Sparse S3 storage: last-known is not current -An empty recent lookup for BucketSizeBytes or NumberOfObjects does not establish zero, deletion, or a stopped stream. These are [daily S3 storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics. AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html) for viewing them. +BucketSizeBytes and NumberOfObjects are [daily storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics; AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html). An empty recent lookup does not establish zero, deletion, or a stopped stream. -Discover the actual bucket and storage-type dimensions. Read bounded history spanning the daily cadence, such as the last three days, and inspect the timestamps of actual samples. Keep this diagnostic history separate from the interval the user asked about. Distinguish a numeric zero, a last known observation, and no observation in the requested interval. Do not claim that a `last_over_time()` result's evaluation timestamp is the original publication time; use raw sample timestamps to establish the age. Preserve the exact returned timestamp: copy raw Unix seconds when needed, or verify the UTC conversion with an available tool. If no reliable conversion tool is available, report the exact epoch only and omit converted UTC dates, times, and computed ages. Do not round it to a date or midnight. +Discover bucket/storage-type dimensions. Inspect bounded raw history spanning the cadence (for example, three days), separately from the user's requested interval. Distinguish numeric zero, last-known value with original timestamp, and no observation in the requested interval. -If expected history is also empty, verify the datasource, filters, source configuration, statistic, and publication cadence before attributing the gap to ingestion. Use available read-only delivery evidence when present. If it is absent, report that whether delivery stopped remains unconfirmed. Do not infer the next publication date or time from “daily” or from the previous sample. +**Daily cadence does not make an old sample current.** Honor a user-specified freshness window; otherwise probe each raw companion at the requested end time, separately from widened history. Do not invent a production freshness threshold. Widened history or a historical `last_over_time()` result cannot establish currentness; its evaluation time is not source publication time. If the observations do not establish currentness, report current **unavailable** and the last-known value separately. -## Metric availability and Discover visibility +Preserve exact returned Unix seconds. Verify UTC conversion and any age calculation with a reliable available tool; otherwise report the epoch only and omit converted dates/times and computed ages. `vector(epoch)` merely echoes a number, not a verified time conversion. Never round a source timestamp to a date or midnight. -A resource can have queryable CloudWatch metrics without appearing in a particular Discover view. Verify ingestion by executing a scoped metric query in the selected datasource. Investigate Discover identity and supported resource discovery separately, using the actual identity dimensions and current integration guidance. Missing application traces or a missing Discover row is not evidence that CloudWatch ingestion failed. +If history is empty, verify datasource, filters, configuration, statistic, and cadence before blaming ingestion. Without read-only delivery evidence, whether delivery stopped remains unconfirmed. Neither “daily” nor the previous sample establishes the next publication time **or that a newer publication has not arrived**. -## Report the evidence +## Report evidence and separate Discover identity -For each requested measurement, preserve the user's requested output name and give the value and unit or **unavailable**, the AWS resource and dimension level, datasource and UTC interval, selected source/statistic, and the exact executed query with its returned evidence or tool-call reference. Cite all operands actually used, including both Sum and SampleCount when computing an average. Separate observations from explanations and unresolved causes. Report partial coverage, last-known sample times, and any unsupported requested statistic explicitly; check that the prose arithmetic agrees with the samples and the stated result. +Keep requested output names. Give value/unit or **unavailable**, resource/dimension level, datasource/UTC interval, source/statistic, and executed queries with evidence/call references. Cite every operand actually used, including Sum and SampleCount for an average. State partial coverage, last-known times, unsupported statistics, and unresolved causes; check prose arithmetic against the samples and result. -For an ingestion-versus-UI question, state what the metric query established and what remains unknown about discovery. Link the [Last9 metric explorer](https://app.last9.io/metrics) when useful. +A resource may have queryable metrics without a Discover row. Establish availability with a scoped metric query; investigate Discover identity and supported resource profiles separately using actual dimensions and current guidance. Missing traces or a Discover row does not prove ingestion failure. Link the [metric explorer](https://app.last9.io/metrics) when useful. ## Related skills -Use `last9-logs` for a separate log investigation or `last9-traces` for application span analysis when relevant and available. Neither is a prerequisite for this CloudWatch workflow. +Use available `last9-logs` or `last9-traces` for separate log/span tasks; neither is required. From 6c9481f4a159c2ef3055313d2ae0229534a1bc97 Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Tue, 8 Sep 2026 08:36:58 +0200 Subject: [PATCH 6/8] fix(skills): distinguish CloudWatch window membership and freshness --- skills/last9-cloudwatch/SKILL.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index d673efc..9a7b658 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -81,7 +81,7 @@ sum(sum_over_time({}[])) The period ratio needs one-to-one companion matching. The window ratio additionally needs disjoint reports, matching populations/periods, and positive total SampleCount. Do not hide unexplained mismatches by dropping labels or convert a missing/zero denominator into zero latency or utilization. -Inspect raw times using `{}[]` through `prometheus_instant_query` at the fixed end. Check result type and timestamp meaning: source observation times need original samples. Vector tuples and chart points can carry evaluation times or repeat earlier observations. Missing periods, duplicates, late data, misaligned boundaries, or unknown timestamp semantics limit totals/averages; report supported coverage instead of filling gaps with zero. +Inspect raw times using `{}[]` through `prometheus_instant_query` at the fixed end. Use verified raw-selector bounds and sample timestamps for window membership; do not contradict them with unverified UTC conversions. Membership alone does not establish complete period coverage. Check result type and timestamp meaning: source observation times need original samples. Vector tuples and chart points can carry evaluation times or repeat earlier observations. Missing periods, duplicates, late data, misaligned boundaries, or unknown timestamp semantics limit totals/averages; report supported coverage instead of filling gaps with zero. ## Worked example: RDS or Aurora CPU and read latency @@ -97,7 +97,7 @@ BucketSizeBytes and NumberOfObjects are [daily storage metrics](https://docs.aws Discover bucket/storage-type dimensions. Inspect bounded raw history spanning the cadence (for example, three days), separately from the user's requested interval. Distinguish numeric zero, last-known value with original timestamp, and no observation in the requested interval. -**Daily cadence does not make an old sample current.** Honor a user-specified freshness window; otherwise probe each raw companion at the requested end time, separately from widened history. Do not invent a production freshness threshold. Widened history or a historical `last_over_time()` result cannot establish currentness; its evaluation time is not source publication time. If the observations do not establish currentness, report current **unavailable** and the last-known value separately. +**Daily cadence does not make an old sample current.** A historical query interval is not a freshness requirement unless the user explicitly makes it one. Honor a user-specified freshness window; otherwise probe each raw companion at the requested end time, separately from widened history. Do not invent a production freshness threshold. Widened history or a historical `last_over_time()` result cannot establish currentness; its evaluation time is not source publication time. If the observations do not establish currentness, report current **unavailable** and the last-known value separately. Preserve exact returned Unix seconds. Verify UTC conversion and any age calculation with a reliable available tool; otherwise report the epoch only and omit converted dates/times and computed ages. `vector(epoch)` merely echoes a number, not a verified time conversion. Never round a source timestamp to a date or midnight. From 34ccb0cbcc6f8d0e9d4409774fc2412f660a3940 Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Tue, 8 Sep 2026 22:15:02 +0200 Subject: [PATCH 7/8] feat(skills): cover all CloudWatch catalog families --- README.md | 2 +- scripts/check-skill-pack-selftest.sh | 67 ++++++++++++++++++- scripts/check-skill-pack.sh | 55 +++++++++++---- skills/last9-cloudwatch/SKILL.md | 37 +++++----- skills/last9-cloudwatch/references/billing.md | 19 ++++++ skills/last9-cloudwatch/references/dms.md | 17 +++++ .../last9-cloudwatch/references/dynamodb.md | 23 +++++++ skills/last9-cloudwatch/references/ec2.md | 22 ++++++ .../references/elasticache.md | 20 ++++++ skills/last9-cloudwatch/references/kms.md | 19 ++++++ skills/last9-cloudwatch/references/msk.md | 19 ++++++ .../last9-cloudwatch/references/rds-aurora.md | 11 +++ skills/last9-cloudwatch/references/s3.md | 21 ++++++ skills/last9-cloudwatch/references/sqs.md | 19 ++++++ 14 files changed, 318 insertions(+), 33 deletions(-) create mode 100644 skills/last9-cloudwatch/references/billing.md create mode 100644 skills/last9-cloudwatch/references/dms.md create mode 100644 skills/last9-cloudwatch/references/dynamodb.md create mode 100644 skills/last9-cloudwatch/references/ec2.md create mode 100644 skills/last9-cloudwatch/references/elasticache.md create mode 100644 skills/last9-cloudwatch/references/kms.md create mode 100644 skills/last9-cloudwatch/references/msk.md create mode 100644 skills/last9-cloudwatch/references/rds-aurora.md create mode 100644 skills/last9-cloudwatch/references/s3.md create mode 100644 skills/last9-cloudwatch/references/sqs.md diff --git a/README.md b/README.md index ccfe66f..7821f37 100644 --- a/README.md +++ b/README.md @@ -16,7 +16,7 @@ MCP gives your agent access. Skills give it judgment. | [`go-agent-install`](skills/go-agent-install/SKILL.md) | Instrument a Go service with Last9 go-agent: detect the stack, wire chi + `database/sql` tracing, promote opt-in body capture, and verify spans land — without double-instrumenting | | [`last9-logs`](skills/last9-logs/SKILL.md) | Log investigation: scope to a service first, attribute filters over body search, aggregate before drilling into raw lines | | [`last9-traces`](skills/last9-traces/SKILL.md) | Trace investigation: a five-question interview that lands on the right tool call, plus a `tracejson` syntax reference card | -| [`last9-cloudwatch`](skills/last9-cloudwatch/SKILL.md) | CloudWatch investigation: discover AWS resource dimensions, interpret period statistics and units, and distinguish sparse data from missing ingestion | +| [`last9-cloudwatch`](skills/last9-cloudwatch/SKILL.md) | CloudWatch investigation across Billing, RDS/Aurora, ElastiCache, MSK, DynamoDB, EC2, SQS, DMS, KMS, and S3; focused family references share discovery, statistic, and evidence rules | ## Installation diff --git a/scripts/check-skill-pack-selftest.sh b/scripts/check-skill-pack-selftest.sh index be61527..8a5ed74 100755 --- a/scripts/check-skill-pack-selftest.sh +++ b/scripts/check-skill-pack-selftest.sh @@ -21,7 +21,7 @@ setup_fixture() { "name": "opencode-last9", "version": "0.0.0-test", "scripts": { - "prepack": "mkdir -p skills && for d in ../../skills/*/; do n=$(basename $d); mkdir -p skills/$n; cp ${d}SKILL.md skills/$n/SKILL.md; done" + "prepack": "mkdir -p skills && cp -R ../../skills/. skills/" } } PKG @@ -99,4 +99,69 @@ cp "$FIX/skills/last9-logs/SKILL.md" "$FIX/plugins/acme/skills/rogue/SKILL.md" git -C "$FIX" add -A && git -C "$FIX" -c user.email=t@t -c user.name=t commit -qm fault expect_fail "committed plugin skill copy" +# References must ship exactly when tracked, without permitting arbitrary payloads. +setup_fixture reference-happy +mkdir -p "$FIX/skills/last9-logs/references" +printf 'focused reference\n' > "$FIX/skills/last9-logs/references/family.md" +commit_fault +if ! run_full sh "$FIX/scripts/check-skill-pack.sh" >/dev/null 2>&1; then + echo "selftest FAILED: tracked Markdown reference expected exit 0" >&2 + exit 1 +fi + +setup_fixture reference-missing +mkdir -p "$FIX/skills/last9-logs/references" +printf 'focused reference\n' > "$FIX/skills/last9-logs/references/family.md" +commit_fault +# Simulate a packer that omits a tracked reference. +jq '.scripts.prepack += " && rm skills/last9-logs/references/family.md"' "$FIX/plugins/opencode-last9/package.json" > "$FIX/package.tmp" +mv "$FIX/package.tmp" "$FIX/plugins/opencode-last9/package.json" +expect_fail "missing packaged reference" + +setup_fixture reference-untracked +mkdir -p "$FIX/skills/last9-logs/references" +printf 'untracked experiment\n' > "$FIX/skills/last9-logs/references/untracked.md" +expect_fail "untracked packaged reference" + +setup_fixture arbitrary-payload +mkdir -p "$FIX/skills/last9-logs/scripts" +printf 'echo unexpected\n' > "$FIX/skills/last9-logs/scripts/run.sh" +commit_fault +expect_fail "arbitrary tracked script" + +setup_fixture reference-symlink +mkdir -p "$FIX/skills/last9-logs/references" +printf 'outside skill\n' > "$FIX/outside.md" +ln -s ../../../outside.md "$FIX/skills/last9-logs/references/escape.md" +commit_fault +expect_fail "tracked reference symlink" + +setup_fixture reference-parent-symlink +mkdir -p "$FIX/skills/last9-logs/references" +printf 'focused reference\n' > "$FIX/skills/last9-logs/references/family.md" +commit_fault +mv "$FIX/skills/last9-logs/references" "$FIX/outside-references" +ln -s ../../outside-references "$FIX/skills/last9-logs/references" +expect_fail "working-tree reference parent symlink" + +setup_fixture reference-without-entrypoint +mkdir -p "$FIX/skills/orphan/references" +printf 'orphan\n' > "$FIX/skills/orphan/references/family.md" +commit_fault +expect_fail "reference without entrypoint" + +# The gate must never delete or inspect another invocation's archive. +setup_fixture archive-isolation +ARCHIVE_TMP="$SANDBOX/archive-temp" +mkdir -p "$ARCHIVE_TMP" +printf 'unrelated archive\n' > "$ARCHIVE_TMP/sentinel.tgz" +if ! TMPDIR="$ARCHIVE_TMP" run_full sh "$FIX/scripts/check-skill-pack.sh" >/dev/null 2>&1; then + echo "selftest FAILED: isolated archive pack expected exit 0" >&2 + exit 1 +fi +if [ ! -f "$ARCHIVE_TMP/sentinel.tgz" ] || [ "$(cat "$ARCHIVE_TMP/sentinel.tgz")" != "unrelated archive" ]; then + echo "selftest FAILED: unrelated archive was changed or deleted" >&2 + exit 1 +fi + echo "check-skill-pack selftests passed" diff --git a/scripts/check-skill-pack.sh b/scripts/check-skill-pack.sh index 907c5ad..91450a0 100755 --- a/scripts/check-skill-pack.sh +++ b/scripts/check-skill-pack.sh @@ -20,6 +20,28 @@ if [ -n "$(git ls-files 'plugins/*/skills/*')" ]; then exit 1 fi +# Canonical payload is deliberately narrow: one entrypoint plus optional direct +# Markdown references. Reject links before prepack can follow them outside skills/. +for payload in $(git ls-files 'skills/**'); do + if ! printf '%s\n' "$payload" | grep -Eq '^skills/[a-z0-9-]+/(SKILL\.md|references/[a-z0-9-]+\.md)$'; then + echo "::error::unsupported canonical skill payload: $payload" >&2 + exit 1 + fi + skill_dir="$(printf '%s\n' "$payload" | cut -d/ -f1,2)" + if ! git ls-files --error-unmatch "$skill_dir/SKILL.md" >/dev/null 2>&1; then + echo "::error::skill reference has no tracked entrypoint: $payload" >&2 + exit 1 + fi + if [ -L skills ] || [ -L "$skill_dir" ] || [ -L "$skill_dir/references" ] || [ -L "$payload" ] || [ ! -f "$payload" ]; then + echo "::error::skill payload must be a regular file without symlink parents: $payload" >&2 + exit 1 + fi + if git ls-files --stage -- "$payload" | grep -q '^120000 '; then + echo "::error::tracked skill symlinks are forbidden: $payload" >&2 + exit 1 + fi +done + # 1. Frontmatter name must equal the skill directory name. Extract from the # YAML frontmatter block only (between the first two --- delimiters), so a # body line starting "name: " cannot false-fail the gate. @@ -90,27 +112,34 @@ jq -e '.skills == "./skills/"' .codex-plugin/plugin.json >/dev/null || { # tarball so npm failures fail fast and listing comes from tar, not logs. cd plugins/opencode-last9 npm run prepack >/dev/null -tgz="$(mktemp "${TMPDIR:-/tmp}/skill-pack.XXXXXX.tgz")" -npm pack --pack-destination "$(dirname "$tgz")" --silent >/dev/null 2>&1 || { +PACK_DIR="$(mktemp -d "${TMPDIR:-/tmp}/skill-pack.XXXXXX")" +trap 'rm -rf "$PACK_DIR"' EXIT +npm pack --pack-destination "$PACK_DIR" --silent >/dev/null 2>&1 || { echo "::error::npm pack failed" >&2 exit 1 } -tar -tzf "$(ls -t "$(dirname "$tgz")"/*.tgz | head -1)" > "$tgz.list" -rm -f "$(dirname "$tgz")"/*.tgz +set -- "$PACK_DIR"/*.tgz +if [ "$#" -ne 1 ] || [ ! -f "$1" ]; then + echo "::error::npm pack did not produce exactly one archive" >&2 + exit 1 +fi +tgz="$1" +tar -tzf "$tgz" > "$tgz.list" cd "$ROOT_DIR" missing=0 -for skill_md in $(git ls-files 'skills/*/SKILL.md'); do - grep -q "^package/skills/${skill_md#skills/}$" "$tgz.list" || { - echo "::error::opencode tarball missing canonical skill: $skill_md" >&2 +for payload in $(git ls-files 'skills/**'); do + grep -Fqx "package/$payload" "$tgz.list" || { + echo "::error::opencode tarball missing canonical skill payload: $payload" >&2 missing=1 } done -extras=$(grep -E "^package/skills/" "$tgz.list" | grep -vE "^package/skills/[^/]+/SKILL\.md$" || true) -if [ -n "$extras" ]; then - echo "::error::opencode tarball ships unexpected skills payload:" >&2 - echo "$extras" >&2 - missing=1 -fi +for packed in $(grep -E '^package/skills/' "$tgz.list" || true); do + payload="${packed#package/}" + if ! printf '%s\n' "$payload" | grep -Eq '^skills/[a-z0-9-]+/(SKILL\.md|references/[a-z0-9-]+\.md)$' || ! git ls-files --error-unmatch -- "$payload" >/dev/null 2>&1; then + echo "::error::opencode tarball ships unexpected or untracked skills payload: $packed" >&2 + missing=1 + fi +done rm -f "$tgz.list" [ "$missing" -eq 0 ] || exit 1 diff --git a/skills/last9-cloudwatch/SKILL.md b/skills/last9-cloudwatch/SKILL.md index 9a7b658..37ba173 100644 --- a/skills/last9-cloudwatch/SKILL.md +++ b/skills/last9-cloudwatch/SKILL.md @@ -1,6 +1,6 @@ --- name: last9-cloudwatch -description: Investigate AWS CloudWatch metrics in Last9 through read-only MCP queries. Use for CloudWatch metric discovery, RDS or Aurora resource metrics, CloudWatch statistics and units, or sparse S3 daily metrics ("CloudWatch in Last9", "RDS CPU", "Aurora latency", "S3 metric is empty"). +description: Investigate AWS CloudWatch metrics in Last9 with read-only MCP queries. Covers Billing, RDS/Aurora, ElastiCache, MSK, DynamoDB, EC2, SQS, DMS, KMS, and S3. Discover resource metrics, interpret statistics and units, and distinguish sparse data from missing delivery. compatibility: Requires the Last9 MCP server connected to the session metadata: author: last9 @@ -18,6 +18,23 @@ Establish the connected organization, selected datasource, AWS account, region, If tools are missing, use sufficient supplied observations or report the capability gap. Mark unexecuted queries; never invent results. +## Choose the AWS family + +Before family-specific discovery or queries, read the matching reference. For a task spanning multiple families, load only those matching references. Namespace names below are discovery hints, not guaranteed ingested names or label spellings. Use the shared rules below for every source. + +| Family | Namespace hint | Read when investigating | +|---|---|---| +| [Billing](references/billing.md) | `AWS/Billing` or a verified cost exporter | Estimated charges, currency, service and linked-account cost scope | +| [RDS / Aurora](references/rds-aurora.md) | `AWS/RDS` | Instance versus cluster/role metrics, CPU, latency, replication | +| [ElastiCache](references/elasticache.md) | `AWS/ElastiCache` | Cache node identity, engine versus host CPU, hits, evictions, lag | +| [MSK](references/msk.md) | `AWS/Kafka` | Broker versus topic/consumer-group metrics, throughput, consumer lag | +| [DynamoDB](references/dynamodb.md) | `AWS/DynamoDB` | Table/index/account scope, consumed capacity, throttling, latency | +| [EC2](references/ec2.md) | `AWS/EC2` | Instance CPU, period network bytes, status checks, credit balances | +| [SQS](references/sqs.md) | `AWS/SQS` | Approximate backlog and age, message-operation counts, inactive queues | +| [DMS](references/dms.md) | `AWS/DMS` | Replication task versus instance, source/target CDC latency | +| [KMS](references/kms.md) | `AWS/KMS` | Operation counts, key material expiration, metric applicability | +| [S3](references/s3.md) | `AWS/S3` | Daily storage versus request metrics, last-known versus current | + ## Tool reference Installed schemas take precedence over this reference: @@ -83,26 +100,10 @@ The period ratio needs one-to-one companion matching. The window ratio additiona Inspect raw times using `{}[]` through `prometheus_instant_query` at the fixed end. Use verified raw-selector bounds and sample timestamps for window membership; do not contradict them with unverified UTC conversions. Membership alone does not establish complete period coverage. Check result type and timestamp meaning: source observation times need original samples. Vector tuples and chart points can carry evaluation times or repeat earlier observations. Missing periods, duplicates, late data, misaligned boundaries, or unknown timestamp semantics limit totals/averages; report supported coverage instead of filling gaps with zero. -## Worked example: RDS or Aurora CPU and read latency - -Resolve account, region, and database dimensions. For Aurora, choose instance, cluster, or role: AWS publishes [different dimension combinations](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/dimensions.html), so selecting every row with a cluster name can overlap resources. - -Discover CPUUtilization and ReadLatency families, batch their companion reads, and verify source, labels, timestamps, and units. For stream summaries, use the latest companion pair for requested period averages; use the weighted template only for a requested window average with verified coverage. Report each metric's aggregation level separately. - -[RDS metrics](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) define CPUUtilization in Percent and ReadLatency in seconds. Confirm ingestion mapping before converting latency to milliseconds by multiplying by 1,000. Enhanced Monitoring and similarly named exporters can differ in units/definitions. Period latency averages or percentiles cannot establish whole-window percentiles without distributions or equivalent raw observations. - -## Sparse S3 storage: last-known is not current - -BucketSizeBytes and NumberOfObjects are [daily storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics; AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html). An empty recent lookup does not establish zero, deletion, or a stopped stream. - -Discover bucket/storage-type dimensions. Inspect bounded raw history spanning the cadence (for example, three days), separately from the user's requested interval. Distinguish numeric zero, last-known value with original timestamp, and no observation in the requested interval. - -**Daily cadence does not make an old sample current.** A historical query interval is not a freshness requirement unless the user explicitly makes it one. Honor a user-specified freshness window; otherwise probe each raw companion at the requested end time, separately from widened history. Do not invent a production freshness threshold. Widened history or a historical `last_over_time()` result cannot establish currentness; its evaluation time is not source publication time. If the observations do not establish currentness, report current **unavailable** and the last-known value separately. +Keep **last-known** and **current** separate. A historical query interval is not a freshness requirement unless the user explicitly makes it one. Honor a user-specified freshness window; otherwise probe each raw companion at the requested end time, separately from widened history. Do not invent a production freshness threshold. Widened history or a historical `last_over_time()` result cannot establish currentness; its evaluation time is not source publication time. If the observations do not establish currentness, report current **unavailable** and the last-known value separately. Preserve exact returned Unix seconds. Verify UTC conversion and any age calculation with a reliable available tool; otherwise report the epoch only and omit converted dates/times and computed ages. `vector(epoch)` merely echoes a number, not a verified time conversion. Never round a source timestamp to a date or midnight. -If history is empty, verify datasource, filters, configuration, statistic, and cadence before blaming ingestion. Without read-only delivery evidence, whether delivery stopped remains unconfirmed. Neither “daily” nor the previous sample establishes the next publication time **or that a newer publication has not arrived**. - ## Report evidence and separate Discover identity Keep requested output names. Give value/unit or **unavailable**, resource/dimension level, datasource/UTC interval, source/statistic, and executed queries with evidence/call references. Cite every operand actually used, including Sum and SampleCount for an average. State partial coverage, last-known times, unsupported statistics, and unresolved causes; check prose arithmetic against the samples and result. diff --git a/skills/last9-cloudwatch/references/billing.md b/skills/last9-cloudwatch/references/billing.md new file mode 100644 index 0000000..e66143d --- /dev/null +++ b/skills/last9-cloudwatch/references/billing.md @@ -0,0 +1,19 @@ +# Billing + +Use the shared scope, statistic, freshness, and evidence rules in [SKILL.md](../SKILL.md). + +## Establish the cost source + +`AWS/Billing` `EstimatedCharges` is an estimated accumulated charge for the current month, not a daily increment, finalized invoice, or amortized cost. AWS publishes billing metrics in US East (N. Virginia), covering worldwide charges, and currently supports USD. Account/resource region and billing-metric region therefore need not be the same. + +Discover the actual `Currency`, `ServiceName`, and `LinkedAccount` dimensions where available. Select total account charges, a service, or a linked-account/service combination deliberately. Do not add a total row to its service or account breakdowns, mix currencies, or infer linked-account coverage from an accessible total. + +Cost exporters can expose names such as `aws_cost_*` with amortized, unblended, or other definitions. These are a separate lineage: inspect their reporting period, currency, scope, and accounting definition. Do not map an exporter gauge to `EstimatedCharges` solely because both are monetary, or describe their difference as a verified discount without accounting evidence. + +## Answer an estimated-charge request + +Discover the requested cost family and scope, then retrieve the latest original observation within an appropriate bounded history. For verified stream summaries, use matching Sum/SampleCount for a period average only when that is the represented statistic; otherwise select the observed AWS statistic. Report the latest estimate with currency, month/period, and original timestamp. Do not sum repeated month-to-date readings or use `rate()` / `increase()` on a charge gauge. A difference between readings is only a change in the estimate unless adjustments and period boundaries are understood. + +Billing updates are sparse. A widened historical result supports a last-known estimate, not currentness; follow the shared current probe and user freshness rules. Missing billing data can reflect configuration or scope, and does not prove zero spend or stopped delivery. Do not change billing preferences as part of this investigation. + +Source: [CloudWatch billing metrics and estimated charges](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/monitor_estimated_charges_with_cloudwatch.html). diff --git a/skills/last9-cloudwatch/references/dms.md b/skills/last9-cloudwatch/references/dms.md new file mode 100644 index 0000000..836ba22 --- /dev/null +++ b/skills/last9-cloudwatch/references/dms.md @@ -0,0 +1,17 @@ +# DMS + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Task and replication instance are separate scopes + +Discover `ReplicationTaskIdentifier` and `ReplicationInstanceIdentifier` for task metrics. Host storage, memory, and network metrics can use the replication instance without a task dimension. Keep the actual account/region/instance scope on host queries; omitting an inapplicable task label does not authorize a fleet-wide query. Serverless replication has its own applicable metric and dimension set. + +## Interpret replication latency before diagnosing + +`CDCLatencySource` and `CDCLatencyTarget` are seconds. Target latency includes source capture delay as well as downstream processing; a high target value alone does not prove the target endpoint is the bottleneck. Retrieve both metrics for the same task, periods, and statistic. Compare their levels/trends and report what is observed before attributing a cause. Full-load metrics describe the initial load, while CDC metrics describe ongoing replication; confirm task phase. + +For a latest-period target/source latency request, batch the four summary companion reads and calculate each matched period mean independently. For a window mean, use total Sum / total SampleCount per metric with matching coverage. Do not add source and target latency or substitute their difference for a directly measured stage duration without a documented model. + +Units vary within DMS: read/write latency is seconds, network throughput is bytes/sec, and CDC/full-load bandwidth is documented in KB/sec. `MemoryUsage` and `MemoryUsageBytes` are not the same unit. Task CPU percentages may exceed 100% when multiple cores are used; do not clamp them or apply an instance CPU interpretation automatically. Preserve the documented source unit unless a verified conversion is requested. + +Sources: [DMS task and replication-instance monitoring](https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Monitoring.html), [CDC latency interpretation](https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Troubleshooting_Latency.html). diff --git a/skills/last9-cloudwatch/references/dynamodb.md b/skills/last9-cloudwatch/references/dynamodb.md new file mode 100644 index 0000000..eb6352d --- /dev/null +++ b/skills/last9-cloudwatch/references/dynamodb.md @@ -0,0 +1,23 @@ +# DynamoDB + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Distinguish table, index, operation, and account + +Discover `TableName`, `GlobalSecondaryIndexName`, and `Operation` on each requested metric. Table-only and GSI rows describe different resources; specify which resource or deliberate combination the request concerns. Account-level metrics must remain account-scoped rather than being filtered by an invented table dimension. Global table replication metrics can add regional dimensions. + +## Capacity totals differ from averages + +`ConsumedReadCapacityUnits` and `ConsumedWriteCapacityUnits` use Sum for total consumed capacity units in a period. A Sum/SampleCount average is a different quantity and does not answer a consumed-capacity total. With complete disjoint period reports, sum their Sums once; divide by elapsed seconds only when average consumed units per second is requested. Provisioned capacity is a configured capacity level, not another consumption event total. On-demand resources need not publish a provisioned-capacity series. + +For example, after verifying the metric and table/index selector, retrieve the raw Sum series over the fixed window or execute: + +```promql +sum(sum_over_time({}[])) +``` + +Confirm full period coverage before reporting a window total. Do not divide by SampleCount, use `increase()` on period summaries, or add repeated chart rollups. + +`SuccessfulRequestLatency` is milliseconds and operation-specific; it excludes unsuccessful requests, so it is not end-to-end latency for all attempts. Throttle events, throttled requests, conditional failures, and transaction conflicts count different things. Publication conditions vary by metric: verify the definition and data path before interpreting an empty event series as zero. Missing data alone is not proof of healthy service. + +Sources: [DynamoDB metrics and dimensions](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/metrics-dimensions.html), [consumed read capacity](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/metrics-dimensions.html#ConsumedReadCapacityUnits). diff --git a/skills/last9-cloudwatch/references/ec2.md b/skills/last9-cloudwatch/references/ec2.md new file mode 100644 index 0000000..1232707 --- /dev/null +++ b/skills/last9-cloudwatch/references/ec2.md @@ -0,0 +1,22 @@ +# EC2 + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Identify instance and reporting period + +For an instance request, discover and select its actual `InstanceId` dimension. AWS also supports aggregation dimensions such as `AutoScalingGroupName`, `ImageId`, and `InstanceType` for applicable metrics; do not assume all rows are per-instance or sum aggregate rows with their instances. Metric availability depends on instance type and monitoring configuration. + +Basic monitoring commonly supplies five-minute periods; detailed monitoring supplies one-minute periods for supported metrics. Some metrics retain their own cadence. Inspect actual timestamps and definitions before assuming sixty-second coverage. + +## Levels, bytes, and credits + +- `CPUUtilization` is percent. A window mean from stream summaries needs matched Sum/SampleCount coverage. +- `NetworkIn` and `NetworkOut` are bytes over the period. For total bytes, add disjoint period Sums once. For average bytes/sec, divide a complete total by elapsed seconds; SampleCount is not a time denominator. +- `StatusCheckFailed*` are status flags. A Maximum of 1 supports an observed failed check, not a count of distinct failures. Do not add the combined flag to its component flags. +- `CPUCreditBalance` is a balance, while `CPUCreditUsage` describes credits spent. Keep the metric's unit and reporting period; applying counter functions to a balance is not usage. Credit metrics are instance-family dependent. + +For a network-total request, read the exact instance's raw Sum reports at the requested bounds, verify that the completed periods cover the interval, then total them or use the shared period-Sum template. Report partial coverage when periods are missing or the window cuts a period. + +Guest memory and filesystem usage generally require agent/custom telemetry, not standard `AWS/EC2` metrics. Preserve a valid CloudWatch-agent or OTel source separately; absence from the EC2 namespace does not establish that memory data is unavailable everywhere. + +Sources: [EC2 metrics and dimensions](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html), [basic and detailed monitoring](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/manage-detailed-monitoring.html). diff --git a/skills/last9-cloudwatch/references/elasticache.md b/skills/last9-cloudwatch/references/elasticache.md new file mode 100644 index 0000000..1e1b734 --- /dev/null +++ b/skills/last9-cloudwatch/references/elasticache.md @@ -0,0 +1,20 @@ +# ElastiCache + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Select engine and entity level + +Establish Valkey, Redis OSS, or Memcached and node-based versus serverless deployment before choosing metrics. Node-based metrics commonly use `CacheClusterId` and `CacheNodeId`; discover their actual ingested keys and combinations. Do not equate a replication group with a cache cluster or add node rows to a cluster aggregate. Serverless and Memcached have different metric inventories; absent Valkey/Redis metrics do not prove a broken stream. + +## Interpret the measurement + +- `EngineCPUUtilization` measures engine CPU; `CPUUtilization` measures the host. Both are percentages with different denominators, so they are not interchangeable or additive. +- `BytesUsedForCache` and `FreeableMemory` are byte levels; `CurrConnections` is a connection level. `Evictions`, `Reclaimed`, `CacheHits`, and `CacheMisses` describe distinct events. Use verified period Sums for event totals, not SampleCount or a peak statistic. +- `CacheHitRate` is a percentage. Do not average node percentages into a fleet hit ratio without request weights; matched hit/miss totals can support the ratio when their populations and periods agree. +- `ReplicationLag` is documented in seconds. `SuccessfulReadRequestLatency`, `SuccessfulWriteRequestLatency`, and command-family latency such as `GetTypeCmdsLatency` are microseconds. `DurabilityLag` and `DB0AverageTTL` are milliseconds. Check the exact metric/engine definition before converting; display requirements do not change source units. + +## Answer an engine CPU or latency request + +Discover the requested node's family and read both summary companions independently at the fixed scope/time. A mean of engine command-latency observations is not automatically request-weighted or end-to-end client latency. For the latest period, use that matched pair. For a window average, use total Sum / total SampleCount with complete matching coverage. Report engine CPU and host CPU separately if both were requested. Keep primary/replica identity for replication metrics; do not turn an absent replica-only metric into zero lag. + +Sources: [Valkey and Redis OSS metrics](https://docs.aws.amazon.com/AmazonElastiCache/latest/dg/CacheMetrics.Redis.html), [Memcached metrics](https://docs.aws.amazon.com/AmazonElastiCache/latest/dg/CacheMetrics.Memcached.html), [node dimensions](https://docs.aws.amazon.com/AmazonElastiCache/latest/dg/CloudWatchMetrics.html), [serverless metrics](https://docs.aws.amazon.com/AmazonElastiCache/latest/dg/serverless-metrics-events-redis.html). diff --git a/skills/last9-cloudwatch/references/kms.md b/skills/last9-cloudwatch/references/kms.md new file mode 100644 index 0000000..0f210bf --- /dev/null +++ b/skills/last9-cloudwatch/references/kms.md @@ -0,0 +1,19 @@ +# KMS + +Use the shared scope, statistic, freshness, and evidence rules in [SKILL.md](../SKILL.md). + +## Discover the applicable metric and key identity + +Do not assume one universal KMS metric or key dimension. `SuccessfulRequest` counts successful cryptographic operations and commonly uses `KeyArn` plus `Operation`; `ReEncrypt` uses source and destination key dimensions. Inspect the returned labels and operation definition before filtering. Do not treat every management API action as a SuccessfulRequest operation, or sum source/destination views without checking overlap. + +For a requested operation total, use the verified CloudWatch Sum for each disjoint period and sum only the requested key/operation population. SampleCount is observations, not automatically successful requests. Quota headroom needs the applicable regional quota and operation grouping; one key's success count cannot establish all attempted traffic, throttling, or remaining quota. + +## Imported key material expiration + +`SecondsUntilKeyMaterialExpiration` is a seconds-remaining measurement for applicable expiring imported key material, with `KeyId` as its key dimension. AWS documents Minimum for this metric. It is not a Unix timestamp, key age, rotation age, or operation count. Select its observed Minimum statistic, preserving the original sample time and applicable key identity. + +For an expiry-horizon request, return the remaining seconds at that observation. If a calendar expiry is requested, add the horizon to the verified source timestamp using an available reliable date tool and identify it as a calculation from that observation. Do not add it to the query evaluation time or label an old horizon current. No observation may mean the key/material is inapplicable, non-expiring, or unavailable in the selected scope; it does not prove no key or no KMS activity. + +Read-only metric investigation does not authorize importing/deleting key material, changing expiration, disabling keys, or altering key policies. + +Source: [KMS CloudWatch metrics and dimensions](https://docs.aws.amazon.com/kms/latest/developerguide/monitoring-cloudwatch.html). diff --git a/skills/last9-cloudwatch/references/msk.md b/skills/last9-cloudwatch/references/msk.md new file mode 100644 index 0000000..8f25e37 --- /dev/null +++ b/skills/last9-cloudwatch/references/msk.md @@ -0,0 +1,19 @@ +# Amazon MSK + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Scope follows the metric + +Discover cluster, broker, topic, and consumer-group dimensions for each metric. AWS dimension names such as `Cluster Name`, `Broker ID`, and `Consumer Group` may be normalized by ingestion; use actual returned keys. A broker filter applied to a consumer-group series can silently remove the requested population. + +Broker metrics, cluster aggregates, and topic breakdowns are not interchangeable. `UnderReplicatedPartitions` has a broker scope; do not classify every health metric as cluster-only. Monitoring level, broker type, cluster mode, and consumer state can affect metric availability. Confirm applicability before treating an absent metric as healthy or an ingestion fault. + +## Throughput and lag + +- `BytesInPerSec`, `BytesOutPerSec`, and `MessagesInPerSec` already describe rates. Use the requested level or average; applying `rate()` again changes the quantity. Combining broker/topic rows needs a verified non-overlapping population. +- `MaxOffsetLag` is the maximum offset lag across the applicable partitions; `SumOffsetLag` is a different aggregate. Neither is a count of unique messages processed during the observation window. +- `EstimatedMaxTimeLag` is in seconds. An offset count cannot be converted to time without additional evidence. + +For a requested worst consumer lag, select the exact cluster, consumer group, and topic. Retrieve the observed Maximum statistic (for a verified summary mapping, the actual maximum quantile) and take the maximum over the requested window. Report it as peak observed maximum offset lag, with partition aggregation and sample coverage. Do not sum per-period maxima or substitute SumOffsetLag. If only average observations exist, say the requested peak is unavailable. + +Sources: [MSK metrics and dimensions](https://docs.aws.amazon.com/msk/latest/developerguide/metrics-details.html), [consumer lag monitoring](https://docs.aws.amazon.com/msk/latest/developerguide/consumer-lag.html). diff --git a/skills/last9-cloudwatch/references/rds-aurora.md b/skills/last9-cloudwatch/references/rds-aurora.md new file mode 100644 index 0000000..4ee1a62 --- /dev/null +++ b/skills/last9-cloudwatch/references/rds-aurora.md @@ -0,0 +1,11 @@ +# RDS and Aurora + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## CPU and read latency + +Resolve account, region, and database dimensions. For Aurora, choose instance, cluster, or role: AWS publishes [different dimension combinations](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/dimensions.html), so selecting every row with a cluster name can overlap resources. + +Discover CPUUtilization and ReadLatency families, batch their companion reads, and verify source, labels, timestamps, and units. For stream summaries, use the latest companion pair for requested period averages; use the weighted template only for a requested window average with verified coverage. Report each metric's aggregation level separately. + +[RDS metrics](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) define CPUUtilization in Percent and ReadLatency in seconds. Confirm ingestion mapping before converting latency to milliseconds by multiplying by 1,000. Enhanced Monitoring and similarly named exporters can differ in units/definitions. Period latency averages or percentiles cannot establish whole-window percentiles without distributions or equivalent raw observations. diff --git a/skills/last9-cloudwatch/references/s3.md b/skills/last9-cloudwatch/references/s3.md new file mode 100644 index 0000000..95d8492 --- /dev/null +++ b/skills/last9-cloudwatch/references/s3.md @@ -0,0 +1,21 @@ +# S3 + +Use the shared scope, statistic, and evidence rules in [SKILL.md](../SKILL.md). + +## Daily storage: last-known is not current + +BucketSizeBytes and NumberOfObjects are [daily storage metrics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html), distinct from request metrics; AWS documents a [daily period and Average statistic](https://docs.aws.amazon.com/AmazonS3/latest/userguide/cloudwatch-monitoring-accessing.html). An empty recent lookup does not establish zero, deletion, or a stopped stream. + +Discover bucket/storage-type dimensions. Inspect bounded raw history spanning the cadence (for example, three days), separately from the user's requested interval. Distinguish numeric zero, last-known value with original timestamp, and no observation in the requested interval. + +**Daily cadence does not make an old sample current.** Apply the shared freshness rules: current probes and widened historical reads answer different questions. Preserve the raw publication timestamp and verify any time conversion or age calculation as described in [SKILL.md](../SKILL.md). + +If history is empty, verify datasource, filters, configuration, statistic, and cadence before blaming ingestion. Without read-only delivery evidence, whether delivery stopped remains unconfirmed. Neither “daily” nor the previous sample establishes the next publication time **or that a newer publication has not arrived**. + +## Storage dimensions and request metrics + +Keep `BucketName` and the observed `StorageType` when selecting daily storage. `NumberOfObjects` uses its documented all-storage-types population; a BucketSizeBytes series for one storage type does not establish the entire bucket size. Before combining storage classes, check their definitions for component/overlap behavior. Do not add daily readings over time as stored bytes, or divide mismatched bucket/storage populations to invent average object size. + +Request metrics are a distinct, optionally configured population and can include `FilterId`. Discover their scope separately; overlapping filters cannot be summed as disjoint traffic. Use verified period Sums for request/error totals and matching populations for error percentages. Request latency metrics are milliseconds; daily storage availability does not prove request-metric availability. Investigating missing request metrics does not authorize changing bucket configuration. + +Source: [S3 metric definitions, dimensions, and statistics](https://docs.aws.amazon.com/AmazonS3/latest/userguide/metrics-dimensions.html). diff --git a/skills/last9-cloudwatch/references/sqs.md b/skills/last9-cloudwatch/references/sqs.md new file mode 100644 index 0000000..884c49f --- /dev/null +++ b/skills/last9-cloudwatch/references/sqs.md @@ -0,0 +1,19 @@ +# SQS + +Use the shared scope, statistic, freshness, and evidence rules in [SKILL.md](../SKILL.md). + +## Queue identity and approximate measurements + +Discover `QueueName` and retain account/region; names need not be globally unique. Establish standard, FIFO, and fair-queue applicability before selecting feature-specific metrics. Quiet-group and noisy-group metrics concern fair queues, not a universal FIFO-only inventory. + +`ApproximateNumberOfMessagesVisible` is waiting backlog, `ApproximateNumberOfMessagesNotVisible` is in-flight messages, and `ApproximateNumberOfMessagesDelayed` is delayed messages. They are approximate levels, not cumulative counters. `ApproximateAgeOfOldestMessage` is seconds; its treatment of repeatedly received or moved messages can affect its meaning. A DLQ age is not automatically total time since original enqueue. + +## Choose the requested statistic + +For peak oldest-message age, select the exact queue and observed Maximum statistic, then take its maximum over the requested interval. If the stream mapping exposes maxima through a quantile, discover that label/value before using `max_over_time()`. Sum/SampleCount answers an average, not peak age. Adding age or backlog observations across time does not produce unique messages or message delay. + +`NumberOfMessagesSent`, `NumberOfMessagesReceived`, and `NumberOfMessagesDeleted` count messages sent, received, or deleted with different semantics. Receives and deletes can include repeated processing; they are not interchangeable with unique messages or exact successfully completed work. Use disjoint period Sums for the requested message total and cite its definition. `NumberOfEmptyReceives` counts receive calls with no returned messages; it is not queue depth. + +Inactive queues can stop emitting metrics and resume with delay. An empty recent query does not prove the queue was deleted, had zero backlog, or stopped delivering telemetry. Separate a last-known observation from current data using the shared freshness rules. + +Sources: [SQS metrics and semantics](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-available-cloudwatch-metrics.html), [CloudWatch monitoring and inactive queues](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/monitoring-using-cloudwatch.html). From 46ac7db00f2d44995d3abfc634f3a17beb32ed61 Mon Sep 17 00:00:00 2001 From: Prathamesh Sonpatki Date: Tue, 8 Sep 2026 23:04:55 +0200 Subject: [PATCH 8/8] docs(last9-cloudwatch): clarify family evidence and period semantics --- skills/last9-cloudwatch/references/dynamodb.md | 2 ++ skills/last9-cloudwatch/references/elasticache.md | 1 + skills/last9-cloudwatch/references/s3.md | 2 +- skills/last9-cloudwatch/references/sqs.md | 2 ++ 4 files changed, 6 insertions(+), 1 deletion(-) diff --git a/skills/last9-cloudwatch/references/dynamodb.md b/skills/last9-cloudwatch/references/dynamodb.md index eb6352d..764bc93 100644 --- a/skills/last9-cloudwatch/references/dynamodb.md +++ b/skills/last9-cloudwatch/references/dynamodb.md @@ -18,6 +18,8 @@ sum(sum_over_time({}[])) Confirm full period coverage before reporting a window total. Do not divide by SampleCount, use `increase()` on period summaries, or add repeated chart rollups. +If the verified data path uses period-end timestamps for period `P`, complete disjoint observations stamped `start + P`, `start + 2P`, through `end` cover the period-aligned interval `(start, end]`. A missing observation stamped exactly at `start` does not by itself leave the first period uncovered. If timestamp-to-period mapping is unverified, state that uncertainty separately. + `SuccessfulRequestLatency` is milliseconds and operation-specific; it excludes unsuccessful requests, so it is not end-to-end latency for all attempts. Throttle events, throttled requests, conditional failures, and transaction conflicts count different things. Publication conditions vary by metric: verify the definition and data path before interpreting an empty event series as zero. Missing data alone is not proof of healthy service. Sources: [DynamoDB metrics and dimensions](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/metrics-dimensions.html), [consumed read capacity](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/metrics-dimensions.html#ConsumedReadCapacityUnits). diff --git a/skills/last9-cloudwatch/references/elasticache.md b/skills/last9-cloudwatch/references/elasticache.md index 1e1b734..fa0bbf3 100644 --- a/skills/last9-cloudwatch/references/elasticache.md +++ b/skills/last9-cloudwatch/references/elasticache.md @@ -12,6 +12,7 @@ Establish Valkey, Redis OSS, or Memcached and node-based versus serverless deplo - `BytesUsedForCache` and `FreeableMemory` are byte levels; `CurrConnections` is a connection level. `Evictions`, `Reclaimed`, `CacheHits`, and `CacheMisses` describe distinct events. Use verified period Sums for event totals, not SampleCount or a peak statistic. - `CacheHitRate` is a percentage. Do not average node percentages into a fleet hit ratio without request weights; matched hit/miss totals can support the ratio when their populations and periods agree. - `ReplicationLag` is documented in seconds. `SuccessfulReadRequestLatency`, `SuccessfulWriteRequestLatency`, and command-family latency such as `GetTypeCmdsLatency` are microseconds. `DurabilityLag` and `DB0AverageTTL` are milliseconds. Check the exact metric/engine definition before converting; display requirements do not change source units. +- `GetTypeCmds` and `GetTypeCmdsLatency` cover read-only commands across data types; AWS examples include `GET`, `HGET`, `SCARD`, and `LRANGE`. Command groups overlap: `LRANGE` also belongs to `ListBasedCmds`. Do not add these groups as disjoint populations or infer command exclusions from a metric name. ## Answer an engine CPU or latency request diff --git a/skills/last9-cloudwatch/references/s3.md b/skills/last9-cloudwatch/references/s3.md index 95d8492..fc03986 100644 --- a/skills/last9-cloudwatch/references/s3.md +++ b/skills/last9-cloudwatch/references/s3.md @@ -10,7 +10,7 @@ Discover bucket/storage-type dimensions. Inspect bounded raw history spanning th **Daily cadence does not make an old sample current.** Apply the shared freshness rules: current probes and widened historical reads answer different questions. Preserve the raw publication timestamp and verify any time conversion or age calculation as described in [SKILL.md](../SKILL.md). -If history is empty, verify datasource, filters, configuration, statistic, and cadence before blaming ingestion. Without read-only delivery evidence, whether delivery stopped remains unconfirmed. Neither “daily” nor the previous sample establishes the next publication time **or that a newer publication has not arrived**. +If history is empty, verify datasource, filters, configuration, statistic, and cadence before blaming ingestion. Daily cadence is one possible explanation for a gap; sample absence does not identify its cause or establish healthy delivery. Without read-only delivery evidence, whether delivery stopped remains unconfirmed. Neither “daily” nor the previous sample establishes the next publication time **or that a newer publication has not arrived**. ## Storage dimensions and request metrics diff --git a/skills/last9-cloudwatch/references/sqs.md b/skills/last9-cloudwatch/references/sqs.md index 884c49f..1b0740e 100644 --- a/skills/last9-cloudwatch/references/sqs.md +++ b/skills/last9-cloudwatch/references/sqs.md @@ -10,6 +10,8 @@ Discover `QueueName` and retain account/region; names need not be globally uniqu ## Choose the requested statistic +For the latest period's Maximum, use the most recent source observation of the verified Maximum statistic and report its timestamp. A maximum over the whole requested interval answers the window peak; it can differ from the latest period's Maximum. + For peak oldest-message age, select the exact queue and observed Maximum statistic, then take its maximum over the requested interval. If the stream mapping exposes maxima through a quantile, discover that label/value before using `max_over_time()`. Sum/SampleCount answers an average, not peak age. Adding age or backlog observations across time does not produce unique messages or message delay. `NumberOfMessagesSent`, `NumberOfMessagesReceived`, and `NumberOfMessagesDeleted` count messages sent, received, or deleted with different semantics. Receives and deletes can include repeated processing; they are not interchangeable with unique messages or exact successfully completed work. Use disjoint period Sums for the requested message total and cite its definition. `NumberOfEmptyReceives` counts receive calls with no returned messages; it is not queue depth.