
Observability
- 3 installs
- 12 repo stars
- Updated June 8, 2026
- aws-samples/sample-claude-code-plugins-for-startups
observability is a Claude Code skill for designing AWS monitoring, logging, and tracing solutions using CloudWatch and X-Ray.
About
This skill helps Claude design AWS observability solutions using CloudWatch and X-Ray. It covers metrics and alarm thresholds by service, log retention and structured logging, ready-made Logs Insights queries, alarms, anomaly detection, dashboards, and tracing. A developer uses it when configuring monitoring or debugging observability gaps.
- Critical CloudWatch metrics and alarm thresholds per AWS service
- CloudWatch Logs retention, structured logging, and Logs Insights query library
- Alarm best practices, anomaly detection, dashboards, and X-Ray tracing
Observability by the numbers
- 3 all-time installs (skills.sh)
- Ranked #1,120 of 1,435 DevOps & CI/CD skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
observability capabilities & compatibility
- Works with
- aws
- Use cases
- devops · debugging
What observability says it does
Design and implement AWS observability solutions. Use when configuring CloudWatch metrics, logs, alarms, dashboards, Logs Insights queries, X-Ray tracing, anomaly detection, or debugging monitoring ga
Set retention on every log group. The default is **never expire** — this gets expensive fast.
npx skills add https://github.com/aws-samples/sample-claude-code-plugins-for-startups --skill observabilityAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 3 |
|---|---|
| repo stars | ★ 12 |
| Last updated | June 8, 2026 |
| Repository | aws-samples/sample-claude-code-plugins-for-startups ↗ |
What it does
Design AWS monitoring: set CloudWatch metrics, alarms, log retention, Logs Insights queries, dashboards, and X-Ray tracing.
Who is it for?
Teams instrumenting AWS workloads who need metrics, alarms, log retention, and tracing set up correctly.
When should I use this skill?
The user configures CloudWatch metrics, logs, alarms, dashboards, Logs Insights queries, X-Ray tracing, or debugs monitoring gaps.
By the numbers
- per-service metric and alarm-threshold table
- recommended log retention 30d dev / 90d prod
Files
You are an AWS observability specialist. Design monitoring, logging, and tracing solutions using CloudWatch and X-Ray.
CloudWatch Metrics
Key Concepts
- Namespace: Grouping for metrics (e.g.,
AWS/EC2,AWS/Lambda, custom) - Metric: Time-ordered set of data points (e.g.,
CPUUtilization) - Dimension: Key-value pair that identifies a metric (e.g.,
InstanceId=i-xxx) - Period: Aggregation interval (60s, 300s, etc.)
- Statistic: Aggregation function (Average, Sum, Min, Max, p99, etc.)
Critical Metrics by Service
| Service | Metric | Alarm Threshold | Notes |
|---|---|---|---|
| Lambda | Errors | > 0 for 1 min | Also alarm on Throttles and Duration p99 |
| Lambda | ConcurrentExecutions | > 80% of account limit | Prevent throttling |
| ALB | HTTPCode_Target_5XX_Count | > 0 for 5 min | Backend errors |
| ALB | TargetResponseTime p99 | > your SLA | Latency SLO |
| ALB | UnHealthyHostCount | > 0 | Failing targets |
| RDS | CPUUtilization | > 80% for 5 min | Sustained high CPU |
| RDS | FreeStorageSpace | < 20% of total | Prevent disk full |
| RDS | DatabaseConnections | > 80% of max | Connection exhaustion |
| DynamoDB | ThrottledRequests | > 0 | Capacity issues |
| SQS | ApproximateAgeOfOldestMessage | > your processing SLA | Queue backlog |
| ECS | CPUUtilization / MemoryUtilization | > 80% for 5 min | Scaling trigger |
Custom Metrics
- Use
PutMetricDataAPI or the CloudWatch Agent - Embedded Metric Format (EMF) for Lambda: log structured JSON that CloudWatch automatically extracts as metrics. Zero API calls, no cost per PutMetricData.
- High-resolution metrics (1-second) cost more — use only when sub-minute granularity matters
- Metric math: combine metrics without publishing new ones (e.g., error rate = Errors / Invocations * 100)
CloudWatch Logs
Log Groups and Retention
- Set retention on every log group. The default is never expire — this gets expensive fast.
- Recommended: 30 days for dev, 90 days for production, archive to S3 for long-term
- Use subscription filters to stream logs to Lambda, Kinesis, or OpenSearch
Structured Logging
Always log in JSON format. This enables Logs Insights queries on fields.
{"level": "ERROR", "message": "Payment failed", "orderId": "123", "errorCode": "DECLINED", "duration_ms": 45}CloudWatch Logs Insights Queries
# Find errors in Lambda functions
fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 100
# P99 latency from structured logs
fields @timestamp, duration_ms
| stats percentile(duration_ms, 99) as p99, avg(duration_ms) as avg_ms by bin(5m)
# Top 10 most frequent errors
fields @timestamp, errorCode, @message
| filter level = "ERROR"
| stats count(*) as error_count by errorCode
| sort error_count desc
| limit 10
# Request rate over time
fields @timestamp
| stats count(*) as requests by bin(1m)
| sort @timestamp desc
# Find slow requests
fields @timestamp, @duration, @requestId
| filter @duration > 5000
| sort @duration desc
| limit 20
# Cold starts in Lambda
filter @type = "REPORT"
| fields @requestId, @duration, @initDuration
| filter ispresent(@initDuration)
| stats count(*) as cold_starts, avg(@initDuration) as avg_init by bin(1h)
# API Gateway latency breakdown
fields @timestamp
| filter @message like /API Gateway/
| stats avg(integrationLatency) as backend_ms, avg(latency) as total_ms by bin(5m)CloudWatch Alarms
Alarm Types
- Static threshold: Fixed value (e.g., CPU > 80%)
- Anomaly detection: ML-based band. Good for metrics with patterns (traffic, latency).
- Composite alarm: Combine multiple alarms with AND/OR logic. Reduces noise.
Alarm Best Practices
- Use 3 out of 5 datapoints evaluation to avoid flapping on transient spikes
- Set
TreatMissingDatatonotBreachingfor low-traffic services (avoids false alarms when no data) - Set
TreatMissingDatatobreachingfor critical health checks (missing data = something is down) - Use composite alarms to create "alarm hierarchies": a top-level alarm that fires only when multiple sub-alarms are in ALARM state
- Always send alarms to SNS. Connect SNS to PagerDuty, Slack, or email.
Anomaly Detection
- Trains on 2 weeks of data. Do not enable during a known-bad period.
- Adjust the band width (number of standard deviations). Start with 2, widen if too noisy.
- Best for: request count, latency, error rate — metrics with daily/weekly patterns.
- Not good for: binary metrics, metrics that are normally zero.
CloudWatch Dashboards
Dashboard Design
- One dashboard per service or domain (not one giant dashboard)
- Top row: key business metrics (request rate, error rate, latency p99)
- Second row: infrastructure health (CPU, memory, connections)
- Third row: dependencies (downstream API latency, queue depth)
- Use metric math to show rates and percentages, not raw counts
- Add text widgets to document what each section monitors and what to do when values are abnormal
Automatic Dashboards
- CloudWatch provides automatic dashboards per service — start there before building custom
- ServiceLens provides an application-centric view combining metrics, logs, and traces
X-Ray Tracing
When to Use X-Ray
- Distributed applications with multiple services
- Debugging latency issues across service boundaries
- Understanding request flow and dependencies
Instrumentation
- AWS SDK automatically instruments calls to AWS services
- Use X-Ray SDK or OpenTelemetry to instrument your application code
- Set sampling rules to control trace volume (default: 1 req/sec + 5% of additional)
Key X-Ray Concepts
- Trace: End-to-end request path
- Segment: A single service's processing of the request
- Subsegment: Detailed breakdown within a segment (DB call, HTTP call)
- Service Map: Visual representation of your architecture based on trace data
- Annotations: Indexed key-value pairs for filtering traces (e.g.,
customerId=123) - Metadata: Non-indexed data attached to segments
X-Ray Best Practices
- Add annotations for business-relevant fields (user ID, order ID) so you can filter traces
- Use groups to define filter expressions for specific trace sets
- Active tracing on API Gateway and Lambda captures the full request lifecycle
- X-Ray daemon runs as a sidecar in ECS or as a DaemonSet in EKS
Contributor Insights
- Identifies top contributors to a metric (e.g., top IPs, top API callers)
- Define rules in JSON that specify log group + fields to analyze
- Good for: identifying noisy neighbors, DDoS sources, hot partition keys in DynamoDB
Common CLI Commands
# Query Logs Insights
aws logs start-query --log-group-name /aws/lambda/my-function \
--start-time $(date -d '1 hour ago' +%s) --end-time $(date +%s) \
--query-string 'fields @timestamp, @message | filter @message like /ERROR/ | limit 20'
# Get query results
aws logs get-query-results --query-id "query-id-here"
# Describe alarms in ALARM state
aws cloudwatch describe-alarms --state-value ALARM --query 'MetricAlarms[*].{Name:AlarmName,Metric:MetricName,State:StateValue}'
# Get metric statistics
aws cloudwatch get-metric-statistics --namespace AWS/Lambda --metric-name Errors \
--start-time 2024-01-01T00:00:00Z --end-time 2024-01-01T01:00:00Z \
--period 300 --statistics Sum --dimensions Name=FunctionName,Value=my-function
# Put custom metric
aws cloudwatch put-metric-data --namespace MyApp --metric-name RequestLatency \
--value 42 --unit Milliseconds --dimensions Name=Environment,Value=prod
# List log groups with retention
aws logs describe-log-groups --query 'logGroups[*].{Name:logGroupName,RetentionDays:retentionInDays,StoredBytes:storedBytes}'
# Set log retention
aws logs put-retention-policy --log-group-name /aws/lambda/my-function --retention-in-days 30
# List X-Ray traces
aws xray get-trace-summaries --start-time $(date -d '1 hour ago' +%s) --end-time $(date +%s)
# Get X-Ray service map
aws xray get-service-graph --start-time $(date -d '1 hour ago' +%s) --end-time $(date +%s)
# List CloudWatch dashboards
aws cloudwatch list-dashboardsOutput Format
| Field | Details |
|---|---|
| Metrics | Critical alarms with thresholds, evaluation periods, and actions |
| Logs | Log groups, retention policy, structured format (JSON), subscription filters |
| Traces | X-Ray or OpenTelemetry, sampling rules, annotations for filtering |
| Dashboards | Dashboard names, key widgets, layout (business/infra/dependencies) |
| Anomaly detection | Metrics with anomaly detection bands, standard deviation config |
| Cost | Estimated monthly cost for logs ingestion, metrics, dashboards, and traces |
Reference Files
references/logs-insights-queries.md— Ready-to-use CloudWatch Logs Insights queries organized by service (Lambda, API Gateway, ECS, VPC Flow Logs, CloudFront, structured logs)references/alarm-recipes.md— Production alarm configurations with thresholds, metric math examples, composite alarm and anomaly detection recipes
Related Skills
lambda— Lambda metrics, Embedded Metric Format, and X-Ray active tracingecs— Container Insights, task-level metrics, and ECS service alarmseks— Control plane logging, Prometheus, and Container Insights for Kubernetescloudfront— CloudFront access logs and cache metricsapi-gateway— API Gateway latency and error monitoringnetworking— VPC Flow Logs, Route53 health checks, and Transit Gateway metrics
Anti-Patterns
- No log retention policy: CloudWatch Logs default to never expire. Costs grow silently. Set retention on every log group.
- Alarming on every metric: Too many alarms leads to alert fatigue. Alarm on symptoms (error rate, latency), not causes (CPU). Use composite alarms to reduce noise.
- Average-based latency alarms: Averages hide tail latency. Use p99 or p95 for latency alarms.
- Missing structured logging: Unstructured logs cannot be queried efficiently with Logs Insights. Always log JSON.
- No tracing in distributed systems: Without X-Ray or OpenTelemetry, debugging cross-service issues requires correlating timestamps across log groups. Enable tracing.
- Sampling rate of 100%: Full tracing in production generates enormous data volume and cost. Use sampling — 1 req/sec + 5% is usually sufficient.
- Not using Embedded Metric Format in Lambda: EMF turns log lines into metrics with zero PutMetricData API calls. It's cheaper and simpler than the alternatives.
- Dashboard without runbook links: A dashboard that shows a problem without explaining what to do about it is only half useful. Add text widgets with runbook links.
- Ignoring CloudWatch anomaly detection: Static thresholds don't work for metrics with daily patterns. Use anomaly detection for request count and latency.
- CloudWatch Agent not installed on EC2: Without the agent, you only get basic metrics (CPU, network, disk I/O). Install the agent for memory utilization, disk space, and custom metrics.
CloudWatch Alarm Recipes
Production-ready alarm configurations organized by service. Each recipe includes the metric, threshold rationale, and recommended settings.
Alarm Configuration Defaults
Unless stated otherwise, all alarms below should use these settings:
| Setting | Value | Rationale |
|---|---|---|
| EvaluationPeriods | 5 | Avoids flapping on transient spikes |
| DatapointsToAlarm | 3 | 3 of 5 datapoints must breach |
| TreatMissingData | notBreaching | Avoids false alarms during low traffic |
| ActionsEnabled | true | Always wire to SNS |
| Period | 60 (seconds) | 1-minute granularity for most metrics |
Override TreatMissingData to breaching for health-check style alarms where missing data means the resource is down.
Lambda
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| Errors | Errors | Sum | > 0 | 60s | Any error is worth knowing about |
| High error rate | Metric math: Errors/Invocations*100 | - | > 5% | 60s | Percentage-based avoids noise on low volume |
| Throttles | Throttles | Sum | > 0 | 60s | Indicates concurrency pressure |
| Duration p99 | Duration | p99 | > 80% of timeout | 60s | Approaching timeout = about to fail |
| Concurrent executions | ConcurrentExecutions | Maximum | > 80% of account limit | 300s | Prevent account-wide throttling |
| Iterator age (streams) | IteratorAge | Maximum | > 60000 ms | 60s | Stream processing falling behind |
Lambda Error Rate Metric Math Example
# CloudFormation snippet
LambdaErrorRateAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "${FunctionName}-error-rate"
Metrics:
- Id: errors
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref MyFunction
Period: 60
Stat: Sum
- Id: invocations
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Invocations
Dimensions:
- Name: FunctionName
Value: !Ref MyFunction
Period: 60
Stat: Sum
- Id: error_rate
Expression: "IF(invocations > 0, errors / invocations * 100, 0)"
Label: "Error Rate %"
ComparisonOperator: GreaterThanThreshold
Threshold: 5
EvaluationPeriods: 5
DatapointsToAlarm: 3
TreatMissingData: notBreaching
AlarmActions:
- !Ref AlertSNSTopicALB / Application Load Balancer
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| 5XX errors | HTTPCode_Target_5XX_Count | Sum | > 0 | 300s | Backend is returning errors |
| High 5XX rate | Metric math: 5XX/RequestCount*100 | - | > 1% | 60s | Percentage-based for noisy services |
| Latency p99 | TargetResponseTime | p99 | > your SLA (e.g., 2s) | 60s | Tail latency breach |
| Unhealthy hosts | UnHealthyHostCount | Maximum | > 0 | 60s | Targets failing health checks |
| Rejected connections | RejectedConnectionCount | Sum | > 0 | 60s | ALB at connection limit |
| Active connections | ActiveConnectionCount | Sum | > 80% of expected max | 60s | Connection exhaustion risk |
RDS / Aurora
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| CPU utilization | CPUUtilization | Average | > 80% | 300s | Sustained high CPU |
| Free storage | FreeStorageSpace | Minimum | < 20% of allocated | 300s | Prevent disk full |
| Connections | DatabaseConnections | Maximum | > 80% of max_connections | 60s | Connection exhaustion |
| Read latency | ReadLatency | p99 | > 20ms | 60s | Disk I/O bottleneck |
| Write latency | WriteLatency | p99 | > 20ms | 60s | Disk I/O bottleneck |
| Replica lag | ReplicaLag | Maximum | > 30s | 60s | Replication falling behind |
| Freeable memory | FreeableMemory | Minimum | < 256 MB | 300s | Instance under memory pressure |
RDS Storage Alarm with Percentage Threshold
RDSStorageAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "${DBInstanceId}-storage-low"
Metrics:
- Id: free
MetricStat:
Metric:
Namespace: AWS/RDS
MetricName: FreeStorageSpace
Dimensions:
- Name: DBInstanceIdentifier
Value: !Ref DBInstance
Period: 300
Stat: Minimum
- Id: threshold
Expression: !Sub "${AllocatedStorageGB} * 1073741824 * 0.2"
Label: "20% of allocated"
ComparisonOperator: LessThanThreshold
Threshold: 0
# Use the expression as the threshold by comparing free < threshold
# Alternative: hardcode the byte value for your instance size
EvaluationPeriods: 3
DatapointsToAlarm: 2
TreatMissingData: breaching
AlarmActions:
- !Ref AlertSNSTopicDynamoDB
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| Throttled requests | ThrottledRequests | Sum | > 0 | 60s | Capacity insufficient |
| Read throttles | ReadThrottleEvents | Sum | > 0 | 60s | Separate from write throttles |
| Write throttles | WriteThrottleEvents | Sum | > 0 | 60s | Separate from read throttles |
| System errors | SystemErrors | Sum | > 0 | 60s | DynamoDB-side errors (rare) |
| User errors | UserErrors | Sum | > 10 | 60s | Conditional check failures, validation |
| Consumed RCU | ConsumedReadCapacityUnits | Sum | > 80% of provisioned | 300s | Provisioned mode only |
| Consumed WCU | ConsumedWriteCapacityUnits | Sum | > 80% of provisioned | 300s | Provisioned mode only |
SQS
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| Queue depth | ApproximateNumberOfMessagesVisible | Maximum | > your processing capacity | 60s | Queue building up |
| Message age | ApproximateAgeOfOldestMessage | Maximum | > your processing SLA | 60s | Messages stuck in queue |
| DLQ depth | ApproximateNumberOfMessagesVisible (DLQ) | Sum | > 0 | 60s | Failed messages accumulating |
| Messages not visible | ApproximateNumberOfMessagesNotVisible | Maximum | > expected in-flight | 60s | Processing bottleneck |
ECS
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| CPU utilization | CPUUtilization | Average | > 80% | 300s | Scaling trigger |
| Memory utilization | MemoryUtilization | Average | > 80% | 300s | Scaling trigger |
| Running task count | RunningTaskCount | Minimum | < desired count | 60s | Tasks crashing |
CloudFront
| Alarm | Metric | Statistic | Threshold | Period | Notes |
|---|---|---|---|---|---|
| 5xx error rate | 5xxErrorRate | Average | > 1% | 300s | Origin errors |
| 4xx error rate | 4xxErrorRate | Average | > 10% | 300s | Client errors (may indicate misconfiguration) |
| Origin latency | OriginLatency | p99 | > 5s | 60s | Slow origin responses |
| Total error rate | TotalErrorRate | Average | > 5% | 300s | Combined error rate |
Composite Alarm Example
Reduce alert fatigue by combining related alarms.
ServiceHealthCompositeAlarm:
Type: AWS::CloudWatch::CompositeAlarm
Properties:
AlarmName: "my-service-unhealthy"
AlarmRule: |
ALARM("my-service-5xx-rate") AND
(ALARM("my-service-latency-p99") OR ALARM("my-service-error-rate"))
AlarmActions:
- !Ref PagerDutySNSTopic
InsufficientDataActions: []
OKActions:
- !Ref PagerDutySNSTopicThis composite alarm fires only when there are 5XX errors AND either high latency or high application error rate. A single noisy metric alone will not page anyone.
Anomaly Detection Alarm Example
LatencyAnomalyAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: "api-latency-anomaly"
Metrics:
- Id: latency
MetricStat:
Metric:
Namespace: AWS/ApiGateway
MetricName: Latency
Dimensions:
- Name: ApiName
Value: !Ref ApiName
Period: 300
Stat: p99
- Id: anomaly_band
Expression: "ANOMALY_DETECTION_BAND(latency, 2)"
Label: "Anomaly Detection Band"
ComparisonOperator: GreaterThanUpperThreshold
ThresholdMetricId: anomaly_band
EvaluationPeriods: 3
DatapointsToAlarm: 2
TreatMissingData: notBreaching
AlarmActions:
- !Ref AlertSNSTopicThe band width of 2 (standard deviations) is a reasonable starting point. Widen to 3 if too noisy. Narrow to 1.5 for critical paths where you want early warning.
CloudWatch Logs Insights Query Examples
Ready-to-use Logs Insights queries organized by use case. Copy and adapt for your log groups.
Lambda Function Debugging
# Find errors in Lambda functions
fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 100
# Cold starts — frequency and duration
filter @type = "REPORT"
| fields @requestId, @duration, @initDuration
| filter ispresent(@initDuration)
| stats count(*) as cold_starts, avg(@initDuration) as avg_init_ms, max(@initDuration) as max_init_ms by bin(1h)
# Lambda timeout detection
filter @message like /Task timed out/
| fields @timestamp, @requestId, @message
| sort @timestamp desc
# Memory usage near limit
filter @type = "REPORT"
| fields @requestId, @maxMemoryUsed, @memorySize
| filter @maxMemoryUsed / @memorySize > 0.8
| sort @maxMemoryUsed desc
| limit 50
# P99 duration over time
filter @type = "REPORT"
| stats percentile(@duration, 99) as p99, percentile(@duration, 95) as p95, avg(@duration) as avg_ms by bin(5m)
# Invocations and errors over time
filter @type = "REPORT"
| stats count(*) as invocations, sum(strcontains(@message, "ERROR")) as errors by bin(5m)Structured Log Analysis
These queries assume JSON-formatted log output with fields like level, message, errorCode, duration_ms, requestId, userId.
# P99 latency from structured logs
fields @timestamp, duration_ms
| stats percentile(duration_ms, 99) as p99, percentile(duration_ms, 95) as p95, avg(duration_ms) as avg_ms by bin(5m)
# Top 10 most frequent errors
fields @timestamp, errorCode, @message
| filter level = "ERROR"
| stats count(*) as error_count by errorCode
| sort error_count desc
| limit 10
# Error rate percentage over time
stats count(*) as total, sum(level = "ERROR") as errors by bin(5m)
| fields @timestamp, total, errors, errors / total * 100 as error_rate_pct
# Find slow requests
fields @timestamp, duration_ms, requestId, userId
| filter duration_ms > 5000
| sort duration_ms desc
| limit 20
# Errors by user
fields @timestamp, userId, errorCode
| filter level = "ERROR"
| stats count(*) as error_count by userId
| sort error_count desc
| limit 20
# Unique users over time
fields userId
| stats count_distinct(userId) as unique_users by bin(1h)API Gateway
# API Gateway latency breakdown
fields @timestamp
| filter @message like /API Gateway/
| stats avg(integrationLatency) as backend_ms, avg(latency) as total_ms by bin(5m)
# 4xx and 5xx error rates
fields @timestamp, status
| stats count(*) as total,
sum(status >= 400 and status < 500) as client_errors,
sum(status >= 500) as server_errors by bin(5m)
| fields @timestamp, total, client_errors, server_errors,
client_errors / total * 100 as client_error_pct,
server_errors / total * 100 as server_error_pct
# Top API paths by request volume
fields path
| stats count(*) as requests by path
| sort requests desc
| limit 20
# Slowest API endpoints
fields path, latency
| stats avg(latency) as avg_ms, percentile(latency, 99) as p99_ms, count(*) as requests by path
| sort p99_ms desc
| limit 20ECS / Container Logs
# OOM kills
fields @timestamp, @message
| filter @message like /OutOfMemory/ or @message like /OOMKilled/ or @message like /oom-kill/
| sort @timestamp desc
# Container restart events
fields @timestamp, @message
| filter @message like /Starting/ or @message like /Stopping/ or @message like /SIGTERM/
| sort @timestamp desc
| limit 50
# Request rate by container
fields @timestamp, containerId
| stats count(*) as requests by containerId, bin(5m)VPC Flow Logs
# Rejected connections (potential security concern)
fields @timestamp, srcAddr, dstAddr, dstPort, action
| filter action = "REJECT"
| stats count(*) as rejected by srcAddr, dstAddr, dstPort
| sort rejected desc
| limit 25
# Top talkers by bytes
fields srcAddr, dstAddr, bytes
| stats sum(bytes) as total_bytes by srcAddr, dstAddr
| sort total_bytes desc
| limit 20
# Traffic to a specific port
fields @timestamp, srcAddr, dstAddr, dstPort, action, bytes
| filter dstPort = 443
| stats sum(bytes) as total_bytes, count(*) as connections by srcAddr
| sort total_bytes desc
| limit 20
# Connections from outside the VPC CIDR
fields @timestamp, srcAddr, dstAddr, dstPort, action
| filter not ispresent(srcAddr like "10.0.")
| filter action = "ACCEPT"
| stats count(*) as connections by srcAddr, dstPort
| sort connections descCloudFront
# Top requested URIs
fields @timestamp, cs-uri-stem, sc-status
| stats count(*) as requests by cs-uri-stem
| sort requests desc
| limit 20
# Cache hit ratio
fields @timestamp, x-edge-result-type
| stats count(*) as total,
sum(x-edge-result-type = "Hit") as hits by bin(5m)
| fields @timestamp, total, hits, hits / total * 100 as hit_rate_pct
# 5xx errors by URI
fields cs-uri-stem, sc-status
| filter sc-status >= 500
| stats count(*) as errors by cs-uri-stem, sc-status
| sort errors desc
| limit 20General Patterns
# Request rate over time
fields @timestamp
| stats count(*) as requests by bin(1m)
| sort @timestamp desc
# Count log volume by log stream
fields @logStream
| stats count(*) as lines by @logStream
| sort lines desc
| limit 20
# Search for a specific request/correlation ID
fields @timestamp, @message
| filter @message like /abc-123-request-id/
| sort @timestamp asc
# Extract and analyze JSON fields dynamically
fields @timestamp, @message
| parse @message '{"action":"*","duration":*}' as action, duration
| stats avg(duration) as avg_ms, count(*) as calls by action
| sort avg_ms descCLI: Running Logs Insights Queries
# Start a query (returns query ID)
aws logs start-query \
--log-group-name /aws/lambda/my-function \
--start-time $(date -d '1 hour ago' +%s) \
--end-time $(date +%s) \
--query-string 'fields @timestamp, @message | filter @message like /ERROR/ | limit 20'
# Get query results (poll until status is "Complete")
aws logs get-query-results --query-id "query-id-here"
# Query multiple log groups at once
aws logs start-query \
--log-group-names /aws/lambda/fn-a /aws/lambda/fn-b /aws/lambda/fn-c \
--start-time $(date -d '6 hours ago' +%s) \
--end-time $(date +%s) \
--query-string 'fields @timestamp, @message | filter @message like /ERROR/ | stats count(*) by @logStream'Related skills
FAQ
How long should I keep CloudWatch logs?
The skill says set retention on every log group because the default is never expire, recommending 30 days for dev, 90 for production, and archiving to S3 long-term.
How do I stop CloudWatch alarms from flapping?
The skill recommends using 3 out of 5 datapoints evaluation to avoid flapping on transient spikes.