Loading...
Loading...
Found 21 Skills
Manage Tencent Cloud CLS alarm policies, notice groups, shields and alarm execution logs. Use when the user asks to: list / create / modify / delete CLS alarms, enable or disable alarms, manage notice recipients (SMS / email / webhook), mute alarms during deploys, or view which alarms fired and when. For searching the underlying log content, use the companion `tencentcloud-cls` skill.
AWS CloudWatch monitoring for logs, metrics, alarms, and dashboards. Use when setting up monitoring, creating alarms, querying logs with Insights, configuring metric filters, building dashboards, or troubleshooting application issues.
Systematic incident investigation methodology. Use when investigating production issues, service degradation, errors, latency spikes, or outages.
DigitalOcean management services for monitoring, uptime checks, and resource organization with Projects. Use when setting up observability, alerts, and operational visibility on DigitalOcean.
Create Alibaba Cloud CMS alert rules via CLI (write-operation skill). Supports CMS 1.0 cloud resource monitoring for ALL CMS-integrated cloud products. This skill performs write operations: creating alert rules, contacts, and contact groups. Use when: creating monitoring alerts, setting up alarm rules, configuring CMS alert policies for any cloud product, or managing cloud monitoring notifications. Triggers: "create alert", "setup monitoring", "configure alarm", "CMS alert", "cloud monitor rule", "告警规则", "创建告警", "监控报警".
Prometheus, Grafana, CloudWatch, Azure Monitor, Stackdriver, logging, alerting, and SRE practices
Triage and manage Coralogix Cases with the `cx cases` CLI — e.g. acknowledge, assign, resolve, or re-prioritize a case, or inspect its event timeline or notification deliveries.
Configures PromQL-based Service Level Objective (SLO) alerting policies for Google Cloud resources registered in App Hub or individually specified. Generates Terraform output. Use when the user asks to configure an SLO or Service Level Objective. Don't use for standard alerting policies.
Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. Use when troubleshooting JobSet restart loops, spot VM preemptions, node readiness failures, host VM issues, or coordinator worker crashes. Don't use for general GKE cluster creation, basic workload deployment, or non-JobSet application issues.