Site Reliability Engineering Bundle
SLO/SLI design, error-budget reviews, runbook authoring, capacity planning, and toil-reduction skills for SREs.
Browse all skill bundles
Included Skills (43)
- SLO Reliability Architect — This tool assists SREs and developers in designing, calculating, and auditing meaningful service level objectives, error budgets, and burn-rate alerts for production systems.
- Resilience Experiment Assistant — Helps SREs and developers plan, execute, and analyze chaos engineering experiments to improve system resilience and reliability.
- Operational Runbook Generator — This tool analyzes codebase and infrastructure configurations to create structured, actionable runbooks that assist on-call engineers during incidents and routine system operations.
- Observability Strategy Architect — This tool helps site reliability engineers design production-ready monitoring strategies, optimize alert configurations, and generate comprehensive dashboards for complex distributed service architectures.
- Runbook Creation Assistant — Generates operational runbooks from a service name, streamlining documentation for on-call procedures and standardizing incident response across teams.
- Incident Response Specialist — This expert SRE tool provides rapid problem resolution, observability guidance, and structured incident management workflows to help engineering teams minimize system downtime.
- Cloud Run Observability — This skill implements production-grade, SRE-focused alerting policies for Cloud Run resources to help engineers maintain high-signal monitoring via Terraform and PromQL.
- Performance Troubleshooting Assistant — This tool enables SREs and developers to perform methodical system performance debugging and root-cause analysis using Brendan Gregg's proven USE and TSA methodologies.
- Incident Response Orchestrator — This framework assists SRE and DevOps teams by managing operational outages through automated severity classification, timeline reconstruction, and structured post-incident review generation.
- PagerDuty Incident Manager — This tool enables developers and SREs to automate incident response, service management, and on-call scheduling directly through the Rube MCP interface.
- Grafana Dashboard Architect — This skill enables engineers to design and implement production-ready Grafana dashboards for effective system observability, infrastructure monitoring, and business metric tracking.
- PagerDuty Configuration — Configure PagerDuty for incident management, on-call scheduling, and alert routing, assisting DevOps engineers and system administrators.
- Pino Logging Integration — Integrate Pino into Node.js projects to enable structured, performant logging for improved observability and easier log aggregation.
- Error Observability Expert — This expert assists developers in implementing robust error tracking, structured logging, and alerting systems to ensure rapid identification and resolution of production issues.
- Distributed Systems Debugger — This expert diagnostic tool assists engineers in performing root-cause analysis, investigating production incidents, and implementing robust observability solutions for complex distributed software systems.
- System Reliability Analyst — This expert assistant helps developers and engineers perform root-cause analysis, investigate production incidents, and design robust observability solutions for complex distributed systems.
- Cloud Monitoring Generator — This tool creates Google Cloud Monitoring dashboard widget textprotos from query data to assist engineers in automating declarative infrastructure and observability configurations.
- Cloud Monitoring Generator — This tool generates precise Cloud Monitoring ListTimeSeries API requests and aggregation configurations to assist developers in building accurate observability queries for GCP.
- Distributed Debugging Expert — This skill assists developers and engineers in configuring distributed tracing, observability tools, and robust debugging workflows for complex multi-service production environments.
- Better Stack Assistant — Provides expert guidance for Better Stack, helping developers configure uptime monitoring, log management, incident response, and status pages.
- Envoy Proxy Expert — This skill helps teams configure Envoy as an API gateway, service mesh, and load balancer for cloud-native applications.
- Error Monitoring Specialist — This skill helps developers implement robust error tracking, structured logging, and automated alerting systems to improve production reliability and incident response efficiency.
- GKE Alerting Architect — This skill assists DevOps engineers in configuring robust Terraform alerting policies for GKE clusters and workloads using PromQL and Google Cloud Managed Service.
- AI Agent Alerting — This tool generates Terraform configurations for OpenTelemetry-based alerting policies to help developers monitor AI agent latency, error rates, token usage, and quality.
- Airflow DAG Troubleshooter — This skill assists data engineers in diagnosing and resolving failed Apache Airflow DAG runs and task instances using cloud command-line diagnostic tools.
- Gke Workload Troubleshooting — Diagnoses GKE workload failures (CrashLoopBackOff, OOMKilled, ImagePullBackOff, Pending, etc.) via logs and events. Use when pods fail to start or crash repeatedly. Don't use for GKE cluster infrastructure provisioning, node pool creation, or non-Kubernetes Google Cloud services.
- Performance Optimization Expert — This skill helps developers identify and resolve application bottlenecks through data-driven profiling to ensure optimal speed, responsiveness, and Core Web Vitals compliance.
- Fluentd Log Management — Configure Fluentd for collecting, filtering, and routing logs across distributed systems, assisting developers and DevOps engineers with log aggregation.
- Automated Debugging Specialist — This skill assists developers by performing automated root cause analysis, stack trace evaluation, and systematic error triage to accelerate complex software debugging workflows.
- DevOps Infrastructure Automator — This tool assists engineers by automating CI/CD pipeline generation, infrastructure as code scaffolding, and deployment management across major cloud platforms.
- Checkly Monitoring Assistant — Provides expert guidance for Checkly, helping developers implement monitoring-as-code, set up checks, configure alerts, and integrate with CI/CD pipelines.
- Security Incident Responder — This skill classifies, triages, and manages declared security incidents, determining severity, escalation paths, and initiating forensic evidence collection for security teams.
- Log Analysis Tool — Analyzes application and server logs to identify root causes, patterns, and anomalies, aiding developers and operations teams.
- Systemd Service Manager — Manage Linux services using systemd to automate application startup, manage background processes, and configure service dependencies for developers.
- System Performance Monitor — This tool diagnoses performance bottlenecks in Claude Code and local systems by analyzing CPU, memory, and API latency to improve development efficiency.
- Server Operations Mentor — This skill provides architectural guidance on process management and monitoring strategies to help developers make informed production infrastructure decisions.
- Error Analysis Specialist — This tool assists developers by parsing logs and stack traces to identify error patterns, correlate system failures, and determine precise root causes.
- Kubernetes DNS Synchronizer — Automatically manages DNS records for Kubernetes Ingress and Service resources, simplifying DNS configuration for DevOps engineers and developers.
- Incident Response Automation — This tool automates production incident management by diagnosing root causes, drafting communications, and generating post-mortem reports to assist SREs during critical system outages.
- Distributed Tracing Architect — This tool helps backend engineers implement Jaeger and Tempo tracing to visualize request flows and diagnose latency across complex microservice architectures.
- API Load Tester — This tool stress-tests API endpoints under progressive concurrency to identify performance bottlenecks and provide actionable optimization recommendations for developers and engineers.
- Ecommerce Performance Monitor — This framework provides proactive monitoring, alerting, and diagnostic tools to help engineering and merchandising teams maintain optimal site speed and uptime.
- Configuration Validation Expert — This expert assistant helps developers and engineers create robust validation schemas and testing strategies to ensure application configurations remain secure, consistent, and error-free.