Site Reliability Engineering Bundle | Prompeteer.ai

SLO/SLI design, error-budget reviews, runbook authoring, capacity planning, and toil-reduction skills for SREs.

Included Skills (38)

  1. SLO Reliability Architect — This tool assists SREs and developers in designing, calculating, and auditing meaningful service level objectives, error budgets, and burn-rate alerts for production systems.
  2. SLO Reliability Architect — This tool assists SREs and developers in designing, calculating, and reviewing meaningful service level objectives, error budgets, and burn-rate alerting strategies.
  3. Resilience Experiment Assistant — Helps SREs and developers plan, execute, and analyze chaos engineering experiments to improve system resilience and reliability.
  4. Resilience Experiment Assistant — Helps SREs and developers plan, execute, and analyze chaos engineering experiments to improve system resilience and reliability.
  5. Operational Runbook Generator — This tool analyzes codebase and infrastructure configurations to create structured, actionable runbooks that assist on-call engineers during incidents and routine system operations.
  6. Observability Strategy Architect — This tool helps site reliability engineers design production-ready monitoring strategies, optimize alert configurations, and generate comprehensive dashboards for complex distributed service architectures.
  7. Runbook Creation Assistant — Generates operational runbooks from a service name, streamlining documentation for on-call procedures and standardizing incident response across teams.
  8. Incident Response Orchestrator — This framework assists SRE and DevOps teams by managing operational outages through automated severity classification, timeline reconstruction, and structured post-incident review generation.
  9. Observability Stack Setup — Sets up end-to-end observability for microservices, helping developers and DevOps engineers monitor and debug production applications.
  10. PagerDuty Configuration — Configure PagerDuty for incident management, on-call scheduling, and alert routing, assisting DevOps engineers and system administrators.
  11. Vector Pipeline Assistant — Provides expert guidance for Vector, a high-performance observability data pipeline, helping developers manage logs, metrics, and traces efficiently.
  12. Pino Logging Integration — Integrate Pino into Node.js projects to enable structured, performant logging for improved observability and easier log aggregation.
  13. Thanos Deployment Manager — Deploys and configures Thanos for scalable Prometheus metric storage, querying, and data management, benefiting DevOps and SRE teams.
  14. Alertmanager Configuration Manager — Configures Prometheus Alertmanager for routing, grouping, and silencing alerts, helping DevOps engineers manage notifications and on-call responsibilities.
  15. Better Stack Assistant — Provides expert guidance for Better Stack, helping developers configure uptime monitoring, log management, incident response, and status pages.
  16. Uptime Kuma Monitor — Monitor service uptime, configure alerts, and create status pages using Uptime Kuma, a self-hosted monitoring tool.
  17. Envoy Proxy Expert — This skill helps teams configure Envoy as an API gateway, service mesh, and load balancer for cloud-native applications.
  18. Detecting Privilege Escalation In Kubernetes Pods — Detect and prevent privilege escalation in Kubernetes pods by combining
  19. Fluentd Log Management — Configure Fluentd for collecting, filtering, and routing logs across distributed systems, assisting developers and DevOps engineers with log aggregation.
  20. Datadog Management Skill — This skill configures and manages Datadog for infrastructure monitoring, APM, log management, and alerting, assisting DevOps engineers.
  21. Detecting AWS Guardduty Findings Automation — Build automated AWS GuardDuty finding response pipelines using EventBridge
  22. Telegraf Metrics Collector — Configure Telegraf to collect system and application metrics, set up input/output plugins, and build custom metric pipelines for infrastructure monitoring.
  23. DevOps Infrastructure Automator — This tool assists engineers by automating CI/CD pipeline generation, infrastructure as code scaffolding, and deployment management across major cloud platforms.
  24. Checkly Monitoring Assistant — Provides expert guidance for Checkly, helping developers implement monitoring-as-code, set up checks, configure alerts, and integrate with CI/CD pipelines.
  25. Load Balancer Configurator — This skill configures and optimizes load balancers and reverse proxies using Nginx, HAProxy, and cloud ALBs for DevOps engineers.
  26. Security Incident Responder — This skill classifies, triages, and manages declared security incidents, determining severity, escalation paths, and initiating forensic evidence collection for security teams.
  27. Google Cloud Global Frontend Configuration — The skill guides agents through a six‑step discovery process to design and deploy Google Cloud global external Application Load Balancers with Cloud CDN, Cloud Armor, and Service Extensions, mapping workload requirements to best‑practice configurations. It generates production‑grade Terraform HCL or gcloud CLI scripts, actuates deployments via Infrastructure Manager or bash scripts, and detects configuration drift. This service benefits architects and DevOps engineers who need to configure, deploy, and maintain secure, globally distributed load balancing solutions on Google Cloud.
  28. Loki Log Aggregator — This skill helps DevOps engineers and SREs deploy and use Loki for cost-effective log aggregation and analysis.
  29. Log Analysis Tool — Analyzes application and server logs to identify root causes, patterns, and anomalies, aiding developers and operations teams.
  30. Error Monitoring — The skill ingests logs from Sentry, Datadog, and Rollbar, automatically detecting patterns, duplicates, and severity levels. It produces categorized reports with root‑cause analysis, enabling teams to audit errors, triage production incidents, and reduce noisy alerts. Developers, operations engineers, and product managers gain clear, actionable insights into application failures.
  31. Systemd Service Manager — Manage Linux services using systemd to automate application startup, manage background processes, and configure service dependencies for developers.
  32. Implementing Web Application Logging With Modsecurity — Configure ModSecurity WAF with the OWASP Core Rule Set (CRS) for web
  33. Implementing Log Forwarding With Fluentd — The skill configures Fluent Bit to forward logs and Fluentd to aggregate, route, filter, and enrich them from syslog, file tails, and application sources. It produces configuration files that deliver logs to Elasticsearch, S3, and Splunk, enabling centralized log collection across distributed infrastructure. This is essential for engineers building or expanding a unified log pipeline.
  34. Nginx Configuration Assistant — This skill helps DevOps engineers configure Nginx for web serving, reverse proxying, load balancing, and security.
  35. Kubernetes DNS Synchronizer — Automatically manages DNS records for Kubernetes Ingress and Service resources, simplifying DNS configuration for DevOps engineers and developers.
  36. Incident Response Automation — This tool automates production incident management by diagnosing root causes, drafting communications, and generating post-mortem reports to assist SREs during critical system outages.
  37. API Load Tester — This tool stress-tests API endpoints under progressive concurrency to identify performance bottlenecks and provide actionable optimization recommendations for developers and engineers.
  38. Ecommerce Performance Monitor — This framework provides proactive monitoring, alerting, and diagnostic tools to help engineering and merchandising teams maintain optimal site speed and uptime.