This course introduces Site Reliability Engineering (SRE), which applies software engineering principles to operating reliable, scalable services. Participants learn to define reliability targets, monitor service health, reduce repetitive operational work, manage incidents, and improve deployment and recovery practices.
The course covers vendor-neutral practices applicable to on-premises, cloud, and hybrid environments. A capstone brings the topics together through a reliability improvement plan for a sample service.
Duration 5 Days – 35 hrs.
Objectives
- Explain SRE principles and their relationship to DevOps and IT operations.
- Define service-level indicators, service-level objectives, and error budgets.
- Design monitoring and alerting around service health and user experience.
- Identify operational toil and opportunities for safe automation.
- Apply structured incident response and blameless postmortem practices.
- Recognize common reliability risks in distributed systems.
- Plan safer deployments, capacity improvements, and service recovery.
- Develop a practical reliability improvement plan.
Target Audience
- Aspiring and practicing Site Reliability Engineers.
- DevOps engineers and platform engineers.
- Cloud engineers and infrastructure engineers.
- System administrators and production support engineers.
- Software engineers responsible for production services.
- Technical leads and operations team leads.
Prerequisites
- Basic knowledge of Linux or Unix operating systems and command-line operations.
- Familiarity with networking concepts, including HTTP, DNS, and TCP/IP.
- Basic scripting or programming experience.
- Understanding of application deployment and version control.
- Familiarity with service architectures and databases.
- Basic exposure to cloud platforms and containers is helpful but not required.
Course Outline
Day 1: SRE Foundations and Reliability Targets
Module 1: Introduction to Site Reliability Engineering
- Purpose, principles, and responsibilities of SRE.
- Relationships among SRE, DevOps, and IT operations.
- Service ownership and collaboration with development teams.
- Reliability, availability, latency, and scalability.
- Balancing reliability, delivery speed, and cost.
Module 2: Service-Level Indicators, Objectives, and Error Budgets
- Critical user journeys and service boundaries.
- Service-level indicators (SLIs), objectives (SLOs), and agreements (SLAs).
- Selecting availability, latency, and correctness indicators.
- Measurement windows and realistic reliability targets.
- Calculating and interpreting error budgets.
- Error budget policies and release decisions.
Day 2: Observability and Operational Efficiency
Module 3: Monitoring, Observability, and Alerting
- Metrics, logs, and distributed traces.
- The four golden signals: latency, traffic, errors, and saturation.
- User-focused dashboards and dependency visibility.
- Symptom-based and cause-based monitoring.
- SLO-based alerts and error budget burn rates.
- Alert routing, actionable notifications, and noise reduction.
Module 4: Toil Reduction and Safe Automation
- Distinguishing toil from valuable operational work.
- Measuring toil and prioritizing reduction opportunities.
- Automating repetitive service maintenance tasks.
- Idempotency, validation, and failure handling.
- Configuration consistency and infrastructure as code concepts.
- Automation safeguards, rollback, and documentation.
Day 3: Incident Management and Learning from Failure
Module 5: On-Call Operations and Incident Response
- On-call responsibilities, handovers, and escalation.
- Incident severity and declaration criteria.
- Incident command, technical response, and communication roles.
- Triage and evidence-based troubleshooting.
- Service mitigation, recovery, and stakeholder updates.
- Runbooks and operational readiness.
Module 6: Blameless Postmortems and Continuous Improvement
- Establishing a blameless learning culture.
- Documenting impact, timelines, and response actions.
- Identifying contributing technical and organizational factors.
- Defining corrective and preventive actions.
- Assigning ownership and tracking improvements.
- Identifying recurring incident patterns.
Day 4: Resilient Systems and Safe Changes
Module 7: Reliability Design, Capacity, and Recovery
- Dependencies, failure domains, and cascading failures.
- Redundancy, failover, and graceful degradation.
- Timeouts, bounded retries, backoff, and circuit breakers.
- Load balancing, rate limiting, and overload protection.
- Capacity forecasting, load testing, and scaling.
- Recovery objectives, backups, and restoration verification.
Module 8: Release Engineering and Production Readiness
- Reliability considerations in continuous integration and delivery.
- Canary, rolling, and blue-green deployments.
- Feature flags and controlled exposure.
- Deployment health signals and rollback criteria.
- Configuration changes and database migration risks.
- Production readiness reviews and service acceptance.
Day 5: SRE Adoption and Capstone
Module 9: Establishing Sustainable SRE Practices
- SRE engagement models and shared service ownership.
- Prioritizing reliability work using service risk and error budgets.
- Balancing engineering improvements with operational responsibilities.
- Sustainable on-call practices and workload management.
- Reliability reporting and stakeholder alignment.
- Building a phased SRE adoption roadmap.
Module 10: Capstone — Service Reliability Improvement Plan
- Map a sample service’s architecture, dependencies, and critical user journeys.
- Define SLIs, SLOs, and an error budget policy.
- Specify dashboards, alerts, and escalation paths.
- Create an incident response runbook and sample postmortem.
- Identify a toil reduction opportunity and design automation safeguards.
- Outline deployment, rollback, capacity, and recovery measures.
- Consolidate recommendations into a prioritized reliability improvement plan.

