Site Reliability Engineering

Inquire now

This course introduces Site Reliability Engineering (SRE), which applies software engineering principles to operating reliable, scalable services. Participants learn to define reliability targets, monitor service health, reduce repetitive operational work, manage incidents, and improve deployment and recovery practices.

The course covers vendor-neutral practices applicable to on-premises, cloud, and hybrid environments. A capstone brings the topics together through a reliability improvement plan for a sample service.

 

Duration 5 Days – 35 hrs.

 

Objectives

  • Explain SRE principles and their relationship to DevOps and IT operations.
  • Define service-level indicators, service-level objectives, and error budgets.
  • Design monitoring and alerting around service health and user experience.
  • Identify operational toil and opportunities for safe automation.
  • Apply structured incident response and blameless postmortem practices.
  • Recognize common reliability risks in distributed systems.
  • Plan safer deployments, capacity improvements, and service recovery.
  • Develop a practical reliability improvement plan.

 

Target Audience

  • Aspiring and practicing Site Reliability Engineers.
  • DevOps engineers and platform engineers.
  • Cloud engineers and infrastructure engineers.
  • System administrators and production support engineers.
  • Software engineers responsible for production services.
  • Technical leads and operations team leads.

  

Prerequisites

  • Basic knowledge of Linux or Unix operating systems and command-line operations.
  • Familiarity with networking concepts, including HTTP, DNS, and TCP/IP.
  • Basic scripting or programming experience.
  • Understanding of application deployment and version control.
  • Familiarity with service architectures and databases.
  • Basic exposure to cloud platforms and containers is helpful but not required.

 

Course Outline

Day 1: SRE Foundations and Reliability Targets

Module 1: Introduction to Site Reliability Engineering

  • Purpose, principles, and responsibilities of SRE.
  • Relationships among SRE, DevOps, and IT operations.
  • Service ownership and collaboration with development teams.
  • Reliability, availability, latency, and scalability.
  • Balancing reliability, delivery speed, and cost.

Module 2: Service-Level Indicators, Objectives, and Error Budgets

  • Critical user journeys and service boundaries.
  • Service-level indicators (SLIs), objectives (SLOs), and agreements (SLAs).
  • Selecting availability, latency, and correctness indicators.
  • Measurement windows and realistic reliability targets.
  • Calculating and interpreting error budgets.
  • Error budget policies and release decisions.

 

Day 2: Observability and Operational Efficiency

Module 3: Monitoring, Observability, and Alerting

  • Metrics, logs, and distributed traces.
  • The four golden signals: latency, traffic, errors, and saturation.
  • User-focused dashboards and dependency visibility.
  • Symptom-based and cause-based monitoring.
  • SLO-based alerts and error budget burn rates.
  • Alert routing, actionable notifications, and noise reduction.

 Module 4: Toil Reduction and Safe Automation

  • Distinguishing toil from valuable operational work.
  • Measuring toil and prioritizing reduction opportunities.
  • Automating repetitive service maintenance tasks.
  • Idempotency, validation, and failure handling.
  • Configuration consistency and infrastructure as code concepts.
  • Automation safeguards, rollback, and documentation.

 

Day 3: Incident Management and Learning from Failure

Module 5: On-Call Operations and Incident Response

  • On-call responsibilities, handovers, and escalation.
  • Incident severity and declaration criteria.
  • Incident command, technical response, and communication roles.
  • Triage and evidence-based troubleshooting.
  • Service mitigation, recovery, and stakeholder updates.
  • Runbooks and operational readiness.

 Module 6: Blameless Postmortems and Continuous Improvement

  • Establishing a blameless learning culture.
  • Documenting impact, timelines, and response actions.
  • Identifying contributing technical and organizational factors.
  • Defining corrective and preventive actions.
  • Assigning ownership and tracking improvements.
  • Identifying recurring incident patterns.

  

Day 4: Resilient Systems and Safe Changes

Module 7: Reliability Design, Capacity, and Recovery

  • Dependencies, failure domains, and cascading failures.
  • Redundancy, failover, and graceful degradation.
  • Timeouts, bounded retries, backoff, and circuit breakers.
  • Load balancing, rate limiting, and overload protection.
  • Capacity forecasting, load testing, and scaling.
  • Recovery objectives, backups, and restoration verification.

 Module 8: Release Engineering and Production Readiness

  • Reliability considerations in continuous integration and delivery.
  • Canary, rolling, and blue-green deployments.
  • Feature flags and controlled exposure.
  • Deployment health signals and rollback criteria.
  • Configuration changes and database migration risks.
  • Production readiness reviews and service acceptance.

 

Day 5: SRE Adoption and Capstone

Module 9: Establishing Sustainable SRE Practices

  • SRE engagement models and shared service ownership.
  • Prioritizing reliability work using service risk and error budgets.
  • Balancing engineering improvements with operational responsibilities.
  • Sustainable on-call practices and workload management.
  • Reliability reporting and stakeholder alignment.
  • Building a phased SRE adoption roadmap.

Module 10: Capstone — Service Reliability Improvement Plan

  • Map a sample service’s architecture, dependencies, and critical user journeys.
  • Define SLIs, SLOs, and an error budget policy.
  • Specify dashboards, alerts, and escalation paths.
  • Create an incident response runbook and sample postmortem.
  • Identify a toil reduction opportunity and design automation safeguards.
  • Outline deployment, rollback, capacity, and recovery measures.
  • Consolidate recommendations into a prioritized reliability improvement plan.

 

Inquire now

Best selling courses

CLOUD COMPUTING

Terraform

Terraform is a configuration orchestration tool for building and managing infrastructure on cloud & data centers. The course is instructor-led, live training (onsite or remote), and is designed for Engineers with little or no previous experience managing infrastructure. The course talks about in-depth Terraform syntax and techniques used to automate the setup and deployment of infrastructure.

Duration  3 days – 21 hrs    Overview    The ITIL Leadership – Digital and IT Strategy training course is designed for senior IT professionals, managers, and leaders who seek to navigate the complex landscape of digital transformation and IT strategy. This course focuses on providing strategic insights, leadership skills, and practical approaches for aligning...

PROGRAMMING / CODING

Spring Architecture and Design

Spring Cloud is a platform for building Java-based distributed systems and microservices. Building complex enterprise applications is challenging. Any change made to a part of the systems could trigger the need for changing the design of the entire system. By the end of this training, participants will have a solid understanding of Service-Oriented Architecture (SOA) and Microservice Architecture as well practical experience using Spring Cloud and related Spring technologies for rapidly developing their own cloud-scale, cloud-ready microservices.

BUSINESS INTELLIGENCE

Dax

Duration 5 days – 35 hrs   Overview The DAX (Data Analysis Expressions) Training Course is designed to provide participants with a comprehensive understanding of DAX, the powerful formula language used in Power BI, Excel, and SQL Server Analysis Services. This course covers the essential concepts, functions, and techniques required to create advanced calculations and...

OPERATING SYSTEMS

Linux Fundamentals

Linux Fundamental provides students a thorough introduction to Linux™ for those who are new to the Linux environment. Delegates will learn how to manage files and directories, utilize the vi editor, work with Linux security mechanisms to protect files and programs, work with the Linux shell to control the flow and processing of data through pipelines, design and write shell programs of moderate complexity, and manage multiple concurrent processes in order to achieve higher utilization of Linux. They will learn how to perform basic operations on the system and how quickly to solve problem.

PROGRAMMING / CODING

Google Apps Script

The Google Apps Script training course give you a detailed knowledge on coding like Automating data calculation, Fetching and sending data from third party software like Trello & Salesforce, connecting different sheets, Documents and other tools, Setting a trigger based on an event. This course is ideal for someone who use google sheets and have no coding background.

This workshop teaches the participants how to design and develop server side applications using the event-driven, non-blocking model framework Node.js. This program inducts the participant in some of the advanced concepts of the JavaScript language so that the participant is well equipped to build end-to-end application using JavaScript.

Duration: 3 days – 21 hrs   Overview This training course is designed to provide participants with a comprehensive understanding of Portfolio Management and Contract Management, focusing on best practices, tools, and techniques. The course covers the strategic alignment of projects within a portfolio, effective management of contracts, risk management, and optimization of resources to...

// BG EARTH WHEN NOT PLAYING

We use cookies on our website to personalize your experience by storing your preferences and recognizing repeat visits. By clicking “Accept”, you agree to the use of all cookies. You can also select “Cookie Settings” to adjust your preferences and provide more specific consent. Cookie Policy