Mastering The Zen of Site Reliability Engineering
Course Overview
In this in-depth training, we will explore the foundational principles of Site Reliability Engineering (SRE) that are essential for balancing risk, minimizing operational toil, and fostering effective collaboration. This course is designed to help you cultivate an SRE mindset that prioritizes learning and adaptability in the face of ever-evolving technological challenges. By the end of this course, you will not only understand SRE best practices but also be able to implement these practices in your organization. Join us on this transformative journey to empower you to drive operational excellence and create resilient systems through collaboration and innovation.
Who is this course for
This course is ideal for system administrators, IT operations professionals, DevOps practitioners, and team leads eager to integrate SRE practices into their roles. Whether transitioning to an SRE role or aiming to enhance your current operations, this course is for you.
Key Benefits
- Position yourself as an SRE expert and boost your career prospects.
- Enhance your ability to manage incidents and conduct postmortem reviews.
- Reduce operational workload through strategic automation.
- Strengthen collaboration and foster a blameless, learning-oriented culture.
- Master essential tools for monitoring, CI/CD, and infrastructure management.
Learning Objectives
By the end of this course, you will be able to:
- Understand and apply key SRE concepts like error budgets and SLOs.
- Develop effective incident response and troubleshooting strategies.
- Implement automation to minimize operational toil.
- Foster a collaborative culture and continuous learning within teams.
- Plan and manage system capacity for future growth and stability.
Format
- Duration: 2 days
- Format: In-person or Remote, Instructor-led training (ILT)
- Course Difficulty: Intermediate
- Experience Level: Intermediate
- Hands-on Activities
Prerequisites
- Linux basics
- Networking basics
- Shell scripting
Course contents
Day 1: SRE Foundations and Core Practices
Module 1: Introduction to SRE, Error Budgets, SLA and SLO
Iliyan Petkov2025-04-14T12:23:27+00:00- SRE role and responsibilities
- Error Budgets and Success Metrics
- Service Level Agreements and Service Level Objectives
- AIOps and the Evolution of Service Management
Module 2: Managing the toil
Iliyan Petkov2025-04-14T12:22:56+00:00- Pros and cons of manual operations
- Reducing the toil
- How much automation
- Securing the Automation
Module 3: Service Monitoring and Service Indicators
Iliyan Petkov2025-04-14T12:22:44+00:00- Monitoring and Observability
- Monitoring Golden signals
- Service Level Indicators
- Preventing Overload
Module 4: Troubleshooting and emergency response
Iliyan Petkov2025-04-14T12:22:33+00:00- Effective problem management workflow
- Triaging, Diagnosing, and fixing a problem
- Handling emergencies
- Managing Incidents
Module 5: Learning from Failures
Iliyan Petkov2025-04-14T12:22:20+00:00- Why and how to learn from past failures
- Incident reviews and postmortems
- Introduction to Antifragility
- Antifragility versus resiliency
Module 6: Testing systems at scale
Iliyan Petkov2025-04-14T12:22:09+00:00- Why testing
- Types of tests
- Using dedicated environments
- Testing requirements
Day 2: SRE Culture, Tooling, and Advanced Practices
Module 7: SRE Mindset
Iliyan Petkov2025-04-14T12:21:55+00:00- SRE guiding principles
- Customer Satisfaction
- The art of change
- Continuous learning and knowledge sharing
Module 8: Developing Collaborative Culture
Iliyan Petkov2025-04-14T12:21:44+00:00- Sharing ownership and responsibility
- Fostering a blameless culture
- Interaction with customers and leadership
- Effective communication strategies
Module 9: Handling load at scale
Iliyan Petkov2025-04-14T12:21:33+00:00- Measuring utilization
- Load balancing strategies
- Addressing capacity limitations
- Reducing the load
Module 10: SRE tooling
Iliyan Petkov2025-04-14T12:21:18+00:00- Version control systems
- Internal Developer Portals and CI/CD
- Monitoring and incident management
- Infrastructure as a code
Module 11: SRE and Release Engineering
Iliyan Petkov2025-04-14T12:21:04+00:00- What is release engineering
- Continuous build and deployment
- Configuration management
- Gatekeeping
Module 12: Capacity Management Best Practices
Iliyan Petkov2025-04-14T12:20:51+00:00- Core principles of capacity management
- Sizing initial resource requirements
- Capacity planning - forecasting the future
- Avoiding failures through redundancies
Related Courses

Linux Fundamentals
The Linux Fundamentals course provides beginners with a comprehensive introduction to Linux operating systems. This course covers essential topics such as ...
VIEW MORE

Docker Foundation
This in-depth Docker Foundation course will give you a solid foundation in Docker and container technology, even if you are new to Docker ...
VIEW MORE

PYTHON FOUNDATION
The Python Foundation course provides beginners with a comprehensive introduction to Python programming. This course covers essential topics such as Python syntax ...
VIEW MORE

KUBERNETES KCNA
This comprehensive Kubernetes training prepares you to excel in the Kubernetes and Cloud Native Associate (KCNA) exam and managing containerized applications ...
VIEW MORE

DEVOPS FOUNDATION
DevOps is a set of practices that combines software development (Dev) and IT operations (Ops) to shorten the development lifecycle and deliver high-quality ...
VIEW MORE
Level up your skills and unlock your potential—contact us today to learn how to enroll and receive special volume discounts!