SRE + Platform Engineering Program

A 16-weekend, hands-on program that takes engineers from SRE fundamentals to designing, building and running an Internal Developer Platform. Every week builds on the same Kubernetes cluster, application and Git repository, so by the final weekend you have a working platform and a portfolio of code and design documents.

Santosh H

Santosh H

Advanced

SRE & Platform Engineering
Our Course Benefits
Cup Icon

Define SLIs, SLOs and error budgets, and turn them into actionable alerts

Debug production issues across Linux, networking, Kubernetes and telemetry

Ship changes safely with canary releases, GitOps and tested rollback

Plan capacity, control cost, and design for disaster recovery

Build an Internal Developer Platform with Terraform, Argo CD and Backstage

Write production automation in Python and Go

Perform with confidence across troubleshooting, system design, and coding interviews

Hands-on labs every weekend run locally; no cloud account required

Career Sectors & Job Roles
Cup Icon

Site Reliability Engineer (SRE)

Platform Engineer

DevOps Engineer

Infrastructure Engineer

Cloud Operations Engineer

Systems Architect

What to expect from this course ?
Book Icon

The SRE + Platform Engineering Program is a four-month, 16-weekend program built around one continuous project: an Internal Developer Platform.

Every weekend follows a structured rhythm: concepts and architecture on Saturday morning, guided hands-on labs on Saturday afternoon, and deliverable review plus interview drills on Sunday.

Phase 1 (Foundations) covers SRE core concepts, where you will learn to define SLIs and SLOs, instrument observability with Prometheus, debug Linux and Kubernetes, and ship safely with Argo Rollouts.

In Phase 2 (Advanced SRE) and Phase 3 (Platform Engineering), you will scale your operations using OpenTelemetry pipelines, build disaster recovery with Velero, automate infrastructure via Terraform and Crossplane, and create a self-service developer portal using Backstage and GitOps.

Phase 4 (Career Readiness) focuses on engineering depth, teaching you production automation in Python and Go, non-abstract system design, and preparing you for senior mock interview loops (Troubleshooting, System Design, Coding, and Behavioral).

The Curriculum
Book Icon

  • How reliability targets are chosen and how they drive engineering decisions
  • Deliverable: SLO worksheet outlining SLIs, targets, and an error budget for a production-style service

  • Metrics and PromQL, the golden signals, labels and cardinality
  • Building dashboards that people actually use
  • Deliverable: PromQL queries and dashboards on the payments application

  • cgroups, processes and memory, packets and connections
  • Log analysis with grep, awk and sed, and distributed traces
  • Deliverable: Debug CPU, memory, disk and log-driven incidents on the payments app

  • Alertmanager routing, inhibition, and silences; SLO burn-rate alerting
  • Page versus ticket, incident roles, and blameless postmortems
  • Deliverable: Alerting labs with an alert console and multi-window burn-rate alerts

  • Pod lifecycle, probes, requests and limits, rollouts, PodDisruptionBudgets, and HPA
  • Service and DNS debugging
  • Deliverable: Three-node kind cluster with Prometheus and Grafana, 9 labs, and a blind incident drill

  • Deployment strategies, canary and blue/green releases, GitOps, progressive delivery
  • Rollback that works, including the database
  • Deliverable: Argo Rollouts canary and blue/green gated by Prometheus analysis

  • Chaos experiments with a hypothesis and game day design with incident roles
  • Interview preparation for SRE roles
  • Deliverable: Three-round game day; interview intensive with live mock and resume clinic

  • OpenTelemetry model (traces, metrics, logs, context propagation) and Collector pipelines
  • Head vs tail sampling, Grafana Alloy, Loki, Tempo, and Cardinality budgets
  • Deliverable: An slo/ folder generating alerts for two services
  • Interview Drill: Design observability for 200 microservices on a fixed budget

  • Capacity planning, load test types (smoke, load, stress, soak), and Little's Law
  • HPA vs VPA vs KEDA, Cluster Autoscaler, and Karpenter
  • FinOps fundamentals with OpenCost (cost per namespace/request)
  • Deliverable: A one-page capacity plan and a cost-per-service report
  • Interview Drill: Prepare for a 10x peak sale event in six weeks

  • RTO and RPO design, backup vs replication, and regional failover patterns
  • Stateful reliability, safe schema migrations, and Chaos Mesh experiments
  • Deliverable: Disaster recovery runbook with measured RTO/RPO and a Production Readiness Review
  • Interview Drill: Your primary region is down. Walk through the first 30 minutes

  • Terraform for SREs: HCL, remote backends, locking, lifecycle rules, OpenTofu
  • Module design, versioning, testing with terratest, and policy checks with Checkov
  • Deliverable: A versioned module repository and a CI pipeline that posts the plan to the PR
  • Interview Drill: Design infrastructure change management for 50 teams sharing one cloud estate

  • Platform as a product, Team Topologies, and IDP reference architecture
  • GitOps at scale with Argo CD (app-of-apps, ApplicationSets) and environment promotion
  • Multi-tenancy (namespaces, RBAC, quotas) and Crossplane compositions/claims
  • Deliverable: A platform charter, a GitOps-managed cluster and a Crossplane claim
  • Interview Drill: Design a multi-tenant Kubernetes platform for 40 teams

  • Backstage software catalog, scaffolder templates, and TechDocs
  • Software supply chain (SBOMs with Syft, image signing with cosign, SLSA levels)
  • Kyverno admission policies, Vault secrets, and cert-manager
  • Deliverable: A working 'create a service' flow, with unsigned images blocked at admission
  • Interview Drill: Design a secure delivery pipeline that produces audit evidence automatically

  • Python for SRE tooling, CLI tools, concurrency, timeouts, and retries
  • Go fundamentals for infrastructure tools (goroutines, channels, context)
  • Kubernetes client libraries and custom controllers
  • Deliverable: A health-checker CLI and a remediation controller with tests
  • Interview Drill: 45-minute practical coding round

  • Distributed systems fundamentals: replication, consensus, backpressure, rate limiting
  • Non-abstract design: capacity estimates with real numbers for requests, storage, bandwidth
  • Deliverable: Two written design documents with capacity estimates
  • Interview Drill: Full 60-minute system design round

  • Running the platform as a product, adoption metrics, and presenting to leadership
  • Behavioral rounds, story banks (ownership, conflict), resume, and GitHub portfolio
  • Deliverable: Capstone demo and mock interview scorecard
  • Interview Drill: Complete mock loop with written feedback

  • Step 1: Clarify users, features, traffic and constraints before drawing
  • Step 2: Agree on availability, latency and durability targets (SLOs)
  • Step 3: Requests per second, storage, bandwidth and machine count estimates
  • Step 4 & 5: High-level architecture, components, and Database/storage choices
  • Step 6 & 7: Failure modes, blast radius, recovery, and Observability signals

  • Traffic & Routing: L4/L7 load balancing, API gateways, service discovery
  • Data & Consistency: SQL vs NoSQL, sharding, Strong vs Eventual consistency, Quorums, Raft
  • Async Processing: Queues, streams, backpressure, retries, and idempotency
  • Protection & Resilience: Rate limiting, load shedding, circuit breakers, graceful degradation

Capstone Project: Mini Internal Developer Platform

Each participant builds and demos a mini Internal Developer Platform for a fictional healthcare or payments company. A new service goes from a portal click to production, then survives a seeded failure.

Terraform Infrastructure Pipeline

Infrastructure provisioned from Terraform modules through a plan-on-PR pipeline with LocalStack and remote state.

Backstage Golden-Path Template

One-click service creation generating repository, CI pipeline, Argo CD application, SLOs, and dashboards automatically.

GitOps Delivery & Admission Control

Application deployment via Argo CD with signed images (cosign) and Kyverno security admission policies enforced.

OpenTelemetry & Automated SLO Alerting

Telemetry pipeline capturing traces, metrics, and logs with automated burn-rate alerts generated via Sloth/OpenSLO.

View More
Get the complete course details in our brochure.

Discover all the essential information about our courses in our detailed brochure. Get insights on curriculum, schedules, and enrollment options to help you make the best choice for your education.

Ready to Master SRE & Platform Engineering?

A 16-weekend, hands-on program that takes engineers from SRE fundamentals to designing, building and running an Internal Developer Platform.

Monthly EMI options upto (24) Months
Monthly EMI options upto (24) Months

Flexible monthly EMI plans available for up to 24 months.

Modes of Payment (UPI, Cards, Wallet, Net Banking)
Modes of Payment (UPI, Cards, Wallet, Net Banking)

Explore various payment modes for secure and convenient transactions.

Course Fees

₹89,999

Final pricing refers to the last and definitive cost of a product or service, including all applicable fees and discounts.

Includes:

  • Live Interactive Classes
  • Lifetime Recorded Sessions
  • Study Material & PDFs
  • Enterprise Assignments
  • Production AI Projects
  • 1:1 Mentorship Sessions
  • Mock Interviews
  • Resume Reviews
  • Portfolio Building
  • Industry Certification
  • Placement Assistance