Practical Data Science For DevOps And SRE is a hands-on series about using simple, explainable analysis methods in operational engineering work.

The goal is not prediction hype, black-box machine learning, or academic statistics. The goal is better awareness: seeing reliability patterns earlier, asking better questions during reviews, and making operational decisions with evidence instead of instinct alone.

Posts In This Series

  1. Using Percentiles Instead Of Averages In Reliability Reviews

Target Topics

  1. Using Simple Trend Analysis For SRE Signals
  2. Alert Fatigue Analysis With Basic Counts And Rates
  3. Detecting Recurring Incident Patterns From Postmortems
  4. Change Failure Rate Analysis Without Overengineering
  5. Capacity Forecasting With Moving Averages
  6. Using Percentiles Instead Of Averages In Reliability Reviews
  7. Correlating Deployments With Incident Windows
  8. Building A Lightweight Reliability Dataset From Tickets And Logs

Each post in this series leverages Data Science and stays close to real DevOps and SRE work: metrics, alerts, incidents, deployments, tickets, logs, postmortems, dashboards, and the operational judgment required to interpret them.