Practical Data Science For DevOps And SRE is a hands-on series about using simple, explainable analysis methods in operational engineering work.
The goal is not prediction hype, black-box machine learning, or academic statistics. The goal is better awareness: seeing reliability patterns earlier, asking better questions during reviews, and making operational decisions with evidence instead of instinct alone.
Posts In This Series
Target Topics
- Using Simple Trend Analysis For SRE Signals
- Alert Fatigue Analysis With Basic Counts And Rates
- Detecting Recurring Incident Patterns From Postmortems
- Change Failure Rate Analysis Without Overengineering
- Capacity Forecasting With Moving Averages
- Using Percentiles Instead Of Averages In Reliability Reviews
- Correlating Deployments With Incident Windows
- Building A Lightweight Reliability Dataset From Tickets And Logs
Each post in this series leverages Data Science and stays close to real DevOps and SRE work: metrics, alerts, incidents, deployments, tickets, logs, postmortems, dashboards, and the operational judgment required to interpret them.