Grafana Alloy Log Collector CPU Guardrails

Log collection is workload traffic. When Grafana Alloy runs as a DaemonSet without a CPU limit, a log storm, destination retry loop, or backlog catch-up can consume most of a worker node. In virtualized clusters, that can also saturate the ESXi hosts underneath the cluster. Symptom The host-level signal may look like a vSphere capacity issue first: esxi-a CPU 95-105% esxi-b CPU 90-100% esxi-c CPU 5% Kubernetes then points at a smaller set of nodes: ...

September 15, 2026 · 5 min · Trinidad Marroquin

Loki Tempo And Incident Correlation Paths

Logs and traces are most useful when they connect to the same incident question as the alert. Loki and Tempo can provide that path, but only if labels, trace IDs, retention, and dashboard links are designed before the outage. Start From The Alert An alert should lead to a dashboard, and the dashboard should lead to logs and traces for the same scope. The path should be obvious: alert -> service dashboard -> Loki logs by service/namespace -> Tempo traces by trace ID or exemplar If responders must invent the LogQL query during every incident, the observability stack is not finished. ...

August 27, 2026 · 2 min · Trinidad Marroquin

Prometheus Operations And Query Safety

Prometheus is easy to start and easy to overload. Most production problems are not caused by one bad alert. They come from unclear scrape ownership, high-cardinality labels, expensive dashboard queries, or rule changes that nobody can explain during an incident. The same ownership rule applies to observability agents. If Grafana Alloy collects logs, give it resource limits and backend-readiness checks too. See Grafana Alloy Log Collector CPU Guardrails. Own The Scrape Path Every scrape target should have an owner and a reason to exist. ...

August 27, 2026 · 3 min · Trinidad Marroquin

Thanos And Mimir Long-Term Metrics Boundaries

Long-term metrics systems solve one problem and create several new ones. Thanos and Mimir can extend Prometheus retention, centralize query, support multi-cluster visibility, and make historical analysis possible. They also introduce object storage, compaction, tenancy, query limits, and failure modes that are easy to miss when the first dashboard loads. Define The Retention Contract Decide retention by use case: recent incident response. capacity planning. SLO reporting. audit or compliance review. seasonal comparison. cost allocation. Do not keep everything forever by default. Long retention with uncontrolled cardinality becomes a storage and query reliability problem. ...

August 27, 2026 · 2 min · Trinidad Marroquin

Service Mesh Telemetry And Debugging Checks

Service mesh incidents are request-path incidents. Debug them like one. The mesh adds useful telemetry, but it also adds another place where a request can fail. Start by proving whether the failure is application, ingress, Service/endpoints, NetworkPolicy, mesh policy, proxy health, or certificate state. Baseline The Path Write the expected path before changing anything: client -> ingress/load balancer -> source workload proxy -> destination service -> destination workload proxy -> application Then check the non-mesh objects first: ...

August 25, 2026 · 2 min · Trinidad Marroquin

Alertmanager Silences Need Independent Maintenance Health Checks

Alertmanager silences are useful during planned infrastructure work. They are also easy to misunderstand. A silence should suppress notification noise. It should not become the health check. During a Kubernetes node power-cycle window, the maintenance plan needed temporary silences for node, kubelet, etcd, and pod-scheduling alerts. That was reasonable: every node reboot would otherwise create expected alert noise. The risk was not the silence itself. The risk was treating the quiet notification path as proof that the cluster was healthy. ...

August 18, 2026 · 5 min · Trinidad Marroquin

Kubernetes Audit Logging Field Note

Kubernetes audit logging is the record of what happened to the cluster API. If it is not configured, enabled, and retained, a suspicious delete, a bad RBAC escalation, or a credential misuse becomes a debate instead of a line in a log. The API server generates the events; the platform team is responsible for making them useful. This note covers operating audit logs as working evidence, not just configuration. What Gets Recorded Audit log lines describe request-level events against the API server: ...

August 10, 2026 · 5 min · Trinidad Marroquin

Burn-Rate Alerting Concepts For Operators

Burn-rate alerting is easier to operate when the team understands the concepts before touching Prometheus rules. The implementation details matter, but the operational idea is simple: alert when a service is consuming its allowed failure budget too quickly. That is different from alerting because a metric crossed a convenient threshold. This note explains the mental model behind burn-rate alerts. For Prometheus examples and alert rule structure, see SLO Burn-Rate Alerting With Prometheus. For dashboard design after the alert fires, see What To Put On An SLO Dashboard. ...

July 28, 2026 · 5 min · Trinidad Marroquin

Incident Review Template For SRE Teams

An incident review should improve the system, not assign blame. The useful output is a better service, a better alert, a better runbook, or a clearer ownership boundary. If the review only produces a timeline and a vague action item, it is documentation theater. This template is designed for platform and SRE teams running production services, Kubernetes platforms, observability stacks, and shared infrastructure. Related notes: SLO Burn-Rate Alerting With Prometheus What To Put On An SLO Dashboard On-Call Escalation Policy For Platform Teams Incident Summary Incident title: Incident date: Severity: Status: Services affected: Primary owner: Incident commander: Review date: Write a short summary in plain language. ...

July 28, 2026 · 6 min · Trinidad Marroquin

On-Call Escalation Policy For Platform Teams

An escalation policy should tell the on-call engineer what to do when ownership, severity, or impact is unclear. It is not only a paging schedule. A schedule says who gets woken up. A policy says when to escalate, who should join, and what decisions are expected. For platform teams, this matters because incidents often cross boundaries: Kubernetes, networking, storage, CI/CD, observability, identity, and application ownership. Related notes: SLO Burn-Rate Alerting With Prometheus What To Put On An SLO Dashboard Incident Review Template For SRE Teams Policy Goals The policy should optimize for clear ownership and fast mitigation. ...

July 28, 2026 · 6 min · Trinidad Marroquin