Alertmanager Silences Need Independent Maintenance Health Checks

Alertmanager silences are useful during planned infrastructure work. They are also easy to misunderstand. A silence should suppress notification noise. It should not become the health check. During a Kubernetes node power-cycle window, the maintenance plan needed temporary silences for node, kubelet, etcd, and pod-scheduling alerts. That was reasonable: every node reboot would otherwise create expected alert noise. The risk was not the silence itself. The risk was treating the quiet notification path as proof that the cluster was healthy. ...

August 18, 2026 · 5 min · Trinidad Marroquin

Kubernetes ContainerCreating: Split Storage Failures From Missing Manifests

Two pods can sit in ContainerCreating for days and need completely different fixes. In one production triage, the first pod was blocked by a CSI volume that Kubernetes could request but the storage array would not present. The second pod was blocked by missing ConfigMap and Secret objects referenced by the pod spec. The visible symptom was the same. The evidence path was not. That distinction matters because ContainerCreating is not a root cause. It is the kubelet saying it cannot finish preparing the container sandbox, volumes, or configuration. ...

August 18, 2026 · 6 min · Trinidad Marroquin

vSphere CD-ROM Cleanup With Power-Cycle Gates

The first version of a vSphere CD-ROM cleanup runbook treated device removal as a live VM reconfigure task. That worked for some VMs, but it also exposed a bad assumption: a CD-ROM change on a running Kubernetes VM is still a VM reconfigure operation, and VM reconfigure operations can disturb guest responsiveness, kubelet status, API availability, and storage controllers. The safer pattern was to stop trying to make CD-ROM removal invisible. Treat it as node maintenance: ...

August 13, 2026 · 5 min · Trinidad Marroquin

CIS Benchmark Review Process

A CIS benchmark review produces a long list of checks. The failure is not having fails; the failure is treating every check with equal weight. CIS profiles are dense and some controls conflict with how a platform is actually operated. This note is about reviewing CIS results in a way that produces evidence, priorities, and defensible decisions instead of a scary PDF. Scope First Before running any scanner, define the scope: ...

August 10, 2026 · 5 min · Trinidad Marroquin

Kubernetes Audit Logging Field Note

Kubernetes audit logging is the record of what happened to the cluster API. If it is not configured, enabled, and retained, a suspicious delete, a bad RBAC escalation, or a credential misuse becomes a debate instead of a line in a log. The API server generates the events; the platform team is responsible for making them useful. This note covers operating audit logs as working evidence, not just configuration. What Gets Recorded Audit log lines describe request-level events against the API server: ...

August 10, 2026 · 5 min · Trinidad Marroquin

Pod Security Standards For Platform Teams

Pod Security Standards are the default answer for “make workloads run non-privileged” in Kubernetes. They replace the removed PodSecurityPolicy admission controller with three policies, Baseline, Restricted, and Privileged, that can be enforced or audited per namespace. The model is simple, but production mistakes come from exemptions, admission confusion, and treating the label as the control. This note covers operating Pod Security Standards on real clusters. Know The Three Policies Privileged: unrestricted, for system workloads that legitimately need it. Baseline: allows the default Kubernetes behavior while preventing known privilege escalation. Good starting point for most non-system namespaces. Restricted: the hardened target. No privileged containers, no host network, strict volume and capability limits. Baseline is the pragmatic default. Restricted is the goal where the workload can satisfy it. Privileged is a documented exception, not a default. ...

August 10, 2026 · 4 min · Trinidad Marroquin

GitOps-Owned vSphere CSI Maintenance Pauses

During a vSphere CNS backlog, stopping the source of new CSI work can be safer than continuing to submit attach, detach, resize, and update requests into a stuck backend queue. The trap is that GitOps may immediately undo the pause. Scaling a vSphere CSI controller Deployment to 0 is not a durable pause if Argo CD, an ApplicationSet, or a higher root Application owns it. The visible command succeeds, then self-heal restores the replicas and new controller pods start submitting tasks again. ...

August 5, 2026 · 4 min · Trinidad Marroquin

Kubernetes Retained PV Missing PVC Rebind

A pod stuck in Pending because a PVC does not exist is not always a provisioning problem. Sometimes the data still exists in a retained PV and only the claim object is missing. The safe recovery is to prove the PV is the intended backing store, then recreate the PVC with an explicit volumeName so it binds to that retained PV instead of provisioning something new. Confirm The Symptom Start with the pod event: ...

August 5, 2026 · 3 min · Trinidad Marroquin

vSphere CSI CNS ExtendVolume Triage

A vSphere CSI controller log that says a CNS ExtendVolume task is pending does not immediately tell you whether the workload is blocked by resize, attachment, mount, or a stale backend task. The first job is to map the CNS volume ID back to Kubernetes state and avoid turning a storage delay into a destructive rollback. For the broader maintenance journal that ties this pattern to GitOps containment, vCenter task queues, VM device cleanup, and stop conditions, see Journal: When Storage Triage Turns Into Platform Maintenance. ...

August 5, 2026 · 7 min · Trinidad Marroquin

vSphere HA Reset Evidence For Kubernetes Nodes

When Kubernetes nodes reboot during a storage incident, the cause matters. An in-guest reboot, a Rancher/system-upgrade action, a human SSH session, and a vSphere HA reset all have different follow-up work. The useful pattern is to prove the reboot from multiple layers before assigning cause. Start With Kubernetes Check node readiness, boot ID, kernel, and recent events: kubectl get nodes -o wide kubectl describe node cp-1 | grep -A5 -E 'Conditions:|Events:' kubectl get events -A --sort-by=.lastTimestamp | grep -i 'reboot\|node' kubectl get node cp-1 -o jsonpath='{.status.nodeInfo.bootID}{"\n"}{.status.nodeInfo.kernelVersion}{"\n"}' Kubernetes can tell you that a node rebooted. It usually cannot tell you why. ...

August 5, 2026 · 3 min · Trinidad Marroquin