Kubernetes OS Maintenance Needs Storage Safety Gates

Updating Kubernetes nodes is not just an apt-get dist-upgrade loop. The safe boundary is different for package download staging, package application, reboot, etcd quorum, CSI detach behavior, and Longhorn replica health. A node can be Ready and still be the wrong node to reboot next. Separate Staging From Applying Use download-only staging when the goal is to warm package caches before a maintenance window: sudo bash -lc ' set -euo pipefail export DEBIAN_FRONTEND=noninteractive apt-get -o DPkg::Lock::Timeout=300 update -qq apt-get -y \ -o DPkg::Lock::Timeout=300 \ --download-only \ -o Dpkg::Options::=--force-confdef \ -o Dpkg::Options::=--force-confold \ dist-upgrade echo "STAGED_DOWNLOAD_ONLY rc=0" echo "upgradable_count=$(apt list --upgradable 2>/dev/null | grep -c upgradable)" echo "cache_mb=$(du -sm /var/cache/apt/archives 2>/dev/null | awk '\''{print $1}'\'')" echo "reboot_required=$(test -f /var/run/reboot-required && echo YES || echo no)" ' After download-only staging, apt list --upgradable can still show the same updates. That is expected. The packages were downloaded, not installed. ...

September 17, 2026 · 5 min · Trinidad Marroquin

Longhorn Free Space Is Not Schedulable Capacity

The disk had space. Longhorn still could not place the replicas. That was the useful part of the incident. The first read looked like an attached volume that had not finished rebuilding. The evidence pointed somewhere else: physical free space was not the same as capacity available to the Longhorn scheduler. Incident Shape The affected Longhorn volume was attached and degraded: volume: pvc-35d1949c-4a61-448d-88b9-59b30c823c81 state: attached robustness: degraded current node: worker-4 size: 53687091200 Its scheduled condition was false: ...

September 17, 2026 · 5 min · Trinidad Marroquin

Grafana Alloy Log Collector CPU Guardrails

Log collection is workload traffic. When Grafana Alloy runs as a DaemonSet without a CPU limit, a log storm, destination retry loop, or backlog catch-up can consume most of a worker node. In virtualized clusters, that can also saturate the ESXi hosts underneath the cluster. Symptom The host-level signal may look like a vSphere capacity issue first: esxi-a CPU 95-105% esxi-b CPU 90-100% esxi-c CPU 5% Kubernetes then points at a smaller set of nodes: ...

September 15, 2026 · 5 min · Trinidad Marroquin

Terraform Refresh-Only Before Narrow vSphere Applies

A Terraform plan can look like a small in-place update and still contain dangerous vSphere operations. When live VMs have been moved by incident response, storage maintenance, DRS, or a CSI controller, Terraform state may lag behind vCenter reality. A normal apply can try to move everything back, detach disks, or rewrite placement while you only meant to update a harmless metadata value. Situation The intended change was narrow: guestinfo.network-config: netmask /16 -> /22 The first plan was not narrow. It also included: ...

September 14, 2026 · 4 min · Trinidad Marroquin

KUBECONFIG Merge Order Can Shadow Fresh Rancher Tokens

A freshly downloaded kubeconfig can be valid and still fail when kubectl uses it. One failure pattern is duplicate user names across a merged KUBECONFIG. Rancher-generated kubeconfigs often use the same user name, such as rancher, across multiple clusters or downloaded files. When several kubeconfig files are merged, the first user with that name can win. A stale token from an older file can shadow the fresh token from the file you just downloaded. ...

September 8, 2026 · 3 min · Trinidad Marroquin

RKE2 System Upgrade Plans Need Taint-Aware Node Selection

An upgrade Plan can be syntactically correct and still create pods that can never schedule. One recurring pattern is a Plan that excludes control-plane and etcd nodes, but unintentionally includes monitor or infrastructure nodes with custom taints. The controller creates Jobs for those nodes. The scheduler rejects the pods. The Jobs age out with DeadlineExceeded. The cluster keeps generating warning events even though the nodes are healthy. The bug is not in the node. It is in the relationship between Plan selection and node taints. ...

September 8, 2026 · 3 min · Trinidad Marroquin

vSphere CSI Mount Loops Can Be Stale Multipath State

Not every FailedMount event means a pod is down. In one cluster, several pods were Running and Ready, but kubelet kept emitting vSphere CSI mount warnings for weeks. The warnings had this shape: MountVolume.MountDevice failed for volume "pvc-<uid>": mount failed: /dev/mapper/mpathX already mounted on /var/lib/kubelet/plugins/kubernetes.io/csi/csi.vsphere.vmware.com/<volume-id>/globalmount The workload was not currently down. The node had stale storage reconciliation state. Split Current Impact From Event Noise Start with the workload, not the event text: ...

September 8, 2026 · 4 min · Trinidad Marroquin

External Secrets Operator Ownership Boundaries

External Secrets Operator adds a useful boundary: applications can consume Kubernetes Secret objects while the source value lives in a dedicated secret manager. It also adds a new failure path. A backend token, SecretStore, controller, target Secret, and consuming workload all have to agree before rotation is actually safe. Ownership Model Write down the owner for each layer: backend secret manager -> source value and access policy SecretStore -> authentication and provider configuration ExternalSecret -> mapping from backend key to Kubernetes Secret target Secret -> Kubernetes object consumed by workloads application workload -> reload behavior and verification platform team -> controller health, RBAC, namespaces, and alerts If an ExternalSecret fails, do not assume it is an application bug or a secret-manager outage. Walk the chain. ...

August 30, 2026 · 3 min · Trinidad Marroquin

Kubernetes Secret Consumer Rotation

A Kubernetes Secret update is not the same thing as a completed rotation. The object can change in etcd while the application still uses an old environment variable, cached file, open connection, or credential loaded at process start. Rotation is complete only when the consumer uses the replacement and the old credential is revoked or made irrelevant. Identify The Consumer Path Start by finding how the workload consumes the secret: ...

August 30, 2026 · 2 min · Trinidad Marroquin

Kubernetes Secret Drift Expiry And Evidence

Secret review should answer operational questions without exposing secret values. The useful evidence is ownership, age, source, consumers, expiry, rollout state, and whether stale credentials still work. Decoding secret data into a ticket or chat channel usually creates a second incident. Inventory Without Values Start with metadata: kubectl get secret -A \ -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,TYPE:.type,AGE:.metadata.creationTimestamp' Then inspect labels and annotations: kubectl -n app-ns get secret app-secret -o yaml Review metadata, not data values: ...

August 30, 2026 · 2 min · Trinidad Marroquin