Fast OS Template Node Replacement Rehearsal

Replacing Kubernetes nodes from a fresh OS template can be faster than repairing legacy VM drift in place, but speed only helps if the risky work is moved out of the maintenance window without creating identity conflicts. The useful rehearsal pattern is to pre-create replacement VMs from the current template, leave them powered off, and join them one at a time during the window. That turns the maintenance window into a controlled cluster-change sequence instead of a race to clone, customize, debug, and drain all at once. ...

August 3, 2026 · 6 min · Trinidad Marroquin

Longhornctl Workstation Install And Operations Boundary

Longhorn has a CLI, longhornctl, but installing it should not change the operational source of truth during maintenance. For Rancher-managed Longhorn clusters, the most useful day-to-day state still usually comes from the Longhorn UI and Kubernetes CRDs such as volumes.longhorn.io, replicas.longhorn.io, nodes.longhorn.io, engines.longhorn.io, and settings.longhorn.io. The practical boundary is simple: use longhornctl as an operator tool, not as a reason to skip in-cluster evidence. Install Locally For an Ubuntu workstation, a user-local install avoids system-wide package changes. Confirm the workstation architecture first: ...

August 3, 2026 · 2 min · Trinidad Marroquin

RKE2 Calico Readiness Failures From Stale Port Owners

After an RKE2 upgrade or node maintenance event, Calico can fail readiness even when the node itself is Ready. One failure pattern is node-local and easy to miss: old Calico or Typha processes survive from a previous RKE2 runtime path and keep owning the ports that the current pods need. Kubernetes shows the new pods as unhealthy, but the real conflict is on the host. This is not a YAML problem first. It is a process ownership problem. ...

August 3, 2026 · 4 min · Trinidad Marroquin

cert-manager Certificate Lifecycle Field Note

Certificate automation is not finished when the first certificate is issued. The operational work is the lifecycle: issuer health, renewal timing, DNS or HTTP challenge reliability, secret ownership, expiration alerts, and safe rotation. This field note focuses on cert-manager as a Kubernetes platform component, not just as an application dependency. Define Certificate Ownership Every production certificate should have a clear owner. Track: hostname. namespace. owning application or platform team. Certificate resource name. Issuer or ClusterIssuer. challenge type. backing secret. renewal policy. alert route. If a certificate expires and nobody knows who owns it, the platform has an ownership problem, not only a TLS problem. ...

July 28, 2026 · 4 min · Trinidad Marroquin

Cluster Autoscaler Operational Review

Cluster autoscaler is capacity automation, not a replacement for capacity ownership. It can add nodes when pods cannot schedule and remove nodes when capacity is unused, but it only works inside the boundaries the platform team gives it: node groups, quotas, labels, taints, pod requests, disruption budgets, and cloud or virtualization capacity. This review is for platform teams that need to know whether autoscaling is safe, predictable, and observable. Start With The Capacity Contract Document the autoscaling boundary. ...

July 28, 2026 · 5 min · Trinidad Marroquin

Kubernetes Ingress Operations Checklist

Ingress is where Kubernetes platform problems become visible to users. The application may be healthy, pods may be ready, and services may have endpoints, but a bad ingress change can still break the user path. Treat ingress as a production edge, not just another manifest. This checklist is for platform teams operating ingress controllers, DNS records, TLS certificates, and application routes across Kubernetes clusters. Define The Ingress Contract Before troubleshooting an ingress issue, write down the expected path. ...

July 28, 2026 · 5 min · Trinidad Marroquin

Vault Kubernetes Auth Method Deep Dive

Vault Kubernetes auth lets workloads authenticate to Vault using Kubernetes service account identity. That makes it a powerful bridge between platform identity and secret access. It also means mistakes in service account binding, namespace scoping, policy mapping, or token lifetime can become production secret exposure. This note focuses on operating the auth method safely. The Auth Contract For every workload using Kubernetes auth, document: cluster. namespace. service account. Vault auth mount. Vault role. attached policies. token TTL. secret paths allowed. owner. Example: ...

July 28, 2026 · 4 min · Trinidad Marroquin

Longhorn No Scheduled Replicas Under Disk Pressure

Longhorn can report no scheduled replicas even when the Kubernetes PVC is still Bound and the data has not obviously disappeared. The message is easy to misread as a corrupt or missing volume. In practice, it often means Longhorn cannot place a replica on any eligible disk because scheduling rules, reservation, or nominal volume size have made the storage pool unschedulable. The important distinction is this: actual filesystem usage != Longhorn scheduled capacity A Longhorn node can have free disk blocks while still refusing new replicas because the declared size of scheduled replicas exceeds the safe scheduling limit. ...

July 24, 2026 · 8 min · Trinidad Marroquin

Temporary Privileged DaemonSets Are Host Access Changes

Sometimes the maintenance path is blocked by host access, not Kubernetes health. In one RKE2 reboot window, SSH worked to the nodes, but noninteractive sudo did not. The maintenance automation needed to reboot hosts and verify node-level state. The workaround was a temporary privileged DaemonSet that wrote a short-lived sudoers rule onto each node, then stayed alive only long enough for the maintenance window. That pattern can be valid in a controlled emergency or tightly scoped window. It is also a host access change. Treat it with the same seriousness as adding an SSH key, changing a sudoers file, or granting a break-glass account. ...

July 23, 2026 · 4 min · Trinidad Marroquin

Downstream Observability Teardown Triage

An observability namespace can look like a platform outage even when customer workloads are healthy. In one downstream-cluster investigation, the noisy symptoms were all in monitoring and alerting: logging resources disappearing, an agent operator stuck during deletion, alert delivery failures, external secret authentication errors, and exporter noise. The first useful move was to split the question in two: Is the cluster currently unhealthy? Is the observability stack being removed, broken, or reconciled by another owner? Those are different incidents. ...

July 22, 2026 · 5 min · Trinidad Marroquin