Kubernetes Maintenance Evidence Bundles Need A Redaction Plan

Good maintenance windows produce evidence. Bad evidence bundles become a new secret store. During RKE2 reboot and upgrade work, the most useful artifacts were not complicated: preflight JSON, per-batch reboot logs, final cluster state captures, and error files. They proved which context was used, which nodes were touched, whether boot IDs changed, whether workers were drained, whether PodDisruptionBudgets blocked eviction, and whether the cluster returned to the expected baseline. That same evidence can expose internal hostnames, IP addresses, SSH usernames, kubeconfig paths, workload names, Vault paths, webhook URLs, and temporary access helpers. Treat maintenance output as operational evidence and sensitive data at the same time. For the access-helper side of this problem, see Temporary Privileged DaemonSets Are Host Access Changes. ...

July 22, 2026 · 5 min · Trinidad Marroquin

Disabling Ubuntu Unattended Upgrades On Kubernetes Nodes

Ubuntu nodes can have the unattended-upgrades package installed without actively running upgrades. The package state alone is not enough. Check the config, service, timers, logs, and package history before deciding whether a node is safe. For production Kubernetes nodes, the desired state is usually: unattended-upgrades package absent or inert unattended-upgrades.service inactive and disabled apt-daily.timer disabled or masked apt-daily-upgrade.timer disabled or masked /etc/apt/apt.conf.d/20auto-upgrades set to 0 patching handled through controlled maintenance windows Audit First Use multiple signals. An empty unattended-upgrades log is useful, but it does not prove the feature is disabled. ...

July 20, 2026 · 6 min · Trinidad Marroquin

Rancher RKE2 Upgrade Pods Mutate The Host Filesystem

Rancher-managed RKE2 upgrades are Kubernetes workflows that intentionally mutate node-local filesystems. That sounds odd until you inspect an upgrade pod. The pod is scheduled onto the node being upgraded, mounts host paths, compares the new RKE2 binary inside the upgrade image with the host binary, replaces the host binary, and restarts the node service. That is the useful mental model: Rancher desired version -> system-upgrade-controller Plan -> upgrade Job/Pod on each selected node -> host filesystem mounted under /host -> RKE2 binary/config/service update on that node -> node restarts components and rejoins The pod is not just a health check. It is the delivery mechanism for node-local changes. ...

July 17, 2026 · 6 min · Trinidad Marroquin

RKE2 Worker Join Failures From Calico Wrong Interface Selection

Replacing an RKE2 worker can look like a token, hostname, or API availability problem when the real issue is lower in the node networking path. One useful pattern is to separate three signals that often appear together: the RKE2 agent cannot reach the supervisor through its local load balancer. the Kubernetes API reports duplicate node identity or stale node password state. Calico advertises an address from the wrong interface or subnet. Those symptoms are related, but they are not all fixed in the same place. ...

July 15, 2026 · 4 min · Trinidad Marroquin

Containerized DNS Drift Detector Boundaries

Containerizing a DNS drift detector is useful when the operator environment is inconsistent. The image can carry the shell runner, Ansible configuration, Python dependencies, and command defaults while the runtime provides the inventory, credentials, and report destination. The important boundary is this: the image should run audits, not own remediation or secrets. Image Contents Keep the image small and boring: FROM python:3.12-slim ENV ANSIBLE_CONFIG=/app/ansible.cfg \ PYTHONDONTWRITEBYTECODE=1 \ PYTHONUNBUFFERED=1 RUN apt-get update \ && apt-get install -y --no-install-recommends \ bash \ ca-certificates \ openssh-client \ sshpass \ bsdextrautils \ && rm -rf /var/lib/apt/lists/* WORKDIR /app COPY requirements.txt ./ RUN pip install --no-cache-dir -r requirements.txt COPY ansible.cfg ./ COPY bin/dns-drift-detector.sh /usr/local/bin/dns-drift-detector RUN chmod +x /usr/local/bin/dns-drift-detector ENTRYPOINT ["dns-drift-detector"] This gives the runner a predictable Bash and Ansible environment without baking in site inventory or operator credentials. ...

July 10, 2026 · 3 min · Trinidad Marroquin

Terraform Targeted Plans For Live Cluster Node Expansion

Targeted Terraform applies are a sharp tool. They are not a normal workflow, but they are sometimes the safer option when a live cluster needs a narrow expansion and the full plan contains unrelated refactor drift. This note covers the pattern for adding a small set of new monitor nodes to an existing RKE2 cluster while avoiding changes to existing etcd, control-plane, worker, and load balancer VMs. Situation The environment already had Terraform-managed vSphere VMs: ...

July 2, 2026 · 5 min · Trinidad Marroquin

SRE Agent Kubernetes Log Access With Namespaced RBAC

An SRE agent can authenticate to Rancher and still fail every useful workload call. Seeing clusters in Rancher does not automatically mean the account can read pods, logs, events, or namespace-scoped workloads in downstream clusters. This pattern applies when an SRE automation identity reports symptoms like: namespaces is forbidden: User "u-example" cannot list resource "namespaces" in API group "" at the cluster scope pods is forbidden: User "u-example" cannot list resource "pods" in namespace "app-namespace" The important distinction is scope. The identity may have management-plane visibility and read-only node visibility, but no downstream project or namespace RBAC. ...

June 29, 2026 · 4 min · Trinidad Marroquin

Concourse Ingress DNS And TLS Cutover

Concourse can be healthy inside the cluster while still failing from the operator’s browser. The common gap is not the web pod. It is the handoff between DNS, ingress controller placement, certificate material, and Concourse’s own external URL. Use this checklist when moving Concourse from a temporary URL or NodePort to a real hostname such as concourse.example.com. Confirm The Ingress Target Start by finding where ingress actually lands. Do not assume every worker should be in DNS until the ingress controller placement confirms it. ...

June 26, 2026 · 4 min · Trinidad Marroquin

Remediation Audits With NotReady Nodes And Calico Checks

Node remediation scripts often fail in boring ways: wrong inventory group, hidden Vault dependency, or one unreachable node that makes the final audit look stuck. Those failures matter because they can turn a clean remediation into a false incident, or worse, hide the one node that still needs attention. Use this workflow when cleaning Kubernetes node drift across a whole RKE2 cluster and validating Calico state at the same time. ...

June 26, 2026 · 5 min · Trinidad Marroquin

DNS Drift Detector Calico Overlay False Positives

DNS drift detectors need to distinguish resolver search domains from interface names. On Kubernetes nodes using Calico, resolvectl domain can include link names such as: Link 3 (caliabc123): Link 5 (vxlan.calico): Those strings can look domain-like to a naive regex. If the detector treats vxlan.calico as an active DNS search domain, clean nodes appear drifted. Symptom The DNS audit shows expected resolver state: resolv_conf_search = . netplan_search = [] resolv_conf_type = symlink:/run/systemd/resolve/stub-resolv.conf But the detector still marks nodes as drift because resolved_domains includes Calico overlay links. ...

June 24, 2026 · 2 min · Trinidad Marroquin