DNS Search Domain Debugging With systemd-resolved

Check Which Layer Is Injecting The Search Domain # 1. Is /etc/resolv.conf managed by systemd-resolved? ls -l /etc/resolv.conf # Expected: /etc/resolv.conf -> /run/systemd/resolve/stub-resolv.conf # 2. Show all search domains (global and per-link) resolvectl domain # Look for "Link <N> (<iface>): <domain>" — not just Global section # 3. What does the active resolver actually show? grep '^search' /etc/resolv.conf || echo "no search line" # 4. Which interface is injecting the domain? resolvectl domain | grep -B1 '\.' # 5. Is DHCP the source? networkctl status eth0 | grep -i domain The Three Layers Layer Inspect With Fix Netplan grep -R "search:" /etc/netplan/ Edit yaml, netplan generate, netplan apply systemd-resolved (global) resolvectl domain (Global:) resolvectl domain "" systemd-resolved (per-link) resolvectl domain (Link N:) resolvectl domain eth0 "" Runtime Fix (Per-Link Domain) Clearing a per-link search domain without restarting systemd-resolved: ...

June 10, 2026 · 3 min · Trinidad Marroquin

etcd Backup, Restore, And Disaster Recovery

Snapshot Manual Snapshot ETCDCTL_ENDPOINTS=https://127.0.0.1:2379 \ ETCDCTL_CACERT=/etc/kubernetes/pki/etcd/ca.crt \ ETCDCTL_CERT=/etc/kubernetes/pki/etcd/server.crt \ ETCDCTL_KEY=/etc/kubernetes/pki/etcd/server.key \ etcdctl snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db Automated Snapshot (RKE2) RKE2 includes etcd-snapshot as a subcommand: rke2 etcd-snapshot save \ --node-name <node-name> \ --s3 \ --s3-bucket=<bucket> \ --s3-region=<region> \ --s3-access-key=<key> \ --s3-secret-key=<secret> Automated snapshots can be configured via the RKE2 config file with etcd-snapshot-schedule-cron and etcd-snapshot-retention. Verify Snapshot ETCDCTL_ENDPOINTS=https://127.0.0.1:2379 \ etcdctl snapshot status /backup/etcd-snapshot-<date>.db -w table Verify the snapshot is not corrupt and check the revision and hash. A snapshot with zero revisions or a mismatched hash indicates corruption. ...

June 10, 2026 · 3 min · Trinidad Marroquin

etcd Member Management And Cluster Health

Check Cluster Health # Endpoint health (all members) etcdctl endpoint health -w table --cluster # Endpoint status with version and DB size etcdctl endpoint status -w table --cluster # Member list etcdctl member list -w table # Leader and term etcdctl endpoint status --cluster -w table | awk '{print $1, $4, $5}' Member Lifecycle Add A New Member # On an existing member, add the new peer etcdctl member add <node-name> --peer-urls=https://<new-peer-ip>:2380 # On the new node, start etcd with the initial cluster state set to "existing" # Then verify etcdctl member list -w table etcdctl endpoint health -w table --cluster Remove A Member etcdctl member remove <member-id> After removal, verify quorum and health. The removed member’s data directory can be cleaned up. ...

June 10, 2026 · 2 min · Trinidad Marroquin

Helm And Terraform Boundary On EKS

The boundary between Terraform and Helm is a common source of confusion. Terraform provisions infrastructure. Helm deploys applications. Terraform’s helm_release resource bridges them, but the chart templates stay in the application repository. For a runnable lab, see the helm-terraform-js-app directory in the IaC repository. The Pattern Terraform manages the Helm release with set blocks that inject environment-specific values: resource "helm_release" "my_app" { name = "my-app" chart = "${path.module}/../helm/myapp" namespace = kubernetes_namespace.my_app.metadata[0].name set { name = "image.repository" value = var.docker_image_repository } set { name = "image.tag" value = var.docker_image_tag } set { name = "replicaCount" value = var.replica_count } } The Helm chart stays portable. Environment-specific values live in Terraform variables. ...

June 10, 2026 · 2 min · Trinidad Marroquin

Longhorn To Enterprise Storage Migration Patterns

Migration Strategies Attach-And-Clone (Lowest Downtime) For platforms that support concurrent attachment: 1. Deploy new StorageClass targeting enterprise storage. 2. Create a clone of the existing PVC on the new StorageClass. 3. Attach the clone to a verification pod and validate data integrity. 4. Scale down the workload, detach old PVC, attach new PVC. 5. Scale up the workload. Downtime is limited to the time between scale-down and attach. Data integrity is verified before the cutover. ...

June 10, 2026 · 3 min · Trinidad Marroquin

vSphere CSI Attach And Mount Checklist

Use this checklist when a Kubernetes workload has a PVC that should use vSphere CSI, but the pod is pending, stuck in init, or failing to mount the volume. For resize-specific CNS task triage, see vSphere CSI CNS ExtendVolume Triage. If the only visible symptom is ContainerCreating, first split storage failures from missing ConfigMap or Secret dependencies; see Kubernetes ContainerCreating: Split Storage Failures From Missing Manifests. If GitOps keeps restarting CSI controllers during a storage freeze, see GitOps-Owned vSphere CSI Maintenance Pauses. If pods are running but kubelet keeps reporting already mounted errors through /dev/mapper/mpath*, inspect host multipath policy before blaming the workload; see vSphere CSI Mount Loops Can Be Stale Multipath State. ...

June 10, 2026 · 5 min · Trinidad Marroquin

iSCSI Bootstrap Readiness For Kubernetes Nodes

A Kubernetes node can be storage-ready before any new disk appears in lsblk. With CSI-backed storage, the disk appears only after the array maps a LUN to the node. Target State For a node with primary and iSCSI networks: ens192 -> service network, default route, DNS iscsi01 -> 169.253.0.x/24 iscsi02 -> 169.253.1.x/24 The node should have: deterministic netplan. reachable storage portals. open-iscsi installed and enabled. an initiator name. successful target discovery. persistent iSCSI node startup. Verify Network First ip -br addr ip route ping -c 3 169.253.0.11 ping -c 3 169.253.1.11 If portal pings fail, stop and fix networking before touching iSCSI. ...

June 9, 2026 · 2 min · Trinidad Marroquin

RKE2 kube-proxy CrashLoopBackOff After Upgrade Due To UFW

After an RKE2/Rancher upgrade, several kube-proxy pods entered CrashLoopBackOff. The initial event stream pointed at liveness probe failures: Warning Unhealthy kubelet Liveness probe failed: HTTP probe failed with statuscode: 503 Warning BackOff kubelet Back-off restarting failed container kube-proxy At first glance, this looked like a Kubernetes 1.31 kube-proxy probe behavior change or an RKE2 manifest mismatch. The real cause was more operational: UFW was active on a subset of Kubernetes nodes and was interfering with kube-proxy’s iptables dataplane programming. ...

February 3, 2026 · 7 min · Trinidad Marroquin