Emergency Stop For Rancher System Upgrade Controller

Use this when a Rancher system-upgrade-controller worker Plan is actively cordoning or draining nodes and a normal live-object patch does not stick because GitOps restores it. Confirm The Controller Is Active kubectl get plan -n system-upgrade kubectl get jobs,pods -n system-upgrade -o wide kubectl get events -n system-upgrade --sort-by=.lastTimestamp Look for jobs like: apply-agent-plan-on-worker-1-... Repeated jobs can mean the upgrade job is timing out while trying to stop rke2-agent or rke2-server. In that case the controller may start over and attempt the shutdown again, keeping the selected worker in the disruption path. ...

June 24, 2026 · 3 min · Trinidad Marroquin

Calico IP Audit Zero Targets Does Not Mean Zero Nodes

An audit script can find contexts and still print zero rows. That may be success, not a parser failure. For Calico node IP audits, many scripts intentionally emit only mismatches between a Kubernetes node’s InternalIP and the Calico node annotation: projectcalico.org/IPv4Address If there are no mismatches, the target files can be empty. Symptom The audit finds contexts: Contexts: site-a-ops-rke2 site-a-prod-rke2 site-a-uat-rke2 But reports zero targets: Counts for DC=site-a: all: 0 prod: 0 uat: 0 qa: 0 dev: 0 unknown: 0 Before assuming the node query broke, inspect the selection logic. ...

June 22, 2026 · 3 min · Trinidad Marroquin

RKE2 Per-Cluster Inventory Generation From Kubeconfig

Inventory generation is only useful if the output matches the repo’s actual inventory contract. A script can successfully read kubectl get nodes and still generate inventory that the playbooks cannot use. The important details are usually group names, expected child groups, node roles, load balancer entries, and local conventions around readiness metadata. Problem A per-cluster RKE2 inventory workflow needed output like: site_a_ops_rke2 rke2_servers rke2_agents api_lb But the generator output did not match the checked-in inventory shape. Common breaks included: ...

June 22, 2026 · 3 min · Trinidad Marroquin

NetBox Cluster Tag Scope Audit

When NetBox inventory moves from DCIM devices to virtualization VMs, tags need their own audit. A tag can exist and still be wrong for the new object model. If a cluster tag is scoped only to DCIM devices or IP addresses, migrated virtualization VMs may not be taggable or discoverable through normal cluster filters. Audit Cluster Tag Scopes Set API variables: NETBOX_URL="${NETBOX_URL:-https://netbox.example.com}" TOKEN="${NETBOX_TOKEN:-}" case "$TOKEN" in nbt_*) AUTH="Authorization: Bearer $TOKEN" ;; *) AUTH="Authorization: Token $TOKEN" ;; esac List cluster-* tags missing virtualization VM scope: ...

June 18, 2026 · 2 min · Trinidad Marroquin

Replacement Node Workflow After Terraform Import Drift

Importing existing vSphere VMs into Terraform can produce a clean source-of-truth checkpoint and still leave a plan that should not be applied. That is common when legacy Kubernetes nodes were built from an older template or outside the current module conventions. Audit Checkpoint A useful checkpoint looks like this: NetBox resources: no-op vSphere resources: update destroy actions: none This means NetBox ownership is reconciled, but Terraform still sees vSphere drift. ...

June 18, 2026 · 3 min · Trinidad Marroquin

Velero Workload Backup And Restore For Kubernetes

Velero has two parts: the CLI is installed on a management machine (your jumpbox or workstation) and issues commands to the cluster via kubectl; the server and node-agent run as pods inside the Kubernetes cluster and carry out the actual backup and restore work. All velero backup, restore, and schedule commands below assume the CLI is installed on a machine with kubectl access to the target cluster. Prerequisites Velero backs up Kubernetes resources and, with the restic/kopia integration, persistent volume data. The backup target must be an S3-compatible object store. ...

June 18, 2026 · 4 min · Trinidad Marroquin

Classifying vSphere Drift After Terraform Import

After existing vSphere VMs are imported into Terraform state, expect drift. The question is not whether drift exists. The question is whether applying that drift is safe. Generate The Audit Plan terraform plan -out=audit.tfplan terraform show -json audit.tfplan \ | jq -r '.resource_changes[]? | [.address, .type, (.change.actions | join(","))] | @tsv' Start with action types: no-op update create delete delete,create For imported production-like nodes, any delete, create, or delete,create action needs explicit review before apply. ...

June 17, 2026 · 3 min · Trinidad Marroquin

NetBox Bulk Import Order For Kubernetes Nodes

NetBox imports are safest when object relationships are created in dependency order. For Kubernetes VM nodes, split the import into devices, interfaces, and IP addresses instead of trying to represent everything as one operation. Import Order Use this order: 1. devices.csv 2. interfaces.csv 3. ip-addresses.csv Why: interfaces need devices to exist first. IP addresses need interfaces to exist before assignment. primary IP selection is a device relationship and may need a later update. Device CSV Example shape: ...

June 12, 2026 · 2 min · Trinidad Marroquin

Network Saturation Evidence Checklist

When a flow collector reports a huge byte count, resist the urge to start with the loudest application log. Start by proving whether the node, interface, and time window support the story. This checklist is useful when Kubernetes or Rancher appears to be involved in a network spike. Identify The Conversation Ask for the flow detail, not only the top-talker summary: source IP destination IP source port destination port protocol bytes sessions start time end time Important ports in Rancher/RKE2 environments include: ...

June 11, 2026 · 3 min · Trinidad Marroquin

Rancher Remotedialer Network Symptoms

Rancher cluster-agent errors can look like Rancher is the root cause. Sometimes Rancher is only the first component sensitive enough to report a degraded network path. Use this note when Rancher-managed clusters show large management-plane traffic, intermittent disconnects, or websocket errors. Common Symptoms In downstream cluster-agent logs: Failed to dial steve aggregation server Remotedialer proxy error websocket: close 1006 (abnormal closure) context deadline exceeded i/o timeout connection reset by peer In Rancher server logs: ...

June 11, 2026 · 3 min · Trinidad Marroquin