Packer SSH Disconnects From Service Restarts

expect_disconnect is useful, but it is not a broom. When a Packer vsphere-iso build fails with Script disconnected unexpectedly, the important question is whether the command already completed before SSH dropped. If the command was interrupted mid-package-install, telling Packer to expect the disconnect can turn a real partial install into a green build step. Symptom The build reaches a shell provisioner and fails after cloud-init finishes: Provisioning with shell script: /tmp/packer-shell123456789 Cloud-init finished... Synchronizing state of open-vm-tools.service with SysV service script... Executing: /usr/lib/systemd/systemd-sysv-install enable open-vm-tools Script disconnected unexpectedly. The template may build successfully in one vSphere environment and fail in another because network convergence and VMXNET3 behavior differ by host, port group, or timing. Do not assume the template is safe just because a sibling environment happened to tolerate the restart. ...

September 17, 2026 · 4 min · Trinidad Marroquin

Terraform Refresh-Only Before Narrow vSphere Applies

A Terraform plan can look like a small in-place update and still contain dangerous vSphere operations. When live VMs have been moved by incident response, storage maintenance, DRS, or a CSI controller, Terraform state may lag behind vCenter reality. A normal apply can try to move everything back, detach disks, or rewrite placement while you only meant to update a harmless metadata value. Situation The intended change was narrow: guestinfo.network-config: netmask /16 -> /22 The first plan was not narrow. It also included: ...

September 14, 2026 · 4 min · Trinidad Marroquin

vSphere CSI Mount Loops Can Be Stale Multipath State

Not every FailedMount event means a pod is down. In one cluster, several pods were Running and Ready, but kubelet kept emitting vSphere CSI mount warnings for weeks. The warnings had this shape: MountVolume.MountDevice failed for volume "pvc-<uid>": mount failed: /dev/mapper/mpathX already mounted on /var/lib/kubelet/plugins/kubernetes.io/csi/csi.vsphere.vmware.com/<volume-id>/globalmount The workload was not currently down. The node had stale storage reconciliation state. Split Current Impact From Event Noise Start with the workload, not the event text: ...

September 8, 2026 · 4 min · Trinidad Marroquin

vSphere CSI VolumeSnapshotClass Validation

Kubernetes snapshot support has several layers. A healthy vSphere CSI deployment does not automatically mean there is a usable snapshot class. The missing object is often small: apiVersion: snapshot.storage.k8s.io/v1 kind: VolumeSnapshotClass metadata: name: vsphere-csi-snapshot-class driver: csi.vsphere.vmware.com deletionPolicy: Delete That object tells Kubernetes which CSI driver should handle a VolumeSnapshot request. Verify The Existing Platform Pieces Start read-only: kubectl api-resources | grep -i snapshot kubectl get volumesnapshotclass kubectl get volumesnapshot -A kubectl get volumesnapshotcontent kubectl get pods -A | grep -Ei 'snapshot-controller|vsphere-csi' kubectl get csidriver csi.vsphere.vmware.com -o yaml You want to confirm: ...

August 28, 2026 · 4 min · Trinidad Marroquin

Full-Site Disaster Recovery Planning For Platform Teams

A full-site disaster recovery plan is not a backup inventory. It is a decision model for operating when the normal platform boundary is gone. Backups answer whether data exists somewhere else. DR answers whether the organization can restore service in the right order, with the right authority, inside an agreed failure window. Start With Service Tiers Do not plan every system as if it has the same recovery objective. Define tiers first: ...

August 23, 2026 · 3 min · Trinidad Marroquin

vSphere Cluster Design Operating Boundaries

A vSphere cluster is not just a folder of hosts. It is an operating boundary. The boundary should tell operators what can fail, what can be maintained, which workloads can coexist, and which automation contracts are safe to depend on. If the cluster boundary is unclear, Terraform, Packer, backup tooling, and Kubernetes node workflows inherit ambiguity. What The Cluster Owns Use the cluster to express infrastructure responsibility: HA and maintenance-mode headroom. host compatibility for the workloads placed there. common datastore and network reachability. DRS behavior and placement assumptions. operational ownership and escalation path. automation target for VM provisioning. Do not let clusters become accidental mixtures of every host that had spare capacity. That makes capacity look larger while making failure behavior harder to reason about. ...

August 23, 2026 · 3 min · Trinidad Marroquin

vSphere DRS And Resource Pool Operational Model

DRS and resource pools are useful when they express an operating model. They are dangerous when they become a second, invisible capacity plan. The mistake is assuming a resource pool is just a folder with limits. It is not. It changes scheduling behavior, and automation that deploys into it inherits those rules. DRS Is A Policy, Not Magic DRS helps place and rebalance workloads inside the cluster’s constraints. It cannot create host capacity, fix datastore contention, or make incompatible workloads safe to colocate. ...

August 23, 2026 · 3 min · Trinidad Marroquin

vSphere Tags And Custom Attributes For Automation Ownership

vSphere metadata becomes operational infrastructure once automation depends on it. Tags and custom attributes are useful because they make inventory searchable and machine-readable. They are risky when teams treat them as decoration. If Terraform, NetBox sync jobs, backup tools, or reports use a tag, that tag is an API contract. Tags Versus Custom Attributes Use tags for categorical membership: environment. workload class. backup policy. lifecycle state. automation owner. Kubernetes cluster or platform grouping. Use custom attributes for small pieces of descriptive data: ...

August 23, 2026 · 2 min · Trinidad Marroquin

Packer Windows Image Pipeline Boundaries

Windows images fail differently than Linux images, but the lifecycle boundary is the same: the template should contain the reusable mechanism, not the clone’s final identity. A Windows Packer build usually has more moving parts before the first provisioner runs: unattended installation or answer-file behavior. VMware Tools installation and reboot timing. WinRM enablement and firewall access. Windows Update or baseline configuration. sysprep and shutdown before template capture. vSphere clone customization or first-boot configuration after deployment. Do not compress those into “Packer built a Windows template.” Ask which stage owns which state. ...

August 21, 2026 · 4 min · Trinidad Marroquin

vSphere CD-ROM Cleanup With Power-Cycle Gates

The first version of a vSphere CD-ROM cleanup runbook treated device removal as a live VM reconfigure task. That worked for some VMs, but it also exposed a bad assumption: a CD-ROM change on a running Kubernetes VM is still a VM reconfigure operation, and VM reconfigure operations can disturb guest responsiveness, kubelet status, API availability, and storage controllers. The safer pattern was to stop trying to make CD-ROM removal invisible. Treat it as node maintenance: ...

August 13, 2026 · 5 min · Trinidad Marroquin