SRE / DevOps / Platform Engineering
Reliable infrastructure, kept understandable.
I am Trinidad Marroquin, an SRE / DevOps / Platform Engineer focused on Kubernetes, infrastructure automation, observability, secrets management, and production operations for systems that need to be operated with confidence. I also bring technical leadership experience from leading a DevOps engineering team for three years.
I work on infrastructure that needs to be understandable, repeatable, and boring in production. My focus areas include Kubernetes, cloud platforms, vCenter, Terraform, Rancher, Packer, Vault, CI/CD, Linux systems, and observability.
Best Fit#
- SRE, DevOps, Platform Engineering, and Infrastructure Automation roles.
- Kubernetes and Rancher platform operations, especially environments that need clear upgrade, recovery, and evidence practices.
- Terraform, Packer, Vault, CI/CD, Linux, vSphere, and observability-heavy infrastructure.
- Teams that value practical runbooks, safer changes, useful alerts, and systems that can be debugged under pressure.
Start Here#
Focus Areas#
Cluster operations, Rancher management, upgrades, access patterns, and production readiness.
Terraform modules, reviewable plans, environment promotion, and safer operational changes.
Cloud platform operations, vCenter administration, VM lifecycle, templates, identity, and network foundations.
Packer image pipelines and CI/CD workflows that keep deployments repeatable and debuggable.
HashiCorp Vault operations, policy design, secrets engines, authentication methods, and operational guardrails.
Metrics, logs, alerts, dashboards, and incident follow-up focused on useful operating signals.
What This Site Is For#
- Practical notes on SRE and DevOps work.
- Project writeups with constraints, tradeoffs, and operational outcomes.
- Field notes for practical command references, troubleshooting checks, and operational judgment.
- Technical blog posts about Kubernetes, cloud platforms, vCenter, Terraform, Rancher, Packer, Vault, CI/CD, and observability.
- A current resume and professional summary.
The headline is remote code execution. The operational lesson is broader: a secrets platform exploit chain is rarely one bug in isolation.
ControlPlane published A Realistic Code Execution Exploit Chain in OpenBao and Vault on September 28, 2026. The write-up describes a path from unauthenticated access to remote code execution by chaining issues across PKI ACME, certificate authentication, policy path canonicalization, namespace policy access, and Raft snapshot restore behavior.
This article does not reproduce the exploit. The public write-up says the full PoC was intentionally withheld while operators patch. That limits what can be proven without access to the real chain.
...
The placement pattern is easy to draw:
ESXi-1 ESXi-2 ESXi-3 ------ ------ ------ Control plane master-1 master-2 master-3 etcd etcd-1 etcd-2 etcd-3 The engineering question is harder:
If I distribute three masters and three etcd members across three ESXi hosts, what failures can the Kubernetes control plane actually survive? This article treats master-1, master-2, and master-3 as example VM names. Functionally, these are RKE2 server nodes running the Kubernetes API server, controller manager, scheduler, and related control-plane components with embedded etcd disabled. The etcd-1, etcd-2, and etcd-3 VMs are RKE2 server nodes running the etcd role with the control-plane components disabled. Current RKE2 documentation describes both dedicated etcd nodes and dedicated control-plane nodes through server role configuration.
...
“Vault is unreachable” was true, but it was not specific enough.
The useful part of this incident was separating two failure domains that overlapped in time: Vault nodes that needed unseal after an infrastructure event, and client traffic that was still being reset after Vault was healthy internally.
Incident Shape The incident context had two independent changes close together:
1. A rack power event affected the ESXi hosts running Vault VMs. 2. Firewall policy was tightened by removing a broad allow rule. After the VMs came back, Vault behaved like Vault should after restart when it uses Shamir sealing: nodes needed to be unsealed before they were usable.
...
Updating Kubernetes nodes is not just an apt-get dist-upgrade loop.
The safe boundary is different for package download staging, package application, reboot, etcd quorum, CSI detach behavior, and Longhorn replica health. A node can be Ready and still be the wrong node to reboot next.
Separate Staging From Applying Use download-only staging when the goal is to warm package caches before a maintenance window:
sudo bash -lc ' set -euo pipefail export DEBIAN_FRONTEND=noninteractive apt-get -o DPkg::Lock::Timeout=300 update -qq apt-get -y \ -o DPkg::Lock::Timeout=300 \ --download-only \ -o Dpkg::Options::=--force-confdef \ -o Dpkg::Options::=--force-confold \ dist-upgrade echo "STAGED_DOWNLOAD_ONLY rc=0" echo "upgradable_count=$(apt list --upgradable 2>/dev/null | grep -c upgradable)" echo "cache_mb=$(du -sm /var/cache/apt/archives 2>/dev/null | awk '\''{print $1}'\'')" echo "reboot_required=$(test -f /var/run/reboot-required && echo YES || echo no)" ' After download-only staging, apt list --upgradable can still show the same updates. That is expected. The packages were downloaded, not installed.
...
The disk had space. Longhorn still could not place the replicas.
That was the useful part of the incident. The first read looked like an attached volume that had not finished rebuilding. The evidence pointed somewhere else: physical free space was not the same as capacity available to the Longhorn scheduler.
Incident Shape The affected Longhorn volume was attached and degraded:
volume: pvc-35d1949c-4a61-448d-88b9-59b30c823c81 state: attached robustness: degraded current node: worker-4 size: 53687091200 Its scheduled condition was false:
...
expect_disconnect is useful, but it is not a broom.
When a Packer vsphere-iso build fails with Script disconnected unexpectedly, the important question is whether the command already completed before SSH dropped. If the command was interrupted mid-package-install, telling Packer to expect the disconnect can turn a real partial install into a green build step.
Symptom The build reaches a shell provisioner and fails after cloud-init finishes:
Provisioning with shell script: /tmp/packer-shell123456789 Cloud-init finished... Synchronizing state of open-vm-tools.service with SysV service script... Executing: /usr/lib/systemd/systemd-sysv-install enable open-vm-tools Script disconnected unexpectedly. The template may build successfully in one vSphere environment and fail in another because network convergence and VMXNET3 behavior differ by host, port group, or timing. Do not assume the template is safe just because a sibling environment happened to tolerate the restart.
...
Log collection is workload traffic.
When Grafana Alloy runs as a DaemonSet without a CPU limit, a log storm, destination retry loop, or backlog catch-up can consume most of a worker node. In virtualized clusters, that can also saturate the ESXi hosts underneath the cluster.
Symptom The host-level signal may look like a vSphere capacity issue first:
esxi-a CPU 95-105% esxi-b CPU 90-100% esxi-c CPU 5% Kubernetes then points at a smaller set of nodes:
...
A Terraform plan can look like a small in-place update and still contain dangerous vSphere operations.
When live VMs have been moved by incident response, storage maintenance, DRS, or a CSI controller, Terraform state may lag behind vCenter reality. A normal apply can try to move everything back, detach disks, or rewrite placement while you only meant to update a harmless metadata value.
Situation The intended change was narrow:
guestinfo.network-config: netmask /16 -> /22 The first plan was not narrow. It also included:
...
A freshly downloaded kubeconfig can be valid and still fail when kubectl uses it.
One failure pattern is duplicate user names across a merged KUBECONFIG. Rancher-generated kubeconfigs often use the same user name, such as rancher, across multiple clusters or downloaded files. When several kubeconfig files are merged, the first user with that name can win. A stale token from an older file can shadow the fresh token from the file you just downloaded.
...
An upgrade Plan can be syntactically correct and still create pods that can never schedule.
One recurring pattern is a Plan that excludes control-plane and etcd nodes, but unintentionally includes monitor or infrastructure nodes with custom taints. The controller creates Jobs for those nodes. The scheduler rejects the pods. The Jobs age out with DeadlineExceeded. The cluster keeps generating warning events even though the nodes are healthy.
The bug is not in the node. It is in the relationship between Plan selection and node taints.
...