Cilium Readiness And CNI Ownership

Cilium is not just a CNI swap. It changes the cluster datapath, policy engine, observability surface, and sometimes kube-proxy ownership. Before treating it as production-ready, prove the network contract rather than stopping at Running pods. Define Ownership First Write down what owns each layer: cluster bootstrap -> installs the selected CNI once Cilium operator -> reconciles Cilium configuration and identities Cilium agents -> enforce datapath and policy on each node kube-proxy -> enabled, replaced, or intentionally absent platform team -> node routing, firewall, MTU, upgrades, and rollback app teams -> service labels, ports, and policy intent Do not let RKE2, Helm, GitOps, and manual manifests all believe they own CNI installation. One owner should install and upgrade the datapath. ...

August 28, 2026 · 2 min · Trinidad Marroquin

Kubernetes Egress Control Operating Model

Egress control is where platform security, application dependencies, DNS, and network operations collide. The goal is not simply to block outbound traffic. The goal is to make allowed outbound paths reviewable, observable, and recoverable when a dependency changes. Define The Egress Path For each cluster, document the path: pod -> CNI policy -> node routing -> NAT gateway or firewall -> proxy, if used -> external service Track ownership for: ...

August 28, 2026 · 2 min · Trinidad Marroquin

Kubernetes NetworkPolicy Rollout Boundaries

NetworkPolicy is production access control. Roll it out like a platform change, not a formatting cleanup. The risk is not only blocking traffic. The risk is blocking traffic without knowing which dependency the workload actually needed. Start With Namespace Ownership Before applying default deny, confirm: namespace owner. workload owner. expected ingress sources. expected egress destinations. DNS dependency. metrics and log scrape paths. admission webhook or sidecar dependencies. emergency rollback path. If the namespace owner cannot name the required network paths, observe first and enforce later. ...

August 28, 2026 · 2 min · Trinidad Marroquin

RKE2 etcd Startup Failures From Expired Peer Certificates

An RKE2 server stuck in activating can look like an API server problem. In one useful failure pattern, the API server is only downstream noise. The real issue is etcd peer TLS. The important clue is the startup sequence: RKE2 starts containerd starts kubelet starts static pods etcd starts or attempts to start RKE2 cannot read local etcd status API server readiness fails Do not begin with cluster-reset, snapshot restore, or deleting /var/lib/rancher/rke2/server/db. First prove what etcd is doing. ...

August 28, 2026 · 4 min · Trinidad Marroquin

vSphere CSI VolumeSnapshotClass Validation

Kubernetes snapshot support has several layers. A healthy vSphere CSI deployment does not automatically mean there is a usable snapshot class. The missing object is often small: apiVersion: snapshot.storage.k8s.io/v1 kind: VolumeSnapshotClass metadata: name: vsphere-csi-snapshot-class driver: csi.vsphere.vmware.com deletionPolicy: Delete That object tells Kubernetes which CSI driver should handle a VolumeSnapshot request. Verify The Existing Platform Pieces Start read-only: kubectl api-resources | grep -i snapshot kubectl get volumesnapshotclass kubectl get volumesnapshot -A kubectl get volumesnapshotcontent kubectl get pods -A | grep -Ei 'snapshot-controller|vsphere-csi' kubectl get csidriver csi.vsphere.vmware.com -o yaml You want to confirm: ...

August 28, 2026 · 4 min · Trinidad Marroquin

ArgoCD Application Ownership And Sync Boundaries

ArgoCD makes Kubernetes drift visible. It does not decide who owns the drift, when to sync, or whether pruning is safe. Those decisions need to be explicit before teams treat the ArgoCD UI as the deployment control plane. Define Application Ownership Every ArgoCD Application should have an owner who can answer: which repo path declares the workload. which team reviews changes. which namespace and cluster are in scope. whether sync is automatic or manual. whether prune and self-heal are enabled. what emergency live changes are allowed. how live fixes get written back to Git. If the owner is “platform,” but the application team controls the chart values, the ownership model is incomplete. ...

August 26, 2026 · 2 min · Trinidad Marroquin

Flux Reconciliation And Promotion Evidence

Flux is strongest when reconciliation evidence is easy to read without guessing which controller owns the change. The operator task is to connect Git revision, source artifact, rendered manifests, apply result, and workload health into one story. Follow The Controller Chain A Flux deployment often crosses several controllers: GitRepository or HelmRepository -> Kustomization or HelmRelease -> Kubernetes objects -> workload readiness When a rollout fails, check each link before changing manifests. ...

August 26, 2026 · 2 min · Trinidad Marroquin

GitOps Drift And Prune Safety Checks

Drift is a signal. Prune is a deletion engine. Treat them differently. A GitOps controller showing drift tells you declared state and live state differ. It does not automatically prove which side is correct. Classify Drift First Before syncing, classify the difference: expected controller mutation. admission webhook defaulting. Helm-rendered value change. manual incident fix. secret or certificate rotation. generated name or checksum annotation. unmanaged object created outside Git. stale manifest that should be removed. The response depends on the class. Syncing everything is fast, but it can erase the clue that explains the incident. ...

August 26, 2026 · 2 min · Trinidad Marroquin

Kubernetes Database Operator Ownership Boundaries

A database operator is not a database team in a container. It can reconcile StatefulSets, users, backups, replicas, failover, and certificates. It cannot decide your recovery objective, storage class risk, schema migration policy, or when data loss is acceptable. Those boundaries need to be explicit before production data lands in the cluster. Draw The Ownership Model Write down what each layer owns: database operator -> database cluster lifecycle and reconciliation storage layer -> persistent volume provisioning and attach/mount behavior backup system -> backup storage, retention, and restore mechanics application team -> schema migrations, connection behavior, and query load platform team -> node capacity, security policy, networking, and observability data owner -> RPO/RTO, retention, and data-loss acceptance If every issue is “the operator is broken,” the ownership model is not useful enough. ...

August 25, 2026 · 2 min · Trinidad Marroquin

MySQL Operator Replication And Restore Checks

MySQL operators can hide a lot of useful machinery behind a simple custom resource. That is convenient until replication or restore becomes an incident. The operating goal is not just a Running primary. It is a database service whose replication, backup, and restore behavior are understood before the outage. Know The Topology Before trusting the operator, record the topology it creates: primary instance. replica instances. writer endpoint. reader endpoint, if used. replication mode. binlog retention or archive behavior. backup schedule and target. failover policy. The topology should be visible to operators without reading the operator source code. ...

August 25, 2026 · 2 min · Trinidad Marroquin