Log collection is workload traffic.

When Grafana Alloy runs as a DaemonSet without a CPU limit, a log storm, destination retry loop, or backlog catch-up can consume most of a worker node. In virtualized clusters, that can also saturate the ESXi hosts underneath the cluster.

Symptom

The host-level signal may look like a vSphere capacity issue first:

esxi-a  CPU 95-105%
esxi-b  CPU 90-100%
esxi-c  CPU 5%

Kubernetes then points at a smaller set of nodes:

kubectl --context cluster-a top nodes | sort -k3 -nr | head

Example shape:

worker-14  7200m  90%
worker-04  7100m  89%
worker-01  7000m  88%
worker-13  6900m  87%

The pod view identifies the workload:

kubectl --context cluster-a top pods -A --no-headers \
  | sort -k3 -hr \
  | head -20

Example shape:

grafana-alloy  alloy-logs-gm6sf   7100m  2800Mi
grafana-alloy  alloy-logs-95xtv   7000m  2800Mi
grafana-alloy  alloy-logs-jhc4w   6900m  2800Mi
grafana-alloy  alloy-logs-k2gmc   5800m  3000Mi

If the hot Alloy pods sit on the hot nodes, the issue is not abstract host pressure. It is a log collection workload consuming node CPU.

Check The Resource Policy

Inspect the DaemonSet resources before changing anything:

kubectl --context cluster-a -n grafana-alloy get ds alloy-logs \
  -o jsonpath='{range .spec.template.spec.containers[*]}name={.name}{"\n"}resources={.resources}{"\n---\n"}{end}'

A risky shape is:

name=alloy
resources={"limits":{"memory":"3Gi"},"requests":{"cpu":"200m","memory":"256Mi"}}
---
name=config-reloader
resources={"requests":{"cpu":"10m","memory":"50Mi"}}

The Alloy container requests 200m, but has no CPU limit. It can consume the whole node if the workload allows it.

Save A Rollback

Save the current object before patching:

kubectl --context cluster-a -n grafana-alloy get ds alloy-logs -o yaml \
  > alloy-logs.before.yaml

Rollback is then simple:

kubectl --context cluster-a -n grafana-alloy apply -f alloy-logs.before.yaml

Apply A Conservative Limit

Patch only the Alloy container. Leave sidecars alone unless they are part of the problem.

kubectl --context cluster-a -n grafana-alloy patch ds alloy-logs \
  --type=json \
  -p='[
    {
      "op":"replace",
      "path":"/spec/template/spec/containers/0/resources",
      "value":{
        "requests":{"cpu":"200m","memory":"256Mi"},
        "limits":{"cpu":"1000m","memory":"3Gi"}
      }
    }
  ]'

This is a mitigation, not a permanent capacity model. The goal is to stop the collector from starving the node while the root cause is investigated.

Restart Old Uncapped Pods

The patch affects only newly created pods. Existing pods may keep running without the CPU limit until replaced.

Check rollout state:

kubectl --context cluster-a -n grafana-alloy get ds alloy-logs -o wide

Find pods that have not picked up the limit:

kubectl --context cluster-a -n grafana-alloy get pods \
  -l app.kubernetes.io/name=alloy-logs \
  -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.spec.nodeName}{" cpuLimit="}{.spec.containers[0].resources.limits.cpu}{"\n"}{end}' \
  | sort

Delete hot uncapped pods so the DaemonSet recreates them with the new limit:

kubectl --context cluster-a -n grafana-alloy delete pod \
  alloy-logs-gm6sf \
  alloy-logs-95xtv \
  alloy-logs-jhc4w \
  alloy-logs-k2gmc

Then verify:

kubectl --context cluster-a -n grafana-alloy get ds alloy-logs -o wide
kubectl --context cluster-a top nodes | sort -k3 -nr | head
kubectl --context cluster-a -n grafana-alloy top pods | sort -k2 -hr | head

A successful mitigation looks like this:

DaemonSet: desired=16 current=16 ready=16 up-to-date=16 available=16
hot workers: 85-90% CPU -> 2-20% CPU
top Alloy pod: 6000-7000m CPU -> less than 200m CPU
ESXi hosts: 95-105% CPU -> about 45% CPU

The exact numbers do not matter. The shape does: capped Alloy pods should stop dominating node CPU.

Inspect Alloy Logs For Cause

After the immediate pressure is controlled, inspect recent Alloy logs:

kubectl --context cluster-a -n grafana-alloy logs ds/alloy-logs \
  -c alloy \
  --since=30m \
  | grep -Ei 'error|dropping|retry|batch|queue|backpressure|tailing|loki'

Two patterns are especially useful.

Missing Loki endpoints:

failed to list rules from loki
lookup loki-ruler.loki.svc.cluster.local: no such host
lookup loki-gateway.loki.svc.cluster.local: no such host
final error sending batch, no retries left, dropping data

Broad file tailing or backlog catch-up:

start tailing file path=/var/log/pods/<namespace>_<pod>_<uid>/<container>/0.log

Repeated missing-service errors do not prove they caused all CPU burn by themselves, but they are a strong configuration clue. Alloy should not repeatedly try to use a Loki service that does not exist in that environment.

Decide Ownership

The CPU cap protects the cluster. It does not decide the observability architecture.

If Loki is expected to exist soon, keep the cap and deploy the missing services. If Loki is not part of the environment yet, disable the Alloy components that depend on it until the backend exists.

Suggested handoff statement:

Alloy log collection was repeatedly trying to use Loki services that do not exist in this environment. Several uncapped Alloy DaemonSet pods consumed multiple CPU cores each and saturated worker nodes. A temporary 1 CPU limit reduced node and host pressure. Please either deploy the expected Loki services or remove the Loki rule/write components from this environment until Loki exists.

Watch For Falling Behind

A CPU limit can reveal under-capacity if log volume is genuinely high. Watch for:

kubectl --context cluster-a -n grafana-alloy top pods | sort -k2 -hr | head

kubectl --context cluster-a -n grafana-alloy logs ds/alloy-logs -c alloy --since=10m \
  | grep -Ei 'dropping|backpressure|retry|batch|queue|too old|error'

Warning signs:

many Alloy pods pinned near 1000m for long periods
dropping data
queue or batch retry errors
growing node disk usage under /var/log/pods
missing logs in the destination once Loki exists

If the cap is too low, raise it gradually, for example from 1000m to 1500m or 2000m. Do not remove the limit entirely unless the backend and log volume problem have been fixed.

Operating Rule

Observability agents need resource guardrails like any other workload.

For Alloy log collection on Kubernetes:

1. Correlate vSphere host CPU to Kubernetes nodes.
2. Correlate hot nodes to Alloy pods.
3. Save the DaemonSet as rollback evidence.
4. Add a conservative CPU limit.
5. Restart old uncapped pods so the limit actually applies.
6. Verify node, pod, and host CPU.
7. Inspect Alloy logs for missing backends, retry loops, and broad tailing.
8. Fix the backend/configuration owner, then tune the limit intentionally.

For broader Loki incident-readiness design, see Loki Tempo And Incident Correlation Paths. For general metrics-system safety boundaries, see Prometheus Operations And Query Safety.