/writing/a-nodeselector-is-a-capacity-request

A nodeSelector Is a Capacity Request

Somewhere in a cluster I look after, an autoscaler had provisioned a full compute node to run two small controller pods. Nothing else would ever schedule onto it. Nothing else could.

The cause was two lines of YAML that had been copied, correctly, from a cluster where they cost nothing.

Two lines, copied correctly

We run the same ingress controller on four clusters. Two of them run a meaningful amount of arm64 workload and keep arm64 capacity around permanently. On those two, someone had pinned the controller to arm64 — a nodeSelector on the architecture label, plus the matching toleration for the arm64 taint. Sensible: the capacity is there, it’s cheaper per unit, the pods fit in the gaps.

When the controller was rolled out to the other two clusters, the values file was copied. Of course it was. That’s how you get consistency, and the diff between the four files looked right in review precisely because the lines matched.

But those two clusters have no arm64 workloads. The node pool exists — the platform configuration is uniform — and until that moment nothing had ever asked it for anything.

What a scheduling constraint means to an autoscaler

Here’s the part that’s easy to say and easy to forget under review.

To the Kubernetes scheduler, a nodeSelector is a filter: it narrows the set of existing nodes a pod may land on. If none match, the pod stays Pending. That’s the mental model most people carry, and in a fixed-size cluster it’s the whole story — you notice immediately, because the pod doesn’t run.

To a provisioning autoscaler, the same nodeSelector is an order. A pending pod with unsatisfiable constraints is not an error condition; it’s a work item. The autoscaler reads the constraint, computes the cheapest instance that satisfies it, and buys one.

So the failure mode inverts. In a fixed cluster, an over-constrained pod is loudly broken. In an autoscaled cluster, it is silently expensive. Everything comes up green. The deployment reports Available, the controller serves traffic, no alert fires. The only evidence is a node in the list with one workload on it and an instance-hours line on the bill.

In our case one cluster got a general-purpose arm64 node and the other got a memory-optimized one — roughly four vCPU and thirty-odd gigabytes of RAM — to host two pods whose combined requests were a rounding error against that.

Nothing needed it

The controller image publishes both linux/amd64 and linux/arm64 in its manifest list. There is no native extension, no compiled sidecar, no architecture-specific behaviour. The pin wasn’t buying compatibility or performance; it was an artifact of where the manifest was first written.

Deleting the two lines rescheduled both replicas onto existing nodes, and the autoscaler deprovisioned the now-empty node in each cluster. That’s the whole fix. No image change, no version bump, no migration.

Why review didn’t catch it

This is the bit worth generalizing, because the review was not lazy.

The correctness of a scheduling constraint is a property of the cluster, not of the manifest. You cannot evaluate nodeSelector: {kubernetes.io/arch: arm64} by reading it. The same characters are free on one cluster and cost you a node on another, and the file gives you no way to tell which. Review compares the change against its neighbours, the neighbours match, and matching is exactly the wrong signal here.

The related conflation, which I’ve now watched several people make including myself:

  • A toleration is permission. It says “I’m allowed onto nodes with this taint.” Adding one to a pod that never encounters the taint changes nothing and costs nothing.
  • A nodeSelector (or a required node affinity) is demand. It says “I will not run anywhere else.” Adding one to a cluster that can’t satisfy it obliges someone to go get capacity.

They’re usually written adjacent to each other, in the same block, by the same person, in a single copied hunk. One of them is inert and one of them spends money.

The same asymmetry applies to anything else that narrows placement: zone pins, dedicated node-pool labels, GPU or accelerator selectors, instance-family requirements, and topologySpreadConstraints with whenUnsatisfiable: DoNotSchedule — which can force a new node purely to satisfy a spread rule.

The question to ask

When copying a workload’s manifest to another cluster, go through every placement constraint and ask one thing:

Does this cluster already run something that satisfies this? If not, I am not writing a preference — I am placing an order.

If the answer is no and the workload doesn’t actually require the constraint, delete it. If the workload does require it, that’s fine, but now you know you’re adding a node and you can size the decision honestly.

And after you remove one: check that the node actually goes away. Consolidation isn’t instant, and anything else that drifted onto that node — or any pod on it carrying a “don’t disrupt me” annotation — will keep it alive and keep you paying for a fix you think you already landed.

Takeaways

  • In an autoscaled cluster, a scheduling constraint is a purchase order. The scheduler filters; the provisioner procures. Same YAML, opposite consequences.
  • Over-constrained pods fail loudly on fixed clusters and quietly on elastic ones. If your intuition was formed on fixed clusters, it is now inverted and nothing will tell you.
  • Tolerations are free, selectors are not. They travel together in copied blocks. Treat them differently.
  • A constraint can’t be reviewed in isolation from the cluster it lands on. “It matches the other environments” is the specific reasoning that produces this bug.
  • Copying manifests between clusters needs a per-constraint pass, not a per-file comparison — the file being identical is the thing that hides it.
  • Verify the node disappears after you unpin. Otherwise you’ve paid for the fix twice: once in the incident and once in the node that never left.