Removing idle capacity is an attractive platform goal. The more important test is what happens when work arrives after that capacity has disappeared.
What changed
Kubernetes 1.37, released on August 26, 2026, moves HorizontalPodAutoscaler scale-to-zero support to Beta and enables it by default. The release announcement specifies object or external metrics for this behavior; CPU and memory measurements alone cannot reactivate a workload without running Pods. [1]
A September 2 technical article explains the operational details. A suitable independent signal, such as queue depth, lets the controller detect returning demand. The post also distinguishes HPA-managed zero replicas from an operator’s manual pause through the ScaledToZero condition. Manually setting a workload to zero is not a substitute for exercising the HPA’s own scale-down path. [2]
This is a Beta capability, not a claim that every managed cluster or application is ready to use it. Confirm your platform’s version and configuration before planning a rollout.
Evaluate the waiting work
My recommendation is to start by describing how much delay a workload can tolerate. A background conversion job and an interactive customer request can have very different expectations even when their containers use similar resources.
Consider an illustrative internal report-generation service. A durable queue holds requests, and users expect the output after a short wait. That may be a useful evaluation candidate. Write down the acceptable completion time and backlog limit before deciding whether zero idle workers is desirable.
The Kubernetes guidance notes that Services do not buffer requests while no Pods are ready. Request-driven systems need a separate buffering mechanism, and cold-start delay remains part of the trade-off. [2]
Test recovery as a user would experience it
A proposed staging exercise should cover more than a replica-count chart:
- Let the controller reduce an idle test workload to zero.
- Submit representative work and measure time to the first completed result.
- Increase the batch size and check completion times across the whole batch.
- Exercise your application’s retry and duplicate-handling behavior.
- Rehearse the operator action that restores service if automatic recovery fails.
The article’s upgrade guidance also matters: both relevant control-plane components must support and enable the feature. Before a downgrade or disabling it, restore a positive minimum replica count and bring zero-replica workloads back up. [2]
Ask where the saving appears
For the business case, I would compare a measured baseline with the trial rather than translating fewer Pods directly into an assumed bill reduction. Include any extra monitoring, buffering, and operational work in the evaluation. Record what actually changes in the environment’s billable resource usage.
Track completion time, backlog age, recovery success, and operator intervention alongside resource consumption. An improvement is convincing when the service continues to meet its delivery expectations and the operational benefit is visible.
For platform teams, this is an opportunity to offer a tested operating pattern with a clear recovery procedure. Application teams then have a concrete choice based on their workload’s tolerance for waiting.
Discussion question: Which background workload could tolerate a cold start—and how would you prove it still meets its service target?
Sources
- Kubernetes: v1.37 release announcement, August 26, 2026.
- Kubernetes: Scale Workloads to Zero with HorizontalPodAutoscaler, September 2, 2026.
The report-generation scenario and evaluation plan are illustrative recommendations. No production results or cost savings are claimed.

