What Is the Kubernetes Cluster Autoscaler?
The Kubernetes Cluster Autoscaler is an open-source component designed to automatically adjust the size of a Kubernetes cluster based on the current workload demands. It works by monitoring pods that fail to schedule due to insufficient resources and then scales up the cluster by adding more nodes when necessary. Conversely, it identifies underutilized nodes and scales down the cluster by removing nodes that are not needed, ensuring optimal resource usage and cost efficiency.
This dynamic scaling helps maintain high availability and performance without manual intervention. The Cluster Autoscaler supports multiple cloud providers, including AWS, Azure, and Google Cloud, integrating seamlessly with their managed Kubernetes offerings. It uses provider-specific APIs to provision and decommission nodes as required, allowing organizations to respond to changing traffic patterns in real time.
This is part of a series of articles about Kubernetes autoscaling
In this article:
- What Are Helm Charts?
- Tutorial: Getting Started with the Cluster Autoscaler with Helm Chart
- Cluster Autoscaler Helm Chart Best Practices
What Are Helm Charts?
Helm charts are standardized packages for Kubernetes applications, containing all the necessary YAML manifests and templates to define, configure, and deploy complex workloads. They simplify the deployment process by bundling application resources, dependencies, and configuration values into a single, reusable artifact. Helm charts can be:
- Versioned
- Shared
- Stored in repositories
Using Helm charts, teams can automate the installation, upgrade, and rollback of applications, reducing the risk of human error and ensuring consistency across environments. Helm’s templating engine allows for customizable deployments, enabling users to override default values through a values file or command-line arguments.
Quick Start: Running Cluster Autoscaler with Helm Chart
Cluster Autoscaler Helm Chart Prerequisites
Before installing the Cluster Autoscaler chart, ensure you are using Helm 3+ and Kubernetes 1.35.x or later. Azure AKS requires Kubernetes 1.10 or later with RBAC enabled.
The Cluster Autoscaler internally simulates Kubernetes scheduler behavior. Although other Kubernetes versions might work by overriding the container image, mismatched versions can introduce subtle scheduling problems. The current chart uses Helm chart API version v2, so Helm versions earlier than 3 are not supported.
If you are migrating from a 1.X version of cluster-autoscaler-chart, uninstall that release before installing cluster-autoscaler version 9.0.0 or later. The older 1.X chart releases are deprecated.
Step 1: Add the Cluster Autoscaler Helm Repository
The Cluster Autoscaler chart is installed from the autoscaler/cluster-autoscaler chart path. Before installation, make sure the repository containing this chart is configured in your Helm environment.
The chart does not create a functional autoscaling deployment with its defaults alone. During installation, you must configure either node-group auto-discovery or static node groups. Using both approaches together is not recommended.
With auto-discovery, set autoDiscovery.clusterName and any required provider-specific values. For static configuration, define one or more groups through autoscalingGroups or autoscalingGroupsnamePrefix.
Step 2: How to Install Cluster Autoscaler with Helm
Install the chart with helm install, supplying the configuration required by your cloud provider. For configurations stored in a values file, use:
helm install my-release autoscaler/cluster-autoscaler -f myvalues.yaml
Your values must identify the node groups that the autoscaler can manage. You can configure auto-discovery with autoDiscovery.clusterName, or explicitly define groups and their minimum and maximum sizes.
After installation, check the autoscaler logs to confirm that its main loop is running. If it is not, inspect the generated pod and verify the arguments passed to the cluster-autoscaler command.
Step 3: Install Cluster Autoscaler on AWS Using Helm
On AWS, the Cluster Autoscaler can automatically discover Auto Scaling Groups (ASGs). Tag each managed ASG with the keys k8s.io/cluster-autoscaler/enabled and k8s.io/cluster-autoscaler/<CLUSTER NAME>. Only the tag keys are significant.
Then install the chart with the cluster name and AWS region:
helm install my-release autoscaler/cluster-autoscaler \
--set autoDiscovery.clusterName=<CLUSTER NAME> \
--set awsRegion=<YOUR AWS REGION>
You can also provide awsAccessKeyID and awsSecretAccessKey during installation. For Amazon EKS, an alternative is to associate the autoscaler's service account with an IAM role and pass its ARN through the service account annotation.
If you do not want auto-discovery, specify ASGs manually:
helm install my-release autoscaler/cluster-autoscaler \
--set "autoscalingGroups[0].name=your-asg-name" \
--set "autoscalingGroups[0].maxSize=10" \
--set "autoscalingGroups[0].minSize=1"
The worker running the autoscaler must have the required IAM permissions to inspect and modify the relevant AWS Auto Scaling resources.
Related content: Read our comparison of Karpenter vs. Cluster Autoscaler
Step 4: Upgrade the Cluster Autoscaler Helm Chart
When upgrading older installations, account for chart-version changes. Releases from version 9.0.0 onward use the cluster-autoscaler chart name. To move from a deprecated 1.X cluster-autoscaler-chart release, first uninstall the existing 1.X installation and then install version 9.0.0 or later.
Version 9.1.0 also changes the meaning of envFromConfigMap. It must contain the name of a ConfigMap referenced by envFrom. Configurations that depend on the earlier envFromConfigMap behavior should rename that setting to extraEnvConfigMaps.
For reference, a release can be removed with:
helm uninstall my-release
This deletes the Kubernetes components associated with that Helm release. Before changing versions, review configuration values that may have changed between the installed and target chart versions.
Cluster Autoscaler Helm Chart Best Practices
Here are some important practices to keep in mind when working with the Cluster Autoscaler for Helm charts.
1. Right-Size CPU and Memory Requests
Right-sizing CPU and memory requests for your workloads is essential when using the Cluster Autoscaler. If requests are too high, the scheduler may over-provision nodes, leading to wasted resources and increased costs. Conversely, setting requests too low risks resource contention and potential application instability.
Key actions:
- Analyze historical resource usage and set requests and limits based on actual needs rather than defaults or guesses.
- Regularly review and adjust these settings as workloads evolve.
- Use Kubernetes monitoring tools to observe real-time resource consumption and identify opportunities for optimization.
Accurate resource requests improve autoscaler efficiency, ensuring new nodes are only added when genuinely necessary and that existing resources are utilized effectively.
2. Treat Workload Rightsizing and Node Autoscaling as One System
Workload rightsizing and node autoscaling should not be managed in isolation. Changes in application resource requests directly impact how the autoscaler scales the cluster. If workloads are consistently over-provisioned, the autoscaler will unnecessarily add nodes, while under-provisioning can lead to unschedulable pods and degraded performance. Coordinating these processes helps maintain balance between resource availability and cost.
Key actions:
- Implement automated tools or policies that adjust both workload requests and node group sizes in tandem.
- Ensure regular communication between development and operations teams to ensure that scaling decisions reflect real-world application needs.
By treating rightsizing and autoscaling as interconnected, you can achieve better resource utilization, application stability, and cost efficiency.
Related content: Read our guide to the Kubernetes Vertical Pod Autoscaler
3. Review Resource Usage Continuously
Continuous monitoring of resource usage is vital for maintaining an efficient autoscaling setup. Static resource allocations quickly become outdated as application demands change, leading to inefficiencies or performance issues. Regularly analyzing usage data allows you to make informed adjustments to workload requests and autoscaler settings.
Key actions:
- Use Kubernetes dashboards and monitoring platforms to track CPU, memory, and node utilization over time.
- Establish a process for periodic review, such as monthly audits or automated alerts for anomalous usage patterns.
- Quickly respond to changes in resource consumption to prevent over-provisioning and ensure that the cluster scales appropriately.
Ongoing review supports a proactive approach, reducing operational costs and maintaining a responsive environment.
4. Design Node Groups for Efficient Bin Packing
Efficient bin packing refers to fitting workloads onto nodes in a way that maximizes resource utilization and minimizes waste. When designing node groups, it’s important to select instance types and sizes that align with your workload profiles.
Key actions:
- Avoid overly large or small nodes that lead to resource fragmentation or underutilization.
- Group similar workloads together to promote more predictable usage patterns and simplify autoscaler configuration.
- Adjust node group configurations as your applications evolve.
- Monitor how workloads are distributed and identify opportunities to consolidate or split node groups for better packing.
Efficient bin packing reduces the total number of nodes required, optimizes autoscaler performance, and lowers infrastructure costs by making the most of available resources.
5. Make Sure HPA and Cluster Autoscaler Work Together
The Horizontal Pod Autoscaler (HPA) and Cluster Autoscaler must be configured to work in harmony to ensure optimal scaling. The HPA adjusts the number of pods based on workload metrics, while the Cluster Autoscaler manages the underlying nodes. If the HPA scales pods beyond the available node capacity, the Cluster Autoscaler must be able to add nodes quickly enough to accommodate them. Misaligned configurations can result in failed pod scheduling or unnecessary resource allocation.
Key actions:
- Coordinate scaling policies and thresholds between the HPA and Cluster Autoscaler.
- Monitor the interaction between the two, using metrics and logs to identify bottlenecks or delays in scaling actions.
- Adjust HPA targets and Cluster Autoscaler parameters to achieve balanced, responsive scaling.
Ensuring both components work together reduces downtime, improves application availability, and supports efficient resource usage.
Maximize Cluster Autoscaler Efficiency with PerfectScale
Node autoscalers like Cluster Autoscaler and Karpenter scale nodes up and down based on demand, improving cluster availability and reducing idle costs. But even a well-configured autoscaler runs into limits when the workloads underneath it are misconfigured: over-provisioned containers waste capacity and force the autoscaler to add more nodes than necessary, under-provisioned containers cause pod evictions and node pressure, and inefficient bin-packing leaves resources underutilized while triggering unnecessary scaling events. PerfectScale addresses these problems at the workload and node level so your autoscaling configuration actually delivers the efficiency it promises.
Key capabilities of PerfectScale for Kubernetes autoscaling:
- Proactive configuration error detection: Continuously analyzes workloads to identify and address configuration errors such as CPU Request Not Set, Memory Request Not Set, and Memory Limit Not Set, preventing unpredictable evictions, node over-commitment, and inefficient autoscaling.
- Autonomous workload right-sizing: Instantly adjusts workload resources based on actual utilization and improves pod bin-packing to boost autoscaling efficiency, without manual interaction.
- Autoscaling group and node pool visibility: Provides visibility into the costs and utilization of your autoscaling groups and node pools so you can pinpoint inefficiencies across the cluster.
- Optimal node type selection: Helps you select node types that improve utilization, achieve precise resource distribution, and maximize cost efficiency.
- Broad autoscaler support: Works alongside Karpenter, Cluster Autoscaler, and Node Auto Provisioning, combining node-level optimization with pod-level rightsizing for better bin packing and faster scaling decisions.
Learn more about maximizing your Kubernetes node autoscaling with PerfectScale