⎈ k8s knowledge compiler

Horizontal Pod Autoscaling [page]deterministic

concepts

In Kubernetes, a _HorizontalPodAutoscaler_ automatically updates a workload resource (such as a [Deployment](#gloss:deployment) or [StatefulSet](#gloss:statefulset)), with the aim of automatically scaling capacity to match demand.

Horizontal scaling means that the response to increased load is to deploy more [Pods](#gloss:pod). This is different from _vertical_ scaling, which for Kubernetes would mean assigning more resources (for example: memory or CPU) to the Pods that are already running for the workload.

If the load decreases, and the number of Pods is above the configured minimum, the HorizontalPodAutoscaler instructs the workload resource (the Deployment, StatefulSet, or other similar resource) to scale back down.

Horizontal pod autoscaling does not apply to objects that can't be scaled (for example: a [DaemonSet](#gloss:daemonset).)

The HorizontalPodAutoscaler is implemented as a Kubernetes API resource and a [controller](#gloss:controller). The resource determines the behavior of the controller. The horizontal pod autoscaling controller, running within the Kubernetes [control plane](#gloss:control-plane), periodically adjusts the desired scale of its target (for example, a Deployment) to match observed metrics such as average CPU utilization, average memory utilization, or any other custom metric you specify.

There is [walkthrough example](/docs/tasks/run-application/horizontal-pod-autoscale-walkthrough/) of using horizontal pod autoscaling.

## How does a HorizontalPodAutoscaler work?

```mermaid graph BT

hpa[HorizontalPodAutoscaler] --> scale[Scale]

subgraph rc[Deployment] scale end

scale -.-> pod1[Pod 1] scale -.-> pod2[Pod 2] scale -.-> pod3[Pod N]

classDef hpa fill:#D5A6BD,stroke:#1E1E1D,stroke-width:1px,color:#1E1E1D; classDef rc fill:#F9CB9C,stroke:#1E1E1D,stroke-width:1px,color:#1E1E1D; classDef scale fill:#B6D7A8,stroke:#1E1E1D,stroke-width:1px,color:#1E1E1D; classDef pod fill:#9FC5E8,stroke:#1E1E1D,stroke-width:1px,color:#1E1E1D; class hpa hpa; class rc rc; class scale scale; class pod1,pod2,pod3 pod ```

Figure 1. HorizontalPodAutoscaler controls the scale of a Deployment and its ReplicaSet

Kubernetes implements horizontal pod autoscaling as a control loop that runs intermittently (it is not a continuous process). The interval is set by the `--horizontal-pod-autoscaler-sync-period` parameter to the [`kube-controller-manager`](/docs/reference/command-line-tools-reference/kube-controller-manager/) (and the default interval is 15 seconds).

Once during each period, the controller manager queries the resource utilization against the metrics specified in each HorizontalPodAutoscaler definition. The controller manager finds the target resource defined by the `scaleTargetRef`, then selects the pods based on the target resource's `.spec.selector` labels, and obtains the metrics from either the resource metrics API (for per-pod resource metrics), or the custom metrics API (for all other metrics).

  • For per-pod resource metrics (like CPU), the controller fetches the metrics from the resource metrics API for each Pod targeted by the HorizontalPodAutoscaler. Then, if a target utilization value is set, the controller calculates the utilization value as a percentage of the equivalent [resource request](/docs/concepts/configuration/manage-resources-containers/#requests-and-limits) on the containers in each Pod. If a target raw value is set, the raw metric values are used directly. The controller then takes the mean of the utilization or the raw value (depending on the type of target specified) across all targeted Pods, and produces a ratio used to scale the number of desired replicas.

Please note that if some of the Pod's containers do not have the relevant resource request set, CPU utilization for the Pod will not be defined and the autoscaler will not take any action for that metric. See the [algorithm details](#algorithm-details) section below for more information about how …(trimmed)

Sources

concepts/workloads/autoscaling/horizontal-pod-autoscale.md · docHorizontal Pod Autoscaling

Related (25)

references DeploymentDeployment conf=1
references StatefulSetStatefulSet conf=1
references PodPods conf=1
references DaemonSetDaemonSet conf=1
references Controllercontroller conf=1
references Control Planecontrol plane conf=1
references Aggregation Layeraggregated APIs conf=1
references Manifestmanifest(s) conf=1
part_of How does a HorizontalPodAutoscaler work?describes conf=1
part_of Pod readiness and autoscaling metricsdescribes conf=1
part_of API Objectdescribes conf=1
part_of Stability of workload scale {#flapping}describes conf=1
part_of Autoscaling during rolling updatedescribes conf=1
part_of Support for resource metricsdescribes conf=1
part_of Scaling on custom metricsdescribes conf=1
part_of Scaling on multiple metricsdescribes conf=1
part_of Support for metrics APIsdescribes conf=1
part_of Configurable scaling behaviordescribes conf=1
part_of Implicit maintenance-mode deactivationdescribes conf=1
part_of {{% heading "whatsnext" %}}describes conf=1
part_of Algorithm detailsdescribes conf=1
part_of Container resource metricsdescribes conf=1

← all Docs