Running an EKS cluster is straightforward when every application has predictable resource requirements. The difficulty begins when a deployment suddenly needs more capacity, a batch job requests a different instance shape, or quiet periods leave you paying for mostly empty nodes.

In our Amazon EKS Cluster Autoscaler guide, we explored scaling existing node groups. Karpenter offers another approach: provision EC2 instances around the requirements of pending workloads.

This tutorial builds a separate demo cluster, configures Karpenter, and proves that application demand can create and remove worker nodes. We will start with On-Demand capacity, then enable Spot with On-Demand fallback.

Scope: This is a self-managed Karpenter walkthrough using the stable v1 resource APIs. It is a documentation-based example, not a claim of a completed live AWS test. Review the configuration in a sandbox account before adapting it to production. EKS, EC2, storage, and networking resources incur charges until removed.

What You Will Learn

  • How Karpenter differs from Cluster Autoscaler
  • How to create an EKS cluster with Karpenter using eksctl
  • How NodePool, EC2NodeClass, and NodeClaim work together
  • How to configure networking and pin an EKS-optimized AMI
  • How to test provisioning and consolidation
  • How to enable Spot capacity without requiring every workload to use Spot
  • How to troubleshoot failed provisioning and plan a migration

Karpenter Compared with Cluster Autoscaler

Both tools help when Kubernetes cannot place Pods on existing capacity. Their infrastructure models differ.

Area Cluster Autoscaler on AWS Karpenter on AWS
Capacity model Changes the size of existing Auto Scaling Groups Provisions EC2 capacity under NodePool constraints
Configuration Node groups define available instance options NodePool requirements define eligible capacity
Workload fit Depends on the shapes configured in node groups Evaluates eligible instance offerings for workloads
Operational choice Useful when node groups suit your needs Useful when workloads need more flexible capacity

Karpenter is not an application autoscaler. An HPA or KEDA can create more application replicas; Karpenter can provide the nodes those replicas need. Raising actual CPU usage alone does not necessarily cause node provisioning.

See AWS’s Karpenter best practices for deployment guidance.

Architecture Used in This Guide

We will keep two small managed nodes for system capacity and let Karpenter provision application nodes separately.

Component Responsibility
EKS control plane Runs the Kubernetes API and scheduler
Managed node group named system Provides initial capacity for the controller and system services
Karpenter controller Evaluates provisioning and disruption decisions
EC2NodeClass named applications Defines AWS networking, AMI, and node identity
NodePool named applications Defines allowed application capacity and disruption settings
Demo Deployment Creates schedulable demand specifically for this NodePool

Keep the controller on managed capacity so it remains available when application capacity reaches zero. AWS recommends using a node group or Fargate for the controller rather than nodes it manages itself.

Prerequisites and Version Selection

Use Bash on Linux, macOS, or Windows through WSL. Install AWS CLI v2, eksctl, kubectl, and Helm from their official distributions. Use an AWS identity authorized to create the required EKS, VPC, EC2, IAM, and CloudFormation resources.

1
2
3
4
5
aws sts get-caller-identity
aws --version
eksctl version
kubectl version --client
helm version

The examples pin Karpenter 1.12.1 and Kubernetes 1.35, a pair reflected in the versioned getting-started documentation. This is an example baseline, not a recommendation to ignore newer security fixes. Check the compatibility matrix and regional EKS availability before creating a cluster.

1
2
3
4
5
6
7
8
export AWS_DEFAULT_REGION="eu-west-2"
export CLUSTER_NAME="codingtricks-karpenter-demo"
export K8S_VERSION="1.35"
export KARPENTER_VERSION="1.12.1"

aws eks describe-cluster-versions \
--query 'clusterVersions[].clusterVersion' \
--output table

Update both version choices deliberately if necessary. Keep the same shell open throughout the walkthrough.

Step 1: Create the Demo Cluster and Install Karpenter

For a new cluster, eksctl supports installing Karpenter and creating its prerequisites. We use that route to avoid maintaining a second hand-written IAM bootstrap alongside the installation.

Create cluster.yaml:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
cat > cluster.yaml <<EOF
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
name: ${CLUSTER_NAME}
region: ${AWS_DEFAULT_REGION}
version: "${K8S_VERSION}"
tags:
karpenter.sh/discovery: ${CLUSTER_NAME}
iam:
withOIDC: true
karpenter:
version: "${KARPENTER_VERSION}"
createServiceAccount: true
withSpotInterruptionQueue: true
managedNodeGroups:
- name: system
instanceType: m5.large
amiFamily: AmazonLinux2023
privateNetworking: true
desiredCapacity: 2
minSize: 2
maxSize: 2
EOF

eksctl create cluster -f cluster.yaml

This command creates a new cluster. It is not an installation command for an existing cluster. The installation route is documented in eksctl Karpenter support.

Inspect the result:

1
2
3
4
5
6
aws eks update-kubeconfig --name "$CLUSTER_NAME"
kubectl config current-context
kubectl get nodes
helm list --all-namespaces
kubectl get deployments --all-namespaces \
-l app.kubernetes.io/name=karpenter

Determine the installed controller namespace rather than assuming it:

1
2
3
4
5
6
export KARPENTER_NAMESPACE="$(kubectl get deployments -A \
-l app.kubernetes.io/name=karpenter \
-o jsonpath='{.items[0].metadata.namespace}')"

test -n "$KARPENTER_NAMESPACE"
kubectl -n "$KARPENTER_NAMESPACE" rollout status deployment/karpenter

Stop if the installation failed or the namespace is empty. Do not proceed simply because worker nodes exist.

Step 2: Keep the Controller on Managed Capacity

For this fresh installation, update the installed Helm release to select the system node group:

1
2
3
4
5
6
7
8
9
helm upgrade karpenter oci://public.ecr.aws/karpenter/karpenter \
--namespace "$KARPENTER_NAMESPACE" \
--version "$KARPENTER_VERSION" \
--reuse-values \
--set-string 'nodeSelector.eks\.amazonaws\.com/nodegroup=system' \
--wait

kubectl -n "$KARPENTER_NAMESPACE" get pods \
-l app.kubernetes.io/name=karpenter -o wide

Confirm the release name in the earlier Helm output; substitute it if necessary. The managed nodes must have enough allocatable resources for the controller and other system Pods.

For production, capture the installation values in your infrastructure or GitOps repository. Treat this imperative upgrade as a lab convenience.

Step 3: Verify Networking and the Node Instance Profile

Karpenter must select subnets and security groups that let new nodes reach the EKS API and required AWS services.

1
2
3
4
5
6
7
8
9
10
11
12
export VPC_ID="$(aws eks describe-cluster --name "$CLUSTER_NAME" \
--query 'cluster.resourcesVpcConfig.vpcId' --output text)"

aws ec2 describe-subnets \
--filters "Name=vpc-id,Values=$VPC_ID" \
--query 'Subnets[].{ID:SubnetId,AZ:AvailabilityZone,PublicIP:MapPublicIpOnLaunch,Tags:Tags}' \
--output json

aws ec2 describe-security-groups \
--filters "Name=vpc-id,Values=$VPC_ID" \
--query 'SecurityGroups[].{ID:GroupId,Name:GroupName,Tags:Tags}' \
--output json

Review subnet route tables to identify the intended private worker subnets. If their discovery tag is missing, apply it only to those reviewed subnets:

1
2
3
4
# Replace these values with the private worker subnet IDs you reviewed.
aws ec2 create-tags \
--resources subnet-REPLACE_FIRST subnet-REPLACE_SECOND \
--tags "Key=karpenter.sh/discovery,Value=$CLUSTER_NAME"

The eksctl integration automatically tags its shared node security group when configured with the discovery tag. Check that the selected group has the required rules; tagging does not create connectivity.

Verify the instance profile created by this installation route:

1
2
export NODE_INSTANCE_PROFILE="eksctl-KarpenterNodeInstanceProfile-${CLUSTER_NAME}"
aws iam get-instance-profile --instance-profile-name "$NODE_INSTANCE_PROFILE"

We will reference that existing profile. Do not substitute the controller’s IAM role: controller permissions and worker-node identity serve different purposes.

Step 4: Select an EKS-Optimized AMI

Resolve the recommended AL2023 image for this Kubernetes version and region, then save its exact ID in the manifest:

1
2
3
4
5
6
7
export AMI_ID="$(aws ssm get-parameter \
--name "/aws/service/eks/optimized-ami/${K8S_VERSION}/amazon-linux-2023/x86_64/standard/recommended/image_id" \
--query 'Parameter.Value' --output text)"

aws ec2 describe-images --image-ids "$AMI_ID" \
--query 'Images[].{ID:ImageId,Name:Name,Architecture:Architecture}' \
--output table

Resolving a recommended image does not establish that it works with your application. Test it first. Pinning the resulting ID makes subsequent provisioning reproducible within this region; changing Kubernetes versions requires reviewing the image selection again.

See AWS’s recommended AMI lookup for the SSM parameter format and Managing AMIs for staged image updates.

Step 5: Create the EC2NodeClass

Create ec2nodeclass.yaml:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
cat > ec2nodeclass.yaml <<EOF
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
name: applications
spec:
amiFamily: AL2023
amiSelectorTerms:
- id: ${AMI_ID}
instanceProfile: ${NODE_INSTANCE_PROFILE}
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: ${CLUSTER_NAME}
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: ${CLUSTER_NAME}
tags:
Project: codingtricks-karpenter-demo
EOF
1
2
3
kubectl apply -f ec2nodeclass.yaml
kubectl wait --for=condition=Ready ec2nodeclass/applications --timeout=180s
kubectl describe ec2nodeclass applications

Check resolved AMIs, subnets, security groups, and readiness conditions. Fix any unresolved dependency before introducing workload demand. The EC2NodeClass reference explains these fields.

Step 6: Create an Application NodePool

Create nodepool.yaml:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: applications
spec:
template:
metadata:
labels:
workload.codingtricks.io/pool: applications
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: applications
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: kubernetes.io/os
operator: In
values: ["linux"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["4"]
expireAfter: 720h
terminationGracePeriod: 1h
limits:
cpu: "32"
memory: 128Gi
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 2m
budgets:
- nodes: "1"
1
2
3
kubectl apply -f nodepool.yaml
kubectl wait --for=condition=Ready nodepool/applications --timeout=180s
kubectl describe nodepool applications

This pool allows several compute, general-purpose, and memory-optimized families. Its aggregate capacity limits are deliberately small for the lab, but they are not a spending cap. Parallel provisioning can temporarily exceed resource limits.

The label lets our demo request this capacity explicitly. The one-node disruption budget limits applicable voluntary disruptions; it does not promise protection from EC2 failures, Spot interruption, or expiration.

See the NodePool reference.

Step 7: Prove That Workloads Create Nodes

Create scaling-demo.yaml:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
apiVersion: apps/v1
kind: Deployment
metadata:
name: karpenter-demo
spec:
replicas: 0
selector:
matchLabels:
app: karpenter-demo
template:
metadata:
labels:
app: karpenter-demo
spec:
nodeSelector:
workload.codingtricks.io/pool: applications
containers:
- name: capacity-demo
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: "1"
memory: 256Mi
limits:
cpu: "1"
memory: 256Mi
1
2
3
kubectl apply -f scaling-demo.yaml
kubectl scale deployment karpenter-demo --replicas=6
kubectl get pods -l app=karpenter-demo -w

In another terminal:

1
2
3
kubectl get nodeclaims
kubectl get nodes -L karpenter.sh/nodepool,karpenter.sh/capacity-type
kubectl get events --sort-by=.metadata.creationTimestamp

The managed nodes do not carry the requested workload label. The demo therefore requires application capacity even if the managed group has spare CPU.

Expect pending Pods, new NodeClaims, registered nodes, and eventually running Pods. The replica count does not dictate an equal number of nodes: several Pods may fit on one instance. This test creates resource requests; the pause container is not a CPU stress test.

Inspect a NodeClaim if provisioning stops:

1
kubectl describe nodeclaim NODECLAIM_NAME

NodeClaim conditions help distinguish launch, registration, and initialization problems.

Step 8: Observe Consolidation

Reduce demand:

1
2
kubectl scale deployment karpenter-demo --replicas=1
kubectl get nodes -l karpenter.sh/nodepool=applications -w

Then remove it:

1
kubectl scale deployment karpenter-demo --replicas=0

Karpenter evaluates whether application nodes can be removed or replaced. consolidateAfter: 2m introduces a delay after relevant Pod changes; it is not an exact termination timer.

Other workloads, scheduling constraints, and disruption controls can prevent consolidation. The managed node group remains because Karpenter does not own it.

Step 9: Enable Spot with On-Demand Fallback

Edit only the capacity requirement in nodepool.yaml:

1
2
3
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
1
2
3
kubectl apply -f nodepool.yaml
kubectl scale deployment karpenter-demo --replicas=6
kubectl get nodes -L karpenter.sh/capacity-type

Allowing both types gives Karpenter a Spot-first provisioning option with On-Demand fallback when eligible Spot capacity is unavailable. It does not guarantee a particular Spot percentage or instantly convert every existing node.

If the account has never used Spot, verify the service-linked role:

1
2
3
4
aws iam get-role --role-name AWSServiceRoleForEC2Spot

# Run only when the role is absent, rather than after any arbitrary error.
aws iam create-service-linked-role --aws-service-name spot.amazonaws.com

Use Spot for applications that tolerate interruption and spread replicas across failure domains. A workload requiring On-Demand should specify:

1
2
3
nodeSelector:
workload.codingtricks.io/pool: applications
karpenter.sh/capacity-type: on-demand

See AWS compute and autoscaling guidance and Karpenter scheduling.

Disruption Settings You Must Understand

Consolidation, drift, and expiration can all replace nodes for different reasons. NodePool budgets and application PodDisruptionBudgets operate at different levels.

Setting Meaning in this example
consolidationPolicy Allows evaluating empty and underutilized capacity
consolidateAfter Delays consolidation evaluation after Pod changes
budgets Restricts applicable voluntary node disruption
expireAfter Starts expiration handling after the configured lifetime
terminationGracePeriod Bounds node draining; forced cleanup can follow

The one-hour draining limit is a lab choice. Evaluate application shutdown, jobs, and storage before using it elsewhere. PDBs do not prevent every forced termination, and an interruption handler cannot manufacture more notice than EC2 provides.

Read the Disruption reference before changing these controls.

Troubleshooting

Start with evidence from all three resource layers:

1
2
3
4
5
6
kubectl describe pod POD_NAME
kubectl describe nodepool applications
kubectl describe ec2nodeclass applications
kubectl get nodeclaims
kubectl -n "$KARPENTER_NAMESPACE" logs deployment/karpenter \
-c controller --tail=100
Symptom What to investigate
Pending Pods, no NodeClaims Pool readiness, selectors, taints, requests, and pool limits
NodeClass not ready Discovery tags, AMI selection, instance profile, and IAM access
Instance launches but node never registers EKS node authentication, API connectivity, DNS, and bootstrap
No eligible instance offering Conflicting requirements, architecture, zones, and requested resources
Image pull failure after node registration Registry access and networking rather than provisioning
Idle nodes remain Other Pods, PDBs, disruption budgets, and consolidation events

Do not resolve a permission error by granting administrator access to the controller. Find the failing AWS action and compare the installed policy with the approved release’s requirements.

Use the official troubleshooting guide for deeper diagnosis.

Moving from Cluster Autoscaler

Use a staged migration for an existing cluster:

  1. Keep stable managed capacity for system services.
  2. Install Karpenter with the correct IAM, network, node-authentication, and interruption prerequisites.
  3. Create a narrowly scoped pool and direct a test workload to it.
  4. Validate provisioning, application behavior, and draining.
  5. Move additional workloads deliberately, then retire superseded autoscaling configuration.

Avoid competing provisioning decisions for the same pending workloads during the transition. The eksctl create cluster command above should never be pointed at an existing production cluster as a migration shortcut.

Follow Migrating from Cluster Autoscaler for the existing-cluster installation procedure.

Clean Up the Demo

Delete application demand first, then delete the demo pool while the controller is still running so it can process finalizers:

1
2
3
4
kubectl delete deployment karpenter-demo
kubectl delete nodepool applications
kubectl get nodeclaims
kubectl get nodes -l karpenter.sh/nodepool=applications

Wait until the corresponding NodeClaims and EC2 instances are gone. If removal stalls, inspect events and controller logs rather than deleting finalizers blindly.

1
2
kubectl delete ec2nodeclass applications
eksctl delete cluster --name "$CLUSTER_NAME" --region "$AWS_DEFAULT_REGION"

Review CloudFormation deletion results and check for remaining EC2 instances, volumes, load balancers, and NAT gateways associated with the demo. Removing local YAML files does not remove AWS resources.

Conclusion

In this article, we learned how to configure Karpenter in EKS. If you have any issue regarding this tutorial, mention your issue in the comment section or reach me through my E-mail.

Happy Coding