Showing posts with label AWS. Show all posts
Showing posts with label AWS. Show all posts

May 7, 2026

AWS Compute Services Explained: EC2 vs ECS vs EKS

Introduction

Choosing the right AWS compute service isn't just a technical decision - it directly affects your team's velocity, your operational costs, and how quickly your business can respond to change. Pick the wrong one and you'll either pay for complexity you don't need, or find yourself locked into a setup that can't scale with you.

EC2, ECS, and EKS are not competing services. They solve fundamentally different problems at different levels of abstraction. EC2 gives you a raw virtual machine with full infrastructure control. ECS is a managed container platform built natively on AWS - no Kubernetes required. EKS brings the full power of Kubernetes to AWS for teams that need portability and advanced orchestration at scale.

Here is what we cover:

  • What EC2, ECS, and EKS actually are - in plain terms
  • Key differences across abstraction, scaling, cost, and operational overhead
  • What each service costs, with real numbers from AWS's official pricing pages
  • Real-world use cases and business implications for each
  • A decision guide to help you pick the right one

Understanding the Basics

Amazon EC2 - Virtual Machines

EC2 is AWS's virtual machine service. You choose the operating system, CPU, memory, and storage - AWS provisions the server. Think of it as renting a physical server in the cloud. Your team owns everything from that point on: patching, scaling, monitoring, and security. Maximum control, maximum responsibility.

Business implication: EC2 requires engineering time to manage and maintain. That time has a cost. It is best suited for workloads where your team needs deep infrastructure control, or for migrating existing applications that weren't built for containers.

Amazon ECS - Managed Containers

ECS is AWS's native container orchestration service. You define your Docker image, CPU and memory requirements, and scaling rules. AWS handles the rest - scheduling, cluster management, health checks, and integrations with AWS services like Application Load Balancer, IAM, and CloudWatch. No Kubernetes knowledge required.

Business implication: ECS reduces the operational surface area significantly. Smaller teams can run production container workloads without dedicated DevOps headcount. It is AWS's recommended starting point for teams new to containers.

Amazon EKS - Managed Kubernetes

EKS runs Kubernetes on AWS. AWS manages the control plane - the brain of the cluster - while you manage the worker nodes, or offload them to AWS Fargate for a fully serverless setup. Kubernetes is the industry-standard container orchestration platform, and EKS makes it available as a managed service.

Business implication: EKS is the most powerful and flexible option, but it brings the highest operational complexity and cost. It pays off at scale, for teams running complex microservices architectures, or for organisations that want the option to run workloads across multiple cloud providers.

Key Differences at a Glance

Feature EC2 ECS EKS
Abstraction Level Low (Infrastructure) Medium (Platform) High (Ecosystem)
Primary Unit VM / Server Container Task Pod
Setup Complexity High Low Very High
Scaling Speed Slow (VM boot) Fast Fast
OS / Host Access Full None Limited
Kubernetes Support No No Yes
Multi-cloud Portability Low Low High
Operational Overhead High Medium Very High
Learning Curve Low–Medium Low High
Relative Cost (entry) Low–Medium Medium High
Best for Legacy apps, full control AWS-native container apps Complex microservices, multi-cloud

Cost Breakdown - With Real Numbers

Cost is one of the most common decision factors, and it is also one of the most misunderstood. The sticker price of compute is only part of the picture. Operational overhead - the engineering time spent managing infrastructure - is a real cost that doesn't show up on your AWS bill but does show up on your payroll.

EC2

With EC2, you pay for instance uptime. AWS offers four main pricing models: On-Demand (pay by the second, no commitment), Reserved Instances (commit to 1 or 3 years for up to 72% off On-Demand rates, per official AWS pricing), Savings Plans (flexible commitment-based discounts, AWS's currently recommended approach over Reserved Instances), and Spot Instances (spare capacity at up to 90% off, but AWS can reclaim with two minutes notice). Most production workloads that run around the clock should be on Reserved Instances or a Savings Plan - running purely On-Demand for steady-state workloads leaves significant money on the table.

Business implication: EC2 can be the most cost-efficient option for stable, predictable workloads - but only if your team actively manages instance sizing and pricing commitments. Teams that over-provision and forget tend to pay more than they should.

ECS

ECS has no additional management fee. You pay for the underlying compute - either EC2 instances you manage yourself, or AWS Fargate (fully serverless, where AWS manages the infrastructure). With Fargate on ECS, billing is per second based on the vCPU and memory your containers actually use. As a reference, in the US East (N. Virginia) region, Fargate charges approximately $0.04048 per vCPU-hour and $0.004445 per GB-hour for memory, per AWS's official Fargate pricing page. A container running 2 vCPUs and 4 GB of RAM costs roughly $0.099 per hour, or around $72 per month running continuously.

Business implication: Fargate on ECS is often the most cost-effective choice for bursty, unpredictable, or variable traffic patterns - you pay only for what you use, and there is no idle compute cost. For high-volume, steady-state workloads, EC2-backed ECS can be cheaper.

EKS

EKS has a fixed control plane fee of $0.10 per cluster per hour - approximately $73 per month per cluster - regardless of cluster size or workload, per AWS's official EKS pricing page. This is just the management fee; worker node compute (EC2 or Fargate), storage, load balancers, and data transfer are all billed separately. Teams running dev, staging, and production environments on separate clusters pay this fee three times minimum. Additionally, if a Kubernetes version ages past its 14-month standard support window without being upgraded, the fee jumps to $0.60 per hour - a 6x increase that catches many teams by surprise.

Business implication: EKS is expensive at small scale and cost-efficient at large scale, where advanced resource scheduling and bin-packing justify the overhead. For small to medium workloads, ECS on Fargate will almost always be the cheaper option. The $73/month control plane fee is a fixed entry cost per cluster - for multi-cluster strategies, it adds up quickly.

Real-World Use Cases

Use EC2 when:

  • You are migrating a legacy or monolithic application that was not built for containers
  • Your application requires custom OS-level configuration, kernel parameters, or specific hardware access
  • Your team needs full visibility and control over the underlying server environment
  • You have predictable, steady-state workloads and want to maximise savings through Reserved Instances or Savings Plans

Business implication: EC2 is the right lift-and-shift vehicle. It minimises application changes during migration and gives your team time to modernise at their own pace. The trade-off is that it demands the most ongoing engineering attention.

Use ECS when:

  • You want to run containers on AWS without learning Kubernetes
  • You need fast, reliable deployments with minimal DevOps overhead
  • Your workloads are bursty or variable and Fargate's pay-per-use model suits your traffic patterns
  • Your team is AWS-first and wants tight, native integration with IAM, CloudWatch, and ALB

Business implication: ECS on Fargate is the fastest path to production for container workloads. It requires the least infrastructure expertise, no cluster maintenance, and scales automatically. It is a strong default choice for startups, product teams, and organisations without dedicated platform engineering.

Use EKS when:

  • Your team already uses Kubernetes and has the expertise to operate it
  • You need multi-cloud portability - the ability to run the same workloads on AWS, GCP, or Azure
  • You are running complex microservices that benefit from Kubernetes-native features like Horizontal Pod Autoscaler, custom scheduling, or service meshes
  • You are operating at a scale where advanced resource optimisation (bin-packing, Spot node groups, Karpenter) delivers meaningful cost savings

Business implication: EKS is a long-term platform investment. It requires Kubernetes expertise to operate well - either in-house or through a managed services partner. The payoff is flexibility, portability, and the ability to run sophisticated workloads that outgrow what ECS can offer.

Quick Decision Guide

Two questions cut through most of the noise:

  1. Do you need Kubernetes?
    • Yes - your team uses it already, or you need multi-cloud portability → EKS
    • No → Move to question 2
  2. Do you want containers?
    • Yes → ECS
    • No, or you have a legacy app to migrate → EC2
Situation Recommended Service Why
Small team, new to containers ECS (Fargate) Lowest ops overhead, no cluster to manage
Legacy app migration EC2 Minimal app changes, full control
Startup scaling fast on AWS ECS (Fargate) Fast deployments, pay-as-you-go, AWS-native
Enterprise microservices EKS Advanced orchestration, multi-team platform
Existing Kubernetes users EKS Familiar tooling, avoid re-platforming cost
Multi-cloud strategy EKS Kubernetes runs anywhere - avoids AWS lock-in
Predictable, high-volume compute EC2 (Reserved / Savings Plan) Up to 72% savings over On-Demand with commitment

Conclusion

EC2, ECS, and EKS are tools for different jobs at different stages of growth. The right choice depends on your team's skills, your application's architecture, your traffic patterns, and your budget - both the AWS bill and the engineering time to manage it.

  • Want simplicity and speed? → ECS on Fargate
  • Have a legacy app to migrate? → EC2
  • Need Kubernetes power and portability? → EKS

ECS on Fargate handles a huge range of production workloads reliably and cost-effectively - and it is far easier to migrate to EKS later than to unwind unnecessary Kubernetes complexity from day one.

If you have any questions, you can reach out to our AWS Cloud Consulting team here.

Event-Driven Autoscaling on EKS with KEDA and Karpenter

Introduction

It's 2 a.m. Your e-commerce platform just hit the front page of a major news site. Your SQS order queue skyrockets from 200 to 200,000 messages in minutes. Pods are overwhelmed. Customers see errors. Your on-call engineer is manually scaling deployments.

This is exactly the failure mode that event-driven scaling prevents. By combining KEDA (Kubernetes Event-Driven Autoscaler) and Karpenter, you can build a platform that reacts to demand automatically - scaling pods and nodes in seconds, then returning to zero when load disappears - all without human intervention.

Why Kubernetes HPA Falls Short for Event-Driven Workloads

Kubernetes' built-in Horizontal Pod Autoscaler (HPA) is often configured with CPU and memory metrics. While HPA does technically support custom and external metrics through the Kubernetes metrics API, wiring it up to event sources like SQS queue depth requires additional metrics adapters and careful configuration. Even then, a deeper problem remains: HPA reacts to metrics after the fact. For event-driven workloads, this creates a dangerous lag:

  • An SQS queue floods with messages; pods start struggling
  • CPU climbs - HPA detects the spike after a 15s scrape and sync cycle
  • New pods are scheduled, but nodes are full - they sit Pending
  • Cluster Autoscaler requests EC2 nodes - taking 3–5 minutes to arrive
  • By then, the queue has grown by tens of thousands of messages

KEDA: Scale Pods on Real Events

KEDA is a CNCF Graduated project that extends Kubernetes to scale workloads based on external event sources - SQS, Kafka, Prometheus, DynamoDB Streams, and 70+ built-in scalers. It installs as a lightweight operator and works alongside your existing setup, connecting workloads directly to event sources without requiring custom metrics adapters.

KEDA introduces two core resources:

  • ScaledObject - scales long-running Deployments/StatefulSets (APIs, background workers)
  • ScaledJob - spawns individual Kubernetes Jobs per event (ETL, ML inference, video transcoding)

Installing KEDA via Helm

helm repo add kedacore https://kedacore.github.io/charts
helm repo update
helm install keda kedacore/keda \
  --namespace keda --create-namespace

Configuring AWS Authentication via TriggerAuthentication

The recommended approach for AWS authentication is a TriggerAuthentication resource using EKS Pod Identity or IRSA. The older identityOwner field on the scaler itself was deprecated in KEDA v2.13 and will be removed in v3 - avoid teaching or using it in new deployments.

First, create a TriggerAuthentication that references your KEDA operator's IAM role via pod identity:

apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
  name: keda-aws-auth
  namespace: order-processing
spec:
  podIdentity:
    provider: aws               # Uses EKS Pod Identity (recommended)
    # provider: aws-eks         # Use this if still on IRSA

ScaledObject Example: SQS Queue Depth

This scales an order-processing Deployment to maintain 10 messages per pod. With 500 messages, KEDA targets 50 pods - capped at maxReplicaCount.

Production note on in-flight messages: By default, KEDA's SQS scaler counts both ApproximateNumberOfMessages (queued) and ApproximateNumberOfMessagesNotVisible (in-flight / being processed). This means pods processing messages are included in the scaling calculation, which is usually the right behaviour. However, if your workers have long processing times or you see unexpected scale-down events mid-processing, tune scaleOnInFlight, your SQS visibility timeout, and your worker shutdown handling carefully - and ensure a Dead Letter Queue is configured to catch messages that fail repeatedly.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: order-processor-scaledobject
spec:
  scaleTargetRef:
    name: order-processor
  pollingInterval: 15       # Check queue every 15s
  cooldownPeriod: 60        # Wait 60s before scaling down
  minReplicaCount: 0        # Scale to zero when idle
  maxReplicaCount: 100
  triggers:
  - type: aws-sqs-queue
    authenticationRef:
      name: keda-aws-auth   # References TriggerAuthentication above
    metadata:
      queueURL: https://sqs.us-east-1.amazonaws.com/123456/orders
      queueLength: '10'     # Target messages-per-pod ratio
      awsRegion: us-east-1
      scaleOnInFlight: 'true'   # Default: true. Set false to exclude in-flight messages

Karpenter: Right-Sized Nodes, Right Now

When KEDA scales your pods, those pods need nodes to land on. Karpenter watches for Pending pods, then automatically provisions the optimal EC2 instance type to satisfy them - typically in under 60 seconds. It also continuously bin-packs workloads and terminates underutilized nodes.

Karpenter vs. Cluster Autoscaler

Feature Cluster Autoscaler Karpenter
Provisioning Speed 3–5+ minutes Typically 30–60 seconds
Instance Selection Pre-configured ASG groups Dynamic - picks optimal type per workload
Spot Support Manual node group setup Native, single NodePool
Node Consolidation Limited Automatic bin-packing

NodePool Configuration

The NodePool resource defines what Karpenter is allowed to provision. The example below targets the stable karpenter.sh/v1 API (available from Karpenter v1.0+) and configures a mixed Spot/On-Demand pool for batch workloads. Note that in v1, the consolidation policy is named WhenEmptyOrUnderutilized (renamed from WhenUnderutilized in v1beta1), and consolidateAfter is now supported alongside it and is required.

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: batch-workers
spec:
  template:
    spec:
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: batch-workers
      requirements:
      - key: karpenter.sh/capacity-type
        operator: In
        values: ['spot', 'on-demand']    # Spot-first, On-Demand fallback
      - key: node.kubernetes.io/instance-category
        operator: In
        values: ['c', 'm', 'r']          # General, compute, memory families
      - key: kubernetes.io/arch
        operator: In
        values: ['amd64', 'arm64']       # Support Graviton for savings
  limits:
    cpu: 1000                            # Safety cap on total cluster cost
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized   # v1 name (was WhenUnderutilized in v1beta1)
    consolidateAfter: 30s                            # Required in v1; set to 0s for immediate

End-to-End Architecture Flow

When an SQS burst hits, the full scale-up sequence - from event arrival to active pod processing - completes in roughly one to two minutes in a well-tuned cluster. Actual time depends on image pull speed, node bootstrap, daemonset startup, and workload readiness. Here is the sequence:

01 Amazon SQS queue depth spikes (e.g., 200,000 messages)
02 KEDA polls queue every 15s, calculates required pod count, updates Deployment replica target
03 New pods are created - many land in Pending state (no capacity yet)
04 Karpenter detects Pending pods, selects optimal EC2 instance types, launches Spot nodes - typically in under 60s
05 Nodes join the cluster; pods are scheduled and begin processing messages
06 Queue drains → KEDA scales pods to 0 → Karpenter terminates idle nodes → worker compute cost drops to zero

Production Best Practices

KEDA

  • Always set maxReplicaCount to guard against runaway scaling from a misconfigured scaler
  • Use cooldownPeriod: 60–120s to prevent scale-down thrashing near zero
  • Authenticate via TriggerAuthentication with podIdentity.provider: aws - the identityOwner field on the scaler is deprecated since v2.13 and will be removed in KEDA v3
  • Set scaleDown.stabilizationWindowSeconds to smooth out spiky workloads
  • For SQS workers, configure visibility timeout, scaleOnInFlight, and graceful shutdown carefully - and always attach a Dead Letter Queue to catch failed messages
  • Test scale-to-zero in staging - some apps have cold-start latency that affects first-message SLA

Karpenter

  • Use the stable karpenter.sh/v1 API - v1beta1 is supported but planned for deprecation
  • Use consolidationPolicy: WhenEmptyOrUnderutilized (the v1 name; WhenUnderutilized from v1beta1 is renamed)
  • Specify multiple instance families (c, m, r) so Karpenter can find available Spot capacity
  • Set consolidateAfter explicitly - it is required in v1 when using WhenEmptyOrUnderutilized; use 0s for the same behaviour as v1beta1
  • Include arm64 (Graviton) in your NodePool - AWS Graviton instances cost up to 20% less per hour than comparable x86 instances, with equal or better performance for most cloud-native workloads
  • Set cpu and memory limits on the NodePool as a hard cost cap
  • Tag all EC2NodeClass nodes with environment, team, and cost-center for AWS Cost Explorer analysis

Observing the Stack in Production

With two autoscalers operating in tandem, visibility across KEDA, Karpenter, SQS, and EC2 simultaneously is what separates a smooth on-call experience from a painful one. When something goes wrong - pods not scaling, nodes not terminating, queue backing up - you need correlated signals from all layers at once.

  • Expose KEDA's /metrics endpoint to Prometheus - scaler values, replica counts, and error rates are all there
  • Use CloudWatch Container Insights for correlated node + pod metrics
  • Alert on SQS ApproximateAgeOfOldestMessage to catch backlogs before they compound
  • Dashboard pod count (KEDA) and node count (Karpenter) together - a node spike without pods often means a misconfigured NodePool
  • Monitor SQS NumberOfMessagesMoved on your Dead Letter Queue - a rising DLQ count signals worker failures that scaling alone cannot fix

Conclusion

KEDA and Karpenter together eliminate the manual scaling work that falls on your on-call engineer at the worst possible moment - scaling pods from real event signals, provisioning the right nodes in seconds, and returning to zero when load clears. Getting the details right (authentication, API versions, SQS in-flight behaviour, consolidation policy) is what makes this stack hold up under pressure in production.

If you have any questions or need help implementing this on your platform, you can reach out to our DevOps & Cloud Engineering team here.

April 25, 2025

Precision Scaling in Kubernetes with KEDA: ScaledObject vs. ScaleJob

Introduction

  • Kubernetes scaling is an art that requires precision, especially when working with event-driven workloads. I recently fine-tuned an AWS EKS workload using KEDA (Kubernetes Event-Driven Autoscaling). KEDA provides two primary scaling mechanisms: ScaledObject and ScaleJob.
  • ScaledObject dynamically adjusts pod replicas for Deployments or StatefulSets based on metrics like SQS queue length.
  • ScaleJob is ideal for short-lived tasks, spinning up Jobs to process a single message before terminating.

What situation we encountered?

  • Initially, we faced challenges in achieving precise autoscaling for an event-driven SQS message processing workload on AWS EKS. The key requirements were:
    • A minimum of one pod should always be running.
    • If there were fewer than four messages, they should be processed by a single pod.
    • Scaling should start when messages exceeded five.
  • We explored KEDA’s ScaleJob, but it terminated pods after processing individual messages. Next, we tried KEDA’s built-in SQS trigger, which lacked precise control, specifically:
  • It scaled down too soon, ignoring in-flight messages.
  • It didn’t allow fine-tuned scaling logic based on both visible and in-flight messages.
  • For my use case—an SQS message processing application—ScaledObject was the right choice, keeping pods alive instead of terminating them.

Why we took this approach?

  • To overcome these limitations, we implemented a custom scaler using Flask and Prometheus, allowing us to:
  • Continuously monitor both visible and in-flight SQS messages.
  • Maintain pod count until all messages were processed.
  • Dynamically scale based on real-time queue depth.
  • Prevent premature scale-downs using a stateful tracking mechanism.

Scaling Requirement: Precise Autoscaling with SQL Messages

  • Target Scaling Behavior
    • Scale-up Logic:
      • 1 pod for 0–5 messages
      • 2 pods for 6 messages
      • 3 pods for 7 messages
      • 10 pods for 14+ messages
    • Scale-down Logic:
      • Maintain current pod count until both visible + in-flight messages hit zero.
      • After 120 seconds of inactivity, scale down to 1 pod.
      • To reduce costs, these pods were deployed on AWS Spot Instances, ensuring resilience while taking advantage of lower pricing.

Initial Approach: Using KEDA's Built-in SQS Trigger

  • KEDA provides an aws-sqs-queue trigger to auto-scale based on queue length:
  • queueLength: Defines how many messages in the queue correspond to one pod.
    • Formula: Desired Pods = Total Messages / queueLength
  • targetValue: The threshold metric value at which scaling occurs (e.g., desired messages per pod).
    • Formula: Desired Pods = Total Messages / targetValue
  • activationQueueLength: Minimum number of messages required in the queue before scaling starts.
    • Formula: Scaling Triggers If Total Messages ≥ activationQueueLength
Aspect ScaledObject ScaledJob
Workload Long-running (e.g., Deployments) Short-lived (e.g., Jobs)
Scaling Adjusts pod replicas (0 to N) Creates new Job instances
Use Case Continuous services (e.g., web apps) One-off tasks (e.g., batch processing)
Execution Pods stay active, process multiple events Jobs run once per event, then terminate
Concurrency Multiple pods run in parallel Multiple Jobs run independently

Custom Scaler Solution: Achieving Precision

  • To overcome this, I built a Flask-based custom scaler, exposing Prometheus metrics via /api/v1/query. The logic ensures pods remain active until both visible and in-flight messages reach zero.
  • Python Code:
    last_replicas = 1 def calculate_replicas(visible, not_visible): global last_replicas total_messages = visible + not_visible app.logger.info(f"Visible: {visible}, NotVisible: {not_visible}, Total: {total_messages}") if total_messages == 0: app.logger.info("Queue empty, returning: 1") last_replicas = 1 return 1 else: if total_messages <= 5: desired_replicas = 1 app.logger.info(f"Total <= 5, desired replicas: {desired_replicas}") else: desired_replicas = min(total_messages - 4, 10) app.logger.info(f"Total > 5, calculated desired replicas: {desired_replicas}") if desired_replicas > last_replicas: last_replicas = desired_replicas app.logger.info(f"Scaling up, updating last_replicas to: {last_replicas}") else: app.logger.info(f"Maintaining last_replicas at: {last_replicas} (no scale-down until queue is empty)") return last_replicas
  • Key Features:
    • Maintains pod count until the queue is fully processed.
    • Prevents premature scale-down using last_replicas.
    • Scales dynamically based on visible + in-flight messages.

Deployment: Integrating with KEDA ScaledObject

  • To use this custom scaler in KEDA, we configured a ScaledObject pointing to the Flask service:
  • KEDA Configuration:
    apiVersion: keda.sh/v1alpha1 kind: ScaledObject metadata: name: custom-sqs-scaler spec: scaleTargetRef: kind: Deployment name: app pollingInterval: 10 # Every 10 seconds cooldownPeriod: 120 # Wait 120s before scaling down minReplicaCount: 1 maxReplicaCount: 10 triggers: - type: prometheus metadata: serverAddress: "http://custom-sqs-scaler.non-prod.svc.cluster.local:8080" metricName: queue_messages query: "queue_messages" threshold: "1"
  • Why This Works?
    • The Prometheus trigger queries the Flask app for real-time SQS message count.
    • The custom logic ensures pods remain active until all messages are processed.
    • Scaling down is prevented until the queue is empty.

Results: Efficient, COst-effective Scaling

  • Here’s proof that the scaler correctly handled a 12-message workload:
  • Log Output:
    INFO:scaler:Visible: 2, In-flight: 10, Total: 12 INFO:scaler:Total > 5, calculated desired replicas: 8 INFO:scaler:Maintaining last_replicas at: 10 (no scale-down until queue is empty)
  • Key Takeaways:
    • Maintained optimal pod count dynamically.
    • Ensured cost-effective scaling with spot instances.
    • Eliminated premature scale-down issues.

How this has helped?

  • By using a Prometheus-based custom scaler, we achieved:
  • Full control over scaling behavior: Pods scale up accurately based on total messages (visible + in-flight).
  • No premature scale-down: The last known pod count is maintained until the queue is fully processed.
  • Cost-effective scaling: By deploying on AWS Spot Instances, we optimized costs while ensuring workload resilience.
  • Seamless workload management: The system efficiently handled varying message loads without delays or bottlenecks.

Final Thoughts

  • KEDA’s built-in triggers provide a great starting point but can fall short in handling complex scaling scenarios. By implementing a custom scaler, we achieved:
  • Precision scaling based on visible + in-flight messages.
  • No premature scale-down, ensuring messages don’t pile up.
  • Cost-optimized scaling with AWS Spot Instances.
  • For workloads requiring precise autoscaling, implementing a Prometheus-based custom scaler significantly enhances efficiency and cost optimization.

March 21, 2025

Setting Up a Local Kubernetes Cluster with Minikube.

Kubernetes has become the go-to container orchestration platform for deploying, scaling, and managing applications. However, setting up a full-scale Kubernetes cluster can be complex, especially for local development. That’s where Minikube comes in! Minikube allows you to run a lightweight Kubernetes cluster locally, perfect for development and testing purposes.

In this Blog, I’ll walk through the steps to set up a local Kubernetes cluster using Minikube, ensuring that you can start experimenting with Kubernetes in no time.


What is Minikube?

Minikube is a tool that sets up a single-node Kubernetes cluster on your local machine. It supports multiple container runtimes like Docker, containerd, and CRI-O, and it’s an excellent option for developers who want to test Kubernetes deployments before pushing them to production.


Prerequisites

Before we dive into the setup process, you’ll need:

  • A machine with at least 2 CPUs and 2GB of RAM
  • A hypervisor like VirtualBox or Hyper-V (if using Windows)
  • kubectl (Kubernetes CLI tool)
  • Minikube

Step 1: Install Minikube

First, you need to install Minikube on your machine. The installation process varies depending on your operating system. Follow these instructions based on your platform:

For macOS (via Homebrew):

brew install minikube

For Linux (via curl):

curl -LO https://storage.googleapis.com/minikube/releases/latest/minikube-linux-amd64
sudo install minikube-linux-amd64 /usr/local/bin/minikube

For Windows (via Chocolatey):

choco install minikube





Step 2: Install kubectl

The Kubernetes command-line tool, kubectl, is essential for interacting with the cluster.

For macOS (via Homebrew):

brew install kubectl

For Linux:

sudo apt-get install -y kubectl

For Windows (via Chocolatey):

choco install kubernetes-cli









Step 3: Start Minikube

Once Minikube is installed, start it with the following command:

minikube start

This command automatically sets up a Kubernetes cluster using the hypervisor installed on your machine (VirtualBox, Hyper-V, Docker, etc.). You can also specify the driver using the --driver flag, like so:

minikube start --driver="ANY REQUIRED"

Note: By default, Minikube will use Docker as the container runtime. If you prefer containerd or CRI-O, you can specify it with the required flag





Step 4: Verify the Setup

After Minikube has started, you can verify that your cluster is running by checking the nodes in the cluster:

kubectl get nodes






Step 5: Deploy an Application on Minikube

A deployment in Kubernetes is a higher-level abstraction that manages the rollout and scaling of applications. It defines how to create and update instances of the application (called pods) consistently across a cluster. Deployments ensure that the desired number of pod replicas are running, and they automatically handle updates, rollbacks, and scaling based on user-defined conditions.

The main purpose of Deployments is because they are essential for handling production workloads and managing containerized apps in a reliable, automated way. This ensures high availability by running multiple instances of an application and scale the application dynamically in response to traffic or resource usage.

Now that Minikube is running, let’s deploy a simple application. We’ll use a sample NGINX deployment to demonstrate.

First, create a Kubernetes deployment:

kubectl create deployment nginx --image=nginx

Verify that the deployment has been created:

kubectl get deployments





Step 6: Expose the Application

By default, the NGINX deployment is not accessible from outside the cluster. To expose it, we’ll create a service:

kubectl expose deployment my-nginx --port=80






Step 7: Access the Application

Now that the service is exposed, you can access the NGINX web server using Minikube’s IP. To get the Minikube IP, run:

minikube ip: 19X.XXX.XXX.XXX:PORT

Combine this IP with the NodePort value from the previous step to access the application in your browser:







Step 8: Stop the Cluster

Once you’re done experimenting, you can stop the Minikube cluster using the following command:

minikube stop

If you want to delete the cluster entirely, run:

minikube delete


Conclusion

Minikube is a fantastic tool for local Kubernetes development, offering a quick and easy way to spin up a local cluster. In this Blog, we went through the setup process, deployed a simple application, and exposed it for external access. Now you can start experimenting with Kubernetes features and workflows in a local environment before deploying them to a production environment.

Start your Kubernetes journey with Minikube today, and happy developing!


Reference Links: