Modern software delivery is no longer about running a script or clicking a deploy button. It's a disciplined, automated chain — where every code change flows through security checks, automated tests, container builds, and policy-gated deployments before a single pod restarts in production.
This post walks through building that exact chain using AWS-native tools. We'll deploy a simple Flask REST API called TaskFlow — but the architecture and every command here applies to any real production service.
"Every artifact is scanned, validated, and deployed in a controlled, auditable, and secure manner — that's the standard this pipeline holds itself to."
— Design principle behind this pipelineWhat You'll Build
One git push triggers a pipeline that:
- Pulls Python dependencies from a private CodeArtifact package registry
- Runs all tests and builds a Docker image in CodeBuild
- Pushes the image to ECR tagged with the git commit SHA
- Waits for a manual approval in CodePipeline before deploying
- Uses CodeDeploy lifecycle hooks to safely roll out to EKS with auto-rollback
- Meanwhile, ArgoCD watches the k8s/ manifests in Git and auto-syncs any infrastructure changes
TaskFlow API
A minimal Flask REST API with proper Kubernetes health probes, configmap-driven config, and a Dockerfile wired to accept build metadata from CodeBuild.
The app itself is intentionally simple — a task manager with CRUD endpoints and a /health, /ready, and /live endpoint. Those last three are what Kubernetes uses to decide whether to send traffic to a pod and whether to restart it.
Follow this sequence: EKS → ECR → CodeArtifact → CodeCommit → CodeBuild → CodePipeline → CodeDeploy → ArgoCD. Each service depends on the one before it being ready. Don't skip ahead.
AWS CodeArtifact
Your builds should never depend on the public internet. CodeArtifact proxies PyPI and caches every package your project uses — giving you full control, auditability, and the ability to ban vulnerable versions.
A domain is the top-level namespace for your organization. Inside it, you create repositories. The pattern here is: one upstream pypi-store that talks to public PyPI and caches everything, and one taskflow-repo that your projects actually pull from — with pypi-store as its upstream.
Create Domain and Repositories
# 1. Create the domain (one per AWS account/org)
aws codeartifact create-domain \
--domain taskflow-domain \
--region ap-south-1
# 2. Create the upstream PyPI proxy store
aws codeartifact create-repository \
--domain taskflow-domain \
--repository pypi-store
aws codeartifact associate-external-connection \
--domain taskflow-domain \
--repository pypi-store \
--external-connection public:pypi
# 3. Create your project repo with pypi-store as upstream
aws codeartifact create-repository \
--domain taskflow-domain \
--repository taskflow-repo \
--upstreams repositoryName=pypi-store
Test pip Login Locally
Before wiring this into CodeBuild, verify it works end-to-end from your local machine. The auth token is short-lived (12 hours) — in CI, CodeBuild generates it fresh at the start of every build.
# Configure pip to point at CodeArtifact
aws codeartifact login \
--tool pip \
--domain taskflow-domain \
--domain-owner $(aws sts get-caller-identity --query Account --output text) \
--repository taskflow-repo
# Install — packages now come from CodeArtifact (cached from PyPI)
pip install flask==3.0.3 gunicorn==22.0.0
# Confirm packages were cached in your repo
aws codeartifact list-packages \
--domain taskflow-domain \
--repository taskflow-repo \
--format pypi
If PyPI goes down — or if a malicious package gets published to PyPI — your builds are unaffected because they're hitting your cached registry. You can also pin exact approved versions per project and block packages that fail your security review.
AWS CodeBuild
Serverless build runner. No infrastructure to manage. Every build starts in a fresh container, runs your tests, builds and pushes the Docker image, and writes deployment metadata for downstream stages.
Understanding buildspec.yml
The buildspec.yml defines four phases. Each phase is a sequence of shell commands. Failure in any command fails the entire build — which is exactly what you want.
| Phase | What Happens |
|---|---|
| install | Sets up Python 3.12 runtime, upgrades pip |
| pre_build | Authenticates to ECR, gets CodeArtifact token, installs deps from CodeArtifact, sets IMAGE_TAG = first 8 chars of git commit SHA |
| build | Runs pytest with coverage, then docker build tagged with git SHA and latest |
| post_build | Pushes both tags to ECR, updates k8s/deployment.yaml image tag via sed, writes imagedefinitions.json for CodeDeploy |
phases:
pre_build:
commands:
# Authenticate to ECR
- aws ecr get-login-password --region $AWS_REGION |
docker login --username AWS --password-stdin
$AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com
# Authenticate to CodeArtifact and configure pip
- aws codeartifact login --tool pip
--domain $DOMAIN_NAME --repository $REPO_NAME
- pip install -r requirements.txt
# Tag = first 8 chars of git commit SHA
- export IMAGE_TAG=${CODEBUILD_RESOLVED_SOURCE_VERSION:0:8}
build:
commands:
- pytest tests/ -v --cov=app --cov-report=xml
- docker build --platform linux/amd64
--build-arg BUILD_NUMBER=$CODEBUILD_BUILD_NUMBER
--build-arg IMAGE_TAG=$IMAGE_TAG
-t $ECR_URI:$IMAGE_TAG -t $ECR_URI:latest .
post_build:
commands:
- docker push $ECR_URI:$IMAGE_TAG
- docker push $ECR_URI:latest
# Update deployment.yaml with the new image tag
- sed -i "s|image:.*taskflow-api.*|image: $ECR_URI:$IMAGE_TAG|g"
k8s/deployment.yaml
# Required by CodeDeploy
- printf '[{"name":"taskflow-api","imageUri":"%s"}]'
$ECR_URI:$IMAGE_TAG > imagedefinitions.json
Create and Test the Project
ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
# Create ECR repository first
aws ecr create-repository --repository-name taskflow-api
# Create the CodeBuild project
aws codebuild create-project \
--name taskflow-api-build \
--source '{"type":"CODECOMMIT","location":"https://git-codecommit.ap-south-1.amazonaws.com/v1/repos/taskflow-api"}' \
--artifacts '{"type":"NO_ARTIFACTS"}' \
--environment '{
"type":"LINUX_CONTAINER",
"image":"aws/codebuild/standard:7.0",
"computeType":"BUILD_GENERAL1_SMALL",
"privilegedMode":true,
"environmentVariables":[
{"name":"AWS_ACCOUNT_ID","value":"'$ACCOUNT_ID'"},
{"name":"AWS_REGION","value":"ap-south-1"},
{"name":"DOMAIN_NAME","value":"taskflow-domain"},
{"name":"REPO_NAME","value":"taskflow-repo"}
]
}' \
--service-role arn:aws:iam::${ACCOUNT_ID}:role/CodeBuildTaskflowRole
# Trigger a test build and follow the logs
BUILD_ID=$(aws codebuild start-build \
--project-name taskflow-api-build --query build.id --output text)
aws logs tail /aws/codebuild/taskflow-api-build --follow
CodeBuild needs this flag to run docker build inside a container (Docker-in-Docker). Without it, the Docker daemon is unavailable and the build fails at the image build step. Only enable this when you actually need Docker.
AWS CodePipeline
The conductor. Listens for every push to main on CodeCommit, triggers CodeBuild, enforces a human approval gate, then hands the artifact to CodeDeploy. All automated — except the approval you choose to keep manual.
The Four Stages
| Stage | Provider | What It Does |
|---|---|---|
| Source | CodeCommit | Detects push to main via EventBridge. Zips the repo and stores it in S3 as SourceOutput artifact. |
| Build | CodeBuild | Takes SourceOutput, runs buildspec.yml, produces BuildOutput artifact (imagedefinitions.json, updated k8s/ manifests). |
| Approve | Manual | Sends SNS notification to your email. Pipeline pauses. You review build output and approve or reject via link in email or CLI. |
| Deploy | CodeDeploy | Takes BuildOutput, triggers CodeDeploy deployment group, runs lifecycle hooks, validates health. |
# Create artifact bucket (versioning required)
BUCKET=taskflow-pipeline-artifacts-$(aws sts get-caller-identity --query Account --output text)
aws s3 mb s3://$BUCKET --region ap-south-1
aws s3api put-bucket-versioning --bucket $BUCKET \
--versioning-configuration Status=Enabled
# Create SNS approval topic + subscribe your email
TOPIC_ARN=$(aws sns create-topic --name taskflow-approvals \
--query TopicArn --output text)
aws sns subscribe --topic-arn $TOPIC_ARN \
--protocol email --notification-endpoint you@email.com
# Create the pipeline (fill in ACCOUNT_ID in pipeline.json first)
aws codepipeline create-pipeline \
--cli-input-json file://pipeline/pipeline.json
Approving (and Rejecting) via CLI
# Get the approval token (only valid while stage is In_Progress)
TOKEN=$(aws codepipeline get-pipeline-state \
--name taskflow-api-pipeline \
--query "stageStates[?stageName=='Approve'].actionStates[0].latestExecution.token" \
--output text)
# Approve
aws codepipeline put-approval-result \
--pipeline-name taskflow-api-pipeline \
--stage-name Approve \
--action-name ManualApproval \
--token $TOKEN \
--result '{"status":"Approved","summary":"Tests passed, LGTM"}'
# Or reject (Deploy stage never runs)
aws codepipeline put-approval-result \
--pipeline-name taskflow-api-pipeline \
--stage-name Approve \
--action-name ManualApproval \
--token $TOKEN \
--result '{"status":"Rejected","summary":"Performance regression found"}'
AWS CodeDeploy
Controlled deployment execution with lifecycle hook scripts, health validation, and automatic rollback — so a failed deployment never stays failed.
The Four Lifecycle Hooks
CodeDeploy calls these shell scripts at specific points in the deployment lifecycle. Each script is a checkpoint — if any exits with a non-zero code, CodeDeploy marks the deployment as failed and rolls back.
before_install.shkubectl is installed and the EKS cluster is reachable before any manifests are applied. A fast early exit if the cluster is unhealthy.after_install.shkubectl apply on the namespace, ConfigMap, Deployment, and Service. The Deployment's image tag has already been updated by CodeBuild.validate_service.shkubectl rollout status to confirm all pods are healthy. If it times out, automatically runs kubectl rollout undo and exits with code 1 — triggering CodeDeploy's rollback.after_allow_traffic.shSetting Up Auto Rollback
ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
# Create CodeDeploy application for EKS
aws deploy create-application \
--application-name taskflow-api \
--compute-platform EKS
# Create deployment group with auto-rollback on failure
aws deploy create-deployment-group \
--application-name taskflow-api \
--deployment-group-name taskflow-api-eks-dg \
--service-role-arn arn:aws:iam::${ACCOUNT_ID}:role/CodeDeployTaskflowRole \
--deployment-config-name CodeDeployDefault.EKSLinear10PercentEvery1Minutes \
--auto-rollback-configuration '{"enabled":true,"events":["DEPLOYMENT_FAILURE"]}'
# Test rollback: break validate_service.sh, push, watch it auto-recover
echo "exit 1 # simulated failure" >> scripts/validate_service.sh
git add scripts/validate_service.sh
git commit -m "test: simulate deployment validation failure"
git push origin main
# CodeDeploy will fail the deployment and roll back to the previous revision
EKSLinear10PercentEvery1Minutes — shifts 10% of pods to the new version every minute. Slow and safe for production.
EKSCanary10Percent5Minutes — 10% for 5 minutes, then 100%. Good middle ground.
EKSAllAtOnce — replaces all at once. Fast but no incremental validation.
Amazon EKS
Managed Kubernetes. AWS owns the control plane. You manage the nodes. The app runs as 2 pods across 2 AZs with topology spread constraints, HPA autoscaling, and liveness/readiness probes.
Cluster Setup
# Create cluster — 2 nodes, 2 AZs, managed node group (~15 mins)
eksctl create cluster \
--name taskflow-eks \
--region ap-south-1 \
--version 1.32 \
--nodegroup-name taskflow-ng \
--node-type t3.small \
--nodes 2 \
--nodes-min 2 \
--nodes-max 4 \
--managed \
--with-oidc \
--asg-access
# Confirm nodes are in different AZs
kubectl get nodes --show-labels | grep topology.kubernetes.io/zone
# Deploy the application
kubectl apply -f k8s/namespace.yaml
kubectl apply -f k8s/configmap.yaml
kubectl apply -f k8s/deployment.yaml
kubectl apply -f k8s/service.yaml
kubectl apply -f k8s/hpa.yaml
# Verify
kubectl get all -n taskflow
kubectl get pods -n taskflow -o wide
Least-Privilege RBAC
The CodeDeploy role needs to call kubectl apply — but it should never be able to read secrets, list all nodes, or modify RBAC rules. We map the IAM role in aws-auth with only system:authenticated, then create a targeted ClusterRole.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: codedeploy-deployer
rules:
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list", "watch", "create", "update", "patch"]
- apiGroups: [""]
resources: ["pods", "services", "namespaces", "configmaps"]
verbs: ["get", "list", "watch", "create", "update", "patch"]
# Notably absent: secrets, nodes, clusterroles, namespaces (delete)
system:masters is full cluster-admin with no restrictions. If a Jenkins credential, CodeDeploy role, or build agent is ever compromised, an attacker with system:masters owns your entire cluster. Always use a scoped ClusterRole.
ArgoCD
Pull-based GitOps. ArgoCD watches the k8s/ directory in CodeCommit and continuously syncs the cluster to match. Change a YAML, push it — the cluster updates itself. No kubectl from CI needed.
CodeDeploy vs ArgoCD — Two Philosophies
This pipeline shows both deployment patterns. They're not mutually exclusive — many teams use both: CodeDeploy for application releases (driven by the pipeline) and ArgoCD for infrastructure changes (Kubernetes config, HPA settings, ConfigMaps).
- CI pipeline drives deployment
- Explicit approval gates
- Lifecycle hooks for custom logic
- Built-in canary / linear strategies
- Pipeline credentials in CI
- Cluster watches Git continuously
- Automatic drift correction
- No cluster creds in CI
- Rollback = revert a git commit
- Single source of truth in Git
Install ArgoCD on EKS
# Install ArgoCD into its own namespace
kubectl create namespace argocd
kubectl apply -n argocd \
-f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml
kubectl wait --for=condition=Ready pods --all -n argocd --timeout=180s
# Get the initial admin password
kubectl get secret argocd-initial-admin-secret -n argocd \
-o jsonpath="{.data.password}" | base64 -d && echo
# Expose the UI (for lab access)
kubectl patch svc argocd-server -n argocd \
-p '{"spec":{"type":"LoadBalancer"}}'
# Install argocd CLI and login
ARGOCD_HOST=$(kubectl get svc argocd-server -n argocd \
-o jsonpath='{.status.loadBalancer.ingress[0].hostname}')
argocd login $ARGOCD_HOST --username admin --insecure \
--password $(kubectl get secret argocd-initial-admin-secret \
-n argocd -o jsonpath="{.data.password}" | base64 -d)
The GitOps Workflow
Once the Application is created, the entire workflow is: edit YAML → git push → ArgoCD syncs. That's it. No pipeline step. No manual kubectl.
# Connect CodeCommit repo to ArgoCD
argocd repo add \
https://git-codecommit.ap-south-1.amazonaws.com/v1/repos/taskflow-api \
--username YOUR_IAM_USER --password YOUR_CODECOMMIT_HTTPS_PASS
# Create the Application (auto-sync + self-heal enabled)
argocd app create taskflow-api \
--repo https://git-codecommit.ap-south-1.amazonaws.com/v1/repos/taskflow-api \
--path k8s \
--dest-server https://kubernetes.default.svc \
--dest-namespace taskflow \
--revision main \
--sync-policy automated \
--auto-prune \
--self-heal
# === Practice: GitOps Update ===
# Change replica count, push → ArgoCD auto-applies in ~3 min
sed -i 's/replicas: 2/replicas: 3/' k8s/deployment.yaml
git add k8s/deployment.yaml && git commit -m "scale: 3 replicas" && git push origin main
argocd app sync taskflow-api # force immediate sync if you don't want to wait
# === Practice: Drift Detection ===
# Manually scale to 1 — ArgoCD self-heals back to 3 within seconds
kubectl scale deployment taskflow-api -n taskflow --replicas=1
kubectl get pods -n taskflow --watch # watch it snap back
# === Practice: Rollback ===
argocd app history taskflow-api
argocd app rollback taskflow-api <REVISION_ID>
Putting It All Together
The first end-to-end run is always the most satisfying. You push one commit to CodeCommit, and over the next few minutes you can watch: CodePipeline light up, CodeBuild run your tests, an image appear in ECR with a git SHA tag, an approval email hit your inbox, CodeDeploy lifecycle hooks fire one by one, and finally kubectl get pods -n taskflow showing the new version running.
ArgoCD adds a separate layer on top — watching the same repo and ensuring the cluster never drifts from what Git says. Try manually editing a replica count with kubectl and watch ArgoCD snap it back within seconds. That's self-healing infrastructure.
EKS control plane costs ~$0.10/hr. Two t3.small nodes add ~$0.04/hr each. When you're done: eksctl delete cluster --name taskflow-eks and delete the ECR images, CodePipeline, S3 bucket, and SNS topic.