The Situation
A European B2B SaaS company with ~200 employees and 5,000+ enterprise customers came to us with a problem: their AWS bill had grown from EUR 18,000/month to EUR 38,000/month over 18 months, but their customer count had only grown 40%. Something was wrong, but their team of 25 engineers was too busy shipping features to investigate.
Their stack: 3 EKS clusters (production, staging, development), 15 RDS instances, 4 ElastiCache clusters, S3 for file storage (~80 TB), CloudFront CDN, and a collection of Lambda functions for async processing. A typical modern SaaS setup.
Week 1: Discovery and Quick Wins
We started with a full AWS Cost Explorer analysis segmented by service, tag, and usage type. The breakdown told the story immediately:
| Service | Monthly Cost | % of Total |
|---|---|---|
| EC2 (EKS nodes) | EUR 14,200 | 37% |
| RDS | EUR 8,100 | 21% |
| ElastiCache | EUR 3,400 | 9% |
| S3 + Data Transfer | EUR 4,800 | 13% |
| CloudFront | EUR 2,900 | 8% |
| Lambda + SQS + SNS | EUR 1,600 | 4% |
| Other (NAT GW, EBS, etc.) | EUR 3,000 | 8% |
Quick Win #1: Shut Down the Dev Cluster Nights and Weekends
The development EKS cluster ran 24/7 but was only used Monday-Friday, 8:00-19:00 CET. We implemented a scheduled scaling policy using Karpenter to scale dev nodes to zero outside business hours.
Savings: EUR 1,900/month
Quick Win #2: Delete Unused EBS Snapshots and Volumes
We found 340 orphaned EBS snapshots totalling 12 TB and 18 unattached EBS volumes from terminated instances. A simple cleanup script:
# Find and delete unattached EBS volumes
aws ec2 describe-volumes \
--filters Name=status,Values=available \
--query 'Volumes[*].[VolumeId,Size,CreateTime]' \
--output table
# Delete snapshots older than 90 days not attached to any AMI
aws ec2 describe-snapshots --owner-ids self \
--query 'Snapshots[?StartTime<`2026-02-01`].[SnapshotId,VolumeSize,StartTime]' \
--output table
Savings: EUR 680/month
Quick Win #3: S3 Lifecycle Policies
Of the 80 TB in S3, 60 TB was log data and old backups that had not been accessed in 90+ days. We implemented lifecycle rules:
- Move to S3 Infrequent Access after 30 days (40% cheaper)
- Move to S3 Glacier Instant Retrieval after 90 days (68% cheaper)
- Delete logs older than 365 days
Savings: EUR 1,800/month (phased in over 90 days as objects transition)
Week 2: Right-Sizing Compute
EKS Node Right-Sizing
Using Kubecost data, we found the production cluster was running at 28% average CPU utilisation and 42% memory utilisation across 18 m6i.2xlarge nodes. Kubernetes resource requests were inflated — most pods requested 2x-4x their actual usage.
Actions taken:
- Ran Goldilocks for 2 weeks to get per-deployment resource recommendations
- Adjusted resource requests to P95 actual usage across all deployments
- Enabled Karpenter with mixed instance types including Graviton (m7g) for 20% better price-performance
- Moved stateless workloads to spot instances (3 node groups with 4 instance type diversification each)
Result: production cluster went from 18 nodes to 11 nodes with the same capacity headroom.
Savings: EUR 4,600/month
RDS Right-Sizing
Of the 15 RDS instances:
- 3 were db.r6g.2xlarge running at 8-12% CPU — downgraded to db.r6g.xlarge
- 4 were development databases on Multi-AZ — switched to Single-AZ (development does not need automatic failover)
- 2 were legacy databases for features that had been deprecated — consolidated into the main cluster
Savings: EUR 3,200/month
Week 3: Commitment-Based Discounts
With the right-sized infrastructure as the new baseline, we could safely commit to reserved capacity.
Compute Savings Plan
We purchased a 1-year Compute Savings Plan covering 70% of the right-sized compute baseline. This applies to EC2, EKS, Lambda, and Fargate — providing flexibility to shift between services.
Savings: EUR 3,100/month (38% discount on covered compute)
RDS Reserved Instances
1-year All Upfront Reserved Instances for the 5 production database instances (the right-sized configurations).
Savings: EUR 1,400/month (42% discount)
ElastiCache Reserved Nodes
1-year reservation for the 2 production ElastiCache clusters after confirming the instance types were appropriate.
Savings: EUR 850/month
Week 4: Network and Transfer Optimization
NAT Gateway Costs
NAT Gateway was costing EUR 1,200/month, mostly from container image pulls and S3 access traversing the NAT. We added:
- S3 VPC Gateway Endpoint (free — eliminates NAT charges for S3 traffic)
- ECR VPC Interface Endpoint (EUR 20/month — eliminates NAT charges for image pulls)
- DynamoDB VPC Gateway Endpoint (free)
Savings: EUR 750/month
CloudFront Optimisation
Increased cache TTLs for static assets from 24 hours to 30 days, reducing origin requests by 60%. Enabled CloudFront Origin Shield to reduce cache misses to the origin.
Savings: EUR 520/month
Results Summary
| Category | Monthly Savings |
|---|---|
| Dev cluster scheduling | EUR 1,900 |
| EBS cleanup | EUR 680 |
| S3 lifecycle policies | EUR 1,800 |
| EKS right-sizing + spot | EUR 4,600 |
| RDS right-sizing | EUR 3,200 |
| Compute Savings Plan | EUR 3,100 |
| RDS Reserved Instances | EUR 1,400 |
| ElastiCache reservations | EUR 850 |
| NAT Gateway optimisation | EUR 750 |
| CloudFront optimisation | EUR 520 |
| Total monthly savings | EUR 18,800 |
New monthly bill: EUR 20,100 (down from EUR 38,000). That is a 47% reduction with zero performance degradation. Annualised savings: EUR 225,600.
What We Did Not Do
Equally important is what we did not change:
- We did not reduce production redundancy or availability
- We did not switch to cheaper, unproven services
- We did not force application code changes — all optimisations were infrastructure-level
- We did not use 3-year commitments — 1-year gives flexibility to re-optimise
Lessons for Your AWS Bill
- Start with visibility — You cannot optimise what you cannot see. Enable Cost Explorer, set up cost allocation tags, and install Kubecost for Kubernetes.
- Quick wins first — Unused resources, lifecycle policies, and dev environment scheduling deliver immediate savings with zero risk.
- Right-size before committing — Never buy reserved capacity for oversized instances. Right-size first, observe for 2 weeks, then commit.
- Make it ongoing — This client now runs a monthly cost review with their platform team. Costs have stayed flat despite 25% customer growth in the 6 months since our engagement.
If your AWS bill feels too high, it probably is. Our DevOps consulting team runs a 2-week cost optimisation sprint that typically delivers 30-50% savings. Book a free consultation — we will review your cost explorer data and tell you where the savings are before any engagement begins.