This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.

Kubernetes Cost Optimization: From Roadblocks to Results

How Trax Retail safely cut Kubernetes costs by up to 75% and improved mission-critical business metrics with PerfectScale

PerfectScale
Trax Retail

The Challenge

At the start of the year, Trax's CFO set aggressive cost-savings goals across the organization. While Mark Serdze, Director of Cloud Infrastructure, and his team quickly reduced costs outside of Kubernetes, they hit roadblocks optimizing their large-scale, multi-cloud, multi-cluster Kubernetes environment. Manual optimization using Vertical Pod Autoscaler (VPA), cluster logs, and existing monitoring solutions didn't provide proper clarity and couldn't scale efficiently without significant development effort. The team was forced into ad-hoc, reactive actions that had minimal impact on their goals. Gaps in their tooling created time-consuming back-and-forth debate, leading to hunch-based decisions that often introduced risk to service resiliency.

The Solution

Trax deployed PerfectScale by DoiT to gain AI-guided cost visibility and optimization recommendations across their 200+ microservice Kubernetes environment. The platform delivered clear visibility into resource needs per service and surfaced the biggest waste-elimination opportunities, while balancing cost savings against resilience. PerfectScale's unique ability to consolidate every replica of a service into a single view enabled Trax to understand heterogeneous utilization across replicas — particularly valuable for ephemeral workloads like Spark and Flink jobs — and rearchitect services for further cost savings without impacting availability.

Results

  • Reduced Kubernetes costs by up to 75% in one cluster, saving over six figures in yearly expenses
  • Gained clear cost visibility across a 200+ microservice, multi-cloud, multi-cluster environment
  • Replaced an existing FinOps tool with a solution built for engineering teams, at virtually no budget impact
  • Improved the business-critical 'cost per processing' metric by rearchitecting services based on replica utilization insights
  • Eliminated countless engineering hours that would have been spent manually evaluating hundreds of replicas
  • Established a strategic partnership with hands-on POC support that drove quick, measurable results

The cost optimization recommendations were key for us, telling us what actions to take with a clear understanding of the impact each change would have. In one of our clusters, we were able to reduce cost by 75%, saving us over 6-figures in yearly expenses.

Mark Serdze, Director of Cloud Infrastructure, Trax

About Trax

Trax's mission is to enable brands and retailers to harness the power of digital technologies to produce the best shopping experiences imaginable. Their industry-leading innovations and excellence through the development of advanced technologies and autonomous data collection methods drive positive shopper experiences and unlock revenue opportunities at all points of sale. Trax's solution portfolio provides mission-critical metrics, analytics, and services that help customers save time and money by improving their shopping experience. Kubernetes is a key component of their infrastructure, supporting a large-scale multi-cloud, multi-cluster environment serving customers in over 90 countries, including some of the globe's largest enterprises.

Manual Optimizations Led to Minimal Impact on Kubernetes Cost Reduction Goals

At the beginning of the year, the CFO laid out aggressive cost savings goals throughout multiple aspects of the organization. For Mark Serdze, Director of Cloud Infrastructure, and his team, this meant quickly taking action to optimize cloud cost. They achieved fast results outside of Kubernetes, but inside their clusters they hit roadblocks. 'We started optimizing manually with available metrics, using Vertical Pod Autoscaler (VPA), cluster logs, and our monitoring solutions,' explained Serdze. 'This approach didn't give us proper clarity and would be challenging to scale efficiently without big development needs. This left us taking ad-hoc, reactive actions that were having minimal impact on our goals.' Gaps in their toolset created time-consuming debate over which actions to take, often resulting in hunch-based decisions, additional rework, and risk to service resiliency.

Quickly Optimizing K8s Cost by up to 75%

Shortly after deploying PerfectScale by DoiT, Serdze and the team gained the cost visibility they were missing across their Kubernetes environment. Across their vast 200-plus microservice environment, they had clear visibility into what resources each service needed and where the biggest waste-elimination opportunities existed. The platform's AI-guided intelligence allowed them to quickly take action, comparing cost savings with overall resilience to adjust resources safely and efficiently without compromising performance. 'The cost optimization recommendations were key for us, telling us what actions to take with a clear understanding of the impact each change would have,' Serdze explained. 'In one of our clusters, we were able to reduce cost by 75%, saving us over 6-figures in yearly expenses.' The comprehensive data and intelligence also let Trax replace an existing FinOps tool that lacked granular cost details and optimization guidance — with virtually no impact on their budget.

Kubernetes Optimization to Improve Business Metrics

After eliminating wasted resources, the Trax team turned to identifying additional cost optimization opportunities tied to business outcomes. 'A key metric for us is cost per processing, which is heavily affected by our Kubernetes efficiency,' said Serdze. 'If it gets over a certain amount, we are under a lot of pressure to figure out why and to take actions to reduce it.' PerfectScale's unique feature that consolidates every replica of a service into a single view gave the team clarity on heterogeneous utilization across replicas — particularly useful for ephemeral workloads like Spark or Flink jobs. 'We were able to build multiple flavors of the service with different levels of resources and route the incoming requests to the proper service based on the size of the data,' explained Serdze. 'This made a big impact on our cost per processing metric. PerfectScale surfaced this data instantly, and without them, we would have spent countless hours evaluating hundreds of replicas to generate the same results.'

Finding a Partner in Optimization

Beyond the quick and effective cost-optimization results, the Trax team attributed much of their success to the strategic partnership and support they received from PerfectScale. 'The support during the proof of concept (POC) made a big impact in helping us drive quick results,' said Serdze. 'The PerfectScale team sat with us, helped us optimize, and ensured our success in using the platform. I have not seen this level of commitment from other vendors, and I am glad we found a partner we can rely on to help us keep our Kubernetes cost in check as we continue to scale.'

Learn how PerfectScale improves Kubernetes efficiency

Explore how PerfectScale helps teams right-size clusters, reduce waste, and improve performance without manual tuning.

More customer stories

OneFootball

PerfectScale by DoiT helps OneFootball optimize Kubernetes for global football traffic at scale

25%
reduction in Kubernetes infrastructure costs
80%
reduction in engineering effort spent on Kubernetes cost optimization and resiliency tuning
Luma Health

Luma Health cuts EKS costs by 40% while freeing up thousands of engineering hours a year

90%
Less time spent on manual Kubernetes rightsizing
40%
Reduction in Amazon EKS costs
+1,700h/year
Engineering hours redirected to reliability and performance
PlayHQ

PlayHQ Optimizes Multi-Tenant Kubernetes Costs

40%
Reduction in non-production Kubernetes costs
40%
Reduction in non-production K8s costs
20%
Reduction in production K8s costs
SNCF

How SNCF Cut K8s Waste and Increased Reliability at Scale

30%
More Workloads at Flat Cost
30%
More Workloads Absorbed at Flat Cost
~€500K
Estimated Annualized Savings
NOS

NOS Cuts Kubernetes Costs in Half and Rebuilds Trust in Optimization

50%
Cost Reduction
50%
Cost reduction on largest cluster
0%
Idle resources on main node pools
K1x

K1x Slashes Cloud Costs and Streamlines Kubernetes Operations

Thousands
Saved Per Month
Thousands
Saved per month on cloud spend
<10%
Resource utilization on over-provisioned nodes before optimization
Trax

How Trax Cut 75% of Kubernetes spend with PerfectScale

75%
Decrease in K8s costs
Riftweaver

How PerfectScale Helps a DevOps Team of One Scale Riftweaver

59%
Reduction in CPU Throttling
6
Critical APIs Improved
80,000+
Game Downloads