This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.

PlayHQ Optimizes Multi-Tenant Kubernetes Costs

How PlayHQ automated Kubernetes rightsizing and gained real-time per-tenant cost visibility with PerfectScale by DoiT

PerfectScale
PlayHQ

The Challenge

PlayHQ migrated from ECS to Kubernetes and achieved significant initial savings, but as its multi-tenant customer base expanded, usage patterns became more complex and unpredictable. With about 90% of compute running on Kubernetes and hundreds of small, variable workloads across tenants, infrastructure-level autoscaling with Karpenter, HPA, and KEDA ensured capacity but could not accurately right-size CPU and memory requests at the workload level. Manual resource optimization did not scale for a lean platform team, and an evaluation of OpenCost revealed too much operational overhead to run and configure.

The Solution

PlayHQ adopted PerfectScale through DoiT to automate Kubernetes rightsizing and improve EKS cost efficiency across its multi-tenant clusters. The team was up and running with automation in a dev cluster within an hour of the initial demo. PerfectScale provided automated workload rightsizing that continuously adjusts CPU and memory requests, maintenance windows that prevent changes during high-traffic Saturday game days, namespace-level exclusions for sensitive workloads, and balanced optimization policies for production clusters that protect stability while reducing cost.

Results

  • 40% reduction in non-production Kubernetes costs
  • 20% reduction in production Kubernetes costs
  • Tens of thousands of dollars saved annually
  • Improved performance consistency across workloads
  • Real-time visibility into per-tenant Kubernetes cost allocation
  • Faster detection and remediation of resiliency issues, including OOM alerts
  • Automation set up in a dev cluster within an hour of the initial demo

Every customer has unique usage patterns. Manual resource optimization simply didn't scale—we needed automation to ensure every customer, regardless of size, had right-sized infrastructure without consuming our team's capacity.

Brad Quinn, Lead Platform Engineer, PlayHQ

Running Community Sports on Kubernetes

Every day across the globe, PlayHQ runs critical digital infrastructure for thousands of community sports competitions. Leading sports organizations—including junior Australian rules football clubs, basketball leagues, and local netball and football associations—depend on the platform during peak match windows when usage spikes sharply. These organizations use PlayHQ to manage registrations, scoring, payments, and competition operations across hundreds of teams. With demand varying widely by sport, season, and region, scaling infrastructure efficiently and controlling costs is a constant challenge. At the center of this operation is PlayHQ's platform engineering group, led by Lead Platform Engineer Brad Quinn, which manages reliability, performance, and cost of the Kubernetes environment. 'Kubernetes is very important to us: It's the day-to-day workflow of our engineers,' Quinn says. With about 90 percent of PlayHQ's compute running on Kubernetes, the team needed a more automated and intelligent approach to cost optimization as their customer base expanded.

Why PlayHQ Prioritized Kubernetes Cost Optimization

PlayHQ migrated from ECS to Kubernetes to reduce infrastructure costs and speed up tenant deployments. While the move delivered early savings, costs began to rise again as more organizations joined the platform and usage patterns became more unpredictable. 'We moved from ECS to Kubernetes and achieved significant initial savings,' Quinn says. 'As our customer base expanded and usage patterns became more complex, we needed to evolve our optimization approach to maintain that efficiency.' PlayHQ was already using Karpenter to scale nodes efficiently and relied on HPA and KEDA to adjust capacity in response to demand, particularly during peak weekend traffic. But workload demands varied widely between tenants. Infrastructure-level autoscaling ensured capacity, but it did not solve the challenge of accurately right-sizing CPU and memory requests across hundreds of small, variable workloads. Quinn also evaluated OpenCost but found the operational overhead too high for a lean platform team.

How PerfectScale Automated Kubernetes Optimization

PerfectScale quickly emerged as the ideal solution for automating Kubernetes rightsizing and improving EKS cost efficiency across PlayHQ's multi-tenant clusters. Speed was a key part of the equation: 'We had the initial demo, got access to the platform, and had automation set up in a dev cluster within an hour,' Quinn says. PerfectScale's capabilities aligned well with PlayHQ's environment and operational rhythms, allowing the team to automate optimization safely while maintaining full control over when and where changes occurred. Key capabilities included automated workload rightsizing to continuously adjust CPU and memory requests, maintenance windows that prevent changes during high-traffic sports periods such as Saturday game days, namespace-level exclusions to avoid modifying sensitive workloads, and balanced optimization policies for production clusters to ensure stability alongside cost reduction.

Results: Lower Costs and Faster Issue Resolution

After adopting PerfectScale through DoiT, PlayHQ achieved significant improvements across both cost and operational performance: a 40% reduction in non-production Kubernetes costs, a 20% reduction in production Kubernetes costs, tens of thousands of dollars saved annually, improved performance consistency across workloads, and greater visibility into per-tenant cost allocation. The team also saw faster detection and remediation of resiliency issues. 'The resilience stuff, like OOM alerts, makes the time to resolution much faster,' Quinn says. Real-time insight into Kubernetes costs per tenant gives Quinn a clearer view of seasonality across different tenants and strengthens how he communicates the value of optimization to leadership. 'We can quickly get the compute costs per tenant on the fly,' Quinn says. 'I can then share the total saved and issues solved with the CTO.'

Building a Cost-Aware Engineering Culture

With automated Kubernetes optimization now in place, PlayHQ plans to extend PerfectScale access beyond the platform team. The goal is to help product engineering squads manage their own SLOs and understand how design decisions influence Kubernetes resource usage and infrastructure costs. This supports PlayHQ's long-term goal of aligning engineering practices with sustainable growth, while continuing to deliver reliable digital experiences to clubs, leagues, and families.

Partnering With DoiT for Kubernetes Optimization at Scale

Quinn underscored the value of DoiT and PerfectScale as long-term partners. During onboarding, the teams identified and resolved an issue within hours, further strengthening PlayHQ's confidence in the solution. With automated optimization, meaningful cost savings, and improved reliability, the platform team has the operational efficiency and visibility needed to support PlayHQ's continued expansion.

Learn how PerfectScale improves Kubernetes efficiency

Explore how PerfectScale helps teams right-size clusters, reduce waste, and improve performance without manual tuning.

More customer stories

OneFootball

PerfectScale by DoiT helps OneFootball optimize Kubernetes for global football traffic at scale

25%
reduction in Kubernetes infrastructure costs
80%
reduction in engineering effort spent on Kubernetes cost optimization and resiliency tuning
Luma Health

Luma Health cuts EKS costs by 40% while freeing up thousands of engineering hours a year

90%
Less time spent on manual Kubernetes rightsizing
40%
Reduction in Amazon EKS costs
+1,700h/year
Engineering hours redirected to reliability and performance
SNCF

How SNCF Cut K8s Waste and Increased Reliability at Scale

30%
More Workloads at Flat Cost
30%
More Workloads Absorbed at Flat Cost
~€500K
Estimated Annualized Savings
NOS

NOS Cuts Kubernetes Costs in Half and Rebuilds Trust in Optimization

50%
Cost Reduction
50%
Cost reduction on largest cluster
0%
Idle resources on main node pools
K1x

K1x Slashes Cloud Costs and Streamlines Kubernetes Operations

Thousands
Saved Per Month
Thousands
Saved per month on cloud spend
<10%
Resource utilization on over-provisioned nodes before optimization
Trax

How Trax Cut 75% of Kubernetes spend with PerfectScale

75%
Decrease in K8s costs
Riftweaver

How PerfectScale Helps a DevOps Team of One Scale Riftweaver

59%
Reduction in CPU Throttling
6
Critical APIs Improved
80,000+
Game Downloads
Rapyd

How Rapyd Solved Observability Gaps to Cut K8s Costs by 40%

35-40%
Cloud Cost Reduction
35-40%
Projected Cloud Cost Reduction
15+
AWS EKS Clusters Optimized