PerfectScale by DoiT helps OneFootball optimize Kubernetes for global football traffic at scale
- 25%
- reduction in Kubernetes infrastructure costs
- 80%
- reduction in engineering effort spent on Kubernetes cost optimization and resiliency tuning
As NOS's Kubernetes footprint grew across a hybrid on-prem and Google Cloud environment, optimization became a major pain point. Observability tools like Prometheus and Grafana provided raw metrics but no actionable recommendations, forcing engineers to guess at rightsizing. Earlier manual optimization attempts led to crashes and SLA breaches, breaking trust in optimization efforts. FinOps teams also spent days each month stitching together fragmented cost data, while overprovisioning remained the default to avoid risk.
NOS adopted PerfectScale, first rolling it out in development clusters where automation ran quietly for four months with costs going down and no incidents. With that success, NOS expanded automation to platform-level components including the ingress controller, cert manager, and observability stack. In production, SREs apply recommendations manually with contextual evidence from PerfectScale to ensure changes won't impact performance. InfraFit identifies optimal node types so Cluster Autoscaler scales efficiently in sync with user traffic.
I believed in the product the first time I saw it. I still show it to everyone. It was the only solution that combined smart automation with real cost savings, without putting performance at risk.
Joao Soares, Platform Engineering Lead, NOS
NOS is one of Portugal's leading telecom providers, serving millions with network, internet, and entertainment services. Almost a decade ago, NOS committed to containerization to reduce overhead and move faster, evolving from virtual machines to Docker and Rancher, and ultimately to Kubernetes rolled out across the company. Today, they operate a hybrid setup: telco-specific systems remain on-prem, while scalable workloads run on Google Cloud. The goal has remained constant—deliver agility and performance while keeping infrastructure spend under control.
As NOS's Kubernetes architecture evolved, complexity grew with it. Teams used Prometheus and Grafana, but these tools only provided raw metrics without the visibility or recommendations needed for safe, confident decisions—especially when SLAs were on the line. Without clear guidance, engineers had to guess. After several failed manual rightsizing attempts that caused crashes and SLA breaches, teams pulled back. 'Developers tried to make changes based on what they thought made sense, and it backfired,' said Joao Soares, Platform Engineering Lead at NOS. Compounding this, FinOps teams spent days each month piecing together fragmented data to build reports, making it hard to spot issues or prepare for budget reviews.
At KubeCon Europe 2024 in Paris, Soares was looking for one thing—a safer way to cut Kubernetes costs. PerfectScale stood out immediately. 'I believed in the product the first time I saw it. I still show it to everyone,' said Soares. 'It was the only solution that combined smart automation with real cost savings, without putting performance at risk.'
PerfectScale was first rolled out in development clusters, where automation ran quietly for four months—with costs going down and without a single issue or complaint. With that success, NOS began automating platform-level components like the ingress controller, cert manager, and observability stack. In production, SREs still apply recommendations manually, but trust is growing as PerfectScale provides the contextual evidence they need to feel confident that changes won't cause performance issues.
NOS reduced costs by over 50% on their largest and most critical cluster—the main API cluster—which had been heavily overprovisioned to avoid risk. PerfectScale uncovered overlooked issues like out-of-memory kills and CPU throttling, providing targeted recommendations to resolve them. InfraFit helped NOS identify optimal node types, enabling Cluster Autoscaler to scale more efficiently. 'We're seeing the sine wave we wanted—scaling up and down perfectly with user traffic. Paired with automated workload rightsizing, several main node pools now run with 0% idle resource,' said Soares. FinOps reporting also accelerated dramatically, reclaiming two to three days every month. 'Now I walk into meetings knowing everything's fine,' said Soares.
PerfectScale now powers day-to-day decisions across NOS, replacing scattered tools and guesswork with a single trusted platform. Platform Engineers use it to automate cluster and resource management; SREs apply recommendations to business-critical services; Developers use it to provision appropriately from dev to prod; Financial Controllers monitor budget and track spend; and Management oversees efficiency and cross-team collaboration. This shared visibility has improved collaboration and made Kubernetes optimization part of day-to-day operations across the company.
Explore how PerfectScale helps teams right-size clusters, reduce waste, and improve performance without manual tuning.