This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.

PicnicAI automates 85-90% of Kubernetes workload and improves reliability with PerfectScale by DoIT

How PicnicAI used PerfectScale by DoiT to automate workload optimization, reduce operational overhead and give engineers greater confidence in their infrastructure

PerfectScale
PicnicAI

Meet PicnicAI

PicnicAI is a health-tech company building intelligence to accelerate human health. Its technology processes and validates large volumes of medical data, helping pharmaceutical researchers run observational clinical trials more efficiently and helping patients manage their care.

The company's cloud infrastructure plays a critical role in that mission. PicnicAI operates eight or nine Kubernetes clusters supporting automated workloads, with CPU and memory requirements that can vary significantly depending on the data being processed. For PicnicAI, reliable infrastructure is ultimately about enabling its engineers to focus on healthcare innovation rather than spending their time manually managing Kubernetes.

The Challenge

As PicnicAI's Kubernetes environment evolved, the team faced an increasingly familiar challenge: understanding how much CPU and memory individual workloads actually needed.

Historically, resource requirements were largely estimated when workloads were created and could remain unchanged until an issue occurred. That approach was particularly problematic for data-processing workloads, where requirements could vary significantly between jobs.

“The way we did scaling before was essentially somebody would guess how much memory and CPU a job would need when they made it, and then it would just never be touched until there were many OOMs (out-of-memory errors)’,” says Josh Zarrabi, Infrastructure Engineer at PicnicAI.

When workloads were under-provisioned, jobs could repeatedly fail and retry, creating queues and putting additional pressure on shared databases. Conversely, over-provisioning created unnecessary infrastructure capacity and cloud spend.

The challenge was compounded by PicnicAI's lean engineering team. Infrastructure decisions had often been made by engineers who were no longer with the company, meaning the team did not always have the context behind existing configurations.

While reducing cloud costs was one motivation for exploring optimization, reliability quickly became the more important objective. “The reliability gains have been really important for us, and the cost savings have been a nice additional benefit,” says Zarrabi.

The Solution

Moving from manual rightsizing to automation

PicnicAI began working with DoiT as part of a broader cloud cost optimization initiative. DoiT introduced PerfectScale to help the team understand resource utilization at workload level and automate Kubernetes optimization.

PerfectScale combines granular Kubernetes visibility with automated, workload-aware rightsizing, helping teams optimize CPU and memory without relying on constant manual monitoring.

For PicnicAI, the implementation was deliberately phased. The team initially focused on improving reliability and establishing confidence in the recommendations before progressively expanding automation.

Identifying inefficient configurations

PerfectScale’s InfraFit capabilities helped PicnicAI identify inefficient infrastructure configurations and understand where workloads were not aligned with the resources being provisioned.

“With PerfectScale’s InfraFit, it’s much easier to be like, okay, we are actually being wasteful here because the shape is incorrect for our workloads,” says Zarrabi. “It gives me more confidence to make changes when we don’t always have full context on the decisions behind them.”

Combining automation with expert support

Technology was only part of the solution. DoiT's Field Engineering and Customer Success teams worked alongside PicnicAI to help the team understand Kubernetes behavior and determine where automation could safely be introduced.

This included reviewing the interaction between HPA (horizontal pod autoscaling) and the cluster autoscaler, adjusting thresholds and introducing controlled policies around resource requirements.

PerfectScale's policy and guardrail capabilities allowed PicnicAI to control where and how optimization was applied, rather than simply switching automation on across every workload.

More sensitive workloads were treated cautiously, while automation was progressively expanded across production, staging and development environments.

For Zarrabi, having ongoing technical support alongside the platform was an important part of the engagement. “It's nice to have somebody who looks over and keeps me accountable,” he says. “Just to make sure that we're getting the most out of the product is helpful.”

Results

  • 85–90% of workloads now automated
  • 50% approximate reduction in infrastructure team time spent on infrastructure
  • Improved reliability and fewer infrastructure distractions
  • Cost efficiency as an added benefit
  • More engineering capacity for higher-value work

“PerfectScale gives us the confidence to automate Kubernetes optimization while letting our engineers focus on building what matters.”

Josh Zarrabi, Infrastructure Engineer

Outcomes

85–90% of workloads now automated

The biggest operational change has been the move away from manual resource management. Before PerfectScale, PicnicAI had no automated workload resizing in place. Today, automated resizing covers approximately 85–90% of workloads across its application clusters.

This means resource requirements can be continuously optimized without engineers having to manually review and adjust individual workloads.

Improved reliability and fewer infrastructure distractions

The benefits extend beyond optimization itself. PicnicAI has experienced fewer memory-related issues and less of the operational disruption that previously accompanied incorrectly sized workloads.

"It's just one less annoying thing to worry about," says Zarrabi.

That change is particularly significant for a lean engineering organization. By automating more of its infrastructure management, PicnicAI has been able to increase the team's capacity to deliver more work across the business without adding operational overhead.

Instead of repeatedly investigating why jobs are failing, queues are building or databases are being placed under unnecessary load, engineers can spend more time on strategic infrastructure and product initiatives.

Cost efficiency as an added benefit

Reliability was the primary objective behind the engagement, but it wasn't the only outcome. Because PerfectScale's rightsizing is workload-aware, it has also helped PicnicAI eliminate the over-provisioning that previously built up when resource limits were guessed once and never revisited - reducing unnecessary infrastructure spend alongside the reliability gains.

"The reliability gains have been really important for us, and the cost savings have been a nice additional benefit," says Zarrabi.

For a lean team already stretched across priorities, that combination has made automation an easy case to keep investing in: it doesn't force a trade-off between running a tighter infrastructure budget and keeping systems reliable — it delivers both.

More engineering capacity for higher-value work

The reduction in manual Kubernetes management has allowed PicnicAI's engineers to focus on longer-term projects, including moving continuous integration away from Jenkins and exploring new AI initiatives.

The value of the engagement, then, goes well beyond the reliability problem it was brought in to solve. By automating a task that previously required engineering judgement and intervention, PicnicAI has also reduced infrastructure costs, freed up scarce technical capacity, and given engineers more confidence working with historical configurations they didn't originally set.

"You want to spend your time focusing on things that are useful," says Zarrabi. "Trying to guess your memory and CPU limits is not."

What's Next?

PicnicAI plans to continue expanding PerfectScale's role as its Kubernetes environment evolves, increasing automation across additional workloads while progressively addressing more sensitive production infrastructure.

The company will also continue to draw on DoiT's broader cloud expertise as its platform and AI initiatives develop. For PicnicAI, the goal is straightforward: automate infrastructure optimization wherever possible, maintain the reliability needed for critical healthcare workloads and allow engineers to spend less time managing Kubernetes and more time building technology that can improve patient outcomes.

Stop guessing Kubernetes CPU and memory limits for every workload

Automate Kubernetes rightsizing with confidence

See what PerfectScale by DoiT finds in your clusters. Get workload-level visibility, automated resizing and policy guardrails you control, with DoiT engineers alongside your team.

More customer stories

Stefanini

Stefanini Cuts Kubernetes Compute Costs 26% with PerfectScale by DoiT

26%
Cut in Kubernetes compute costs
>60%
In infrastructure cost reduction in the last year
OneFootball

PerfectScale by DoiT helps OneFootball optimize Kubernetes for global football traffic at scale

25%
reduction in Kubernetes infrastructure costs
80%
reduction in engineering effort spent on Kubernetes cost optimization and resiliency tuning
Luma Health

Luma Health cuts EKS costs by 40% while freeing up thousands of engineering hours a year

90%
Less time spent on manual Kubernetes rightsizing
40%
Reduction in Amazon EKS costs
+1,700h/year
Engineering hours redirected to reliability and performance
PlayHQ

PlayHQ Optimizes Multi-Tenant Kubernetes Costs

40%
Reduction in non-production Kubernetes costs
40%
Reduction in non-production K8s costs
20%
Reduction in production K8s costs
SNCF

How SNCF Cut K8s Waste and Increased Reliability at Scale

30%
More Workloads at Flat Cost
30%
More Workloads Absorbed at Flat Cost
~€500K
Estimated Annualized Savings
NOS

NOS Cuts Kubernetes Costs in Half and Rebuilds Trust in Optimization

50%
Cost Reduction
50%
Cost reduction on largest cluster
0%
Idle resources on main node pools
K1x

K1x Slashes Cloud Costs and Streamlines Kubernetes Operations

Thousands
Saved Per Month
Thousands
Saved per month on cloud spend
<10%
Resource utilization on over-provisioned nodes before optimization
Riftweaver

How PerfectScale Helps a DevOps Team of One Scale Riftweaver

59%
Reduction in CPU Throttling
6
Critical APIs Improved
80,000+
Game Downloads