PerfectScalePerfectScale

PerfectScale

Kubernetes Multi-Tenancy: Models, Isolation Layers, and Best Practices

How Kubernetes multi-tenancy actually works: soft vs. hard isolation models, the layers (RBAC, NetworkPolicy, ResourceQuotas) that hold it together, and where it breaks in practice.

This page is also available in Deutsch, Español, Français, Italiano, 日本語, and Português.

Sep 14, 20269 min read
Josh Palmer

About Josh Palmer

Head of Content

I'm Josh Palmer, Head of Content at DoiT, where I split my time across multiple business units including DoiT Cloud Intelligence, PerfectScale (Kubernetes cost optimization), and SELECT (Snowflake, Databricks, and BigQuery cost optimization). Before DoiT, I spent four and a half years at OnBoard building content for a board intelligence platform used by 6,000+ organizations, and before that, two years as Content Marketing Manager at Zylo, a SaaS management platform.

My personal page

TL;DR

  • Kubernetes multi-tenancy means running multiple teams, projects, or customers on one shared cluster instead of a dedicated cluster each, and none of the isolation happens automatically.
  • Four core layers hold it together: namespaces, RBAC, network policy, and resource quotas (LimitRanges included), plus Pod Security Standards protecting the workload level. Skip one, and a tenant that should stay isolated starts affecting everyone else.
  • Three isolation models exist: soft (namespace-based, cheapest, logical isolation only), hard (virtual clusters like vCluster or Capsule), and physical (dedicated node pools or clusters, strongest isolation, highest cost). Most organizations mix all three rather than standardizing on one.
  • Most failures trace back to drift, not bad initial setup: teams configure quotas and RBAC correctly once, then never revisit them as workloads change.

Put ten teams on one Kubernetes cluster, and by default nothing stops team A from reading team B's secrets, eating team B's CPU, or taking down a control plane team B also depends on. Kubernetes multi-tenancy closes that gap through four configurable layers: namespaces, RBAC, network policy, and resource quotas.

Multi-tenancy lets organizations stop paying for dozens of idle control planes and start sharing infrastructure across teams, projects, or customers ("tenants"). It only works if every one of those layers actually gets configured and stays configured. Kubernetes doesn't ship multi-tenant by default. Namespaces draw a logical boundary, not a security one. Skip a layer, and a tenant that should have stayed isolated starts affecting everyone else on the cluster, whether that means a security boundary or just one workload eating another's CPU.

Why teams share a cluster in the first place

Three reasons show up in almost every case: cost, operational load, and developer speed.

Running ten clusters means ten control planes, ten sets of over-provisioned node pools, ten load balancers idling most of the day. Consolidating into two or three shared clusters cuts that overhead directly. It also cuts operational load: one cluster to patch, one place to watch etcd health, one set of policies to keep consistent, instead of the same job repeated across every cluster you run.

Developer speed matters just as much, though people notice it less. Teams waiting on a platform team to provision a new cluster wait days. Teams requesting a namespace on an existing multi-tenant cluster can get one in minutes, self-service, with quotas and network policy already attached. Platform teams call this "namespace-as-a-service," and it often drives why they build multi-tenancy in the first place.

The three multi-tenancy models

Not every tenant needs the same isolation. The model you pick trades blast radius against cost and the operational complexity you choose to own.

multi-tenancy models

Model How it isolates Isolation strength Overhead Best for
Soft (namespace-based) Namespaces + RBAC + NetworkPolicy + ResourceQuota, shared control plane and kernel Logical only. A kernel or API server exploit can still cross tenants Lowest cost, easiest to operate Internal teams that already trust each other
Hard (virtual clusters) Each tenant gets its own virtual API server (vcluster, Capsule, Kiosk) layered over a shared set of nodes Strong control-plane isolation; workloads can still share a kernel unless you pair them with sandboxed runtimes Moderate: more moving parts, still one physical cluster to manage Platform teams selling "cluster-like" environments to many internal or external tenants
Physical (dedicated nodes/clusters) Node taints and tolerations, or fully separate clusters, per tenant Strongest, with a separate kernel and a separate blast radius Highest cost, most infrastructure to run Regulated workloads (HIPAA, PCI-DSS), GPU tenants, or anyone who can't share a kernel for compliance reasons

Most organizations don't pick one model and stop. They run soft multi-tenancy for trusted internal teams, and reserve hard or physical isolation for the handful of tenants (an external customer, a GPU-heavy ML team, a regulated workload) that need it. Treating "multi-tenancy" as a single on/off setting over-engineers or under-engineers a lot of these projects from the start.

The isolation layers that make soft multi-tenancy hold

Most teams start with soft multi-tenancy by default. Four layers build it, and you have to configure every one; none come switched on out of the box.

soft multitenancy layer stack

Namespaces give you the boundary everything else attaches to: a place to scope RBAC, quotas, and network policy. On their own they isolate names, not behavior. A pod in namespace A can still talk to a pod in namespace B, and a workload with no resource limits can still eat a node that namespace B's pods also use.

RBAC decides who can act on what. Start deny-by-default, grant the minimum roles each tenant's team actually needs, and bind to groups instead of individual users so access doesn't rot as people join and leave teams. Tenant RBAC sprawl (dozens of near-duplicate Roles and RoleBindings drifting out of sync) ranks among the top reasons multi-tenant clusters get harder to audit over time.

NetworkPolicy stops tenants from reaching each other over the network, which RBAC and namespaces alone don't do:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-cross-namespace
namespace: team-a
spec:
podSelector: {}
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: team-a

Most teams start with a default-deny policy per namespace, then add explicit allow rules for what actually needs to talk cross-namespace. Without it, any pod on the cluster can reach any other pod by default. Namespaces don't enforce network boundaries on their own.

ResourceQuotas and LimitRanges stop one tenant from starving the rest of the cluster on CPU, memory, or object count. A ResourceQuota caps what a namespace can consume in total; a LimitRange sets sane per-container defaults and maximums inside it. Multi-tenancy and the noisy neighbor problem meet directly here: a quota only protects the cluster when teams set the requests and limits behind it correctly, and close to real usage. Quotas set too loose don't stop noisy neighbors; quotas set too tight just push teams to over-request and waste the consolidation savings multi-tenancy promised.

Pod Security Standards round this out at the workload level, restricting privileged containers, host networking, and root access per namespace so an attacker who compromises a pod in one tenant can't escalate into the node or the rest of the cluster.

Where soft multi-tenancy breaks in practice

Two things go wrong more than anything else.

The first carries the name "blast radius": multi-tenancy isolates tenants from each other, but everyone still shares the control plane. A namespace with no object-count quota can flood etcd with CRDs or watch requests and slow down the API server for every tenant on the cluster. No amount of per-tenant RBAC or NetworkPolicy prevents this, because the failure doesn't originate at tenant-to-tenant. It originates at the tenant-to-control-plane.

blast radius comparison

The second problem: drift. Quotas, LimitRanges, and RBAC get configured correctly once, at rollout, and then nobody revisits them as workloads change. Six months later, half the quotas have gone stale: either too tight (teams padding requests to get past them) or too loose (no longer providing real protection). The noisy neighbor post covers the same failure mode: isolation that held on day one silently stops holding, and the first sign usually shows up as an eviction or an OOM kill landing on a workload that did nothing wrong.

Kubernetes multi-tenancy best practices

  • Set a default-deny network policy per namespace; add explicit allows only for what needs cross-namespace traffic
  • Scope RBAC to groups, not individual users, and review it on a schedule instead of setting it once and forgetting it
  • Attach a ResourceQuota and LimitRange to every tenant namespace, not just the ones that already caused a problem
  • Enforce Pod Security Standards per namespace, not as a cluster-wide afterthought
  • Cap object counts (not just CPU/memory) to protect the control plane from any single tenant
  • Isolate any tenant with compliance, GPU, or performance requirements on dedicated node pools (taints/tolerations) when soft multi-tenancy can't satisfy them
  • Track actual usage against requested, per namespace and per team; without that visibility, teams end up guessing which tenant causes problems

FAQ

What is the difference between hard and soft multi-tenancy in Kubernetes? Soft multi-tenancy isolates tenants logically: namespaces, RBAC, network policy, and quotas, while sharing one control plane and kernel. Hard multi-tenancy gives each tenant its own virtual API server (via tools like vCluster, Capsule, or Kiosk) or dedicated nodes, so a control-plane or kernel-level issue in one tenant can't reach another.

Does Kubernetes support multi-tenancy natively? Not out of the box. Kubernetes provides the primitives (namespaces, RBAC, NetworkPolicy, ResourceQuota, Pod Security Standards), but it doesn't wire any of them together by default, and none of them alone deliver full isolation. You build multi-tenancy from those primitives; you don't switch it on as a mode.

Do I need a separate cluster per tenant? Only for tenants that specifically require it: regulated workloads, GPU-heavy workloads with strict performance isolation needs, or external customers where a shared-kernel risk is unacceptable. Soft multi-tenancy on a shared cluster serves most internal teams well; reserve dedicated clusters or node pools for the exceptions rather than the default.

Keeping a multi-tenant cluster honest over time

Every layer above (quotas, LimitRanges, RBAC) holds up only as well as it stays current. PerfectScale by DoiT keeps requests and limits matched to real usage as workloads change, and gives you per-namespace and per-team visibility into what each team actually consumes, so you catch quota drift and noisy neighbors before they turn into an incident instead of after.

Want to see it on a real cluster? Join the live multitenancy workshop on September 29, or book a technical session.