platform-eng

InfiniBand Kubernetes Performance Meets Tenant Isolation

InfiniBand Kubernetes delivers bare-metal speed for AI workloads. vCluster Platform adds fully isolated tenant clusters on top, with Private Nodes as the production default, without sacrificing GPU performance or network throughput.

Trusted by the fastest-growing AI cloud providers
Problem

The InfiniBand Kubernetes Tradeoff

High-performance fabrics demand infrastructure that keeps pace without compromising tenant security.

Isolation Undermines Throughput

Standard Kubernetes isolation mechanisms add latency and overhead, negating the raw throughput advantages InfiniBand networking provides.

Namespace Isolation Is Too Weak

Tenants can see platform internals they should not, including cluster-wide agents and other tenants' nodes and pods.

Separate Clusters Are Too Expensive

Provisioning full physical clusters per tenant destroys the economics of shared InfiniBand fabric and GPU infrastructure.

Solution

Control Plane Virtualization on Bare Metal Fabric

vCluster Platform virtualizes the Kubernetes control plane itself, running CNCF-certified tenant clusters as lightweight processes directly on bare metal. For production deployments, Private Nodes, dedicated worker nodes with per-tenant CNI and storage, deliver hardware-level isolation. Every tenant gets their own API server, etcd, and RBAC, while InfiniBand networking and GPU hardware remain directly accessible and unthrottled.

Built for InfiniBand Kubernetes at GPU Scale

A complete stack from bare metal provisioning through tenant cluster orchestration to workload-level isolation on high-performance fabrics.

Hardware Isolation

Private Nodes for InfiniBand Tenants

Each tenant gets fully dedicated physical nodes with their own CNI and CSI configuration. No workloads from other tenants share the same hardware, preserving InfiniBand fabric integrity and GPU performance.

  • No cross-tenant hardware sharing
  • Own CNI and CSI per tenant
  • Full InfiniBand fabric access preserved
Tenant Clusters

Isolated Control Planes on Bare Metal

Tenant Kubernetes control planes run as lightweight pods on the control plane cluster, each with a dedicated API server, etcd, and scheduler. Spins up in seconds with near-zero overhead on the underlying InfiniBand fabric.

  • Own API server and etcd per tenant
  • Spins up in seconds
  • Near-zero control plane overhead
Network Isolation

Hardware-Enforced Per-Tenant Network Boundaries

Hardware-enforced VLANs, VXLANs, VRFs, ACLs, and DPU policies deliver per-tenant network boundaries without sacrificing the low-latency, high-bandwidth characteristics of InfiniBand Kubernetes environments.

  • VLANs and VRFs per tenant
  • DPU policy enforcement
  • Low-latency boundaries preserved
Workload Security

Kernel-Native Isolation Without Hypervisor Tax

vNode uses seccomp, cgroups, namespaces, and AppArmor to prevent container breakout and limit blast radius at the kernel level, preserving near-bare-metal performance. No hypervisor overhead means bare metal GPU and InfiniBand throughput are preserved across tenant boundaries.

  • No VM or hypervisor overhead
  • Container breakout protection included
  • Bare metal GPU performance preserved
Bare Metal Provisioning

Zero-Touch GPU Server Provisioning

vMetal handles PXE boot, OS installation, machine registration, and network automation for GPU servers. Provisions InfiniBand-connected GPU racks from zero to production-ready bare metal without manual intervention. vCluster Standalone then runs CNCF-certified tenant clusters directly on that provisioned hardware.

  • PXE boot and OS install automated
  • Full machine lifecycle management
  • Network automation including DPU policies

Why vCluster

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.

100K+
GPU Nodes Powered
50+
GPU Clouds & F500s
<45
Days to Launch
30K
GitHub Stars

Get Started in 3 Steps

1
Schedule a Demo

Talk to our team about your stack

2
Deploy vCluster

Deploy vCluster on your infra in minutes

3
Onboard Your Tenants

Go live with a hyperscaler-grade tenant experience in days

FAQs

How does vCluster preserve InfiniBand performance while adding tenant isolation?

vCluster virtualizes the Kubernetes control plane itself rather than inserting a hypervisor between workloads and hardware. Each tenant gets their own dedicated API server and etcd running as lightweight processes, while the underlying InfiniBand fabric and GPU hardware remain directly accessible. Private Nodes assign dedicated physical servers per tenant, and vNode prevents container breakout and limits blast radius using seccomp and cgroups without introducing VM overhead. The result is full tenant separation with minimal to no measurable impact on InfiniBand throughput or GPU performance.

Can vCluster run directly on bare metal without an existing Kubernetes cluster?

Yes. vCluster Standalone runs as a single binary directly on bare metal Linux with no external Kubernetes dependency. There is no need for k3s or kubeadm as a base layer. This is particularly relevant for InfiniBand Kubernetes deployments where the base infrastructure is GPU racks rather than a pre-existing managed cluster. vMetal handles the full provisioning path from PXE boot and OS install; vCluster Standalone then runs CNCF-certified tenant clusters directly on that provisioned hardware.

What tenant isolation options are available for InfiniBand Kubernetes environments?

vCluster Platform offers a flexible isolation spectrum. Private Nodes are the production default, giving each tenant fully dedicated physical servers with their own CNI and CSI. Private Nodes (production default) give each tenant fully dedicated physical servers with their own CNI and CSI, eliminating noisy-neighbor contention on high-performance fabrics. Shared Nodes are suited for dev, test, CI/CD, and trusted-team workloads, placing multiple tenants on the same physical hardware with namespace and resource quota boundaries. vNode prevents container breakout and limits blast radius on top of any of these configurations for the strongest workload-level isolation. AI cloud providers typically use Private Nodes for InfiniBand workloads.

Is the Kubernetes control plane per tenant CNCF certified?

Yes. Every tenant cluster created by vCluster is a fully CNCF-conformant Kubernetes control plane with 100 percent API compatibility. Tenants receive full cluster-admin access and can install their own CRDs, configure RBAC, and deploy any Kubernetes-native tooling without affecting neighboring tenants or the control plane cluster. This is relevant for AI teams who have built workflows on standard Kubernetes APIs and need compatibility guarantees on InfiniBand-backed GPU infrastructure.

What GPU cloud providers use vCluster in production today?

vCluster powers more than 100,000 GPU nodes across 50+ GPU clouds and Fortune 500 customers. Named customers include CoreWeave and Nscale. Boost Run launched a managed Kubernetes offering in less than 45 days using vCluster. Lintasarta launched a production GPU cloud in Indonesia in 90 days using vCluster. vCluster is also named in the NVIDIA DGX SuperPOD reference architecture

Does vCluster support network automation for InfiniBand and high-performance fabrics?

Yes. Through the Netris integration, vMetal automates VLANs, VXLANs, VRFs, ACLs, and DPU policies to enforce per-tenant network boundaries at the hardware level. This ensures that even on a shared InfiniBand fabric, network traffic between tenants is strictly isolated. The network automation is part of the full provisioning flow managed by vMetal, so tenant network configuration happens automatically as part of bare metal node lifecycle management.

Launch Isolated Kubernetes on Your InfiniBand Fabric

See how GPU cloud providers deploy tenant isolation on bare metal without sacrificing performance.