Skip to main content

GPU security threat model

Status: approved by Security and Product, 2026-08-07. The residual risks in Approval gates and residual risks are accepted as documented; each is a property of the design and travels with the capability wherever it is described. Shared GPU and Dedicated GPU VM creation are enabled per cluster once its qualification record is complete.

Scope: Shared GPU containers, Dedicated GPU VMs, catalog/discovery, quota, billing add-ons, operator installation, node-mode transitions, monitoring, and the tenant/admin user interfaces.

This model complements the platform security model. It does not turn a cooperative software-sharing mechanism into a hardware security boundary.

Security objectives

The GPU design must preserve these invariants:

  1. A tenant cannot gain GPU entitlement by creating or editing ordinary Kubernetes objects.
  2. A GPU request is accepted only for an enabled stable catalog profile and within organization plus project quota.
  3. A Shared GPU Pod cannot bypass the selected scheduler, request bounds, or admission-visible runtime restrictions.
  4. Failure of the Shared GPU mutating webhook fails closed for GPU Pods without blocking ordinary Pods.
  5. A Dedicated GPU VM can receive only a catalogued whole device through the KubeVirt controller and cannot be live-migrated or directly node-steered.
  6. Shared-container and VM device plugins never own the same physical device at the same time.
  7. Tenant responses never expose node names, native device-resource names, PCI selectors, device UUIDs, or other tenants' holders.
  8. Privileged reads or writes happen only after authorization with the caller's identity and organization/project scope.
  9. Add-on, annotation, HRQ, status, UI, and invoice quantity converge before a GPU product is shown as ready.
  10. Entitlement reductions, cancellation, and suspension never undercut active authoritative usage.
  11. Mode changes and upgrades stop on holders, unsupported version tuples, or incomplete rollback/monitoring evidence.
  12. Every failure either denies a new allocation, leaves an existing holder intact for controlled release, or keeps the affected node cordoned.

Trust boundaries and actors

Actor/componentTrustRelevant authority
Project member/adminUntrusted for platform integrityCreates namespaced workloads and some ordinary ResourceQuotas/ConfigMaps
Organization adminPartially trustedRequests project-cap and add-on changes only inside the JWT-bound organization
Browser/UIUntrusted clientPresents stable profile IDs; never supplies authoritative entitlement or device identity
kube-dc backendPrivileged after caller checksReads aggregate cluster inventory and writes reserved project quotas
kube-dc managerPrivileged controllerReconciles billing annotations, HRQ, status, and release state
HNC and KubeVirt controllersTrusted system identitiesPropagate HRQ; create VMIs and launcher Pods
GPU scheduler/webhooks/pluginsHighly privilegedMutate/schedule GPU Pods and access host devices/runtime state
Fleet/GitOps operatorTrusted administrative boundaryOwns drivers, plugins, monitoring, modes, upgrades, and rollback
Billing provider/webhookExternal partially trusted inputSupplies provider item IDs/quantities; cannot directly write quota
Guest/container imageUntrusted workload codeRuns inside the selected Pod or VM isolation boundary

The strongest host-risk boundary is the privileged GPU software supply chain. The strongest tenant boundary offered in the first release is the whole-device VM. Shared GPU is appropriate only for approved cooperative code.

Threat register

IDThreat and impactControls and verificationResidual/owner
GPU-T01A compromised GPU Operator, scheduler, webhook, device plugin, or driver DaemonSet takes over a node through host mounts/privilegeComponents are GitOps-owned and version-pinned; nodes/modes are explicit; Project Roles cannot modify them; upgrade and transition commands require reviewed state and rollback evidenceImage provenance/SBOM/CVE policy; Platform + Security
GPU-T02A tenant bypasses Shared GPU memory/compute policy through alternate resources, scheduler, node steering, host access, runtime class, privilege, envFrom, or CUDA_DISABLE_CONTROLCatalog-derived Pod admission validates request/limit equality and steps, exact scheduler, profile, node/runtime/security fields, volumes/devices, environment and ephemeral-container attach; the live matrix includes these reject pathsShell/arguments and CUDA-library behavior cannot be completely inspected; Product + Security decision
GPU-T03The mutating webhook or scheduler fails and an unmodified GPU Pod reaches the default schedulerThe webhook failure policy is fail-closed only for profile-labelled GPU Pods; admission requires the injected GPU scheduler; ordinary Pods bypass the selector. Controlled outage and recovery passed liveFrozen-but-listening plugin requires monitoring; the allocation-canary alert is mandatory
GPU-T04A tenant requests a native whole-device resource directly or spoofs a KubeVirt launcher ownerVM/VMI admission requires stable profile propagation and exact catalog device mapping; launcher admission allows the native resource only for the configured KubeVirt controller identity with a real VMI owner; generic hostDevices are deniedKubeVirt controller compromise is cluster-admin impact; monitor controller/supply chain
GPU-T05VM attachment weakens the promised boundary through migration, node steering, or uncatalogued devicesEffective eviction strategy must be non-migrating; node name/selector/affinity and generic host devices are denied; catalog, KubeVirt allowlist, external provider, PCI selector and node capacity must agreeGuest driver remains tenant-controlled; no device identity stability promise
GPU-T06Shared and VM plugins claim one GPU concurrently, corrupting allocation or exposing a device twiceOne expected/active mode label selects exactly one plugin; conflict/wrong-mode alerts, holder-safe transitions, exact plugin checks and rollback are implemented; live wrong-mode and both-direction driver handoff were exercisedA transition after cordon failure stays cordoned for operator recovery
GPU-T07A Project admin creates a ResourceQuota that inflates entitlement or changes controller-owned quotaOnly exact reserved names are trusted; Kubernetes-effective hard is the minimum and used is the maximum; a VAP denies reserved-name writes except backend/HNC and deletion controllers; Project RBAC is read-only for quotaBackend/controller service-account compromise can change quota; audit those identities
GPU-T08An Organization admin uses the backend service account to edit another Organization or ProjectEvery request first checks the caller JWT role, reads the Project with the caller token inside the token-derived organization, validates identifiers, then performs the narrow privileged write; route tests cover cross-Organization rejectionCluster-wide core RBAC cannot restrict create-by-name; authorization ordering is the compensating control
GPU-T09Discovery leaks node, PCI, UUID, native resource, or other-tenant holder dataTenant access uses per-request SSAR and Organization annotation scope; only a field-allowlisted aggregate is cached/returned; admin details use a separate superadmin route; errors are static; URL segments are encodedTenant dashboards/recording rules must be separately reviewed
GPU-T10A tenant amplifies cluster-wide Node/Pod/KubeVirt discovery into API-server denial of service or cache poisoningDiscovery uses a 30-second single-flight refresh and stores only the immutable redacted aggregate; authorization stays outside the cache; namespace quota reads remain scopedAdd endpoint request metrics/rate policy if pilot load shows abuse; API/SRE
GPU-T11Catalog drift maps a stable profile to the wrong native resource or unsafe request boundsGo, backend, Helm and admission validate profile shape, unique resource names, numeric ranges/steps, billing eligibility and passthrough consistency; invalid catalogs fail render/reconcile/discovery closedRules exist across languages; hardware-free fixture CI and G6 consistency tests are drift controls
GPU-T12Billing provider quantity or webhook replay grants unintended quotaProvider price mapping is catalog-owned; placeholder/missing IDs fail before mutation; one subscription item quantity maps to one stable add-on; controller, not provider, reconciles HRQ; readiness requires annotation/HRQ/status agreementExternal Stripe acceptance and provider audit
GPU-T13Reduction/cancel/suspend removes quota under active Pods/VMIs, causing accounting or service inconsistencyAuthoritative HRQ usage and named Pod/VMI blockers fail reductions closed across quota-only, Stripe and WHMCS; unsafe trial expiry defers release, emits warnings and retriesLive holder lifecycle acceptance
GPU-T14A user mistakes quota for reserved capacity or software compute percentage for hard performance/isolationUI separates entitlement, project cap, and physical capacity; copy says quota is not reservation; Shared GPU disclosure describes cooperative isolation and compute convergence/library limitations; whole-device VM is the hard-boundary optionProduct approval required before reservation sales
GPU-T15Pending/error UI shows another project's cached data or stale data as liveProject identity keys the hook state; project changes clear attribution; polling preserves stale/error state instead of relabelling it live; unknown reason codes have safe copyBrowser RBAC/accessibility E2E
GPU-T16Operator upgrade/mode commands disrupt holders or leave two owners activeCommands require exact live/fleet agreement, clean Git, both creation gates off, zero holders before and after cordon, qualification tuple/canary evidence, atomic Git operations, target plugin/label/Ready checks, and explicit cordoned resumeReal Wave-0-gated transition and selected upgrade canary
GPU-T17A failed device/node/plugin continues accepting new work or silently loses allocationsUnhealthy/conflicting/stale states remove readiness or fail discovery closed; allocation canary, plugin/scheduler/webhook/mode alerts and quota guards cover known pathsECC/XID/thermal and complete failure matrix
GPU-T18Admin monitoring data leaks physical or cross-tenant identity into a tenant Grafana organizationAdmin inventory and tenant capability contracts are separate; public fields are allowlisted; tenant metrics must use Organization/Project scope and omit node/device labelsRecording rules and cross-tenant dashboard proof
GPU-T19Secrets or licensed vGPU credentials enter ordinary config, logs, plans, or tenant APIsCurrent installer accepts only a secret-readiness boolean and rejects secret/license/UUID-shaped output; vGPU is deferredSOPS wiring and licensed vGPU review/G11

Control ownership

Control planeAuthoritative controls
AdmissionShared Pod VAP; VM/VMI VAP; launcher Pod rules; reserved ResourceQuota VAP
AuthorizationTenant RBAC; caller-token SSAR; JWT organization scope; superadmin inventory role
EntitlementBilling catalog/annotation; manager reconciliation; organization HRQ; optional project cap
DiscoveryValidated catalog; redacted single-flight aggregate; exact KubeVirt/resource consistency
RuntimeOne GPU mode/plugin owner; scheduler/webhook; whole-device VM attachment
OperationsFlux ownership; allocation canary/alerts; upgrade and transition gates; rollback runbooks
ProductIndependent discovery/billing/shared-create/VM-create flags; isolation and capacity disclosure

No browser decision, Pod annotation, tenant-created quota, provider callback, or Node label alone is sufficient to grant entitlement or prove readiness.

Required security evidence

Before opening tenant GPU creation on a cluster, its release record must retain:

  • rendered admission policies and their reject/allow matrix, including live webhook failure and controller-created launcher behavior;
  • caller-token authorization, cross-Organization rejection, reserved-quota protection, and tenant/admin redaction tests;
  • exact image/chart/driver provenance plus scan disposition;
  • zero-entitlement baseline, one controlled grant, HRQ/status convergence, and holder-safe release evidence;
  • firing and recovery evidence for frozen plugin, wrong mode, health, and quota conditions;
  • a whole-device VM create/run/stop/start/migration-denial and guest-driver qualification record;
  • manual review of Shared GPU isolation copy and its acceptance;
  • cross-tenant monitoring/dashboard tests with no node/device identity leakage;
  • rollback timing ending with zero holders and one authoritative plugin owner.

Approval gates and residual risks

Security approval is a recorded decision, not an inference from passing tests. Security and Product accepted the following residual risks on 2026-08-07:

  1. Shared compute is cooperative. Memory and control injection is not a hostile-tenant boundary: startup can exceed a requested compute percentage, and some CUDA library paths can remain above it. Workloads that must not share silicon with another tenant belong on a Dedicated GPU VM.
  2. Privileged supply chain. GPU drivers and operators carry a node-compromise blast radius. Installer pins and the runtime digest audit are in the GPU supply-chain policy; SBOM and scan disposition are part of the release record.
  3. Capacity is not reservation. Quota entitlement does not guarantee a free physical device — a reserved product requires the separate operational contract in GPU capacity reservations.
  4. Guest support is qualified per combination. The VM attachment lifecycle is proven; each guest OS, driver and CUDA combination is qualified before it is offered, per the guest and driver matrix.
  5. Observability coverage is defined per deployment. Hardware health recording rules, tenant dashboards and cross-tenant leakage proof are completed as part of a cluster's qualification record.

Each cluster's qualification record is what opens tenant Shared GPU and Dedicated GPU VM creation on that cluster.