NVIDIA GPU Operator
When you enable the NVIDIA GPU Operator cluster add-on, you can pass the following key/value pairs as arguments.
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
affinity |
affinity |
A group of affinity scheduling rules. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
nodeSelectors |
node selectors |
You can use node selectors and node labels to control the worker nodes on which add-on pods run. For a pod to run on a node, the pod's node selector must have the same key/value as the node's label. Set JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | {"foo":"bar", "foo2": "bar2"}The pod will only run on nodes that have the |
numOfReplicas |
numOfReplicas | The number of replicas of the add-on deployment. Not used by:
Possible equivalents:
|
Required | 1Creates one replica of the add-on deployment per cluster. |
2Creates two replicas of the add-on deployment per cluster. |
rollingUpdate |
rollingUpdate |
Controls the desired behavior of rolling update by maxSurge and maxUnavailable. JSON format in plain text or Base64 encoded. Not used by:
Possible equivalents:
|
Optional | null | null |
tolerations |
tolerations |
You can use taints and tolerations to control the worker nodes on which add-on pods run. For a pod to run on a node that has a taint, the pod must have a corresponding toleration. Set JSON format in plain text or Base64 encoded. Possible equivalents:
|
Optional | null | [{"key":"tolerationKeyFoo", "value":"tolerationValBar", "effect":"noSchedule", "operator":"exists"}]Only pods that have this toleration can run on worker nodes that have the |
topologySpreadConstraints |
topologySpreadConstraints |
How to spread matching pods among the given topology. JSON format in plain text or Base64 encoded. Not used by:
|
Optional | null | null |
| Key (API and CLI) | Key's Display Name (Console) | Description | Required/Optional | Default Value | Example Value |
|---|---|---|---|---|---|
skipNodeFeatureDiscoveryDependencyCheck |
skipNodeFeatureDiscoveryDependencyCheck | Skip Node Feature Discovery dependency check. | Optional | false |
|
disableNvidiaGpuPlugin |
disableNvidiaGpuPlugin | Disable the existing NvidiaDevicePlugin add-on and related resources. | Optional | false |
|
cdi.enabled |
Enable Container Device Interface | Flag to enable Container Device Interface. | Optional |
|
|
cdi.default |
Enable CDI as default | Flag to enable the container runtime to use CDI as the default option to perform device injection. | Optional |
Deprecated in version 25.10.x and later. |
|
daemonsets.labels |
DaemonSets Labels | Labels for all GPU Operator DaemonSets. JSON text or Base64-encoded JSON map. | Optional | null |
|
daemonsets.annotations |
DaemonSets Annotations | Annotations for all GPU Operator DaemonSets. JSON text or Base64-encoded JSON map. | Optional | null |
|
daemonsets.tolerations |
DaemonSets Tolerations | Tolerations for all GPU Operator DaemonSets. JSON text or Base64-encoded JSON array. | Optional | null |
|
daemonsets.rollingUpdate.maxUnavailable |
DaemonSets MaxUnavailable | Maximum number of unavailable pods for rolling update on GPU Operator DaemonSets. | Optional | 1 |
|
dcgm.enabled |
Enable DCGM | Enable NVIDIA DCGM Hostengine as a separate pod. | Optional | false |
|
dcgm.args |
DCGM Args | Arguments for DCGM Hostengine. JSON plain text or Base64-encoded JSON array. | Optional | null |
|
dcgm.env |
DCGM Env | Environment variables for DCGM Hostengine. JSON plain text or Base64-encoded JSON. | Optional | null |
|
dcgm.containerResources |
DCGM Resources | Resource requests and limits for the DCGM Hostengine pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
dcgmExporter.enabled |
Enable DCGM Exporter | Enable the DCGM Exporter DaemonSet for metrics. | Optional | true |
|
dcgmExporter.env |
DCGM Exporter Env | Environment variables for DCGM Exporter. JSON plain text or Base64-encoded JSON. | Optional | null |
|
dcgmExporter.resources |
DCGM Exporter Resources | Resource requests and limits for the DCGM Exporter pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
dcgmExporter.service.internalTrafficPolicy |
Exporter Service Traffic Policy | InternalTrafficPolicy for the DCGM Exporter ClusterIP service. | Optional | Cluster |
|
dcgmExporter.serviceMonitor.enabled |
Enable ServiceMonitor | Enable a ServiceMonitor for NVIDIA DCGM Exporter. | Optional | false |
|
dcgmExporter.serviceMonitor.interval |
ServiceMonitor Interval | Scrape interval for NVIDIA DCGM Exporter. Supports y, w, d, h, m, s, and ms. |
Optional | 15s |
|
dcgmExporter.serviceMonitor.honorLabels |
ServiceMonitor Honor Labels | Prefer target metric labels on collisions. | Optional | false |
|
dcgmExporter.serviceMonitor.additionalLabels |
ServiceMonitor Additional Labels | Additional labels for the ServiceMonitor. JSON plain text or Base64-encoded JSON map. | Optional | null |
|
dcgmExporter.serviceMonitor.relabelings |
ServiceMonitor Relabelings | Relabeling rules for the ServiceMonitor. JSON plain text or Base64-encoded JSON. | Optional | null |
|
dcgmExporter.config.name |
Exporter ConfigMap Name | Name of the ConfigMap that provides custom metrics configuration for DCGM Exporter. | Optional | null |
|
migManager.enabled |
Enable MIG Manager | Enable deployment of NVIDIA MIG Manager. | Optional | false |
|
migManager.resources |
MIG Manager Resources | Resource requests and limits for each MIG Manager pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
migManager.env |
MIG Manager Env | Environment variables for MIG Manager. JSON plain text or Base64-encoded JSON. | Optional | null |
|
migManager.config.name |
MIG ConfigMap Name | ConfigMap name with mig-parted configuration for MIG Manager. |
Optional | null |
|
migManager.gpuClientsConfig.name |
MIG GPU Clients Config | ConfigMap name with gpu-clients configuration for MIG Manager. |
Optional | null |
|
devicePlugin.enabled |
Enable Device Plugin | Enable NVIDIA Kubernetes Device Plugin. | Optional | true |
|
devicePlugin.args |
Device Plugin Args | Arguments for the Device Plugin. JSON plain text or Base64-encoded JSON array. | Optional | null |
|
devicePlugin.env |
Device Plugin Env | Environment variables for the Device Plugin. JSON plain text or Base64-encoded JSON. | Optional | null |
|
devicePlugin.resources |
Device Plugin Resources | Resource requests and limits for each Device Plugin pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
devicePlugin.config.name |
Device Plugin ConfigMap Name | ConfigMap name that provides Device Plugin configuration, shared with GFD when applicable. | Optional | null |
|
devicePlugin.mps.root |
MPS Root | MPS root path on the host. | Optional | /run/nvidia/mps |
|
toolkit.enabled |
Enable Toolkit | Enable deployment of NVIDIA Container Toolkit. | Optional | true |
|
toolkit.env |
Toolkit Env | Environment variables for the Container Toolkit. JSON plain text or Base64-encoded JSON. | Optional | null |
|
toolkit.resources |
Toolkit Resources | Resource requests and limits for each Container Toolkit pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
toolkit.installDir |
Toolkit Install Dir | Install directory of Container Toolkit utilities on the host. | Optional | /usr/local/nvidia |
|
mig.strategy |
MIG Strategy | MIG strategy for GFD and Device Plugin. Valid values are none, single, and mixed. |
Optional | single |
|
operator.logging.level |
Operator Log Level | Zap logging level: debug, info, error, or an integer > 0 for custom verbosity. |
Optional | info |
|
operator.resources |
Operator Resources | Resource requests and limits for the GPU Operator pod. JSON plain text or Base64-encoded JSON. | Optional | null |
|
hostPaths.rootFS |
Host RootFS | Host root filesystem path. The path must be chroot-able. | Optional | / |
|
hostPaths.driverInstallDir |
Driver Install Dir | Root directory for NVIDIA driver files on the host. | Optional | /run/nvidia/driver |
|
driver.rdma.enabled |
Enable RDMA in Driver | Enable RDMA support for the NVIDIA driver. | Optional | false |
|
driver.rdma.useHostMofed |
Use Host MOFED | Use host-provided Mellanox OFED instead of container-provided Mellanox OFED. | Optional | false |
The following table describes considerations for configuring this cluster add-on in large clusters.
| Argument Name | Roles wrt number of Nodes | Description | As cluster size increases | Risks | Recommendation |
|---|---|---|---|---|---|
devicePlugin.*
|
Y |
Controls the NVIDIA device plugin that exposes GPUs to Kubernetes. |
The device plugin runs on every GPU node and is responsible for:
The configuration directly affects scheduler decisions and GPU workload placement. |
If this argument is misconfigured:
|
|
toolkit.*
|
Y |
Controls the NVIDIA Container Toolkit, which provides GPU access inside containers. |
The toolkit runs on every GPU node and is required for GPU workloads to access GPUs inside containers. |
If this argument is misconfigured, pods might be scheduled successfully but:
|
|
daemonsets.tolerations
|
Y |
Controls whether GPU Operator components can run on GPU nodes. |
GPU nodes are often tainted. All GPU Operator DaemonSets must tolerate those taints to run on the GPU nodes. |
If the required tolerations are missing:
|
|
daemonsets.rollingUpdate.maxUnavailable
|
Y |
Controls how many nodes are updated simultaneously. |
The value affects the number of GPU nodes that are temporarily unavailable during an update. In a large cluster, an update might affect many nodes. |
If the value is too high:
If the value is too low, rolling out the update across thousands of nodes might take a long time. |
Use a balanced value that is not overly aggressive. Set the value based on the cluster size and the amount of disruption that workloads can tolerate. |
operator.resources
|
Y |
Defines the CPU and memory resources for the GPU Operator controller. |
The operator manages components including:
As the number of nodes increases, the operator manages more objects and requires more resources. |
If the resources are undersized:
|
|
cdi.enabled
|
Y |
Enables the Container Device Interface. |
|
If the Container Device Interface is disabled, the older runtime integration might:
|
|
mig.strategy
|
Y |
Controls Multi-Instance GPU partitioning. |
Multi-Instance GPU partitioning enables workloads to share GPUs and affects scheduler behavior. |
If this argument is misconfigured:
|
|
dcgmExporter.*
|
Y |
Controls GPU metrics collection. |
The exporter generates GPU metrics for each node, increasing Prometheus load and network traffic as the number of GPU nodes increases. |
If this argument is misconfigured:
|
|
dcgm.*
|
Y |
Controls the NVIDIA Data Center GPU Manager host engine. |
The optional component tracks GPU health across the cluster. |
|
|
skipNodeFeatureDiscoveryDependencyCheck
|
N |
Controls the dependency check for Node Feature Discovery. |
The dependency check affects node labeling and GPU detection. |
If Node Feature Discovery is missing:
|
|
disableNvidiaGpuPlugin
|
Y |
Disables the standalone NVIDIA GPU device plugin. |
Disabling the standalone plugin prevents duplicate GPU device plugins from running in the cluster. |
If this argument is configured incorrectly:
|
Disable the standalone NVIDIA GPU device plugin when using the NVIDIA GPU Operator. |