Repairing GPU Hosts
Find out how to configure and use GPU host repair for managed GPU worker nodes in clusters created with Kubernetes Engine (OKE).
GPU host repair helps you respond to specified hardware-related conditions on eligible managed bare metal GPU worker nodes in managed node pools. When a node is eligible for repair, OKE coordinates the repair workflow. The workflow first makes the node unavailable for new workloads and attempts to drain workloads from it before it performs the applicable Compute action.
GPU host repair supports the following repair types:
- User-initiated repair: Use this repair type when you identify a condition that requires a GPU worker node to be replaced. OKE coordinates draining the node, reporting the condition to Compute, terminating the affected instance, and restoring managed node pool capacity.
- GPU instance recovery: Use this repair type when GPU health monitoring reports an eligible critical GPU health signal. OKE coordinates draining the node and reboots the affected instance.
GPU host repair is available for eligible GPU managed worker nodes in clusters running Kubernetes version 1.32 or later. A node is eligible only when its GPU shape is supported and available in the selected region. The functionality does not apply to self-managed nodes, virtual nodes, or non-GPU worker nodes.
GPU host repair uses Kubernetes resources and standard Kubernetes tools. A NodeRepairConfig is a Kubernetes custom resource that selects the managed GPU worker nodes to which a repair policy applies. A NodeHealthReport is a Kubernetes custom resource that records health signals and repair progress for a node. You can inspect a NodeHealthReport, but you must not edit its status. The Node Repair Controller coordinates the repair workflow. For GPU instance recovery, the GPU Fault Monitor reports GPU health signals. Repair coordination uses a Kubernetes Lease to prevent concurrent processing of the same repair request.
Use dry-run mode to verify configuration and observe a proposed repair without cordoning or draining the node, applying a Compute tag, terminating or rebooting the instance, or replacing the node.
GPU host repair can disrupt workloads running on an affected node. Before configuring GPU host repair, design workloads for interruption. Review pod disruption budgets, replica configuration, checkpointing requirements, topology constraints, and available spare GPU capacity.
A repair request with a Completed state indicates that the repair workflow completed. It does not confirm that the underlying GPU fault was corrected. After any repair, verify the node's current health signals and application health.
Configuring User-Initiated Repair for Managed GPU Worker Nodes
Before configuring user-initiated repair, identify an eligible managed GPU node pool and a controlled test node. Ensure that workloads on the selected node can be interrupted.
User-initiated repair requires access control in both OCI IAM and Kubernetes RBAC. In OCI IAM, the OKE cluster resource principal requires permission to apply the defined tag used to report the host condition to Compute. In Kubernetes, only a trusted operator should be able to initiate the protected repair request. The trusted operator requires standard Kubernetes RBAC permissions to update the target Node resource and the OKE-specific authorization delegated through a ClusterRoleBinding.
When you request user-initiated repair, OKE coordinates the following actions:
- Ensures only one repair request is processed at a time.
- Cordons the node so that Kubernetes does not schedule new pods on it.
- Attempts to drain workloads from the node.
- Reports the condition to Compute and terminates the affected instance.
- Restores the managed node pool’s configured capacity.
To configure and test user-initiated repair:
-
Verify that the cluster is an enhanced cluster, and is running Kubernetes version 1.32 or later. Verify that the node is an eligible managed GPU worker node. Verify that the OKE-managed Node Repair Controller and
NodeHealthReportCRDs, theoke-customer-node-repair-triggerClusterRole, and the trigger-protection admission policy and binding are installed on the cluster. If any OKE-managed resource is missing, contact Oracle Support. Do not recreate OKE-managed resources. Before using live repair, open a Support request to enable user-reported hardware-fault handling for your tenancy. Include the tenancy, region, cluster OCID, and GPU node pool in the request.
-
Create or reuse the defined tag that the Node Repair Controller uses to report the host condition to Compute.
Create or reuse a tag namespace named
ComputeInstanceHostActions. In that namespace, create or reuse a tag key definition namedCustomerReportedHostStatus. Do not create a free-form tag with the same text. For more information, see Creating a Tag Namespace and Creating a Tag Key Definition.Before terminating the original instance, the Node Repair Controller writes the value
unhealthyto theComputeInstanceHostActions.CustomerReportedHostStatusdefined tag on the instance. Create an OCI IAM policy that grants the OKE cluster resource principal the narrowly scoped permission to apply the defined tag:
Allow any-user to use tag-namespace in tenancy where all {request.principal.type = 'cluster', request.principal.id = '<cluster-ocid>', target.tag-namespace.name = 'ComputeInstanceHostActions'}IAM changes can take time to propagate. Wait for propagation before performing destructive validation.
-
Delegate the protected user-initiated repair trigger to trusted operators.
Create a
ClusterRoleBindingthat binds a trusted OCI IAM group OCID to the OKE-ownedoke-customer-node-repair-triggerClusterRole. Operators also require standard Kubernetes RBAC permissions to update the targetNoderesource.Create a file named
cir-trigger-operators.yamlwith the following content:apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: oke-node-repair-trigger-operators roleRef: apiGroup: rbac.authorization.k8s.io kind: ClusterRole name: oke-customer-node-repair-trigger subjects: - apiGroup: rbac.authorization.k8s.io kind: Group name: <trusted-oci-iam-group-ocid>Replace
<trusted-oci-iam-group-ocid>with the OCID of the OCI IAM group that contains the trusted operators.Apply and verify the binding by entering:
kubectl apply -f cir-trigger-operators.yamlkubectl get clusterrolebinding \ oke-node-repair-trigger-operators \ -o yamlThis binding does not grant general permission to update
Noderesources. Use your existing Kubernetes authorization model to grant only the standard Kubernetes RBAC permissions required by the trusted operators. -
Create a non-overlapping, cluster-scoped
NodeRepairConfigcustom resource in dry-run mode.Select nodes by using a durable node-pool label that you manage and that is present on replacement nodes. The selector label is user-defined. Configure the label on the intended managed node pool so that replacement nodes retain it.
A node must match exactly one
NodeRepairConfigresource. Requests with no matching configuration or overlapping configurations are rejected without repair side effects. You can use the same or separate configurations for user-initiated repair and GPU instance recovery. If you use separate configurations, do not include the same nodes in multiple configurations; otherwise, the repair request will be rejected.The following example starts in dry-run mode, allows one repair at a time, sets the unhealthy-node threshold to 10%, uses a 60-minute drain grace period, and disables force-after-grace behavior:
apiVersion: oci.oraclecloud.com/v1beta1 kind: NodeRepairConfig metadata: name: gpu-node-repair spec: nodeSelector: matchLabels: example.com/gpu-repair: "enabled" policy: dryRun: true evictionGracePeriodMinutes: 60 forceActionAfterGracePeriod: false maxUnhealthyNodeThreshold: "10%" maximumParallelNodeRepairs: 1In the
policysection,dryRunandforceActionAfterGracePeriodare Boolean values.evictionGracePeriodMinutesis a non-negative integer.maximumParallelNodeRepairsis a positive integer.maxUnhealthyNodeThresholdaccepts a positive node count or a quoted percentage, such as"10%".Start with
dryRun: true. After validating the configuration, update onlydryRuntofalsefor live operation. -
Request a dry-run repair by applying the protected label
oci.oraclecloud.com/customer-report-host-status=unhealthyto the selected node. Inspect theNodeHealthReportand related node state. A dry-run request creates a completedNodeHealthReportrecord, but does not cordon, drain, apply the Compute defined tag, terminate, reboot, or replace the node. Trigger a live repair by applying the protected trigger label together with the one-shot live override
oci.oraclecloud.com/node-repair-dry-run=false. The live override is valid only with an active trigger, and is sampled only when a new trigger is created.To retry a user-initiated repair request, remove and reapply the
oci.oraclecloud.com/customer-report-host-status=unhealthylabel. Reapplying the same label value does not start another request. Removing the label does not cancel an active repair.Important
A live user-initiated repair can drain workloads and terminates the original Compute instance. Draining uses the Kubernetes Eviction API and honors pod disruption budgets. DaemonSet pods and static pods are skipped. If force-after-grace is enabled, the Node Repair Controller does not force-delete remaining pods, but the subsequent host action can still disrupt them.Verify the progress of the repair action, the Compute defined tag, termination of the original instance, managed-node-pool capacity restoration, and creation of a replacement node with a new UID and instance OCID. The
NodeHealthReportfor terminated instances is retained temporarily after the workflow completes.A
Deferredresult means that the workflow is temporarily blocked. AFailedresult means that the attempt stopped. ASupersededresult means that a higher-priority OKE workflow took over. ACompletedresult means that OKE finished the requested termination workflow. It does not mean that physical repair is complete or that replacement capacity isReady. After the workflow completes, verify that the managed node pool has the expected capacity and that the replacement node reports current healthy status.
For user-initiated repair, creation of a replacement worker node depends on managed-node-pool reconciliation and available GPU capacity. A completed workflow does not confirm that the underlying GPU fault was corrected.
Configuring GPU Instance Recovery for Managed GPU Worker Nodes
Before configuring GPU instance recovery, identify an eligible managed GPU node pool and confirm that its workloads can be interrupted. GPU instance recovery can cordon and drain an affected node before rebooting its instance.
Before configuring GPU instance recovery, create the IAM policies required by the GPU Health Monitoring add-on to monitor GPU hosts:
Allow dynamic-group <dg-name> to manage instance-family in compartment id <compartment-ocid>
Allow dynamic-group <dg-name> to read instance-agent-plugins in compartment id <compartment-ocid>
where:
-
<dg-name>is the name of the dynamic group that includes the Compute instances for the GPU worker nodes. For example, using a rule such asALL {instance.compartment.id = '<compartment-ocid>'}. By applying tags to node pools, you can create more restrictive rules such as:ALL { instance.compartment.id = 'ocid1.compartment.oc1..examplecompartment', tag.Operations.Workload.value = 'gpu-monitoring', tag.Operations.Environment.value = 'prod' }For more information, see Applying Tags to Node Pools.
<compartment-ocid>is the OCID of the compartment that contains the GPU worker nodes.
GPU instance recovery is supported only on eligible managed GPU worker nodes using the following GPU shapes:
BM.GPU.A10.4BM.GPU.A100-v2.8BM.GPU.B200.8BM.GPU.B300.8BM.GPU.B300.HS.8BM.GPU.B4.8BM.GPU.GB200-v2.4BM.GPU.GB200-v3.4BM.GPU.GB200.4BM.GPU.GB300.4BM.GPU.H100.8BM.GPU.H100T.8BM.GPU.H200-NC.8BM.GPU.H200.8BM.GPU.L40S-NC.4BM.GPU.L40S.4BM.GPU.MI300X.8BM.GPU.MI355X-v1.8BM.GPU.MI355X.8
GPU shape availability varies by region. Confirm that the selected GPU shape is available in the region where the cluster runs.
Oracle Cloud Agent version 1.64.0 or later is required by the Compute RDMA GPU Monitoring plug-in and must be installed on the selected nodes.
When OKE receives an eligible GPU health signal, it coordinates the following actions:
- Ensures only one recovery request is processed at a time.
- Cordons the node so that Kubernetes does not schedule new pods on it.
- Attempts to drain workloads from the node.
- Reboots the affected Compute instance.
To configure GPU instance recovery:
-
Verify that the cluster is running Kubernetes version 1.32 or later. Verify that the node is an eligible managed GPU worker node. Verify that the OKE-managed Node Repair Controller and
NodeHealthReportCRDs are installed on the cluster, and verify that the GPU Health Monitoring add-on is available for the cluster. Verify that the GPU Health Monitoring add-on has an
ACTIVEoption for the Kubernetes version running on the cluster by entering:oci ce addon-option list \ --kubernetes-version <version> \ --addon-name GpuHealthMonitoringWhen you install the add-on, install the latest
ACTIVEoption rather than pinning a version. The installed version must bev1.1.0or later.Install the OKE-managed GPU Health Monitoring add-on by entering:
oci ce cluster install-addon \ --addon-name GpuHealthMonitoring \ --cluster-id $CLUSTER_ID \ --region $REGIONThe command installs the latest available version of the add-on for the cluster.
The supported add-on configuration keys are
nodeSelectors,tolerations, androllingUpdate. TherollingUpdatevalue controls rolling update behavior bymaxSurgeandmaxUnavailable, using JSON format in plain text.-
Create a
GpuHealthMonitoringConfigresource to enable GPU health monitoring for the selected GPU worker nodes. The configuration requiresspec.enabledand a non-empty node selector.apiVersion: oci.oraclecloud.com/v1beta1 kind: GpuHealthMonitoringConfig metadata: name: gpu-health-monitoring spec: enabled: true nodeSelector: matchLabels: nvidia.com/gpu.present: "true"Save the configuration as
gpu-health-monitoring.yaml, and apply it by entering:kubectl apply -f gpu-health-monitoring.yamlThe
spec.enabledfield enables or disables monitoring for the selected nodes. Thespec.nodeSelectorfield selects the GPU nodes on which the Compute RDMA GPU Monitoring plug-in and GPU Fault Monitor run.Avoid creating multiple enabled
GpuHealthMonitoringConfigresources whose selectors match the same node. Overlapping policies are rejected with aReadycondition that hasstatusset toFalseandreasonset toMultipleMatchingPolicies. Verify GPU health monitoring configuration on the selected nodes.
The controller publishes the
oci.oraclecloud.com/GpuHealthMonitoringConfiguredNode condition. A value ofTrueindicates that the Compute RDMA GPU Monitoring plug-in was configured as enabled. A value ofFalseindicates that it was configured as disabled. A value ofUnknownindicates that the configured state cannot currently be trusted.The GPU Fault Monitor publishes the
oci.oraclecloud.com/GpuHealthMonitoringUnavailableNode condition. A value ofTrueindicates that one or more required GPU health-result files are unavailable. Possible causes include an outdated Oracle Cloud Agent version, unsupported shape, disabled or unhealthy plug-in, or misconfiguration. A value ofFalseindicates that the required result files are present.-
Create a non-overlapping, cluster-scoped
NodeRepairConfigcustom resource in dry-run mode.Select nodes by using a durable node-pool label that you manage. The selector label is user-defined.
A node must match exactly one
NodeRepairConfigresource. Requests with no matching configuration or overlapping configurations are rejected without repair side effects. You can use the same or separate configurations for user-initiated repair and GPU instance recovery. If you use separate configurations, do not include the same nodes in multiple configurations; otherwise, the repair request will be rejected.The following example starts in dry-run mode, allows one repair at a time, sets the unhealthy-node threshold to 10%, uses a 60-minute drain grace period, and disables force-after-grace behavior:
apiVersion: oci.oraclecloud.com/v1beta1 kind: NodeRepairConfig metadata: name: gpu-node-repair spec: nodeSelector: matchLabels: example.com/gpu-repair: "enabled" policy: dryRun: true evictionGracePeriodMinutes: 60 forceActionAfterGracePeriod: false maxUnhealthyNodeThreshold: "10%" maximumParallelNodeRepairs: 1In the
policysection,dryRunandforceActionAfterGracePeriodare Boolean values.evictionGracePeriodMinutesis a non-negative integer.maximumParallelNodeRepairsis a positive integer.maxUnhealthyNodeThresholdaccepts a positive node count or a quoted percentage, such as"10%".Start with
dryRun: true. After validating the configuration, update onlydryRuntofalsefor live operation.Note that GPU instance recovery is health-signal-driven. Do not apply the user-initiated repair label
oci.oraclecloud.com/customer-report-host-status=unhealthyfor GPU instance recovery. Verify that an eligible GPU health signal is reported for the node.
A node is eligible for automated reboot only when all of the following conditions are met:
- The Node condition
oci.oraclecloud.com/GpuHealthMonitoringFaultDetectedisTrue. - The
NodeHealthReportcontains an entry understatus.signals.clusterHealthCheck[]. - The signal has a non-empty
id. - The signal has
action=REBOOT. The value is case-insensitive. - The signal has a non-null
firstObservedAtvalue. - The signal has a non-null
lastObservedAtvalue, indicating that the signal is active. - The signal has remained active for at least 10 minutes.
- Node Repair Controller retry and cooldown rules permit another attempt. The repeat reboot cooldown is 60 minutes.
- The effective dry-run value is
falsefor an actual reboot. - Subsequent admission and disruption guardrails pass.
Signals with
action=CUSTOMER_ACTION_REQUIREDdo not trigger an automated reboot. If you see one of these signals, review theNodeHealthReportand node conditions, investigate the reported GPU health issue, and contact Support if you need help resolving it.- The Node condition
-
Enable live repair in the
NodeRepairConfigresource.GPU instance recovery starts only after the eligible signal has remained active for at least 10 minutes, and the repeat reboot cooldown is 60 minutes. Drain behavior follows the repair policy specified in the matching
NodeRepairConfigresource. Draining uses the Kubernetes Eviction API and honors pod disruption budgets. The time required to complete GPU instance recovery can vary, depending on workload disruption settings, node drain behavior, and Compute instance state. -
Verify that GPU instance recovery completed successfully.
Confirm that the
NodeHealthReportshows aRebootaction in live mode and has a terminal status ofCompleted. Confirm that the Compute instance is in theRunningstate, the Kubernetes node isReady, and the current GPU health result no longer reports the fault. Review the GPU health result after the workflow completes.
If the GPU health result still reports the same fault after the reboot, or reports the fault again, the fault has not been resolved. A completed recovery workflow does not mean that the underlying GPU fault was corrected. GPU instance recovery uses cooldown rules to prevent immediate reboot loops. After the GPU health result stops reporting the fault, a later GPU health result that reports a new fault can make the node eligible for another repair.
The generated
GpuHealthMonitoringAction resource is controller-owned. Do not create or edit GpuHealthMonitoringAction resources.Monitoring and Troubleshooting GPU Host Repair for GPU Worker Nodes
Use Kubernetes resources and standard Kubernetes monitoring tools to inspect the NodeHealthReport, Node conditions, affected workloads, and repair workflow.
A
NodeHealthReport with a Completed state indicates that the repair workflow completed. It does not indicate that the underlying GPU fault was fixed. Verify current health signals, the worker node's readiness, and application health after the workflow completes.To monitor a GPU host repair workflow:
Inspect the
NodeHealthReportand the affected worker node before, during, and after repair.Review the
NodeHealthReportfor the repair request, repair progress, action, dry-run or live mode, and workflow result. Review the affected KubernetesNoderesource for node labels and conditions. Review Kubernetes Events for recent node and repair activity. For GPU instance recovery, check theGpuHealthMonitoringConfigured,GpuHealthMonitoringUnavailable, andGpuHealthMonitoringFaultDetectedconditions on the affected KubernetesNoderesource.Treat controller-owned resources as read-only. Do not create or edit generated actions or controller-owned resources.
Review the supported repair status and node state.
Use the
NodeHealthReportto review the repair policy, effective dry-run value, admission result, terminal result, and last repair attempt. Use the affected KubernetesNoderesource to review node conditions. Where applicable, verify the Compute instance lifecycle state, the Compute defined tag used for user-initiated repair, and whether replacement capacity is ready.Do not rely on internal leases, request IDs, file paths, polling intervals, or intermediate workflow phases as status indicators.
-
For user-initiated repair, investigate rejected requests, deferred repair, failed repair, and superseded repair.
A request can be rejected when no
NodeRepairConfigresource matches the node or when more than one configuration applies.Use the
NodeHealthReportresult to decide what to investigate:- If the
NodeHealthReportshowsAdmissionRejected, verify theNodeRepairConfigresource, trigger authorization, and label configuration. - If the
NodeHealthReportshowsDeferred, wait for the blocking condition to clear, or adjust capacity if replacement capacity is unavailable. - If the
NodeHealthReportshowsFailed, inspect theNodeHealthReport, Kubernetes Events, pod disruption budgets, IAM policy, Compute defined tag, and Compute instance state. - If the
NodeHealthReportshowsSuperseded, follow the OKE workflow that took over the repair. - If the
NodeHealthReportshowsCompleted, verify that the original instance was terminated and that replacement capacity is ready.
- If the
-
For GPU instance recovery, investigate GPU health monitoring configuration, unavailable health results, reboot signals, signals that require action, incomplete drains, completed reboots with an ongoing fault, and continuing signals that do not trigger repeated reboots.
Use node conditions, GPU health signals, and the
NodeHealthReportresult to decide what to investigate:- If the
GpuHealthMonitoringConfigurednode condition isFalseorUnknown, verify the Oracle Cloud Agent configuration and theGpuHealthMonitoringadd-on configuration. - If the
GpuHealthMonitoringUnavailablenode condition isTrue, verify that the GPU shape is supported, Oracle Cloud Agent version1.64.0or later is installed, the Compute RDMA GPU Monitoring plug-in is enabled, and the GPU Health Monitoring add-on is installed and configured. - If an active GPU health signal has
action=REBOOT, monitor theNodeHealthReportfor the GPU instance recovery workflow. - If a GPU health signal has
action=CUSTOMER_ACTION_REQUIRED, investigate the reported GPU health issue. This signal does not trigger an automated reboot. - If the
NodeHealthReportshows an incomplete drain or failed repair, inspect Kubernetes Events, pod disruption budgets, workload state, and node conditions. - If the reboot completed but the GPU health result still reports the fault, investigate the ongoing fault.
- If the same signal continues after reboot, GPU instance recovery does not immediately reboot the node again. Cooldown rules prevent immediate reboot loops.
- If the
-
Before opening a service request, collect the information needed to investigate the repair workflow.
Include the tenancy OCID, cluster OCID, region, Kubernetes version, GPU Health Monitoring add-on version, Oracle Cloud Agent version, affected worker node name, Compute instance OCID, relevant
NodeHealthReportandNodeRepairConfigresources, node labels and conditions, Kubernetes Events, pod disruption budget and workload state, timestamps, and current GPU health results.Do not modify OKE-owned CRDs, controller resources, admission objects, generated actions, or coordination resources.
Troubleshooting GPU Host Repair
| Issue | Action |
|---|---|
| No repair starts after a user-initiated request. | Check that exactly one NodeRepairConfig selects the node, confirm that the request was made by an authorized operator, and inspect the NodeHealthReport admission result. |
| The node remains cordoned or workloads are not drained. | Review pod disruption budgets, pods that cannot be evicted, workload scheduling constraints, and available replacement capacity. |
| The managed node pool does not return to its configured capacity after user-initiated repair. | Check GPU shape capacity in the selected region, service limits, compartment quotas, and the managed node pool's state. |
| A GPU health signal does not start GPU instance recovery. | Confirm that the signal and node are eligible, and verify the GPU Health Monitoring add-on and NodeRepairConfig. |
| The GPU worker node does not become ready after GPU instance recovery. | Check Node conditions, current GPU health signals, and application health. A completed workflow does not confirm that the underlying GPU fault was resolved. |
Use Kubernetes resources, node conditions, and Kubernetes Events as troubleshooting evidence. If you open a service request, Oracle Support can retrieve service-side logs as needed.