Repairing GPU Hosts

Find out how to configure and use GPU host repair for managed GPU worker nodes in clusters created with Kubernetes Engine (OKE).

GPU host repair helps you respond to specified hardware-related conditions on eligible managed bare metal GPU worker nodes in managed node pools. When a node is eligible for repair, OKE coordinates the repair workflow. The workflow first makes the node unavailable for new workloads and attempts to drain workloads from it before it performs the applicable Compute action.

GPU host repair supports the following repair types:

  • User-initiated repair: Use this repair type when you identify a condition that requires a GPU worker node to be replaced. OKE coordinates draining the node, reporting the condition to Compute, terminating the affected instance, and restoring managed node pool capacity.
  • GPU instance recovery: Use this repair type when GPU health monitoring reports an eligible critical GPU health signal. OKE coordinates draining the node and reboots the affected instance.

GPU host repair is available for eligible GPU managed worker nodes in clusters running Kubernetes version 1.32 or later. A node is eligible only when its GPU shape is supported and available in the selected region. The functionality does not apply to self-managed nodes, virtual nodes, or non-GPU worker nodes.

GPU host repair uses Kubernetes resources and standard Kubernetes tools. A NodeRepairConfig is a Kubernetes custom resource that selects the managed GPU worker nodes to which a repair policy applies. A NodeHealthReport is a Kubernetes custom resource that records health signals and repair progress for a node. You can inspect a NodeHealthReport, but you must not edit its status. The Node Repair Controller coordinates the repair workflow. For GPU instance recovery, the GPU Fault Monitor reports GPU health signals. Repair coordination uses a Kubernetes Lease to prevent concurrent processing of the same repair request.

Use dry-run mode to verify configuration and observe a proposed repair without cordoning or draining the node, applying a Compute tag, terminating or rebooting the instance, or replacing the node.

Important

GPU host repair can disrupt workloads running on an affected node. Before configuring GPU host repair, design workloads for interruption. Review pod disruption budgets, replica configuration, checkpointing requirements, topology constraints, and available spare GPU capacity.

A repair request with a Completed state indicates that the repair workflow completed. It does not confirm that the underlying GPU fault was corrected. After any repair, verify the node's current health signals and application health.

Configuring User-Initiated Repair for Managed GPU Worker Nodes

Before configuring user-initiated repair, identify an eligible managed GPU node pool and a controlled test node. Ensure that workloads on the selected node can be interrupted.

User-initiated repair requires access control in both OCI IAM and Kubernetes RBAC. In OCI IAM, the OKE cluster resource principal requires permission to apply the defined tag used to report the host condition to Compute. In Kubernetes, only a trusted operator should be able to initiate the protected repair request. The trusted operator requires standard Kubernetes RBAC permissions to update the target Node resource and the OKE-specific authorization delegated through a ClusterRoleBinding.

When you request user-initiated repair, OKE coordinates the following actions:

  1. Ensures only one repair request is processed at a time.
  2. Cordons the node so that Kubernetes does not schedule new pods on it.
  3. Attempts to drain workloads from the node.
  4. Reports the condition to Compute and terminates the affected instance.
  5. Restores the managed node pool’s configured capacity.

To configure and test user-initiated repair:

  1. Verify that the cluster is an enhanced cluster, and is running Kubernetes version 1.32 or later. Verify that the node is an eligible managed GPU worker node. Verify that the OKE-managed Node Repair Controller and NodeHealthReport CRDs, the oke-customer-node-repair-trigger ClusterRole, and the trigger-protection admission policy and binding are installed on the cluster. If any OKE-managed resource is missing, contact Oracle Support. Do not recreate OKE-managed resources.

  2. Before using live repair, open a Support request to enable user-reported hardware-fault handling for your tenancy. Include the tenancy, region, cluster OCID, and GPU node pool in the request.

  3. Create or reuse the defined tag that the Node Repair Controller uses to report the host condition to Compute.

    Create or reuse a tag namespace named ComputeInstanceHostActions. In that namespace, create or reuse a tag key definition named CustomerReportedHostStatus. Do not create a free-form tag with the same text. For more information, see Creating a Tag Namespace and Creating a Tag Key Definition.

    Before terminating the original instance, the Node Repair Controller writes the value unhealthy to the ComputeInstanceHostActions.CustomerReportedHostStatus defined tag on the instance.

  4. Create an OCI IAM policy that grants the OKE cluster resource principal the narrowly scoped permission to apply the defined tag:

    Allow any-user to use tag-namespace in tenancy where all {request.principal.type = 'cluster', request.principal.id = '<cluster-ocid>', target.tag-namespace.name = 'ComputeInstanceHostActions'}

    IAM changes can take time to propagate. Wait for propagation before performing destructive validation.

  5. Delegate the protected user-initiated repair trigger to trusted operators.

    Create a ClusterRoleBinding that binds a trusted OCI IAM group OCID to the OKE-owned oke-customer-node-repair-trigger ClusterRole. Operators also require standard Kubernetes RBAC permissions to update the target Node resource.

    Create a file named cir-trigger-operators.yaml with the following content:

    apiVersion: rbac.authorization.k8s.io/v1
    kind: ClusterRoleBinding
    metadata:
      name: oke-node-repair-trigger-operators
    roleRef:
      apiGroup: rbac.authorization.k8s.io
      kind: ClusterRole
      name: oke-customer-node-repair-trigger
    subjects:
    - apiGroup: rbac.authorization.k8s.io
      kind: Group
      name: <trusted-oci-iam-group-ocid>

    Replace <trusted-oci-iam-group-ocid> with the OCID of the OCI IAM group that contains the trusted operators.

    Apply and verify the binding by entering:

    kubectl apply -f cir-trigger-operators.yaml
    kubectl get clusterrolebinding \
      oke-node-repair-trigger-operators \
      -o yaml

    This binding does not grant general permission to update Node resources. Use your existing Kubernetes authorization model to grant only the standard Kubernetes RBAC permissions required by the trusted operators.

  6. Create a non-overlapping, cluster-scoped NodeRepairConfig custom resource in dry-run mode.

    Select nodes by using a durable node-pool label that you manage and that is present on replacement nodes. The selector label is user-defined. Configure the label on the intended managed node pool so that replacement nodes retain it.

    A node must match exactly one NodeRepairConfig resource. Requests with no matching configuration or overlapping configurations are rejected without repair side effects. You can use the same or separate configurations for user-initiated repair and GPU instance recovery. If you use separate configurations, do not include the same nodes in multiple configurations; otherwise, the repair request will be rejected.

    The following example starts in dry-run mode, allows one repair at a time, sets the unhealthy-node threshold to 10%, uses a 60-minute drain grace period, and disables force-after-grace behavior:

    apiVersion: oci.oraclecloud.com/v1beta1
    kind: NodeRepairConfig
    metadata:
      name: gpu-node-repair
    spec:
      nodeSelector:
        matchLabels:
          example.com/gpu-repair: "enabled"
      policy:
        dryRun: true
        evictionGracePeriodMinutes: 60
        forceActionAfterGracePeriod: false
        maxUnhealthyNodeThreshold: "10%"
        maximumParallelNodeRepairs: 1

    In the policy section, dryRun and forceActionAfterGracePeriod are Boolean values. evictionGracePeriodMinutes is a non-negative integer. maximumParallelNodeRepairs is a positive integer. maxUnhealthyNodeThreshold accepts a positive node count or a quoted percentage, such as "10%".

    Start with dryRun: true. After validating the configuration, update only dryRun to false for live operation.

  7. Request a dry-run repair by applying the protected label oci.oraclecloud.com/customer-report-host-status=unhealthy to the selected node. Inspect the NodeHealthReport and related node state. A dry-run request creates a completed NodeHealthReport record, but does not cordon, drain, apply the Compute defined tag, terminate, reboot, or replace the node.

  8. Trigger a live repair by applying the protected trigger label together with the one-shot live override oci.oraclecloud.com/node-repair-dry-run=false. The live override is valid only with an active trigger, and is sampled only when a new trigger is created.

    To retry a user-initiated repair request, remove and reapply the oci.oraclecloud.com/customer-report-host-status=unhealthy label. Reapplying the same label value does not start another request. Removing the label does not cancel an active repair.

    Important

    A live user-initiated repair can drain workloads and terminates the original Compute instance. Draining uses the Kubernetes Eviction API and honors pod disruption budgets. DaemonSet pods and static pods are skipped. If force-after-grace is enabled, the Node Repair Controller does not force-delete remaining pods, but the subsequent host action can still disrupt them.
  9. Verify the progress of the repair action, the Compute defined tag, termination of the original instance, managed-node-pool capacity restoration, and creation of a replacement node with a new UID and instance OCID. The NodeHealthReport for terminated instances is retained temporarily after the workflow completes.

    A Deferred result means that the workflow is temporarily blocked. A Failed result means that the attempt stopped. A Superseded result means that a higher-priority OKE workflow took over. A Completed result means that OKE finished the requested termination workflow. It does not mean that physical repair is complete or that replacement capacity is Ready. After the workflow completes, verify that the managed node pool has the expected capacity and that the replacement node reports current healthy status.

For user-initiated repair, creation of a replacement worker node depends on managed-node-pool reconciliation and available GPU capacity. A completed workflow does not confirm that the underlying GPU fault was corrected.

Configuring GPU Instance Recovery for Managed GPU Worker Nodes

Before configuring GPU instance recovery, identify an eligible managed GPU node pool and confirm that its workloads can be interrupted. GPU instance recovery can cordon and drain an affected node before rebooting its instance.

Before configuring GPU instance recovery, create the IAM policies required by the GPU Health Monitoring add-on to monitor GPU hosts:

Allow dynamic-group <dg-name> to manage instance-family in compartment id <compartment-ocid>
Allow dynamic-group <dg-name> to read instance-agent-plugins in compartment id <compartment-ocid>

where:

  • <dg-name> is the name of the dynamic group that includes the Compute instances for the GPU worker nodes. For example, using a rule such as ALL {instance.compartment.id = '<compartment-ocid>'}. By applying tags to node pools, you can create more restrictive rules such as:

    ALL {
      instance.compartment.id = 'ocid1.compartment.oc1..examplecompartment',
      tag.Operations.Workload.value = 'gpu-monitoring',
      tag.Operations.Environment.value = 'prod'
    }

    For more information, see Applying Tags to Node Pools.

  • <compartment-ocid> is the OCID of the compartment that contains the GPU worker nodes.

GPU instance recovery is supported only on eligible managed GPU worker nodes using the following GPU shapes:

  • BM.GPU.A10.4
  • BM.GPU.A100-v2.8
  • BM.GPU.B200.8
  • BM.GPU.B300.8
  • BM.GPU.B300.HS.8
  • BM.GPU.B4.8
  • BM.GPU.GB200-v2.4
  • BM.GPU.GB200-v3.4
  • BM.GPU.GB200.4
  • BM.GPU.GB300.4
  • BM.GPU.H100.8
  • BM.GPU.H100T.8
  • BM.GPU.H200-NC.8
  • BM.GPU.H200.8
  • BM.GPU.L40S-NC.4
  • BM.GPU.L40S.4
  • BM.GPU.MI300X.8
  • BM.GPU.MI355X-v1.8
  • BM.GPU.MI355X.8

GPU shape availability varies by region. Confirm that the selected GPU shape is available in the region where the cluster runs.

Oracle Cloud Agent version 1.64.0 or later is required by the Compute RDMA GPU Monitoring plug-in and must be installed on the selected nodes.

When OKE receives an eligible GPU health signal, it coordinates the following actions:

  1. Ensures only one recovery request is processed at a time.
  2. Cordons the node so that Kubernetes does not schedule new pods on it.
  3. Attempts to drain workloads from the node.
  4. Reboots the affected Compute instance.

To configure GPU instance recovery:

  1. Verify that the cluster is running Kubernetes version 1.32 or later. Verify that the node is an eligible managed GPU worker node. Verify that the OKE-managed Node Repair Controller and NodeHealthReport CRDs are installed on the cluster, and verify that the GPU Health Monitoring add-on is available for the cluster.

  2. Verify that the GPU Health Monitoring add-on has an ACTIVE option for the Kubernetes version running on the cluster by entering:

    oci ce addon-option list \
      --kubernetes-version <version> \
      --addon-name GpuHealthMonitoring

    When you install the add-on, install the latest ACTIVE option rather than pinning a version. The installed version must be v1.1.0 or later.

  3. Install the OKE-managed GPU Health Monitoring add-on by entering:

    oci ce cluster install-addon \
      --addon-name GpuHealthMonitoring \
      --cluster-id $CLUSTER_ID \
      --region $REGION

    The command installs the latest available version of the add-on for the cluster.

    The supported add-on configuration keys are nodeSelectors, tolerations, and rollingUpdate. The rollingUpdate value controls rolling update behavior by maxSurge and maxUnavailable, using JSON format in plain text.

  4. Create a GpuHealthMonitoringConfig resource to enable GPU health monitoring for the selected GPU worker nodes. The configuration requires spec.enabled and a non-empty node selector.

    apiVersion: oci.oraclecloud.com/v1beta1
    kind: GpuHealthMonitoringConfig
    metadata:
      name: gpu-health-monitoring
    spec:
      enabled: true
      nodeSelector:
        matchLabels:
          nvidia.com/gpu.present: "true"

    Save the configuration as gpu-health-monitoring.yaml, and apply it by entering:

    kubectl apply -f gpu-health-monitoring.yaml

    The spec.enabled field enables or disables monitoring for the selected nodes. The spec.nodeSelector field selects the GPU nodes on which the Compute RDMA GPU Monitoring plug-in and GPU Fault Monitor run.

    Avoid creating multiple enabled GpuHealthMonitoringConfig resources whose selectors match the same node. Overlapping policies are rejected with a Ready condition that has status set to False and reason set to MultipleMatchingPolicies.

  5. Verify GPU health monitoring configuration on the selected nodes.

    The controller publishes the oci.oraclecloud.com/GpuHealthMonitoringConfigured Node condition. A value of True indicates that the Compute RDMA GPU Monitoring plug-in was configured as enabled. A value of False indicates that it was configured as disabled. A value of Unknown indicates that the configured state cannot currently be trusted.

    The GPU Fault Monitor publishes the oci.oraclecloud.com/GpuHealthMonitoringUnavailable Node condition. A value of True indicates that one or more required GPU health-result files are unavailable. Possible causes include an outdated Oracle Cloud Agent version, unsupported shape, disabled or unhealthy plug-in, or misconfiguration. A value of False indicates that the required result files are present.

  6. Create a non-overlapping, cluster-scoped NodeRepairConfig custom resource in dry-run mode.

    Select nodes by using a durable node-pool label that you manage. The selector label is user-defined.

    A node must match exactly one NodeRepairConfig resource. Requests with no matching configuration or overlapping configurations are rejected without repair side effects. You can use the same or separate configurations for user-initiated repair and GPU instance recovery. If you use separate configurations, do not include the same nodes in multiple configurations; otherwise, the repair request will be rejected.

    The following example starts in dry-run mode, allows one repair at a time, sets the unhealthy-node threshold to 10%, uses a 60-minute drain grace period, and disables force-after-grace behavior:

    apiVersion: oci.oraclecloud.com/v1beta1
    kind: NodeRepairConfig
    metadata:
      name: gpu-node-repair
    spec:
      nodeSelector:
        matchLabels:
          example.com/gpu-repair: "enabled"
      policy:
        dryRun: true
        evictionGracePeriodMinutes: 60
        forceActionAfterGracePeriod: false
        maxUnhealthyNodeThreshold: "10%"
        maximumParallelNodeRepairs: 1

    In the policy section, dryRun and forceActionAfterGracePeriod are Boolean values. evictionGracePeriodMinutes is a non-negative integer. maximumParallelNodeRepairs is a positive integer. maxUnhealthyNodeThreshold accepts a positive node count or a quoted percentage, such as "10%".

    Start with dryRun: true. After validating the configuration, update only dryRun to false for live operation.

    Note that GPU instance recovery is health-signal-driven. Do not apply the user-initiated repair label oci.oraclecloud.com/customer-report-host-status=unhealthy for GPU instance recovery.

  7. Verify that an eligible GPU health signal is reported for the node.

    A node is eligible for automated reboot only when all of the following conditions are met:

    • The Node condition oci.oraclecloud.com/GpuHealthMonitoringFaultDetected is True.
    • The NodeHealthReport contains an entry under status.signals.clusterHealthCheck[].
    • The signal has a non-empty id.
    • The signal has action=REBOOT. The value is case-insensitive.
    • The signal has a non-null firstObservedAt value.
    • The signal has a non-null lastObservedAt value, indicating that the signal is active.
    • The signal has remained active for at least 10 minutes.
    • Node Repair Controller retry and cooldown rules permit another attempt. The repeat reboot cooldown is 60 minutes.
    • The effective dry-run value is false for an actual reboot.
    • Subsequent admission and disruption guardrails pass.

    Signals with action=CUSTOMER_ACTION_REQUIRED do not trigger an automated reboot. If you see one of these signals, review the NodeHealthReport and node conditions, investigate the reported GPU health issue, and contact Support if you need help resolving it.

  8. Enable live repair in the NodeRepairConfig resource.

    GPU instance recovery starts only after the eligible signal has remained active for at least 10 minutes, and the repeat reboot cooldown is 60 minutes. Drain behavior follows the repair policy specified in the matching NodeRepairConfig resource. Draining uses the Kubernetes Eviction API and honors pod disruption budgets. The time required to complete GPU instance recovery can vary, depending on workload disruption settings, node drain behavior, and Compute instance state.

  9. Verify that GPU instance recovery completed successfully.

    Confirm that the NodeHealthReport shows a Reboot action in live mode and has a terminal status of Completed. Confirm that the Compute instance is in the Running state, the Kubernetes node is Ready, and the current GPU health result no longer reports the fault.

  10. Review the GPU health result after the workflow completes.

    If the GPU health result still reports the same fault after the reboot, or reports the fault again, the fault has not been resolved. A completed recovery workflow does not mean that the underlying GPU fault was corrected. GPU instance recovery uses cooldown rules to prevent immediate reboot loops. After the GPU health result stops reporting the fault, a later GPU health result that reports a new fault can make the node eligible for another repair.

Important

The generated GpuHealthMonitoringAction resource is controller-owned. Do not create or edit GpuHealthMonitoringAction resources.

Monitoring and Troubleshooting GPU Host Repair for GPU Worker Nodes

Use Kubernetes resources and standard Kubernetes monitoring tools to inspect the NodeHealthReport, Node conditions, affected workloads, and repair workflow.

Important

A NodeHealthReport with a Completed state indicates that the repair workflow completed. It does not indicate that the underlying GPU fault was fixed. Verify current health signals, the worker node's readiness, and application health after the workflow completes.

To monitor a GPU host repair workflow:

  1. Inspect the NodeHealthReport and the affected worker node before, during, and after repair.

    Review the NodeHealthReport for the repair request, repair progress, action, dry-run or live mode, and workflow result. Review the affected Kubernetes Node resource for node labels and conditions. Review Kubernetes Events for recent node and repair activity. For GPU instance recovery, check the GpuHealthMonitoringConfigured, GpuHealthMonitoringUnavailable, and GpuHealthMonitoringFaultDetected conditions on the affected Kubernetes Node resource.

    Treat controller-owned resources as read-only. Do not create or edit generated actions or controller-owned resources.

  2. Review the supported repair status and node state.

    Use the NodeHealthReport to review the repair policy, effective dry-run value, admission result, terminal result, and last repair attempt. Use the affected Kubernetes Node resource to review node conditions. Where applicable, verify the Compute instance lifecycle state, the Compute defined tag used for user-initiated repair, and whether replacement capacity is ready.

    Do not rely on internal leases, request IDs, file paths, polling intervals, or intermediate workflow phases as status indicators.

  3. For user-initiated repair, investigate rejected requests, deferred repair, failed repair, and superseded repair.

    A request can be rejected when no NodeRepairConfig resource matches the node or when more than one configuration applies.

    Use the NodeHealthReport result to decide what to investigate:

    • If the NodeHealthReport shows AdmissionRejected, verify the NodeRepairConfig resource, trigger authorization, and label configuration.
    • If the NodeHealthReport shows Deferred, wait for the blocking condition to clear, or adjust capacity if replacement capacity is unavailable.
    • If the NodeHealthReport shows Failed, inspect the NodeHealthReport, Kubernetes Events, pod disruption budgets, IAM policy, Compute defined tag, and Compute instance state.
    • If the NodeHealthReport shows Superseded, follow the OKE workflow that took over the repair.
    • If the NodeHealthReport shows Completed, verify that the original instance was terminated and that replacement capacity is ready.
  4. For GPU instance recovery, investigate GPU health monitoring configuration, unavailable health results, reboot signals, signals that require action, incomplete drains, completed reboots with an ongoing fault, and continuing signals that do not trigger repeated reboots.

    Use node conditions, GPU health signals, and the NodeHealthReport result to decide what to investigate:

    • If the GpuHealthMonitoringConfigured node condition is False or Unknown, verify the Oracle Cloud Agent configuration and the GpuHealthMonitoring add-on configuration.
    • If the GpuHealthMonitoringUnavailable node condition is True, verify that the GPU shape is supported, Oracle Cloud Agent version 1.64.0 or later is installed, the Compute RDMA GPU Monitoring plug-in is enabled, and the GPU Health Monitoring add-on is installed and configured.
    • If an active GPU health signal has action=REBOOT, monitor the NodeHealthReport for the GPU instance recovery workflow.
    • If a GPU health signal has action=CUSTOMER_ACTION_REQUIRED, investigate the reported GPU health issue. This signal does not trigger an automated reboot.
    • If the NodeHealthReport shows an incomplete drain or failed repair, inspect Kubernetes Events, pod disruption budgets, workload state, and node conditions.
    • If the reboot completed but the GPU health result still reports the fault, investigate the ongoing fault.
    • If the same signal continues after reboot, GPU instance recovery does not immediately reboot the node again. Cooldown rules prevent immediate reboot loops.
  5. Before opening a service request, collect the information needed to investigate the repair workflow.

    Include the tenancy OCID, cluster OCID, region, Kubernetes version, GPU Health Monitoring add-on version, Oracle Cloud Agent version, affected worker node name, Compute instance OCID, relevant NodeHealthReport and NodeRepairConfig resources, node labels and conditions, Kubernetes Events, pod disruption budget and workload state, timestamps, and current GPU health results.

    Do not modify OKE-owned CRDs, controller resources, admission objects, generated actions, or coordination resources.

Troubleshooting GPU Host Repair

IssueAction
No repair starts after a user-initiated request.Check that exactly one NodeRepairConfig selects the node, confirm that the request was made by an authorized operator, and inspect the NodeHealthReport admission result.
The node remains cordoned or workloads are not drained.Review pod disruption budgets, pods that cannot be evicted, workload scheduling constraints, and available replacement capacity.
The managed node pool does not return to its configured capacity after user-initiated repair.Check GPU shape capacity in the selected region, service limits, compartment quotas, and the managed node pool's state.
A GPU health signal does not start GPU instance recovery.Confirm that the signal and node are eligible, and verify the GPU Health Monitoring add-on and NodeRepairConfig.
The GPU worker node does not become ready after GPU instance recovery.Check Node conditions, current GPU health signals, and application health. A completed workflow does not confirm that the underlying GPU fault was resolved.

Use Kubernetes resources, node conditions, and Kubernetes Events as troubleshooting evidence. If you open a service request, Oracle Support can retrieve service-side logs as needed.