Managing Autoscaling

Learn how to validate, monitor, scale, and remove autoscaling nodes.

Control-Plane Scheduling When Static Workers Are Disabled

Autoscaling is supported on OpenShift Container Platform 4.22 and later. Use an RHCOS worker image that matches the OpenShift Container Platform minor version of the cluster.

When a cluster is installed with autoscaling enabled and no static workers, control-plane nodes can receive the worker role and remain schedulable. Workloads scheduled on control-plane nodes can satisfy pending demand and prevent autoscaler scale-up. After dedicated autoscaled workers are ready, follow Red Hat solution 4564851 to remove the worker role from control-plane nodes.

Validating Autoscaling

From the Console, use the following commands to validate that your autoscaling nodes are operating:

oc get ns oci-openshift-autoscaling-operator
oc get pods -n oci-openshift-autoscaling-operator

Monitoring Autoscaling

To investigate or monitor the instance provisioning process:

  1. Export ENV:
    export STACK_NAMESPACE=oci-openshift-autoscaling-operator
    export AUTOSCALER_NAME=ociclusterautoscaler
  2. Check the provisioning chain:
    oc get ociclusterautoscaler -n ${STACK_NAMESPACE} ${AUTOSCALER_NAME} -o yaml
    oc get machinedeployments.cluster.x-k8s.io -n ${STACK_NAMESPACE}
    oc get machinesets.cluster.x-k8s.io -n ${STACK_NAMESPACE}
    oc get machines.cluster.x-k8s.io -n ${STACK_NAMESPACE} -o wide
    oc get ocimachines.infrastructure.cluster.x-k8s.io -n ${STACK_NAMESPACE} -o wide
    oc get nodes -o wide
  3. Check events for provisioning failures:
    oc get events -n ${STACK_NAMESPACE} --sort-by=.lastTimestamp
    
    ## For a specific stuck Machine:
    
    export MACHINE_NAME=<machine-name>
    oc describe machine.cluster.x-k8s.io -n ${STACK_NAMESPACE} ${MACHINE_NAME}
    oc get event -n ${STACK_NAMESPACE} \
      --field-selector involvedObject.name=${MACHINE_NAME} \
      --sort-by=.lastTimestamp
    
    ## Find the related OCIMachine:
    
    oc get machine.cluster.x-k8s.io -n ${STACK_NAMESPACE} ${MACHINE_NAME} \
      -o jsonpath='{.spec.infrastructureRef.name}{"\n"}'
    
    export OCI_MACHINE_NAME=<ocimachine-name>
    oc describe ocimachine.infrastructure.cluster.x-k8s.io -n ${STACK_NAMESPACE} ${OCI_MACHINE_NAME}
    oc get ocimachine.infrastructure.cluster.x-k8s.io -n ${STACK_NAMESPACE} ${OCI_MACHINE_NAME} -o yaml
  4. Check controller logs:
    # OCI CAPI operator logs
    oc logs -n ${STACK_NAMESPACE} \
      deployment/oci-capi-operator-controller-manager \
      -c manager \
      --tail=300
    
    # CAPOCI provider logs: most useful for OCI provisioning errors
    oc logs -n ${STACK_NAMESPACE} \
      deployment/capoci-controller-manager \
      --tail=500
    
    # CAPI controller logs
    oc logs -n ${STACK_NAMESPACE} \
      deployment/capi-manager \
      --tail=500
    
    # Cluster Autoscaler logs: useful if no Machine was created
    oc logs -n ${STACK_NAMESPACE} \
      deployment/oci-cluster-autoscaler \
      --tail=500

These are the most common failure signals:

# No new Machine:
oc logs -n ${STACK_NAMESPACE} deployment/oci-cluster-autoscaler --tail=500

# Machine exists, but OCIMachine has no providerID:
oc describe ocimachine.infrastructure.cluster.x-k8s.io -n ${STACK_NAMESPACE} ${OCI_MACHINE_NAME}
oc logs -n ${STACK_NAMESPACE} deployment/capoci-controller-manager --tail=500

# OCI instance exists, but no OpenShift node joined:
oc get csr
oc get nodes
oc get mcp
oc get co machine-config

If an instance is provisioned but not able to join the cluster, in the OCI Console, select OS Management and check the Console History. Here you can review OS-level error messages.

Scaling Up or Down

Use the following commands to monitor the scaling up and down of your autoscaling resources:

oc get machinedeployments.cluster.x-k8s.io -n oci-openshift-autoscaling-operator
oc get machinesets.cluster.x-k8s.io -n oci-openshift-autoscaling-operator
oc get machines.cluster.x-k8s.io -n oci-openshift-autoscaling-operator -o wide
oc get nodes

To modify the minimum and/or maximum number of nodes in your autoscaling instance:

  1. Verify the current minimum and maximum values:
    oc get ociclusterautoscaler.capi.openshift.io -n oci-openshift-autoscaling-operator ociclusterautoscaler \
    	-o jsonpath='{.spec.autoscaling.minNodes}{" "}{.spec.autoscaling.maxNodes}{"\n"}'
  2. Set new minimum and/or maximum values with this command. In this example, a new minimum of 1 and maximum of 3 nodes is set. Change the values according to your needs:
    oc patch ociclusterautoscaler.capi.openshift.io -n oci-openshift-autoscaling-operator ociclusterautoscaler \
    	--type=merge \
    	-p '{"spec":{"autoscaling":{"minNodes":1,"maxNodes":3}}}'