Configure a Reranking Model

A container administrator can configure a reranking model in the container configuration file.

The following example configures a reranking model that uses a model template:

{
  "models": [
    {
      "name": "mxbai-rerank-base-v2",
      "path": "mixedbread-ai/mxbai-rerank-base-v2",
      "runtime": "vllm",
      "model_template": "mxbai-rerank"
    }
  ]
}

In this case, the model template supplies the runtime arguments and the TEXT_RERANK capability for the model.

The following examples provide sample configurations for reranking models using vLLM, llama.cpp, and ONNX runtimes, respectively:

The following example sends a reranking request to the /v1/rerank API endpoint. It asks the bge-reranker model to score and rank the three listed documents by how relevant they are to the given query.

curl --location 'https://localhost:9091/v1/rerank' \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer $API_KEY" \
  --cacert "$SECRETS_DIR/cert.pem" \
  --data '{
    "model": "bge-reranker",
    "query": "What is the capital of France?",
    "documents": [
      "The capital of Brazil is Brasilia.",
      "The capital of France is Paris.",
      "Horses and cows are both animals."
    ]
}'

For more information about configuring the Private AI Services Container, see Configure the Private AI Services Container.