Configure the Private Large Language Model Service

Models are associated with a runtime engine using the service configuration file, which is config.json by default. Included are examples of configuration files that are set up to use Llama.cpp or vLLM, along with an example chat completion request.

For LLMs not shipped with the container, for example gemma-4-31B-it-GGUF, the model can either be provided in your local file system, as a PAR link, or can be downloaded from Hugging Face. If the model is not found in the local file system or provided as a PAR link, the runtime engine will attempt to download the model from the Hugging Face repository based on the model info in the configuration file.

You can choose to explicitly download a model yourself from Hugging Face rather than have the container download the model for you:

wget https://huggingface.co/unsloth/gemma-3-1b-it-GGUF/resolve/main/gemma-3-1b-it-Q8_0.gguf
cp gemma-3-1b-it-Q8_0.gguf /home/opc/models

In this example, the configuration file would look like the following:

{
  "models": [
    {
      "name":"gemma-3-1b-it-Q8_0",
      "path":"gemma-3-1b-it-Q8_0.gguf",
      "runtime": "llamacpp",
      "context_size":4096,
      "capabilities":["TEXT_GENERATION"]
    }
  ]
}

The following example sends a test text prompt to the local LLM endpoint. A successful JSON response containing an assistant message confirms that the endpoint, TLS certificate, API key, and specified model are working together.

curl --noproxy '*' \
  -X POST \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer $API_KEY" \
  --cacert "$SECRETS_DIR/cert.pem" \
  --data '{
    "model": "gpt-oss-20b-GGUF",
    "messages": [
      {
        "role": "user",
        "content": "What does the SELECT statement in SQL do?"
      }
    ],
    "max_tokens": 256,
    "temperature": 0.7,
    "stream": false
  }' https://localhost:8443/v1/chat/completions

The following example is similar, but includes text and image input and uses a different model:

BASE64_IMG=$(base64 -w 0 dog-kitten-small.jpg)
curl --noproxy '*' \
  -X POST \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer $API_KEY" \
  --cacert "$SECRETS_DIR/cert.pem" \
  --data '{
      "model": "Ministral-3-3B-Reasoning-2512",
      "messages": [
        {
          "role": "user",
          "content": [
            { "type": "text", "text": "Describe the image in one short sentence." },
            { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,'"${BASE64_IMG}"'" } }
      ]
    }
  ],
  "max_tokens": 64,
  "temperature": 0.7
}' https://localhost:9091/v1/chat/completions

For more information about configuring the Private AI Services Container, see Configure the Private AI Services Container.