Models Available to the Private Large Language Model Service

Included are models that are shipped with the container or available to be downloaded from Hugging Face and are known to work with the vLLM and llama.cpp runtimes.

The model Ministral-3-3B-Reasoning-2512 is shipped with the large inference container images and does not need to be separately downloaded. This is in addition to the models listed in Available Embedding Models, which are shipped with all inference containers for CPU hardware.

Note:

The model Llama-Prompt-Guard-2-86M is also shipped with the container as a mechanism to implement guardrails. It cannot be invoked directly and is not listed with the other shipped models. For more information about guardrails, see Guardrails.

Consider the memory and disk space required, along with support for tool calling, when deciding which model to download. The memory requirement listed in the following tables is the minimum. Concurrent requests, longer context windows or less quantization will require more memory.

Note that all llama.cpp GGUF models use Q8_0 quantization for determining the GGUF disk space requirement.

Table 6-1 Available Models for Llama.cpp

Provider GGUF Format LLM Associated Safe Tensor Format LLM GGUF Memory Required (GB) GGUF Disk Space Required (GB) Tool Calling Support?
OpenAI unsloth/gpt-oss-20b-GGUF openai/gpt-oss-20b 25 12 Yes
unsloth/gpt-oss-120b-GGUF openai/gpt-oss-120b 120 60 Yes
unsloth/gpt-oss-safeguard-120b-GGUF openai/gpt-oss-safeguard-120b 118 65 Yes
vonjack/whisper-large-v3-gguf openai/whisper-large-v3 15 9 No
Meta unsloth/Llama-3.1-8B-Instruct-GGUF meta-llama/Llama-3.1-8B-Instruct 20 15 Yes
unsloth/Llama-3.3-70B-Instruct-GGUF meta-llama/Llama-3.3-70B-Instruct 137 132 Yes
unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF meta-llama/Llama-4-Scout-17B-16E-Instruct 112 109 Yes
unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF meta-llama/Llama-4-Maverick-17B-128E-Instruct 402 399 Yes
Mistral.ai unsloth/Ministral-3-3B-Reasoning-2512-GGUF mistralai/Ministral-3-3B-Reasoning-2512 12 8 Yes
unsloth/Ministral-3-8B-Instruct-2512-GGUF mistralai/Ministral-3-8B-Instruct-2512 21 17 Yes
unsloth/Ministral-3-14B-Instruct-2512-GGUF mistralai/Ministral-3-14B-Instruct-2512 31 26 Yes
unsloth/Mistral-Medium-3.5-128B-GGUF mistralai/Mistral-Medium-3.5-128B 135 129 Yes
unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF mistralai/Mistral-Large-3-675B-Instruct-2512 691 384 Yes
Google unsloth/functiongemma-270m-it-GGUF google/functiongemma-270m-it 1 1 No
unsloth/gemma-4-E4B-it-GGUF google/gemma-4-E4B-it 19 15 Yes
unsloth/gemma-4-26B-A4B-it-GGUF google/gemma-4-26B-A4B-it 30 27 Yes
unsloth/gemma-4-31B-it-GGUF google/gemma-4-31B-it 37 31 Yes
bullerwins/translategemma-4b-it-GGUF google/translategemma-4b-it 7 4 No
bullerwins/translategemma-12b-it-GGUF google/translategemma-12b-it 17 12 Yes
bullerwins/translategemma-27b-it-GGUF google/translategemma-27b-it 32 27 Yes
Deepreinforce-ai deepreinforce-ai/Ornith-1.0-9B-GGUF deepreinforce-ai/Ornith-1.0-9B 16 15 Yes
deepreinforce-ai/Ornith-1.0-35B-GGUF deepreinforce-ai/Ornith-1.0-35B 40 35 Yes
bartowski/deepreinforce-ai_Ornith-1.0-397B-GGUF deepreinforce-ai/Ornith-1.0-397B 208 385 Yes
Microsoft unsloth/Phi-4-reasoning-GGUF microsoft/Phi-4-reasoning 18 15 Yes
jamesburton/Phi-4-reasoning-vision-15B-GGUF microsoft/Phi-4-reasoning-vision-15B 17 15 No
IBM unsloth/granite-4.1-3b-GGUF ibm-granite/granite-4.1-3b 11 7 Yes
unsloth/granite-4.1-8b-GGUF ibm-granite/granite-4.1-8b 21 17 Yes
unsloth/granite-4.1-30b-GGUF ibm-granite/granite-4.1-30b 34 29 Yes
ibm-granite/granite-docling-258M-GGUF ibm-granite/granite-docling-258M 1 1 No
NVIDIA unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 12 8 Yes
bartowski/nvidia_NVIDIA-Nemotron-Nano-12B-v2-GGUF nvidia/NVIDIA-Nemotron-Nano-12B-v2 15 13 Yes
unsloth/Nemotron-3-Nano-30B-A3B-GGUF nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 36 32 Yes
unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 37 33 Yes
bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF nvidia/Nemotron-Cascade-2-30B-A3B 35 32 No
unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 124 120 No
unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 551 545 Yes
DeepSeek-ai unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF deepseek-ai/DeepSeek-R1-Distill-Llama-8B 20 15 Yes
unsloth/DeepSeek-R1-Distill-Qwen-7B-GGUF deepseek-ai/DeepSeek-R1-Distill-Qwen-7B 19 15 Yes
TheBloke/deepseek-coder-33B-instruct-GGUF deepseek-ai/deepseek-coder-33b-instruct 38 33 Yes
Qwen unsloth/Qwen3.6-27B-GGUF Qwen/Qwen3.6-27B 32 33 Yes
unsloth/Qwen3.6-35B-A3B-GGUF Qwen/Qwen3.6-35B-A3B 41 36 Yes
unsloth/Qwen3-Coder-Next-GGUF Qwen/Qwen3-Coder-Next 87 43 Yes
Other unsloth/Kimi-K2.6-GGUF moonshotai/Kimi-K2.6 1048 545 Yes
unsloth/Kimi-K2.7-Code-GGUF moonshotai/Kimi-K2.7-Code 628 462 Yes
unsloth/MiniMax-M2.7-GGUF MiniMaxAI/MiniMax-M2.7 232 227 Yes
unsloth/GLM-4.7-Flash-GGUF zai-org/GLM-4.7-Flash 34 30 Yes
unsloth/GLM-5.1-GGUF zai-org/GLM-5.1 670 433 Yes
unsloth/GLM-5.2-GGUF zai-org/GLM-5.2 222 223 Yes
unsloth/Step-3.7-Flash-GGUF stepfun-ai/Step-3.7-Flash 209 200 Yes
LiquidAI/LFM2.5-8B-A1B-GGUF LiquidAI/LFM2.5-8B-A1B-GGUF 13 9 Yes
NousResearch/Hermes-4.3-36B-GGUF NousResearch/Hermes-4.3-36B 40 36 Yes
fdtn-ai/Foundation-Sec-8B-Q8_0-GGUF fdtn-ai/Foundation-Sec-8B 11 8 Yes
TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF TinyLlama/TinyLlama-1.1B-Chat-v1.0 2 2 No
TheBloke/leo-hessianai-13B-chat-GGUF LeoLM/leo-hessianai-13b-chat 116 13 Yes
TheBloke/SauerkrautLM-7B-v1-GGUF TheBloke/SauerkrautLM-7B-v1 73 7 Yes
LSX-UniWue/LLaMmlein_7B_chat-gguf LSX-UniWue/LLaMmlein_7B 79 13 Yes

Table 6-2 Available Models for vLLM on CPU Hardware

Provider Safe Tensor Format LLM Memory Required (GB)
OpenAI openai/whisper-large-v3 9
Meta meta-llama/Llama-3.1-8B-Instruct 15
meta-llama/Llama-3.3-70B-Instruct 132
meta-llama/Llama-4-Scout-17B-16E-Instruct 176
Mistral.ai mistralai/Ministral-3-3B-Reasoning-2512 15
mistralai/Ministral-3-8B-Instruct-2512 20
mistralai/Ministral-3-14B-Instruct-2512 30
Google google/functiongemma-270m-it 1
google/gemma-4-E4B-it 15
google/gemma-4-31B-it 54
Deepreinforce-ai deepreinforce-ai/Ornith-1.0-9B 18
deepreinforce-ai/Ornith-1.0-35B 66
Microsoft microsoft/Phi-4-reasoning 28
IBM ibm-granite/granite-4.1-3b 7
ibm-granite/granite-4.1-8b 7
ibm-granite/granite-4.1-30b 54
DeepSeek-ai deepseek-ai/DeepSeek-R1-Distill-Llama-8B 15
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B 15
deepseek-ai/deepseek-coder-33b-instruct 63
Qwen Qwen/Qwen3.6-27B 52
Qwen/Qwen3.6-35B-A3B 67
Qwen/Qwen3-Coder-Next 149
Other NousResearch/Hermes-4.3-36B 68
WeiboAI/VibeThinker-3B 6
TinyLlama/TinyLlama-1.1B-Chat-v1.0 2
deepreinforce-ai/Ornith-1.0-9B 18
deepreinforce-ai/Ornith-1.0-35B 66

In the following table, in addition to the memory and disk space requirements, columns are provided to indicate whether the particular LLM is known to work using the given number of A100 GPUs. For example, the openai/gpt-oss-120b LLM can be run using 4 or 8 GPUs, but is not supported using only one or two.

Table 6-3 Available Models for vLLM on GPU Hardware

Provider Safe Tensor Format LLM GPU VRAM Required (GB) 1 * A100 2 * A100 4 * A100 8 * A100
OpenAI openai/gpt-oss-20b 26 Y Y Y Y
openai/gpt-oss-120b 122 - - Y Y
openai/gpt-oss-safeguard-120b 61 - - Y Y
openai/whisper-large-v3 9 - - - -
Meta meta-llama/Llama-3.1-8B-Instruct 15 Y Y Y Y
meta-llama/Llama-3.3-70B-Instruct 132 - - Y Y
meta-llama/Llama-4-Scout-17B-16E-Instruct 176 - - - Y
Mistral.ai mistralai/Ministral-3-3B-Reasoning-2512 15 Y - - Y
mistralai/Ministral-3-8B-Instruct-2512 20 Y - - Y
mistralai/Ministral-3-14B-Instruct-2512 30 Y - - Y
Google google/functiongemma-270m-it 1 Y Y Y -
google/gemma-4-E4B-it 15 Y Y Y Y
google/gemma-4-26B-A4B-it 49 - Y Y Y
google/gemma-4-31B-it 54 - Y Y Y
google/translategemma-4b-it 9 - - - -
Deepreinforce-ai deepreinforce-ai/Ornith-1.0-9B 18 Y Y Y Y
deepreinforce-ai/Ornith-1.0-35B 66 - Y Y Y
Microsoft microsoft/Phi-4-reasoning 28 Y Y - -
IBM ibm-granite/granite-4.1-3b 7 Y Y Y Y
ibm-granite/granite-4.1-8b 7 Y Y Y Y
ibm-granite/granite-4.1-30b 54 - Y Y Y
ibm-granite/granite-docling-258M 1 Y - - -
DeepSeek-ai deepseek-ai/DeepSeek-R1-Distill-Llama-8B 15 Y - - Y
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B 15 Y Y Y -
deepseek-ai/deepseek-coder-33b-instruct 63 - Y - Y
Qwen Qwen/Qwen3.6-27B 52 - - Y Y
Qwen/Qwen3.6-35B-A3B 67 - - Y Y
Qwen/Qwen3-Coder-Next 149 - - - Y
Other zai-org/GLM-4.7-Flash 50 - - Y -
LiquidAI/LFM2.5-8B-A1B - Y - - Y
NousResearch/Hermes-4.3-36B 68 - - Y Y
WeiboAI/VibeThinker-3B 6 Y - - Y
TinyLlama/TinyLlama-1.1B-Chat-v1.0 2 Y - - Y
deepreinforce-ai/Ornith-1.0-9B 18 Y Y Y Y
deepreinforce-ai/Ornith-1.0-35B 66 - - Y Y