Models Available to the Private Large Language Model Service
Included are models that are shipped with the container or available to be downloaded from Hugging Face and are known to work with the vLLM and llama.cpp runtimes.
The model Ministral-3-3B-Reasoning-2512 is shipped with the large inference container images and does not need to be separately downloaded. This is in addition to the models listed in Available Embedding Models, which are shipped with all inference containers for CPU hardware.
Note:
The model Llama-Prompt-Guard-2-86M is also shipped with the container as a mechanism to implement guardrails. It cannot be invoked directly and is not listed with the other shipped models. For more information about guardrails, see Guardrails.Consider the memory and disk space required, along with support for tool calling, when deciding which model to download. The memory requirement listed in the following tables is the minimum. Concurrent requests, longer context windows or less quantization will require more memory.
Note that all llama.cpp GGUF models use Q8_0 quantization for determining the GGUF disk space requirement.
Table 6-1 Available Models for Llama.cpp
| Provider | GGUF Format LLM | Associated Safe Tensor Format LLM | GGUF Memory Required (GB) | GGUF Disk Space Required (GB) | Tool Calling Support? |
|---|---|---|---|---|---|
| OpenAI | unsloth/gpt-oss-20b-GGUF | openai/gpt-oss-20b | 25 | 12 | Yes |
| unsloth/gpt-oss-120b-GGUF | openai/gpt-oss-120b | 120 | 60 | Yes | |
| unsloth/gpt-oss-safeguard-120b-GGUF | openai/gpt-oss-safeguard-120b | 118 | 65 | Yes | |
| vonjack/whisper-large-v3-gguf | openai/whisper-large-v3 | 15 | 9 | No | |
| Meta | unsloth/Llama-3.1-8B-Instruct-GGUF | meta-llama/Llama-3.1-8B-Instruct | 20 | 15 | Yes |
| unsloth/Llama-3.3-70B-Instruct-GGUF | meta-llama/Llama-3.3-70B-Instruct | 137 | 132 | Yes | |
| unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF | meta-llama/Llama-4-Scout-17B-16E-Instruct | 112 | 109 | Yes | |
| unsloth/Llama-4-Maverick-17B-128E-Instruct-GGUF | meta-llama/Llama-4-Maverick-17B-128E-Instruct | 402 | 399 | Yes | |
| Mistral.ai | unsloth/Ministral-3-3B-Reasoning-2512-GGUF | mistralai/Ministral-3-3B-Reasoning-2512 | 12 | 8 | Yes |
| unsloth/Ministral-3-8B-Instruct-2512-GGUF | mistralai/Ministral-3-8B-Instruct-2512 | 21 | 17 | Yes | |
| unsloth/Ministral-3-14B-Instruct-2512-GGUF | mistralai/Ministral-3-14B-Instruct-2512 | 31 | 26 | Yes | |
| unsloth/Mistral-Medium-3.5-128B-GGUF | mistralai/Mistral-Medium-3.5-128B | 135 | 129 | Yes | |
| unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF | mistralai/Mistral-Large-3-675B-Instruct-2512 | 691 | 384 | Yes | |
| unsloth/functiongemma-270m-it-GGUF | google/functiongemma-270m-it | 1 | 1 | No | |
| unsloth/gemma-4-E4B-it-GGUF | google/gemma-4-E4B-it | 19 | 15 | Yes | |
| unsloth/gemma-4-26B-A4B-it-GGUF | google/gemma-4-26B-A4B-it | 30 | 27 | Yes | |
| unsloth/gemma-4-31B-it-GGUF | google/gemma-4-31B-it | 37 | 31 | Yes | |
| bullerwins/translategemma-4b-it-GGUF | google/translategemma-4b-it | 7 | 4 | No | |
| bullerwins/translategemma-12b-it-GGUF | google/translategemma-12b-it | 17 | 12 | Yes | |
| bullerwins/translategemma-27b-it-GGUF | google/translategemma-27b-it | 32 | 27 | Yes | |
| Deepreinforce-ai | deepreinforce-ai/Ornith-1.0-9B-GGUF | deepreinforce-ai/Ornith-1.0-9B | 16 | 15 | Yes |
| deepreinforce-ai/Ornith-1.0-35B-GGUF | deepreinforce-ai/Ornith-1.0-35B | 40 | 35 | Yes | |
| bartowski/deepreinforce-ai_Ornith-1.0-397B-GGUF | deepreinforce-ai/Ornith-1.0-397B | 208 | 385 | Yes | |
| Microsoft | unsloth/Phi-4-reasoning-GGUF | microsoft/Phi-4-reasoning | 18 | 15 | Yes |
| jamesburton/Phi-4-reasoning-vision-15B-GGUF | microsoft/Phi-4-reasoning-vision-15B | 17 | 15 | No | |
| IBM | unsloth/granite-4.1-3b-GGUF | ibm-granite/granite-4.1-3b | 11 | 7 | Yes |
| unsloth/granite-4.1-8b-GGUF | ibm-granite/granite-4.1-8b | 21 | 17 | Yes | |
| unsloth/granite-4.1-30b-GGUF | ibm-granite/granite-4.1-30b | 34 | 29 | Yes | |
| ibm-granite/granite-docling-258M-GGUF | ibm-granite/granite-docling-258M | 1 | 1 | No | |
| NVIDIA | unsloth/NVIDIA-Nemotron-3-Nano-4B-GGUF | nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 | 12 | 8 | Yes |
| bartowski/nvidia_NVIDIA-Nemotron-Nano-12B-v2-GGUF | nvidia/NVIDIA-Nemotron-Nano-12B-v2 | 15 | 13 | Yes | |
| unsloth/Nemotron-3-Nano-30B-A3B-GGUF | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | 36 | 32 | Yes | |
| unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 | 37 | 33 | Yes | |
| bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF | nvidia/Nemotron-Cascade-2-30B-A3B | 35 | 32 | No | |
| unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 124 | 120 | No | |
| unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 551 | 545 | Yes | |
| DeepSeek-ai | unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 20 | 15 | Yes |
| unsloth/DeepSeek-R1-Distill-Qwen-7B-GGUF | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 19 | 15 | Yes | |
| TheBloke/deepseek-coder-33B-instruct-GGUF | deepseek-ai/deepseek-coder-33b-instruct | 38 | 33 | Yes | |
| Qwen | unsloth/Qwen3.6-27B-GGUF | Qwen/Qwen3.6-27B | 32 | 33 | Yes |
| unsloth/Qwen3.6-35B-A3B-GGUF | Qwen/Qwen3.6-35B-A3B | 41 | 36 | Yes | |
| unsloth/Qwen3-Coder-Next-GGUF | Qwen/Qwen3-Coder-Next | 87 | 43 | Yes | |
| Other | unsloth/Kimi-K2.6-GGUF | moonshotai/Kimi-K2.6 | 1048 | 545 | Yes |
| unsloth/Kimi-K2.7-Code-GGUF | moonshotai/Kimi-K2.7-Code | 628 | 462 | Yes | |
| unsloth/MiniMax-M2.7-GGUF | MiniMaxAI/MiniMax-M2.7 | 232 | 227 | Yes | |
| unsloth/GLM-4.7-Flash-GGUF | zai-org/GLM-4.7-Flash | 34 | 30 | Yes | |
| unsloth/GLM-5.1-GGUF | zai-org/GLM-5.1 | 670 | 433 | Yes | |
| unsloth/GLM-5.2-GGUF | zai-org/GLM-5.2 | 222 | 223 | Yes | |
| unsloth/Step-3.7-Flash-GGUF | stepfun-ai/Step-3.7-Flash | 209 | 200 | Yes | |
| LiquidAI/LFM2.5-8B-A1B-GGUF | LiquidAI/LFM2.5-8B-A1B-GGUF | 13 | 9 | Yes | |
| NousResearch/Hermes-4.3-36B-GGUF | NousResearch/Hermes-4.3-36B | 40 | 36 | Yes | |
| fdtn-ai/Foundation-Sec-8B-Q8_0-GGUF | fdtn-ai/Foundation-Sec-8B | 11 | 8 | Yes | |
| TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF | TinyLlama/TinyLlama-1.1B-Chat-v1.0 | 2 | 2 | No | |
| TheBloke/leo-hessianai-13B-chat-GGUF | LeoLM/leo-hessianai-13b-chat | 116 | 13 | Yes | |
| TheBloke/SauerkrautLM-7B-v1-GGUF | TheBloke/SauerkrautLM-7B-v1 | 73 | 7 | Yes | |
| LSX-UniWue/LLaMmlein_7B_chat-gguf | LSX-UniWue/LLaMmlein_7B | 79 | 13 | Yes |
Table 6-2 Available Models for vLLM on CPU Hardware
| Provider | Safe Tensor Format LLM | Memory Required (GB) |
|---|---|---|
| OpenAI | openai/whisper-large-v3 | 9 |
| Meta | meta-llama/Llama-3.1-8B-Instruct | 15 |
| meta-llama/Llama-3.3-70B-Instruct | 132 | |
| meta-llama/Llama-4-Scout-17B-16E-Instruct | 176 | |
| Mistral.ai | mistralai/Ministral-3-3B-Reasoning-2512 | 15 |
| mistralai/Ministral-3-8B-Instruct-2512 | 20 | |
| mistralai/Ministral-3-14B-Instruct-2512 | 30 | |
| google/functiongemma-270m-it | 1 | |
| google/gemma-4-E4B-it | 15 | |
| google/gemma-4-31B-it | 54 | |
| Deepreinforce-ai | deepreinforce-ai/Ornith-1.0-9B | 18 |
| deepreinforce-ai/Ornith-1.0-35B | 66 | |
| Microsoft | microsoft/Phi-4-reasoning | 28 |
| IBM | ibm-granite/granite-4.1-3b | 7 |
| ibm-granite/granite-4.1-8b | 7 | |
| ibm-granite/granite-4.1-30b | 54 | |
| DeepSeek-ai | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 15 |
| deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 15 | |
| deepseek-ai/deepseek-coder-33b-instruct | 63 | |
| Qwen | Qwen/Qwen3.6-27B | 52 |
| Qwen/Qwen3.6-35B-A3B | 67 | |
| Qwen/Qwen3-Coder-Next | 149 | |
| Other | NousResearch/Hermes-4.3-36B | 68 |
| WeiboAI/VibeThinker-3B | 6 | |
| TinyLlama/TinyLlama-1.1B-Chat-v1.0 | 2 | |
| deepreinforce-ai/Ornith-1.0-9B | 18 | |
| deepreinforce-ai/Ornith-1.0-35B | 66 |
In the following table, in addition to the memory and disk space requirements, columns are provided to indicate whether the particular LLM is known to work using the given number of A100 GPUs. For example, the openai/gpt-oss-120b LLM can be run using 4 or 8 GPUs, but is not supported using only one or two.
Table 6-3 Available Models for vLLM on GPU Hardware
| Provider | Safe Tensor Format LLM | GPU VRAM Required (GB) | 1 * A100 | 2 * A100 | 4 * A100 | 8 * A100 |
|---|---|---|---|---|---|---|
| OpenAI | openai/gpt-oss-20b | 26 | Y | Y | Y | Y |
| openai/gpt-oss-120b | 122 | - | - | Y | Y | |
| openai/gpt-oss-safeguard-120b | 61 | - | - | Y | Y | |
| openai/whisper-large-v3 | 9 | - | - | - | - | |
| Meta | meta-llama/Llama-3.1-8B-Instruct | 15 | Y | Y | Y | Y |
| meta-llama/Llama-3.3-70B-Instruct | 132 | - | - | Y | Y | |
| meta-llama/Llama-4-Scout-17B-16E-Instruct | 176 | - | - | - | Y | |
| Mistral.ai | mistralai/Ministral-3-3B-Reasoning-2512 | 15 | Y | - | - | Y |
| mistralai/Ministral-3-8B-Instruct-2512 | 20 | Y | - | - | Y | |
| mistralai/Ministral-3-14B-Instruct-2512 | 30 | Y | - | - | Y | |
| google/functiongemma-270m-it | 1 | Y | Y | Y | - | |
| google/gemma-4-E4B-it | 15 | Y | Y | Y | Y | |
| google/gemma-4-26B-A4B-it | 49 | - | Y | Y | Y | |
| google/gemma-4-31B-it | 54 | - | Y | Y | Y | |
| google/translategemma-4b-it | 9 | - | - | - | - | |
| Deepreinforce-ai | deepreinforce-ai/Ornith-1.0-9B | 18 | Y | Y | Y | Y |
| deepreinforce-ai/Ornith-1.0-35B | 66 | - | Y | Y | Y | |
| Microsoft | microsoft/Phi-4-reasoning | 28 | Y | Y | - | - |
| IBM | ibm-granite/granite-4.1-3b | 7 | Y | Y | Y | Y |
| ibm-granite/granite-4.1-8b | 7 | Y | Y | Y | Y | |
| ibm-granite/granite-4.1-30b | 54 | - | Y | Y | Y | |
| ibm-granite/granite-docling-258M | 1 | Y | - | - | - | |
| DeepSeek-ai | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 15 | Y | - | - | Y |
| deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 15 | Y | Y | Y | - | |
| deepseek-ai/deepseek-coder-33b-instruct | 63 | - | Y | - | Y | |
| Qwen | Qwen/Qwen3.6-27B | 52 | - | - | Y | Y |
| Qwen/Qwen3.6-35B-A3B | 67 | - | - | Y | Y | |
| Qwen/Qwen3-Coder-Next | 149 | - | - | - | Y | |
| Other | zai-org/GLM-4.7-Flash | 50 | - | - | Y | - |
| LiquidAI/LFM2.5-8B-A1B | - | Y | - | - | Y | |
| NousResearch/Hermes-4.3-36B | 68 | - | - | Y | Y | |
| WeiboAI/VibeThinker-3B | 6 | Y | - | - | Y | |
| TinyLlama/TinyLlama-1.1B-Chat-v1.0 | 2 | Y | - | - | Y | |
| deepreinforce-ai/Ornith-1.0-9B | 18 | Y | Y | Y | Y | |
| deepreinforce-ai/Ornith-1.0-35B | 66 | - | - | Y | Y |