xAI Grok 4.6
Grok 4.6 (xai.grok-4.6) is designed for coding, agentic tasks, and knowledge work. The model supports long-running agent workflows, research across unfamiliar domains, codebase exploration, and iterative development of applications and work artifacts. It has a 500,000-token context window.
Learn about Grok 4.6.
Regions for this Model
For supported regions, endpoint types (on-demand or dedicated AI clusters), and hosting (OCI Generative AI or external calls) for this model, see the Models by Region page. For details about the regions, see the Generative AI Regions page.
Access this Model
Access this model on demand through supported text API requests in the supported regions.
Key Features
- Model name in OCI
Generative AI:
xai.grok-4.6 - Available On-Demand: Access this model with standard or priority processing through supported text API requests.
- Context Length: 500,000 tokens.
- Cached Input Tokens: Yes. For the number of cached input tokens, see the
cachedTokensattribute in PromptTokensDetails Reference API and for prices, see the Pricing Page.
Limits
- Tokens per minute (TPM)
- For a TPM limit increase, use the limit name
grok-4-6-tokens-per-minute-count. See Creating a Limit Increase Request.
On-Demand Mode
The Grok models are available only in the on-demand mode.
| Model Name | OCI Model Name |
|---|---|
| xAI Grok 4.6 | xai.grok-4.6 |
OCI Release and Retirement Dates
For release and retirement dates and replacement model options, see the Model Retirement Dates (On-Demand Mode).
Pricing
Grok 4.6 has separate rates for standard and priority processing and for short- and long-context requests. When a prompt reaches 200,000 tokens, long-context pricing applies to the entire request, including all input and output tokens. For current rates, see the Pricing Page.
Priority processing is available for supported text API requests only. A priority request is charged at the priority rate only when the response includes
service_tier: priority. Otherwise, standard rates apply.