Route on-demand inference requests across regions with smart model router
- Services: Generative AI
- Release Date: September 23, 2026
OCI Generative AI now supports smart model router for cross-region routing of on-demand inference requests within a user-selected regional scope. Create a routing profile that specifies an on-demand model and select regions within the same realm where that model is available. Then reference the routing profile in a Chat API or Responses API request. OCI Generative AI routes requests only among the regions selected in the profile.
You can create and manage routing profiles by using the Console or API. Regional routing helps applications use on-demand model capacity across approved regions and reduces the need to build application-side cross-region routing. It also keeps inference processing within the user-defined regional scope.
To find the regions available for each on-demand model, see Generative AI Models by Region. For more information about the service, see the Generative AI documentation.