Route Model Requests Across OCI Regions

Create a routing profile for an on-demand model, then use the routing profile OCID in Chat Completions and Responses requests from Python applications.

Overview

A routing profile specifies one supported on-demand model and one or more target OCI regions in the same realm where that model is available. Applications send requests to a regional OCI Generative AI API endpoint and pass the routing profile OCID as the model value. Generative AI serves each request from an eligible region in the profile.

Use one profile when applications need the same model and approved set of serving regions. You can update the profile's configuration in OCI instead of maintaining a region list in each application. A profile with multiple target regions doesn't imply round-robin routing or a fixed region priority.

In this tutorial, two Python applications call the Chicago API endpoint. One uses Chat Completions and the other uses Responses. Both use a routing profile for openai.gpt-oss-120b that allows serving from Chicago, Frankfurt, or Osaka.

Two applications send Chat Completions and Responses requests to the Chicago OCI Generative AI API endpoint. Both pass the same routing profile OCID as the model. The profile selects the on-demand model openai.gpt-oss-120b and permits serving from Chicago, Frankfurt, or Osaka.

You create a profile, install the OpenAI Python SDK, run both examples, and check the x-genai-selected-region response header to see which region served each request.

Note

This release supports routing profiles for on-demand models only. You can use a profile with both APIs only if its selected model supports both OCI Chat Completions and Responses API.

Before You Begin

To complete this tutorial, you need:

  • Access to OCI Generative AI, permission to create a routing profile in the selected compartment, and permission to use it for inference. See API-Level Permissions for Routing Profiles.
  • A supported on-demand model available to your tenancy in the target regions. To run both examples with one profile, select a model that supports both APIs. See Generative AI Models by Region.
  • An OCI Generative AI API key authorized for requests through the Chicago endpoint. Don't put the key in source code or commit it to a repository.
  • Python 3.11 or later, and uv installed on macOS, Linux, or Windows Subsystem for Linux. The commands in this tutorial use a Bash-compatible shell.

The examples use the OpenAI Python SDK with OCI's OpenAI-compatible endpoint. They don't require the OCI Python SDK or OCI CLI.

1. Create a Routing Profile

This example creates a profile in US Midwest (Chicago). Its target regions can differ from the API endpoint region, but they must be available for the selected model and in the same realm. If a model or region in this example isn't available to your tenancy, select an available option.

  1. Sign in to the OCI Console and select US Midwest (Chicago) (us-chicago-1) from the region selector.
  2. Open Analytics & AI, select Generative AI, and then select Routing profiles.
    For more information about the list page, see Listing Routing Profiles.
  3. Select Create routing profile.
  4. Enter a name, such as gpt-oss-multiregion, select a compartment, and optionally add a description or tags.
  5. Under Configuration, select openai.gpt-oss-120b from Model select.

    OpenAI gpt-oss-120b is available on demand in Chicago, Frankfurt, and Osaka and supports both OCI Chat Completions and Responses APIs.

    To select a model, check Models by Region for on-demand availability in your target regions, and Models and Regions for OCI OpenAI-Compatible Endpoints for models supported by the APIs you plan to call.

  6. Under Target OCI regions, select the regions from which the model can be served.

    For this example, select the following regions:

    • US Midwest (Chicago): us-chicago-1
    • Germany Central (Frankfurt): eu-frankfurt-1
    • Japan Central (Osaka): ap-osaka-1
  7. Select Create.
  8. Open the profile's details page. Confirm that its status is Active and review the selected model and target OCI regions. Copy the profile's full OCID.
    For details about this page, see Getting a Routing Profile's Details. You use the OCID as the model value in both scripts.
The profile is ready for supported on-demand inference requests.

2. Prepare the Python Environment

  1. In a working directory, create and activate a virtual environment and install the OpenAI Python SDK:
    uv venv --python 3.11 .venv
    source .venv/bin/activate
    uv pip install --upgrade openai

    If .venv already exists, activate it and run the installation command. The OCI Python SDK isn't needed for these examples.

  2. Set the API endpoint region and the routing profile OCID for the current shell:
    export OCI_REGION=us-chicago-1
    export OCI_ROUTING_PROFILE_OCID='<your-routing-profile-ocid>'

    Replace <your-routing-profile-ocid> with the full OCID shown in the routing profile. OCI_REGION determines which regional API endpoint receives the request. The routing profile specifies the regions where the model can serve it.

  3. Have an OCI Generative AI API key ready.
    Each script prompts for the key without displaying it if OCI_GENAI_API_KEY isn't set by a secret manager or another secure mechanism. Don't paste the key into either script.

3. Run the Chat Completions Example

  1. Save the following code as routing-chat-completion.py in your working directory:
    """Chat with an OCI Generative AI routing profile."""
    
    from __future__ import annotations
    
    import os
    from getpass import getpass
    
    from openai import OpenAI
    
    
    def main() -> None:
        region = os.environ.get("OCI_REGION", "us-chicago-1")
        profile_ocid = os.environ["OCI_ROUTING_PROFILE_OCID"]
        api_key = os.environ.get("OCI_GENAI_API_KEY") or getpass(
            "OCI Generative AI API key: "
        )
        client = OpenAI(
            base_url=f"https://inference.generativeai.{region}.oci.oraclecloud.com/openai/v1",
            api_key=api_key,
        )
    
        messages: list[dict[str, str]] = []
        print("Chat Completions with a routing profile. Type 'exit' to stop.\n")
    
        while True:
            try:
                prompt = input("You: ").strip()
            except (EOFError, KeyboardInterrupt):
                print("\nGoodbye!")
                break
    
            if prompt.lower() in {"exit", "quit"}:
                print("Goodbye!")
                break
            if not prompt:
                continue
    
            messages.append({"role": "user", "content": prompt})
            answer_parts: list[str] = []
    
            try:
                with client.chat.completions.create(
                    model=profile_ocid,
                    messages=messages,
                    stream=True,
                ) as stream:
                    selected_region = stream.response.headers.get(
                        "x-genai-selected-region"
                    )
                    print(f"Served from OCI region: {selected_region or 'not reported'}")
                    print("Assistant: ", end="", flush=True)
    
                    for chunk in stream:
                        if not chunk.choices:
                            continue
                        content = chunk.choices[0].delta.content
                        if content:
                            answer_parts.append(content)
                            print(content, end="", flush=True)
            except Exception as error:
                messages.pop()
                print(f"\nRequest failed: {error}\n")
                continue
    
            answer = "".join(answer_parts)
            if answer:
                messages.append({"role": "assistant", "content": answer})
            else:
                messages.pop()
            print("\n")
    
    
    if __name__ == "__main__":
        main()
  2. With the virtual environment activated, run the script:
    python routing-chat-completion.py
  3. At You:, enter a prompt. After the response, enter a follow-up or exit to stop.
    The script streams text from chunk.choices[0].delta.content and retains the conversation history locally for the current process.

4. Run the Responses Example

Use this example only if the model selected in the routing profile supports the Responses API. It uses the same profile OCID and regional endpoint as the Chat Completions example.

  1. Save the following code as routing-responses-api.py in the working directory:
    """Use the Responses API with an OCI Generative AI routing profile."""
    
    from __future__ import annotations
    
    import os
    from getpass import getpass
    
    from openai import OpenAI
    
    
    def main() -> None:
        region = os.environ.get("OCI_REGION", "us-chicago-1")
        profile_ocid = os.environ["OCI_ROUTING_PROFILE_OCID"]
        api_key = os.environ.get("OCI_GENAI_API_KEY") or getpass(
            "OCI Generative AI API key: "
        )
        client = OpenAI(
            base_url=f"https://inference.generativeai.{region}.oci.oraclecloud.com/openai/v1",
            api_key=api_key,
        )
    
        messages: list[dict[str, str]] = []
        print("Responses with a routing profile. Type 'exit' to stop.\n")
    
        while True:
            try:
                prompt = input("You: ").strip()
            except (EOFError, KeyboardInterrupt):
                print("\nGoodbye!")
                break
    
            if prompt.lower() in {"exit", "quit"}:
                print("Goodbye!")
                break
            if not prompt:
                continue
    
            messages.append({"role": "user", "content": prompt})
            answer_parts: list[str] = []
    
            try:
                with client.responses.create(
                    model=profile_ocid,
                    input=messages,
                    stream=True,
                ) as stream:
                    selected_region = stream.response.headers.get(
                        "x-genai-selected-region"
                    )
                    print(f"Served from OCI region: {selected_region or 'not reported'}")
                    print("Assistant: ", end="", flush=True)
    
                    for event in stream:
                        if event.type == "response.output_text.delta":
                            answer_parts.append(event.delta)
                            print(event.delta, end="", flush=True)
            except Exception as error:
                messages.pop()
                print(f"\nRequest failed: {error}\n")
                continue
    
            answer = "".join(answer_parts)
            if answer:
                messages.append({"role": "assistant", "content": answer})
            else:
                messages.pop()
            print("\n")
    
    
    if __name__ == "__main__":
        main()
  2. With the virtual environment activated, run the script:
    python routing-responses-api.py
  3. Enter a prompt and a follow-up, then enter exit to stop.
    The script streams response.output_text.delta events. It maintains conversation history locally and resends it on each turn; it doesn't use previous_response_id or a server-managed conversation. Other event types are ignored.

Verify the Result and Troubleshoot

In either script, enter a prompt such as Explain model routing in two sentences. The assistant's response should stream to the terminal. If OCI returns x-genai-selected-region, the script displays the region that served the request. It can differ from Chicago, the region of the API endpoint, but must be one of the target regions in the profile.

Both examples skip empty prompts. Enter exit or quit, press Ctrl+C, or send end-of-input at the prompt to stop. History is retained only while the script is running.

SymptomWhat to check
Authentication or authorization error Check the OCI Generative AI API key, its region and access, routing-profile read permission, and the permission for the inference operation.
Profile not found or invalid model value Use the full profile OCID in OCI_ROUTING_PROFILE_OCID. Confirm the profile is Active and that OCI_REGION points to the endpoint region you intend to call.
Unsupported API or request parameter Confirm that the profile's selected model supports the API used by the script. This tutorial doesn't require a model-specific reasoning parameter.
Request can't be served Check that the selected model is available in the profile's target regions and review the returned error.
Serving region is not reported The script didn't find x-genai-selected-region in the response headers. Check the response and service behavior; don't infer the serving region from the API endpoint alone.

The examples show profile-based requests and streaming text. They don't define or verify a routing priority, load-balancing algorithm, or failover policy.