LLMs, Embedders and Rerankers
This page presents the abstract interfaces used to plug LLMs, embedders, and rerankers into Oracle Agent Memory.
LLM Interface
class oracleagentmemory.apis.llms.ILlm
Bases: ABC
Abstract interface for LLM invocation.
method generate (abstract)
Generate a response from an LLM synchronously.
- Parameters:
- prompt
str | Sequence[Message | Mapping[str, str | Sequence[Mapping[str, Any]]]] | ChatMessageTypedDictT– A string, chat-style dictionary sequence, or sequence ofMessageobjects. Strings become user messages; message content may include text and image parts. Image parts are validated before they are sent to the provider. - response_json_schema
dict[str, Any] | None– Optional JSON Schema describing the expected response format. - **kwargs (Any) – Additional call options. The built-in Llm accepts
api_type=LlmApiType.RESPONSESto select the Responses API andimage_input_limit_config=ImageInputLimitConfig(...)to override image limits for that request. Other keyword arguments are forwarded to the underlying backend.
- prompt
- Returns: Normalized LLM output.
- Return type: LlmResponse
method generate_async (abstract, async)
Asynchronously generate a response from an LLM.
- Parameters:
- prompt
str | Sequence[Message | Mapping[str, str | Sequence[Mapping[str, Any]]]] | ChatMessageTypedDictT– A string, chat-style dictionary sequence, or sequence ofMessageobjects. Strings become user messages; message content may include text and image parts. Image parts are validated before they are sent to the provider. - response_json_schema
dict[str, Any] | None– Optional JSON Schema describing the expected response format. - **kwargs (Any) – Additional call options. The built-in Llm accepts
api_type=LlmApiType.RESPONSESto select the Responses API andimage_input_limit_config=ImageInputLimitConfig(...)to override image limits for that request. Other keyword arguments are forwarded to the underlying backend.
- prompt
- Returns: Normalized LLM output.
- Return type: LlmResponse
LLM Responses
class oracleagentmemory.apis.llms.LlmResponse
Bases: object
A small normalized response returned by ILlm.
- Parameters:
text
str
text
The primary generated text content.
- Type: str
Embedder Interface
class oracleagentmemory.apis.IEmbedder
Bases: ABC
Abstract interface for text embedders.
method embed (abstract)
Embed a batch of texts into a 2D float32 NumPy array.
- Parameters:
- texts
list[str]– Batch of texts to embed. - is_query
bool– Whether the batch is being embedded for query-time retrieval.
- texts
- Returns:
A 2D array shaped
(len(texts), dim)withdtype=float32. - Return type: numpy.ndarray
method embed_async (abstract, async)
Embed a batch of texts into a 2D float32 NumPy array.
- Parameters:
- texts
list[str]– Batch of texts to embed. - is_query
bool– Whether the batch is being embedded for query-time retrieval.
- texts
- Returns:
A 2D array shaped
(len(texts), dim)withdtype=float32. - Return type: numpy.ndarray
property embedding_dimension
- Return Type: int
- Description: Return the size of the embeddings produced by this embedder.
Subclasses may override this property when the embedding width is
known from configuration or provider metadata. The default
implementation probes embed() once and caches the result size.
- Returns: Positive number of floating-point values in each embedding vector.
- Return type: int
property max_input_tokens
- Return Type: int
- Description: Return the maximum supported input tokens.
Subclasses may override this property when the model’s input budget is
known from configuration or provider metadata. The default implementation
validates a probe sized to an estimated 512 input tokens once and
caches 512 as a conservative fallback. It does not run a model
tokenizer locally, so callers should set max_input_tokens manually
when the model’s actual input budget is known.
- Returns: Positive maximum input-token count for one text payload.
- Return type: int
Reranker Interface
class oracleagentmemory.apis.IReranker
Bases: ABC
Abstract interface for synchronous and asynchronous document reranking.
Implementations must return one result for every input document. Each
zero-based input index must appear exactly once. relevance_score must be
finite, and higher scores must indicate greater relevance. Results must be
ordered from the highest score to the lowest score.
method rerank
Rank documents synchronously by delegating to rerank_async.
- Parameters:
- query
str– Search query used to compare the documents. - documents
list[str]– Candidate document text in stable input order. Result indexes refer to positions in this list. - **kwargs (Any) – Provider-specific options.
- query
- Returns:
Complete ranking with exactly
len(documents)results. Every input index appears once, ordered from most relevant to least relevant. - Return type: RerankResponse
method rerank_async (abstract, async)
Asynchronously rank documents by relevance to a query.
- Parameters:
- query
str– Search query used to compare the documents. - documents
list[str]– Candidate document text in stable input order. Result indexes refer to positions in this list. - **kwargs (Any) – Provider-specific options.
- query
- Returns:
Complete ranking with exactly
len(documents)results. Every input index appears once, ordered from most relevant to least relevant. - Return type: RerankResponse
class oracleagentmemory.apis.RerankResponse
Bases: object
Complete document ranking ordered by descending relevance score.
- Parameters:
results
list[RerankResponseResult]
class oracleagentmemory.apis.RerankResponseResult
Bases: object
One document’s reranking result.
- Parameters:
- index
int– Zero-based position of the document in the input list. - relevance_score
float– Relevance score for the query-document pair. Higher values mean greater relevance. - document
str | None– The document text when the provider returned it, otherwiseNone.
- index
LiteLLM Adapters
class oracleagentmemory.core.llms.LlmApiType
Bases: str, Enum
Supported OpenAI-compatible API families for Llm.
CHAT_COMPLETIONS = ‘chat_completions’
RESPONSES = ‘responses’
class oracleagentmemory.core.llms.Llm
Bases: ILlm
Adapter for generating model responses.
Create an LLM adapter.
- Parameters:
- model
str– Model identifier sent to the underlying model provider. - api_base
str | None– Optional base URL for an OpenAI-compatible endpoint. - api_key
str | None– Optional API key used when contacting the provider. - api_type
LlmApiType– API family to call. UseLlmApiType.CHAT_COMPLETIONSfor Chat Completions orLlmApiType.RESPONSESfor the Responses API. Defaults toLlmApiType.CHAT_COMPLETIONS. - stream
bool– Whether to request streaming output. The stream is consumed internally and returned as a singleLlmResponse. - temperature
float | None– Optional sampling temperature. - max_tokens
int | None– Optional output token limit. Withapi_type=LlmApiType.CHAT_COMPLETIONSthis is sent asmax_tokens. - reasoning_effort
str | None– Optional reasoning effort. Withapi_type=LlmApiType.CHAT_COMPLETIONSthis is sent asreasoning_effort. Withapi_type=LlmApiType.RESPONSESthis is converted toreasoning={"effort": ...}. - enable_structured_output_reminder
bool– Whether to append a structured-output schema to the prompt when structured output is requested. If omitted, the default follows the model route: it is enabled for hosted vLLM and non-closed models on LiteLLM’s explicit"openai/..."route. Set this explicitly for bare model names or other custom endpoints. - supports_vision
bool– Whether this model accepts image content in generation prompts. When omitted, support is detected lazily: the adapter first checks model metadata and, when metadata is unavailable, sends a fixed red-image probe. Set this toTrueorFalseto skip detection. - image_input_limit_config
ImageInputLimitConfig– Optional raw-image and per-request image limits. Omitted fields use SDK defaults. Validation remains enabled for every image request. - max_concurrent_requests
int | None– Maximum number of asynchronous requests allowed through this LLM instance. Omit the value to use16. PassNoneto disable concurrency limiting. Set this explicitly when the endpoint needs different behavior. - proxy
str | None– Optional proxy URL for provider requests. When provided, this explicit proxy takes precedence over proxy settings from the environment. If omitted, environment proxy settings are used whentrust_envis enabled. - trust_env
bool– Whether provider requests read proxy and TLS settings from environment variables. Defaults toTruewhen omitted. Set toFalseto ignore those environment settings; an explicitproxyis still used. - key_file
str | None– Optional path to the client private key file in PEM format. Provide it together withcert_filewhen the server requires mutual TLS. - cert_file
str | None– Optional path to the client certificate chain file in PEM format. Provide it together withkey_filewhen the server requires mutual TLS. - ca_file
str | None– Optional path to a trusted CA certificate or bundle in PEM format used to verify the server certificate. Use this for a private or otherwise non-system CA. - **default_kwargs (Any) – Advanced default keyword arguments applied to every call. Prefer
the explicit parameters above for common connection and generation
settings. When the same setting is provided both explicitly and in
default_kwargs, the explicit parameter takes precedence.
- model
Examples
OCI Generative AI models use LiteLLM’s "oci/..." model identifiers.
A common setup is to pass OCI API key authentication details from the
standard OCI config file through LiteLLM-specific keyword arguments.
The OCI Python SDK is not installed by this package; applications that
already depend on it may alternatively pass an oci_signer object.
import configparser
from pathlib import Path
parser = configparser.RawConfigParser()
parser.read(Path("~/.oci/config").expanduser())
cfg = parser["DEFAULT"]
key_file = Path(cfg["key_file"]).expanduser()
oci_llm = Llm(
model="oci/openai.gpt-oss-120b",
oci_compartment_id="ocid1.compartment.oc1..example",
oci_region=cfg.get("region", "us-chicago-1"),
oci_user=cfg["user"],
oci_fingerprint=cfg["fingerprint"],
oci_tenancy=cfg["tenancy"],
oci_key_file=str(key_file),
)
oci_llm.generate("Reply with OK.")
OpenAI-hosted models use LiteLLM model identifiers such as
"openai/gpt-5.1" and an OpenAI API key. Chat Completions is the
default API family.
openai_llm = Llm(
model="openai/gpt-5.1",
api_key="sk-example",
temperature=0,
max_tokens=128,
)
openai_llm.model
'openai/gpt-5.1'
openai_llm.generate("Reply with OK.")
Use api_type=LlmApiType.RESPONSES when the target model should be
called through the OpenAI Responses API instead of Chat Completions.
responses_llm = Llm(
model="openai/gpt-5.4",
api_key="sk-example",
api_type=LlmApiType.RESPONSES,
reasoning_effort="high",
stream=True,
)
responses_llm.model
'openai/gpt-5.4'
Self-hosted OpenAI-compatible servers, including vLLM, are called with
an "openai/..." model identifier plus the server’s /v1 base URL.
Pass a nominal api_key such as "none" when the endpoint does not
enforce authentication.
vllm_llm = Llm(
model="openai/openai/gpt-oss-120b",
api_base="http://localhost:8000/v1",
api_key="none",
stream=True,
)
vllm_llm.model
'openai/openai/gpt-oss-120b'
vllm_llm.generate("Reply with OK.")
method generate
Generate a response.
- Parameters:
- prompt
str | Sequence[Message | Mapping[str, str | Sequence[Mapping[str, Any]]]] | ChatMessageTypedDictT– A string, chat-style dictionary sequence, or sequence of Message objects. Strings become user messages; content may include text and image parts. Image parts are validated before they are sent to the provider. - response_json_schema
dict[str, Any] | None– Optional JSON Schema describing the expected response format. When provided, this method uses the provider-native structured output mechanism via OpenAI-compatibleresponse_format. - **kwargs (Any) – Additional call parameters. Pass
api_type=LlmApiType.RESPONSESto route this call through the Responses API. For image prompts, passimage_input_limit_config=ImageInputLimitConfig(...)to override this Llm instance’s image limits for this request; omitted fields inherit the instance configuration. Other keyword arguments are sent with the provider request.
- prompt
- Returns: Normalized LLM output.
- Return type: LlmResponse
method generate_async (async)
Asynchronously generate a response.
- Parameters:
- prompt
str | Sequence[Message | Mapping[str, str | Sequence[Mapping[str, Any]]]] | ChatMessageTypedDictT– A string, Prompt, or chat-style message sequence. Strings become user messages; content may include text and image parts. Image parts are validated before they are sent to the provider. - response_json_schema
dict[str, Any] | None– Optional JSON Schema describing the expected response format. When provided, this method uses the provider-native structured output mechanism via OpenAI-compatibleresponse_format. - **kwargs (Any) – Additional call parameters. Pass
api_type=LlmApiType.RESPONSESto route this call through the Responses API. For image prompts, passimage_input_limit_config=ImageInputLimitConfig(...)to override this Llm instance’s image limits for this request; omitted fields inherit the instance configuration. Other keyword arguments are sent with the provider request.
- prompt
- Returns: Normalized LLM output.
- Return type: LlmResponse
property supports_vision
- Return Type: bool
- Description: Return whether this LLM supports image input, detecting lazily.
class oracleagentmemory.core.embedders.Embedder
Bases: IEmbedder
Provider-backed embedder.
Create a provider-backed embedder.
- Parameters:
- model
str– Model identifier sent to the underlying embedding provider. - api_base
str | None– Optional base URL for an OpenAI-compatible endpoint. - api_key
str | None– Optional API key used when contacting the provider. - embedding_dimension
int | None– Optional embedding vector dimension. When provided, DB-backed clients can create or validate vector schemas without sending a provider probe. When omitted,embedding_dimensioninfers the dimension lazily with a small fallback probe. - max_input_tokens
int– Maximum input-token count supported by the embedding model. When omitted, themax_input_tokensproperty validates a provider probe sized to an estimated512input tokens and caches512as a conservative fallback. It does not run a model tokenizer locally, so setmax_input_tokensmanually based on the model’s documented input budget. - normalize
bool– Whether to L2-normalize embeddings returned by the provider. - query_prefix
str | None– Optional prefix added only when embedding query texts. - document_prefix
str | None– Optional prefix added only when embedding non-query texts. - truncate_prompt_tokens
int | None– Optional input token limit forwarded to providers that support truncating long embedding prompts. - proxy
str | None– Optional proxy URL for provider requests. When provided, this explicit proxy takes precedence over proxy settings from the environment. If omitted, environment proxy settings are used whentrust_envis enabled. - trust_env
bool– Whether provider requests read proxy and TLS settings from environment variables. Defaults toTruewhen omitted. Set toFalseto ignore those environment settings; an explicitproxyis still used. - key_file
str | None– Optional path to the client private key file in PEM format. Provide it together withcert_filewhen the server requires mutual TLS. - cert_file
str | None– Optional path to the client certificate chain file in PEM format. Provide it together withkey_filewhen the server requires mutual TLS. - ca_file
str | None– Optional path to a trusted CA certificate or bundle in PEM format used to verify the server certificate. Use this for a private or otherwise non-system CA. - **default_kwargs (Any) – Advanced default keyword arguments applied to every embedding call. Prefer the explicit parameters above for common settings.
- model
Examples
OCI Generative AI embedding models use "oci/..." model identifiers.
A common setup is to pass OCI API key authentication details from the
standard OCI config file through LiteLLM-specific keyword arguments.
The OCI Python SDK is not installed by this package; applications that
already depend on it may alternatively pass an oci_signer object.
import configparser
from pathlib import Path
parser = configparser.RawConfigParser()
parser.read(Path("~/.oci/config").expanduser())
cfg = parser["DEFAULT"]
key_file = Path(cfg["key_file"]).expanduser()
oci_embedder = Embedder(
model="oci/cohere.embed-english-v3.0",
oci_compartment_id="ocid1.compartment.oc1..example",
oci_region=cfg.get("region", "us-chicago-1"),
oci_user=cfg["user"],
oci_fingerprint=cfg["fingerprint"],
oci_tenancy=cfg["tenancy"],
oci_key_file=str(key_file),
)
oci_embedder.embed(["hello world"])
OpenAI-hosted embedding models use identifiers such as
"openai/text-embedding-3-small" with an OpenAI API key.
openai_embedder = Embedder(
model="openai/text-embedding-3-small",
api_key="sk-example",
truncate_prompt_tokens=8192,
)
openai_embedder.model
'openai/text-embedding-3-small'
openai_embedder.embed(["hello world"])
Self-hosted OpenAI-compatible embedding servers, including vLLM, use
the "hosted_vllm/..." provider prefix with the server’s /v1
base URL.
vllm_embedder = Embedder(
model="hosted_vllm/sentence-transformers/all-MiniLM-L6-v2",
api_base="http://localhost:8000/v1",
)
vllm_embedder.model
'hosted_vllm/sentence-transformers/all-MiniLM-L6-v2'
vllm_embedder.embed(["hello world"])
method embed
Embed a batch of texts using the configured provider.
- Parameters:
- texts
list[str]– Batch of raw text strings to embed. - is_query
bool– Whether the text is a query. Query texts receivequery_prefixand non-query texts receivedocument_prefixwhen configured.
- texts
- Returns:
A two-dimensional
float32matrix with the embedding vectors returned by the provider. - Return type: numpy.ndarray
- Raises: RuntimeError – If the provider response payload does not include embedding data.
method embed_async (async)
Asynchronously embed a batch of texts using the configured provider.
- Parameters:
- texts
list[str]– Batch of raw text strings to embed. - is_query
bool– Whether the text is a query. Query texts receivequery_prefixand non-query texts receivedocument_prefixwhen configured.
- texts
- Returns:
A two-dimensional
float32matrix with the embedding vectors returned by the provider. - Return type: numpy.ndarray
- Raises: RuntimeError – If the provider response payload does not include embedding data.
property embedding_dimension
- Return Type: int
-
Description: Return the configured or inferred embedding dimension.
- Returns: Positive number of dimensions in each embedding vector.
- Return type: int
Notes
A constructor-provided value is returned without contacting the provider. Otherwise the property probes once and caches the result.
property max_input_tokens
- Return Type: int
-
Description: Return the configured or inferred embedding input-token limit.
- Returns: Positive maximum input-token count for one text payload.
- Return type: int
Notes
A constructor-provided value is returned without contacting the
provider. Otherwise, the property validates a provider probe sized to
an estimated 512 input tokens and caches 512 as a conservative
fallback. It does not run a model tokenizer locally, so set
max_input_tokens manually from the model’s documented input budget
when precision matters.
class oracleagentmemory.core.Reranker
Bases: IReranker
Reranker backed by a provider-neutral rerank interface.
- Parameters:
- model
str– Reranker model identifier. Prefix OCI Generative AI models withoci/, for exampleoci/cohere.rerank-v4.0-fast. - api_base
str | None– Optional base URL for an OpenAI-compatible reranking endpoint. - api_key
str | None– Optional API key used by the provider. -
**default_kwargs (Any) –
Provider-specific options applied to every reranking request.
Note: OCI reranking requires
oci_compartment_id,oci_region,oci_user,oci_fingerprint,oci_tenancy, andoci_key_file. Install the optionalrerank-ocidependency group before using an OCI model.
- model
Examples
reranker = Reranker(
model="your-reranker-model",
api_base="https://your-reranker-endpoint/v1",
api_key="your-api-key",
)
reranker.rerank("favorite food", ["The user likes pasta."])
oci_reranker = Reranker(
model="oci/cohere.rerank-v4.0-fast",
oci_compartment_id="ocid1.compartment...",
oci_region="your-region",
oci_user="ocid1.user...",
oci_fingerprint="aa:bb:cc",
oci_tenancy="ocid1.tenancy...",
oci_key_file="~/.oci/oci_api_key.pem",
)
oci_reranker.rerank("favorite food", ["The user likes pasta."])
Create a provider-backed reranker.
method rerank_async (async)
Asynchronously rank documents by relevance to a query.
- Parameters:
- query
str– Search query used to compare the documents. - documents
list[str]– Candidate document text in stable input order. - **kwargs (Any) – Provider-specific options forwarded to the reranking provider.
- query
- Returns: Complete provider ranking ordered from the highest relevance score to the lowest.
- Return type: RerankResponse
Oracle DB Embedders
class oracleagentmemory.core.embedders.OracleDBEmbedder
Bases: IEmbedder
Embed text by invoking Oracle AI Database embedding SQL.
This embedder keeps the package’s existing embedder contract intact while
delegating embedding generation to the database over SQL. Direct embedding
prefers VECTOR_EMBEDDING for database-resident model configurations and
falls back to DBMS_VECTOR_CHAIN.UTL_TO_EMBEDDING when the vectorizer
configuration needs the JSON provider-parameter surface.
Create an embedder backed by Oracle AI Database SQL execution.
- Parameters:
- connection
object– Oracle DB connection or pool-like object with a callablecursor()oracquire()method. - model
str– Model identifier. For the default"database"provider, this must be an unquoted Oracle SQL identifier or schema-qualified identifier for an in-database embedding model. The connected schema must be able to resolve this model name in SQL. For a remote provider, use the model name or provider-specific model identifier expected by that service. - input_name
str– Model input name used byVECTOR_EMBEDDINGwhen the vectorizer configuration targets a database-resident model. Defaults to"DATA", the input name used by Oracle’s DBMS_VECTOR ONNX embedding-model examples and metadata. Pass the actual model input name here if the imported model uses a different attribute. - embedding_dimension
int | None– Optional embedding vector dimension. When provided, DB-backed clients can create or validate vector schemas without sending a dimension-probe query. When omitted, the dimension is inferred lazily with one probe embedding request. - max_input_tokens
int– Maximum input-token budget used by the default store chunker. When omitted, themax_input_tokensproperty validates a database-model probe sized to an estimated512input tokens and caches512as a conservative fallback. It does not run a model tokenizer locally, so setmax_input_tokensmanually based on the model’s documented input budget. - normalize
bool– Whether to L2-normalize embeddings after they are fetched from the database. - query_prefix
str | None– Optional prefix added only when embedding query texts. - batch_size
int– Maximum number of texts grouped into one SQL embedding round-trip. - provider
str– Embedding provider configured in Oracle AI Database. The default,"database", uses an embedding model loaded into Oracle Database, typically in ONNX format. Remote providers include services such as Cohere, OpenAI, Google AI, and Oracle Cloud Infrastructure Generative AI. OpenAI-compatible services such as vLLM use"openai"together with the service endpoint inprovider_options. See Oracle’s DBMS_VECTOR_CHAIN documentation for supported providers and their configuration: https://docs.oracle.com/en/database/oracle/oracle-database/26/vecse/dbms_vector_chain-vecse.html - provider_options
Mapping[str, Any] | None– Optional Oracle vectorizer settings, such asurl,credential_name, orhost="local". Constructor arguments overrideprovider,model, andinput_namein this mapping. Oracle documents provider options and how to create a database credential at: https://docs.oracle.com/en/database/oracle/oracle-database/26/vecse/utl_to_embedding-and-utl_to_embeddings-dbms_vector_chain.html https://docs.oracle.com/en/database/oracle/oracle-database/26/vecse/create_credential-dbms_vector_chain.html
- connection
Examples
Use an Oracle connection pool and a DB-resident embedding model:
import oracledb
pool = oracledb.create_pool(
user="scott",
password="tiger",
dsn="dbhost.example.com/orclpdb",
)
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
embedding_dimension=768,
)
embedder.embed(["hello world"])
Schema-qualified model names can be used when the connected schema has privileges on a model owned by another schema:
shared_embedder = OracleDBEmbedder(
connection=pool,
model="MY_OTHER_SCHEMA.MY_ONNX_MODEL",
embedding_dimension=768,
)
shared_embedder.embed(["hello world"])
The following examples show how to configure Oracle AI Database embeddings with OpenAI, vLLM, Cohere, and other providers:
OpenAI example:
openai_embedder = OracleDBEmbedder(
connection=pool,
provider="openai",
model="text-embedding-3-small",
provider_options={
"credential_name": "OPENAI_CRED",
"url": "https://api.openai.example.com/embeddings",
},
)
openai_embedder.embed(["hello world"])
Cohere example:
cohere_embedder = OracleDBEmbedder(
connection=pool,
provider="cohere",
model="embed-english-v3.0",
provider_options={
"credential_name": "COHERE_CRED",
"url": "https://api.cohere.example.com/embed",
"input_type": "search_document",
},
)
cohere_embedder.embed(["hello world"])
OpenAI-compatible services such as vLLM also use the "openai"
provider. Set host to "local" when the endpoint does not
require an Oracle AI Database credential:
vllm_embedder = OracleDBEmbedder(
connection=pool,
provider="openai",
model="BAAI/bge-small-en-v1.5",
provider_options={
"url": "http://localhost:8080/v1/embeddings",
"host": "local",
},
)
vllm_embedder.embed(["hello world"])
Gemini example:
gemini_embedder = OracleDBEmbedder(
connection=pool,
provider="googleai",
model="gemini-embedding-001",
provider_options={
"credential_name": "GOOGLEAI_CRED",
"url": "https://googleapis.example.com/models/",
},
)
gemini_embedder.embed(["hello world"])
Hugging Face example:
huggingface_embedder = OracleDBEmbedder(
connection=pool,
provider="huggingface",
model=(
"sentence-transformers/all-MiniLM-L6-v2"
),
provider_options={
"credential_name": "HF_CRED",
"url": "https://router.huggingface.example.com/",
},
)
huggingface_embedder.embed(["hello world"])
Query-specific prefixes can be configured without changing the store API:
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
query_prefix="search_document: ",
)
embedder.embed(["pizza"], is_query=True)
method embed
Embed a batch of texts by executing SQL in Oracle AI Database.
- Parameters:
- texts
list[str]– Batch of raw text strings to embed. - is_query
bool– Whether the text is a query. Query texts receivequery_prefixwhen one was configured.
- texts
- Returns:
A two-dimensional
float32matrix with one row per input text. - Return type: numpy.ndarray
Examples
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
)
matrix = embedder.embed(["alpha", "beta"])
matrix.shape[0]
2
method embed_async (async)
Asynchronously embed a batch of texts using Oracle AI Database SQL.
- Parameters:
- texts
list[str]– Batch of raw text strings to embed. - is_query
bool– Whether the text is a query. Query texts receivequery_prefixwhen one was configured.
- texts
- Returns:
A two-dimensional
float32matrix with one row per input text. - Return type: numpy.ndarray
Examples
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
)
matrix = await embedder.embed_async(["hello"])
matrix.shape
(1, 384)
property embedding_dimension
- Return Type: int
-
Description: Return the configured or inferred embedding dimension.
- Returns: Positive number of dimensions in each embedding vector.
- Return type: int
Notes
A constructor-provided value is returned without contacting the database model. Otherwise the property probes once and caches the result for future accesses.
Examples
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
embedding_dimension=768,
)
embedder.embedding_dimension
768
method get_vectorizer_config_json
Return Oracle vectorizer preference JSON for this DB model.
The same model configuration is used by direct embedding and by managed
hybrid indexes. Direct embedding uses it to decide whether
VECTOR_EMBEDDING can represent the configured database model or
whether DBMS_VECTOR_CHAIN.UTL_TO_EMBEDDING is needed for provider
JSON. Hybrid indexing passes it to
DBMS_VECTOR_CHAIN.CREATE_PREFERENCE and then Oracle’s vectorizer
pipeline owns embedding work for that index.
- Returns:
Compact JSON payload suitable for
DBMS_VECTOR_CHAIN.CREATE_PREFERENCEwithDBMS_VECTOR_CHAIN.VECTORIZER. - Return type: str
Examples
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
embedding_dimension=768,
)
embedder.get_vectorizer_config_json()
'{"model":"DOC_MODEL"}'
custom_embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
input_name="TEXT",
embedding_dimension=768,
)
custom_embedder.get_vectorizer_config_json()
'{"model":"DOC_MODEL","input_name":"TEXT"}'
shared_embedder = OracleDBEmbedder(
connection=pool,
model="MY_OTHER_SCHEMA.MY_ONNX_MODEL",
embedding_dimension=768,
)
shared_embedder.get_vectorizer_config_json()
'{"model":"MY_OTHER_SCHEMA.MY_ONNX_MODEL"}'
remote_embedder = OracleDBEmbedder(
connection=pool,
provider="openai",
model="text-embedding-3-small",
provider_options={"host": "local", "url": "http://localhost:8080/v1/embeddings"},
)
remote_embedder.get_vectorizer_config_json()
'{"embedder_spec":{"host":"local","url":"http://localhost:8080/v1/embeddings","provider":"openai","model":"text-embedding-3-small"}}'
property max_input_tokens
- Return Type: int
-
Description: Return the configured or inferred input-token budget for chunking.
- Returns: Positive maximum input-token count for one text payload.
- Return type: int
Notes
A constructor-provided value is returned without contacting the
database model. Otherwise, the property validates a database-model
probe sized to an estimated 512 input tokens and caches 512 as
a conservative fallback. It does not run a model tokenizer locally, so
set max_input_tokens manually from the model’s documented input
budget when precision matters.
Examples
embedder = OracleDBEmbedder(
connection=pool,
model="DOC_MODEL",
max_input_tokens=2048,
)
embedder.max_input_tokens
2048