使用影像與多模態訊息

專員通常需要記住影像和文字。螢幕截圖、文件、圖表和相片可包含無法保留純文字記憶體的詳細資料。

在本指南中,您將學習如何:

提示:如需設定套裝程式,請參閱開始使用代理程式記憶體。如果此範例需要本機 Oracle AI Database,請遵循在本機執行 Oracle AI Database 。

建立 OracleAgentMemory 從屬端

建立本指南中的範例所使用的 OracleAgentMemory 用戶端。它可連線至 Oracle AI Database、設定嵌入器和具備視覺功能的 LLM,並設定影像輸入的限制。請參閱本指南結尾附近的「檢閱格式」、「限制」和「資料處理」,以取得這些限制的說明。

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()

從屬端會在建立 Llm 時設定 supports_vision=True。只有在選取的模型和提供者端點接受影像輸入時,才設定此選項。如果省略,Llm 會檢查模型中繼資料。沒有描述資料時,Llm 會傳送小型測試影像,以判斷端點是否接受影像輸入。設定 supports_vision=True 會略過這兩個檢查;它不會讓只有文字的模型能夠處理影像。

API 參照:Llm 影像輸入限制組態 記憶體擷取組態

儲存獨立影像

使用 OracleAgentMemory.add_image() 來新增獨立影像。至少包含 user_id、agent_id 或 thread_id 其中之一。當您稍後呼叫 list_images() 或 search() 時,必須提供相符的 ID。

若為搜尋,SDK 代表含有文字描述的影像。使用向量搜尋,會在該描述中內嵌與用於純文字搜尋的相同文字內嵌器,但不會內嵌影像位元組。如果您將 description 傳送至 add_image(),則 SDK 會使用該文字。如果省略,設定的支援視覺功能的 LLM 會自動產生描述。下列範例提供描述,因此新增影像不會提出 LLM 要求。

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)

list_images() 會傳回符合提供之 user_id、agent_id 或 thread_id 的獨立影像。每個 ImageRecord.content 值都包含影像描述。

若要將獨立影像與執行緒建立關聯,請呼叫 OracleThread.add_image()。此方法使用執行緒 ID 與任何儲存在執行緒中的使用者或代理程式 ID,因此您不會再次傳送它們。

API 參考: OracleAgentMemory 繁體中文 影像紀錄

將影像儲存為訊息的一部分

當影像提供對話轉換的相關資訊環境且不需要個別管理時,將影像儲存為執行緒訊息的一部分。

唯讀文字訊息可以使用字串作為 content。針對包含文字和影像的訊息,傳送內容部分的排序清單。SDK 會在儲存及擷取訊息以及建立記憶體擷取提示時,保留該順序。

在說明表單中,每個內容部分都有一個 type:

支援的 mime_type 值為 "image/png"、"image/jpeg" 和 "image/webp"。

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]

此範例中的影像字典會省略 description,因此 SDK 會產生具備已設定之具備視覺功能的 LLM 的影像字典。

從屬端使用 BACKGROUND,因此 SDK 會儲存訊息、產生佇列描述以及在產生完成前傳回。當描述必須在 add_messages() 傳回之前就緒時,請使用 INLINE。

請先呼叫 wait_for_memory_extraction(),再讀取或搜尋產生的描述。此方法會等待此從屬端佇列的 image-description 與記憶體擷取工作。

產生成功後,get_message() 會傳回影像部分 ImageContent.description 欄位中的描述:

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)

如果產生失敗或佇列拒絕任務,則影像會保留不包含描述。請先檢查 ImageContent.description,再使用產生的文字。

您也可以使用 TextContent 和 ImageContent 物件而非字典來建構訊息。這兩個表單都儲存相同的訊息。

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)

由於此 ImageContent 包含描述,因此 SDK 不會傳送影像給 LLM 以產生描述。

API 參照:Llm 記憶體擷取組態 訊息與訊息內容 繁體中文

搜尋影像

設定 record_types=["image"] 以搜尋影像描述。這會同時搜尋獨立影像與附加至訊息的影像。每個結果都是不含原始影像位元組的 ImageRecord。

您可以自行提供描述,或讓已設定的支援視覺功能的 LLM 產生描述。SDK 會以與提供的描述相同的方式儲存和搜尋產生的描述。

使用本指南中使用的預設 VECTOR 策略時,SDK 會視需要將影像描述分割成區塊,並將這些區塊內嵌在設定的 Embedder 中。它會內嵌具有相同內嵌器的查詢,並比較產生的向量。此程序與附加至訊息的獨立影像和影像相同。

其他搜尋策略的處理程序描述不同。KEYWORD 會搜尋儲存的描述文字,但不建立內嵌項目。HYBRID 會將文字比對與其設定的 OracleDBEmbedder 所產生的向量結合。

下列範例會搜尋這兩個影像。每個結果都包含影像 ID。附加的影像也包含其父項訊息的 ID。

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)

後續的擷取區段會使用這些 ID 來載入原始位元組。

使用 record_types=["message"] 時,搜尋只會檢查訊息文字。它不會搜尋附加至訊息的影像描述;請在這些描述中使用 record_types=["image"]。

API 參考: OracleAgentMemory 繁體中文 OracleSearch 結果 嵌入器

擷取獨立映像檔位元組

list_images() 預設會傳回影像描述資料和描述,但不會傳回儲存的位元組。若要擷取位元組,請設定 include_bytes=True 並提供搜尋結果中的影像 ID 以及相符的 user_id、agent_id 或 thread_id。這些篩選會將要求限制為一個影像。

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]
API 參考: OracleAgentMemory 影像紀錄

擷取附加至訊息之影像的位元組

依照預設,get_message() 和 get_messages() 傳回的影像部分不包含其位元組。若要擷取影像,請將其 ID 從搜尋結果傳送至 get_message(..., included_image_ids=[...])。搜尋結果的 message_id 會識別要擷取的訊息。

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)
API 參考:訊息與訊息內容 繁體中文

管理附加至訊息的影像

附加的影像屬於其上階訊息。若要取代或移除附加的影像,請以新的訊息內容呼叫 update_message()。刪除此訊息會一併刪除附加至此訊息的每個影像。delete_image() 只會刪除獨立影像。您無法直接更新附加影像的 TTL。

新增附加的影像時,會收到上階訊息的到期時間。如果 update_message() 讓訊息更快到期,SDK 也可縮短影像的到期時間。延長訊息的到期時間,或使用 ttl_days=None 清除訊息,並不會延長或清除影像已儲存的到期時間。讀取和搜尋會在影像或其父項訊息過期後排除影像。

API 參考:訊息與訊息內容 繁體中文

選取何謂自動擷取記憶體

設定 MemoryExtractionConfig.memory_extraction_image_context 以控制包含記憶體擷取 LLM 接收之影像的訊息的哪些部分。此設定只會變更擷取提示。不會變更儲存的訊息,或產生缺少的圖片描述。

自動擷取記憶體的影像相關資訊環境

值 擷取 LLM 收到的內容 選取時機
DISABLED 只有訊息的文字部份。 解開時應該忽略影像 。這是預設值。
CAPTION 文字部分與影像描述,依其原始順序排列。 描述包含擷取所需的視覺資訊,或 LLM 提供者不得接收影像位元組。
IMAGE 文字部分和原始影像的原始順序。 記憶體取決於說明中未包含的視覺詳細資料。LLM 及其端點必須接受影像輸入。

下列範例使用 CAPTION。使用 memory_extraction_frequency=1 時,擷取會在第一個訊息之後執行。使用 MemoryExtractionMode.INLINE 時,擷取會在 add_messages() 傳回之前完成。

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)

若要改為傳送原始影像,請以 MemoryExtractionImageContext.IMAGE 取代 MemoryExtractionImageContext.CAPTION。在 OracleAgentMemory 上設定的 LLM 必須接受影像輸入。

在 CAPTION 模式中,每個傳送供擷取的影像都必須具有非空白的描述。在新增每個影像時提供 description,或設定具備視覺功能的 LLM 以產生描述。在 MemoryExtractionMode.BACKGROUND 中,SDK 會在擷取相同執行緒的記憶體之前產生佇列描述。如果影像在擷取開始時仍沒有描述,SDK 會拒絕要求。

請勿使用 MemoryExtractionImageContext.MEMORY。此值保留供未來使用,而 SDK 拒絕此值。

API 參照:MemoryExtractionImageContext 記憶體擷取組態

更新或刪除獨立映像檔

新增獨立影像之後,請使用 update_image() 來變更其描述、描述資料或位元組。描述會儲存在 ImageRecord.content 中,並編製索引以供搜尋。傳送字串以取代字串,或傳送 None 以使用設定的視覺 LLM 產生取代項目。如果您省略 description,現有描述會維持不變。

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)

此範例會在 update_image() 之後擷取影像,以驗證新的描述和描述資料。然後將 image_id 傳送至 delete_image(),並檢查是否已刪除一個影像。

若要取代位元組,請同時傳送 image 和 mime_type。在同一個呼叫中,您可以保留目前的描述、提供新的描述,或要求使用 description=None 產生的取代項目。與附加至訊息的影像不同,獨立影像可以有自己的描述資料、時戳和 TTL。

API 參考: OracleAgentMemory OracleSearch 結果

複查格式、限制及資料處理

在儲存影像之前,SDK 會解碼其位元組並驗證格式。如果您通過 mime_type,則解碼格式必須與它相符。SDK 接受 PNG、JPEG 與 WebP 圖片,但拒絕使用動畫 PNG 與 WebP。您無法停用此驗證。

依預設,原始影像最多可為 10 MiB。單一 LLM 要求最多可包含 100 個影像和 100 MiB 的影像資料。使用 ImageInputLimitConfig 可降低部署的這些限制,或提高至記載的最大值。本指南開頭的用戶端允許每個影像 10 MiB、每個要求 20 個影像,以及每個要求 50 MiB 的影像資料。

Oracle AI Agent Memory 會將影像位元組儲存在 Oracle AI Database 中。產生描述會將這些位元組傳送給設定的 LLM 提供者。記憶體擷取也會以 IMAGE 模式傳送位元組。在 CAPTION 模式中,記憶體擷取會改為傳送影像描述。描述可以顯示原始影像的資訊。請先複查安全考量,再將機密影像傳送至 LLM 以產生描述或擷取記憶體。

結論

在本指南中,我們學會如何新增獨立影像、將影像附加至訊息、擷取影像位元組、搜尋影像描述,以及設定自動擷取記憶體,以使用訊息文字、影像描述或原始影像。

→ 瞭解如何使用影像和多模型訊息之後,您現在可以繼續使用訊息和記憶體的存留時間。

完整代碼

本指南中包含完整的範例,可供您複製和執行。

#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.

#Oracle Agent Memory Code Example - Use Images and Multimodal Messages
#---------------------------------------------------------------------

##Configure a vision capable memory client

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()



##Store a standalone image

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)



##Add a message with an image dictionary

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]



##Generate an attached image description

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)



##Add typed multimodal message content

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)



##Search for images

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)



##Retrieve standalone image bytes

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]



##Retrieve attached image bytes

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)



##Extract memories from an image description

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)



##Update and delete a standalone image

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)