使用图像和多模式消息

代理通常需要记住图像和文本。屏幕截图、文档、图表和照片可以包含纯文本内存无法保留的详细信息。

在本指南中,您将学习如何:

提示:有关软件包设置,请参见 Get Started with Agent Memory 。如果在此示例中需要本地 Oracle AI Database,请遵循在本地运行 Oracle AI Database 。

创建 OracleAgentMemory 客户机

创建本指南中的示例使用的 OracleAgentMemory 客户机。它连接到 Oracle AI Database,配置嵌入器和具有视觉功能的 LLM,并为图像输入设置限制。有关这些限制的说明,请参见本指南末尾附近的“查看格式、限制和数据处理”。

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()

客户机在创建 Llm 时设置 supports_vision=True。仅当所选模型和提供程序端点接受映像输入时才设置此选项。如果省略它,Llm 将检查模型元数据。当没有元数据可用时,Llm 将发送一个小测试映像以确定端点是否接受映像输入。设置 supports_vision=True 将跳过这两个检查;它不会使纯文本模型能够处理图像。

API 参考:Llm 图像输入限制配置 内存提取配置

存储独立映像

使用 OracleAgentMemory.add_image() 添加独立映像。至少包含 user_id、agent_id 或 thread_id 中的一个。在以后调用 list_images() 或 search() 时,必须提供匹配的标识符。

对于搜索,SDK 表示具有文本说明的图像。通过向量搜索,它嵌入了用于仅文本搜索的相同文本嵌入器的描述;它不嵌入图像字节。如果将 description 传递给 add_image(),则 SDK 将使用该文本。如果省略它,配置的可视化 LLM 会自动生成说明。以下示例提供了说明,因此添加映像不会发出 LLM 请求。

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)

list_images() 返回与提供的 user_id、agent_id 或 thread_id 匹配的独立映像。每个 ImageRecord.content 值都包含图像说明。

要将独立映像与线程关联,请调用 OracleThread.add_image()。此方法使用线程 ID 以及存储在线程中的任何用户或代理 ID,因此您不会再次传递它们。

API 参考:OracleAgentMemory OracleThread 图像记录

将映像存储为消息的一部分

将图像存储为线程消息的一部分,当它为对话转动提供上下文时,不需要单独管理。

仅文本消息可以对 content 使用字符串。对于包含文本和图像的消息,传递内容部分的有序列表。SDK 在存储和检索消息以及构建内存提取提示时保留该顺序。

在字典形式中,每个内容部分都有一个 type:

支持的 mime_type 值包括 "image/png"、"image/jpeg" 和 "image/webp"。

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]

此示例中的图像字典省略 description,因此 SDK 将生成一个配置了支持视觉的 LLM。

客户端使用 BACKGROUND,因此 SDK 会存储消息、队列说明生成,并在生成完成之前返回。如果说明必须在 add_messages() 返回之前就绪,请使用 INLINE。

在读取或搜索生成的说明之前,请先调用 wait_for_memory_extraction()。该方法等待此客户机排队的图像描述和内存提取任务。

生成成功后,get_message() 将返回映像部分的 ImageContent.description 字段中的说明:

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)

如果生成失败或队列拒绝任务,则图像将保留而不提供说明。在使用生成的文本之前,请检查 ImageContent.description。

您还可以使用 TextContent 和 ImageContent 对象而不是字典来构造消息。这两种表单都存储相同的消息。

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)

由于此 ImageContent 包含说明,因此 SDK 不会将映像发送到 LLM 以生成说明。

API 参考:Llm 内存提取配置 消息和消息内容 OracleThread

搜索图像

将 record_types=["image"] 设置为搜索图像说明。这将搜索附加到消息的独立图像和图像。每个结果都是一个没有原始图像字节的 ImageRecord。

您可以自己提供描述,也可以让配置的可视化 LLM 生成描述。SDK 以与提供的说明相同的方式存储和搜索生成的说明。

使用本指南中使用的默认 VECTOR 策略,SDK 在必要时将图像描述拆分为块,并将这些块与配置的 Embedder 嵌入。它使用相同的嵌入器嵌入查询,并比较生成的向量。对于附加到消息的独立映像和映像,此过程相同。

其他搜索策略处理描述的方式不同。KEYWORD 在不创建嵌入的情况下搜索存储的说明文本。HYBRID 将文本匹配与其配置的 OracleDBEmbedder 生成的向量相结合。

以下示例搜索两个图像。每个结果都包含图像 ID。附加的图像还包括其父消息的 ID。

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)

后面的检索部分使用这些 ID 来加载原始字节。

使用 record_types=["message"] 时,搜索仅检查消息文本。它不搜索附加到消息的图像的描述;对这些描述使用 record_types=["image"]。

API 参考:OracleAgentMemory OracleThread OracleSearchResult 嵌入

检索独立映像字节数

缺省情况下,list_images() 返回映像元数据和说明,但不返回存储的字节。要检索字节,请设置 include_bytes=True 并从搜索结果中提供映像 ID 以及匹配的 user_id、agent_id 或 thread_id。这些筛选器将请求限制为一个图像。

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]
API 参考:OracleAgentMemory 图像记录

检索附加到消息的图像的字节数

缺省情况下,get_message() 和 get_messages() 返回的映像部分不包括其字节。要检索图像,请将其 ID 从搜索结果传递到 get_message(..., included_image_ids=[...])。搜索结果的 message_id 标识要检索的消息。

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)
API 参考:消息和消息内容 OracleThread

管理附加到消息的图像

附加的图像属于其父消息。要替换或删除附加的图像,请使用新消息内容调用 update_message()。删除消息还会删除附加到它的每个映像。delete_image() 仅删除独立映像。无法直接更新附加映像的 TTL。

添加附加的图像时,它将收到父消息的到期时间。如果 update_message() 使消息更快过期,则 SDK 还会缩短映像的过期时间。延长消息的失效时间,或者使用 ttl_days=None 清除消息,不会延长或清除已为映像存储的失效时间。在图像或其父消息过期后,读取和搜索将排除该图像。

API 参考:消息和消息内容 OracleThread

选择自动内存提取的含义

将 MemoryExtractionConfig.memory_extraction_image_context 设置为控制包含内存提取 LLM 接收的图像的消息的哪些部分。此设置仅更改提取提示。它不会更改存储的消息或生成缺少的图像描述。

自动内存提取的图像上下文

值 提取 LLM 接收的内容 在以下情况下选择它
DISABLED 仅消息的文本部分。 提取应忽略图像。这是默认设置。
CAPTION 文本部分和图像说明,按原始顺序排列。 这些说明包含提取所需的可视信息,或者 LLM 提供程序不能接收映像字节。
IMAGE 文本部分和原始图像,按原始顺序排列。 记忆取决于描述不包含的视觉细节。LLM 及其端点必须接受映像输入。

以下示例使用 CAPTION。使用 memory_extraction_frequency=1 时,提取将在第一条消息后运行。使用 MemoryExtractionMode.INLINE 时,提取将在 add_messages() 返回之前完成。

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)

要改为发送原始映像,请将 MemoryExtractionImageContext.CAPTION 替换为 MemoryExtractionImageContext.IMAGE。在 OracleAgentMemory 上配置的 LLM 必须接受映像输入。

在 CAPTION 模式下,发送进行提取的每个图像都必须具有非空说明。添加每个映像时提供 description,或配置支持可视的 LLM 以生成说明。使用 MemoryExtractionMode.BACKGROUND 时,SDK 会将描述生成排队,然后再为同一线程提取内存。如果在开始提取时图像仍然没有说明,则 SDK 会拒绝该请求。

请勿使用 MemoryExtractionImageContext.MEMORY。此值保留供将来使用,SDK 将拒绝该值。

API 参考:MemoryExtractionImageContext 内存提取配置

更新或删除独立映像

添加独立映像后,使用 update_image() 更改其说明、元数据或字节。说明存储在 ImageRecord.content 中,并为搜索编制索引。传递一个字符串以替换该字符串,或者传递 None 以使用配置的视觉 LLM 生成替换。如果省略 description,现有说明将保持不变。

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)

该示例在 update_image() 之后检索映像以验证新说明和元数据。然后,它将 image_id 传递到 delete_image() 并检查是否删除了一个映像。

要替换字节,请同时传递 image 和 mime_type。在同一通话中,您可以保留当前描述,提供新描述,或者使用 description=None 请求生成的替换。与附加到消息的映像不同,独立映像可以有自己的元数据、时间戳和 TTL。

API 参考:OracleAgentMemory OracleSearchResult

复核格式、限制和数据处理

在存储映像之前,SDK 会解码其字节并验证格式。如果传递 mime_type,则解码的格式必须与其匹配。SDK 接受 PNG、JPEG 和 WebP 图像,但拒绝动画 PNG 和动画 WebP。您无法禁用此验证。

默认情况下,原始映像最多可以达到 10 MiB。单个 LLM 请求最多可以包含 100 个图像和 100 MiB 图像数据。使用 ImageInputLimitConfig 可降低部署的这些限制,或将其提高到所记录的最大值。本指南开头的客户端允许每个图像 10 MiB、每个请求 20 个图像以及每个请求 50 MiB 的图像数据。

Oracle AI Agent Memory 将映像字节存储在 Oracle AI Database 中。生成说明会将这些字节发送到配置的 LLM 提供程序。内存提取还以 IMAGE 模式发送字节。在 CAPTION 模式下,内存提取将改为发送图像说明。描述可以显示原始图像中的信息。在将敏感映像发送到 LLM 以生成描述或提取内存之前,请查看 Security Considerations 。

结论

在本指南中,我们学习了如何添加独立图像、将图像附加到消息、检索图像字节、搜索图像描述以及配置自动内存提取以使用消息文本、图像描述或原始图像。

→学会了如何使用图像和多模式消息后,现在可以继续执行 Use Time-to-Live for Messages and Memories 。

完整代码

本指南中包括了完整的示例,供您复制和运行。

#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.

#Oracle Agent Memory Code Example - Use Images and Multimodal Messages
#---------------------------------------------------------------------

##Configure a vision capable memory client

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()



##Store a standalone image

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)



##Add a message with an image dictionary

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]



##Generate an attached image description

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)



##Add typed multimodal message content

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)



##Search for images

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)



##Retrieve standalone image bytes

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]



##Retrieve attached image bytes

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)



##Extract memories from an image description

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)



##Update and delete a standalone image

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)