Use Images and Multimodal Messages
Agents often need to remember images as well as text. Screenshots, documents, charts, and photographs can contain details that text-only memory cannot preserve.
In this guide, you will learn how to:
- add standalone images and attach images to thread messages;
- search both types of images;
- retrieve the original image bytes;
- configure automatic memory extraction to use message text, image descriptions, or original images.
Hint: For package setup, see the Get Started with Agent Memory. If you need a local Oracle AI Database for this example, follow Run Oracle AI Database locally.
Create an OracleAgentMemory Client
Create the OracleAgentMemory client used by the examples in this guide. It
connects to Oracle AI Database, configures an embedder and a vision-capable LLM,
and sets limits for image input. See Review Formats, Limits, and Data Handling
near the end of this guide for an explanation of these limits.
from pathlib import Path
import oracledb
from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
ImageInputLimitConfig,
MemoryExtractionConfig,
MemoryExtractionImageContext,
MemoryExtractionMode,
OracleAgentMemory,
SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
model="YOUR_VISION_CAPABLE_MODEL",
supports_vision=True,
)
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"
memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=vision_llm,
memory_extraction_config=MemoryExtractionConfig(
extract_memories=False,
extraction_mode=MemoryExtractionMode.BACKGROUND,
),
image_input_limit_config=ImageInputLimitConfig(
max_raw_image_bytes=10 * 1024 * 1024,
max_images_per_llm_request=20,
max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
),
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
)
image_bytes = Path("sample.png").read_bytes()
The client sets supports_vision=True when it creates Llm. Set this
option only when the selected model and provider endpoint accept image input.
If you omit it, Llm checks the model metadata. When no metadata is
available, Llm sends a small test image to determine whether the endpoint
accepts image input. Setting supports_vision=True skips both checks; it does
not make a text-only model capable of processing images.
| API Reference: Llm | ImageInputLimitConfig | MemoryExtractionConfig |
Store a Standalone Image
Use OracleAgentMemory.add_image() to add a standalone image. Include at
least one of user_id, agent_id, or thread_id. You must provide a
matching identifier when you later call list_images() or search().
For search, the SDK represents an image with a text description. With vector
search, it embeds that description with the same text embedder used for
text-only search; it does not embed the image bytes. If you pass
description to add_image(), the SDK uses that text. If you omit it, the
configured vision-capable LLM generates the description automatically. The
following example provides a description, so adding the image does not make an
LLM request.
image_id = memory.add_image(
image_bytes,
mime_type=ImageMimeType.PNG,
user_id="customer_123",
description="A red sample image used by the image support guide.",
metadata={"source": "profile-photo"},
)
images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)
list_images() returns standalone images that match the supplied
user_id, agent_id, or thread_id. Each ImageRecord.content value
contains the image description.
To associate a standalone image with a thread, call OracleThread.add_image().
This method uses the thread ID and any user or agent IDs stored on the thread,
so you do not pass them again.
| API Reference: OracleAgentMemory | OracleThread | ImageRecord |
Store an Image as Part of a Message
Store an image as part of a thread message when it provides context for a conversation turn and does not need to be managed separately.
A text-only message can use a string for content. For a message that
contains text and images, pass an ordered list of content parts. The SDK
preserves that order when it stores and retrieves the message and when it
builds a memory-extraction prompt.
In dictionary form, each content part has a type:
- A text part is
{"type": "text", "text": "..."}. - An image part contains
"type": "image", rawbytes, andmime_type. It can also contain adescription.
The supported mime_type values are "image/png", "image/jpeg", and
"image/webp".
thread = memory.create_thread(
thread_id="image-support-thread",
user_id="customer_123",
memory_extraction_config=MemoryExtractionConfig(
extract_memories=False,
extraction_mode=MemoryExtractionMode.BACKGROUND,
),
)
message_id = thread.add_messages(
[
{
"role": "user",
"content": [
{"type": "text", "text": "Please remember what is in this image."},
{
"type": "image",
"bytes": image_bytes,
"mime_type": "image/png",
},
],
}
]
)[0]
The image dictionary in this example omits description, so the SDK
generates one with the configured vision-capable LLM.
The client uses BACKGROUND, so the SDK stores the message, queues
description generation, and returns before generation finishes. Use INLINE
when the description must be ready before add_messages() returns.
Call wait_for_memory_extraction() before reading or searching for the
generated description. The method waits for the image-description and
memory-extraction tasks queued by this client.
After generation succeeds, get_message() returns the description in the
image part’s ImageContent.description field:
memory.wait_for_memory_extraction()
stored_message = thread.get_message(message_id)
attached_image = next(
part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)
If generation fails or the queue rejects the task, the image remains stored
without a description. Check ImageContent.description before using the
generated text.
You can also construct the message with TextContent and ImageContent
objects instead of dictionaries. Both forms store the same message.
typed_message_id = thread.add_messages(
[
Message(
role="user",
content=[
TextContent(text="This message uses the typed content API."),
ImageContent(
bytes=image_bytes,
mime_type=ImageMimeType.PNG,
description="A red square on a white background.",
),
],
)
]
)[0]
print(typed_message_id)
Because this ImageContent includes a description, the SDK does not send the
image to an LLM for description generation.
| API Reference: Llm | MemoryExtractionConfig | Messages and message contents | OracleThread |
Search for Images
Set record_types=["image"] to search image descriptions. This searches both
standalone images and images attached to messages. Each result is an
ImageRecord without the original image bytes.
You can provide the description yourself or let the configured vision-capable LLM generate it. The SDK stores and searches generated descriptions in the same way as supplied descriptions.
With the default VECTOR strategy used in this guide, the SDK splits an
image description into chunks when necessary and embeds those chunks with the
configured Embedder. It embeds the query with the same embedder and
compares the resulting vectors. This process is the same for standalone images
and images attached to messages.
Other search strategies process descriptions differently. KEYWORD searches
the stored description text without creating embeddings. HYBRID combines
text matching with vectors produced by its configured OracleDBEmbedder.
The following example searches for both images. Each result includes the image ID. An attached image also includes the ID of its parent message.
standalone_matches = memory.search(
"red sample image",
user_id="customer_123",
record_types=["image"],
max_results=5,
)
for match in standalone_matches:
print(match.record.id, match.content)
attached_matches = thread.search(
"red square on a white background",
record_types=["image"],
max_results=5,
)
for match in attached_matches:
print(match.record.message_id, match.record.id, match.content)
The retrieval sections that follow use these IDs to load the original bytes.
With record_types=["message"], search examines only message text. It does
not search descriptions of images attached to messages; use
record_types=["image"] for those descriptions.
| API Reference: OracleAgentMemory | OracleThread | OracleSearchResult | Embedder |
Retrieve Standalone Image Bytes
By default, list_images() returns image metadata and descriptions, but not
the stored bytes. To retrieve the bytes, set include_bytes=True and provide
the image ID from the search result together with a matching user_id,
agent_id, or thread_id. These filters limit the request to one image.
#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
image_id=standalone_match.record.id,
user_id="customer_123",
include_bytes=True,
)[0]
| API Reference: OracleAgentMemory | ImageRecord |
Retrieve Bytes for an Image Attached to a Message
By default, image parts returned by get_message() and get_messages() do
not include their bytes. To retrieve an image, pass its ID from the search
result to get_message(..., included_image_ids=[...]). The search result’s
message_id identifies the message to retrieve.
#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
match
for match in attached_matches
if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
attached_match.record.message_id,
included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
part for part in message_with_bytes.content if isinstance(part, ImageContent)
)
| API Reference: Messages and message contents | OracleThread |
Manage Images Attached to Messages
An attached image belongs to its parent message. To replace or remove an
attached image, call update_message() with new message content. Deleting the
message also deletes every image attached to it. delete_image() deletes
only standalone images. You cannot update the TTL of an attached image
directly.
When you add an attached image, it receives the parent message’s expiration
time. If update_message() makes the message expire sooner, the SDK also
shortens the image’s expiration time. Extending the message’s expiration time,
or clearing it with ttl_days=None, does not extend or clear the expiration
time already stored for the image. Reads and searches exclude the image after
either the image or its parent message has expired.
| API Reference: Messages and message contents | OracleThread |
Select What Automatic Memory Extraction Sees
Set MemoryExtractionConfig.memory_extraction_image_context to control which
parts of a message containing images the memory-extraction LLM receives. This
setting changes only the extraction prompt. It does not change the stored
message or generate a missing image description.
Image Context for Automatic Memory Extraction
| Value | What the extraction LLM receives | Select it when |
|---|---|---|
DISABLED |
Only the text parts of the message. | Extraction should ignore images. This is the default. |
CAPTION |
Text parts and image descriptions, in their original order. | The descriptions contain the visual information needed for extraction, or the LLM provider must not receive image bytes. |
IMAGE |
Text parts and the original images, in their original order. | The memories depend on visual details that the descriptions do not contain. The LLM and its endpoint must accept image input. |
The following example uses CAPTION. With
memory_extraction_frequency=1, extraction runs after the first message.
With MemoryExtractionMode.INLINE, extraction completes before
add_messages() returns.
caption_thread = memory.create_thread(
thread_id="caption-extraction-thread",
user_id="customer_123",
memory_extraction_config=MemoryExtractionConfig(
extract_memories=True,
extraction_mode=MemoryExtractionMode.INLINE,
memory_extraction_frequency=1,
memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
),
)
caption_thread.add_messages(
[
Message(
role="user",
content=[
TextContent(text="Remember the color and shape in this image."),
ImageContent(
bytes=image_bytes,
mime_type=ImageMimeType.PNG,
description="A red square on a white background.",
),
],
)
]
)
To send the original image instead, replace
MemoryExtractionImageContext.CAPTION with
MemoryExtractionImageContext.IMAGE. The LLM configured on
OracleAgentMemory must accept image input.
In CAPTION mode, every image sent for extraction must have a nonempty
description. Provide description when adding each image, or configure a
vision-capable LLM to generate the descriptions. With
MemoryExtractionMode.BACKGROUND, the SDK queues description generation
before memory extraction for the same thread. If an image still has no
description when extraction begins, the SDK rejects the request.
Do not use MemoryExtractionImageContext.MEMORY. This value is reserved for
future use, and the SDK rejects it.
| API Reference: MemoryExtractionImageContext | MemoryExtractionConfig |
Update or Delete a Standalone Image
After adding a standalone image, use update_image() to change its
description, metadata, or bytes. The description is stored in
ImageRecord.content and indexed for search. Pass a string to replace it, or
pass None to generate a replacement with the configured vision LLM. If you
omit description, the existing description remains unchanged.
memory.update_image(
image_id,
description="A red square used by the image support guide.",
metadata={"source": "reviewed-profile-photo"},
)
updated_image = memory.list_images(
image_id=image_id,
user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)
deleted = memory.delete_image(image_id)
print(deleted)
The example retrieves the image after update_image() to verify the new
description and metadata. It then passes image_id to delete_image() and
checks that one image was deleted.
To replace the bytes, pass image and mime_type together. In the same
call, you can keep the current description, supply a new one, or request a
generated replacement with description=None. Unlike an image attached to a
message, a standalone image can have its own metadata, timestamps, and TTL.
| API Reference: OracleAgentMemory | OracleSearchResult |
Review Formats, Limits, and Data Handling
Before storing an image, the SDK decodes its bytes and verifies the format. If
you pass mime_type, the decoded format must match it. The SDK accepts PNG,
JPEG, and WebP images, but rejects animated PNG and animated WebP. You cannot
disable this validation.
By default, a raw image can be up to 10 MiB. A single LLM request can contain
up to 100 images and 100 MiB of image data. Use ImageInputLimitConfig to
lower these limits for your deployment or raise them up to the documented
maximum values. The client at the start of this guide allows 10 MiB per image,
20 images per request, and 50 MiB of image data per request.
Oracle AI Agent Memory stores image bytes in Oracle AI Database. Generating a
description sends those bytes to the configured LLM provider. Memory extraction
also sends the bytes in IMAGE mode. In CAPTION mode, memory extraction
sends image descriptions instead. A description can reveal information from the
original image. Review the Security Considerations before
sending sensitive images to an LLM for description generation or memory
extraction.
Conclusion
In this guide we learned how to add standalone images, attach images to messages, retrieve image bytes, search image descriptions, and configure automatic memory extraction to use message text, image descriptions, or original images.
→ Having learned how to use images and multimodal messages, you may now proceed to Use Time-to-Live for Messages and Memories.
Full Code
The complete example is included in this guide for you to copy and run.
#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.
#Oracle Agent Memory Code Example - Use Images and Multimodal Messages
#---------------------------------------------------------------------
##Configure a vision capable memory client
from pathlib import Path
import oracledb
from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
ImageInputLimitConfig,
MemoryExtractionConfig,
MemoryExtractionImageContext,
MemoryExtractionMode,
OracleAgentMemory,
SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
model="YOUR_VISION_CAPABLE_MODEL",
supports_vision=True,
)
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"
memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=vision_llm,
memory_extraction_config=MemoryExtractionConfig(
extract_memories=False,
extraction_mode=MemoryExtractionMode.BACKGROUND,
),
image_input_limit_config=ImageInputLimitConfig(
max_raw_image_bytes=10 * 1024 * 1024,
max_images_per_llm_request=20,
max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
),
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
)
image_bytes = Path("sample.png").read_bytes()
##Store a standalone image
image_id = memory.add_image(
image_bytes,
mime_type=ImageMimeType.PNG,
user_id="customer_123",
description="A red sample image used by the image support guide.",
metadata={"source": "profile-photo"},
)
images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)
##Add a message with an image dictionary
thread = memory.create_thread(
thread_id="image-support-thread",
user_id="customer_123",
memory_extraction_config=MemoryExtractionConfig(
extract_memories=False,
extraction_mode=MemoryExtractionMode.BACKGROUND,
),
)
message_id = thread.add_messages(
[
{
"role": "user",
"content": [
{"type": "text", "text": "Please remember what is in this image."},
{
"type": "image",
"bytes": image_bytes,
"mime_type": "image/png",
},
],
}
]
)[0]
##Generate an attached image description
memory.wait_for_memory_extraction()
stored_message = thread.get_message(message_id)
attached_image = next(
part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)
##Add typed multimodal message content
typed_message_id = thread.add_messages(
[
Message(
role="user",
content=[
TextContent(text="This message uses the typed content API."),
ImageContent(
bytes=image_bytes,
mime_type=ImageMimeType.PNG,
description="A red square on a white background.",
),
],
)
]
)[0]
print(typed_message_id)
##Search for images
standalone_matches = memory.search(
"red sample image",
user_id="customer_123",
record_types=["image"],
max_results=5,
)
for match in standalone_matches:
print(match.record.id, match.content)
attached_matches = thread.search(
"red square on a white background",
record_types=["image"],
max_results=5,
)
for match in attached_matches:
print(match.record.message_id, match.record.id, match.content)
##Retrieve standalone image bytes
#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
image_id=standalone_match.record.id,
user_id="customer_123",
include_bytes=True,
)[0]
##Retrieve attached image bytes
#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
match
for match in attached_matches
if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
attached_match.record.message_id,
included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
part for part in message_with_bytes.content if isinstance(part, ImageContent)
)
##Extract memories from an image description
caption_thread = memory.create_thread(
thread_id="caption-extraction-thread",
user_id="customer_123",
memory_extraction_config=MemoryExtractionConfig(
extract_memories=True,
extraction_mode=MemoryExtractionMode.INLINE,
memory_extraction_frequency=1,
memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
),
)
caption_thread.add_messages(
[
Message(
role="user",
content=[
TextContent(text="Remember the color and shape in this image."),
ImageContent(
bytes=image_bytes,
mime_type=ImageMimeType.PNG,
description="A red square on a white background.",
),
],
)
]
)
##Update and delete a standalone image
memory.update_image(
image_id,
description="A red square used by the image support guide.",
metadata={"source": "reviewed-profile-photo"},
)
updated_image = memory.list_images(
image_id=image_id,
user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)
deleted = memory.delete_image(image_id)
print(deleted)