イメージおよびマルチモーダル・メッセージの使用

多くの場合、エージェントは画像とテキストを記憶する必要があります。スクリーンショット、ドキュメント、チャートおよび写真には、テキストのみのメモリーで保持できない詳細を含めることができます。

このガイドでは、次のことを学習します。

ヒント:パッケージの設定については、エージェント・メモリーのスタート・ガイドを参照してください。この例にローカルのOracle AI Databaseが必要な場合は、「Oracle AI Databaseをローカルで実行」に従います。

OracleAgentMemoryクライアントの作成

このガイドの例で使用するOracleAgentMemoryクライアントを作成します。Oracle AI Databaseに接続し、埋込み機能とビジョン対応LLMを構成し、イメージ入力の制限を設定します。これらの制限については、このガイドの最後にあるレビュー形式、制限およびデータ処理を参照してください。

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()

クライアントは、Llmの作成時にsupports_vision=Trueを設定します。このオプションは、選択したモデルおよびプロバイダ・エンドポイントがイメージ入力を受け入れる場合にのみ設定します。省略すると、Llmによってモデル・メタデータがチェックされます。使用可能なメタデータがない場合、Llmは小さいテスト・イメージを送信して、エンドポイントがイメージ入力を受け入れるかどうかを判断します。supports_vision=Trueを設定すると、両方のチェックがスキップされます。イメージを処理できるテキストのみのモデルは作成されません。

APIリファレンス: Llm ImageInputLimitConfig メモリー抽出構成

スタンドアロン・イメージの格納

スタンドアロン・イメージを追加するには、OracleAgentMemory.add_image()を使用します。user_id、agent_idまたはthread_idの少なくとも1つを含めます。後でlist_images()またはsearch()をコールする場合は、一致する識別子を指定する必要があります。

検索の場合、SDKはテキストの説明を含むイメージを表します。ベクトル検索では、テキストのみの検索に使用されるのと同じテキスト・エンベダーでその説明が埋め込まれます。イメージ・バイトは埋め込まれません。descriptionをadd_image()に渡すと、SDKはそのテキストを使用します。これを省略すると、構成済のビジョン対応LLMによって説明が自動的に生成されます。次の例では説明を示しているため、イメージを追加してもLLMリクエストは作成されません。

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)

list_images()は、指定されたuser_id、agent_idまたはthread_idに一致するスタンドアロン・イメージを返します。各ImageRecord.content値には、イメージの説明が含まれます。

スタンドアロン・イメージをスレッドに関連付けるには、OracleThread.add_image()をコールします。このメソッドでは、スレッドIDと、スレッドに格納されているユーザーまたはエージェントIDが使用されるため、再度渡すことはありません。

APIリファレンス: OracleAgentMemory OracleThread イメージレコード

イメージをメッセージの一部として格納

会話のターンにコンテキストを提供し、別々に管理する必要がない場合は、イメージをスレッド・メッセージの一部として格納します。

テキストのみのメッセージでは、contentに文字列を使用できます。テキストおよびイメージを含むメッセージの場合は、コンテンツ・パートの順序付きリストを渡します。SDKは、メッセージを格納および取得するとき、およびメモリー抽出プロンプトを構築するときに、その順序を保持します。

ディクショナリ形式では、各コンテンツ・パートにtypeがあります。

サポートされているmime_type値は、"image/png"、"image/jpeg"および"image/webp"です。

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]

この例のイメージ・ディクショナリはdescriptionを省略するため、SDKは構成されたビジョン対応LLMでイメージ・ディクショナリを生成します。

クライアントはBACKGROUNDを使用するため、SDKはメッセージを格納し、説明の生成をキューに入れて、生成が終了する前に戻ります。INLINEは、add_messages()が返される前に説明の準備が整っている必要がある場合に使用します。

生成された説明を読み取るか検索する前に、wait_for_memory_extraction()をコールします。このメソッドは、このクライアントによってキューに入れられたimage-descriptionおよびmemory-extractionタスクを待機します。

生成が成功すると、get_message()はイメージ・パートのImageContent.descriptionフィールドに説明を返します。

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)

生成が失敗した場合、またはキューがタスクを拒否した場合、イメージは説明なしで保存されたままになります。生成されたテキストを使用する前に、ImageContent.descriptionを確認してください。

また、ディクショナリのかわりにTextContentおよびImageContentオブジェクトを使用してメッセージを作成することもできます。どちらのフォームにも同じメッセージが格納されます。

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)

このImageContentには説明が含まれているため、SDKは説明の生成のためにイメージをLLMに送信しません。

APIリファレンス: Llm メモリー抽出構成 メッセージとメッセージの内容 OracleThread

イメージの検索

イメージの説明を検索するには、record_types=["image"]を設定します。これにより、メッセージに添付されたスタンドアロン・イメージとイメージの両方が検索されます。各結果は、元のイメージ・バイトのないImageRecordになります。

説明は自分で入力するか、構成されたビジョン対応LLMで生成できます。SDKは、提供された説明と同じ方法で生成された説明を格納および検索します。

このガイドで使用されるデフォルトのVECTOR戦略では、SDKは必要に応じてイメージの説明をチャンクに分割し、それらのチャンクを構成済のEmbedderに埋め込みます。問合せを同じ埋込み子で埋め込み、結果のベクトルを比較します。このプロセスは、メッセージに添付されたスタンドアロン・イメージおよびイメージの場合と同じです。

その他の検索戦略では、説明の処理方法が異なります。KEYWORDは、埋込みを作成せずに、格納された説明テキストを検索します。HYBRIDは、テキスト一致を構成済のOracleDBEmbedderによって生成されたベクトルと組み合せます。

次の例では、両方のイメージを検索します。各結果にはイメージIDが含まれます。添付されたイメージには、その親メッセージのIDも含まれます。

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)

後続の取得セクションでは、これらのIDを使用して元のバイトがロードされます。

record_types=["message"]を指定すると、検索ではメッセージ・テキストのみが検査されます。メッセージに添付されたイメージの説明は検索しません。これらの説明にはrecord_types=["image"]を使用します。

APIリファレンス: OracleAgentMemory OracleThread OracleSearch結果 埋込み

スタンドアロン・イメージ・バイトの取得

デフォルトでは、list_images()はイメージ・メタデータおよび説明を返しますが、格納されたバイトは返しません。バイトを取得するには、include_bytes=Trueを設定し、一致するuser_id、agent_idまたはthread_idとともに検索結果からイメージIDを指定します。これらのフィルタは、リクエストを1つのイメージに制限します。

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]
APIリファレンス: OracleAgentMemory イメージレコード

メッセージに添付されたイメージのバイト数の取得

デフォルトでは、get_message()およびget_messages()によって返されるイメージ・パートには、そのバイトは含まれません。イメージを取得するには、検索結果からそのIDをget_message(..., included_image_ids=[...])に渡します。検索結果のmessage_idは、取得するメッセージを識別します。

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)
APIリファレンス: メッセージおよびメッセージ・コンテンツ OracleThread

メッセージに添付されたイメージの管理

添付されたイメージは、その親メッセージに属します。アタッチされたイメージを置換または削除するには、update_message()を新しいメッセージ・コンテンツでコールします。メッセージを削除すると、そのメッセージに添付されているすべてのイメージも削除されます。delete_image()は、スタンドアロン・イメージのみを削除します。アタッチされたイメージのTTLを直接更新することはできません。

添付されたイメージを追加すると、親メッセージの有効期限が届きます。update_message()によってメッセージの有効期限が早くなると、SDKによってイメージの有効期限も短縮されます。メッセージの有効期限を延長したり、ttl_days=Noneでクリアしても、イメージにすでに格納されている有効期限は延長またはクリアされません。イメージまたはその親メッセージの有効期限が切れた後、イメージを読み取りおよび検索で除外します。

APIリファレンス: メッセージおよびメッセージ・コンテンツ OracleThread

自動メモリー抽出の選択

メモリー抽出LLMが受信するイメージを含むメッセージのどの部分を制御するには、MemoryExtractionConfig.memory_extraction_image_contextを設定します。この設定では、抽出プロンプトのみが変更されます。保存されたメッセージは変更されず、欠落しているイメージの説明も生成されません。

自動メモリー抽出のイメージ・コンテキスト

値 抽出LLMが受け取るもの 次の場合に選択します。
DISABLED メッセージのテキスト部分のみ。 抽出ではイメージは無視されます。これがデフォルト値です。
CAPTION テキスト・パーツとイメージの説明を元の順序で記述します。 説明には、抽出に必要な視覚的な情報が含まれているか、LLMプロバイダがイメージ・バイトを受信しないようにする必要があります。
IMAGE テキスト・パーツと元のイメージを元の順序で並べます。 記憶は、説明に含まれていない視覚的な詳細に依存します。LLMとそのエンドポイントはイメージ入力を受け入れる必要があります。

次の例ではCAPTIONを使用しています。memory_extraction_frequency=1では、抽出は最初のメッセージの後に実行されます。MemoryExtractionMode.INLINEでは、add_messages()が返される前に抽出が完了します。

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)

かわりに元のイメージを送信するには、MemoryExtractionImageContext.CAPTIONをMemoryExtractionImageContext.IMAGEに置き換えます。OracleAgentMemoryで構成されたLLMは、イメージ入力を受け入れる必要があります。

CAPTIONモードでは、抽出のために送信されるすべてのイメージに空でない説明が必要です。各イメージを追加するときにdescriptionを指定するか、説明を生成するようにビジョン対応LLMを構成します。MemoryExtractionMode.BACKGROUNDを使用すると、SDKは、同じスレッドのメモリー抽出前に説明の生成をキューに入れます。抽出の開始時にイメージにまだ説明がない場合、SDKはリクエストを拒否します。

MemoryExtractionImageContext.MEMORYを使用しないでください。この値は、将来の使用のために予約されており、SDKによって拒否されます。

APIリファレンス: MemoryExtractionImageContext メモリー抽出構成

スタンドアロン・イメージの更新または削除

スタンドアロン・イメージを追加した後、update_image()を使用してその説明、メタデータまたはバイトを変更します。この説明はImageRecord.contentに格納され、検索用に索引付けされます。文字列を渡して置換するか、Noneを渡して構成済のビジョンLLMとの置換を生成します。descriptionを省略した場合、既存の説明は変更されません。

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)

この例では、update_image()の後にイメージを取得して、新しい説明およびメタデータを検証します。次に、image_idをdelete_image()に渡し、1つのイメージが削除されたことを確認します。

バイトを置換するには、imageとmime_typeを一緒に渡します。同じコールで、現在の説明を保持するか、新しい説明を指定するか、またはdescription=Noneを使用して生成された置換をリクエストできます。メッセージにアタッチされたイメージとは異なり、スタンドアロン・イメージには独自のメタデータ、タイムスタンプおよびTTLを含めることができます。

APIリファレンス: OracleAgentMemory OracleSearch結果

フォーマット、制限およびデータ処理の確認

イメージを格納する前に、SDKはそのバイトをデコードし、フォーマットを検証します。mime_typeを渡す場合、デコードされた形式はそれと一致する必要があります。SDKはPNG、JPEGおよびWebPイメージを受け入れますが、アニメーションPNGおよびアニメーションWebPは拒否されます。この検証を無効化することはできません。

デフォルトでは、RAWイメージは最大10MiBです。単一のLLMリクエストには、最大100個のイメージと100MiBのイメージ・データを含めることができます。ImageInputLimitConfigを使用して、デプロイメントのこれらの制限を下げるか、文書化された最大値まで上げます。このガイドの開始時のクライアントは、イメージごとに10MiB、リクエストごとに20イメージ、およびリクエストごとに50MiBのイメージ・データを許可します。

Oracle AI Agent Memoryでは、イメージ・バイトがOracle AI Databaseに格納されます。説明を生成すると、それらのバイトが構成済みのLLMプロバイダに送信されます。メモリー抽出では、バイトもIMAGEモードで送信されます。CAPTIONモードでは、かわりにメモリー抽出によってイメージの説明が送信されます。説明は、元のイメージから情報を表示できます。説明の生成またはメモリーの抽出のために機密イメージをLLMに送信する前に、セキュリティに関する考慮事項を確認してください。

まとめ

このガイドでは、スタンドアロン・イメージの追加、メッセージへのイメージのアタッチ、イメージ・バイトの取得、イメージの説明の検索、およびメッセージ・テキスト、イメージの説明または元のイメージを使用するように自動メモリー抽出を構成する方法を学習しました。

→ イメージおよびマルチモーダル・メッセージの使用方法を学習した後、「メッセージおよびメモリーに存続時間を使用」に進むことができます。

完全コード

このガイドには、コピーして実行するための完全な例が含まれています。

#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.

#Oracle Agent Memory Code Example - Use Images and Multimodal Messages
#---------------------------------------------------------------------

##Configure a vision capable memory client

from pathlib import Path

import oracledb

from oracleagentmemory.apis.message import ImageContent, ImageMimeType, Message, TextContent
from oracleagentmemory.core import (
    ImageInputLimitConfig,
    MemoryExtractionConfig,
    MemoryExtractionImageContext,
    MemoryExtractionMode,
    OracleAgentMemory,
    SchemaPolicy,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm


embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
vision_llm = Llm(
    model="YOUR_VISION_CAPABLE_MODEL",
    supports_vision=True,
)
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_IMAGE_GUIDE"

memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=vision_llm,
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
    image_input_limit_config=ImageInputLimitConfig(
        max_raw_image_bytes=10 * 1024 * 1024,
        max_images_per_llm_request=20,
        max_total_raw_image_bytes_per_llm_request=50 * 1024 * 1024,
    ),
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
)

image_bytes = Path("sample.png").read_bytes()



##Store a standalone image

image_id = memory.add_image(
    image_bytes,
    mime_type=ImageMimeType.PNG,
    user_id="customer_123",
    description="A red sample image used by the image support guide.",
    metadata={"source": "profile-photo"},
)

images = memory.list_images(user_id="customer_123", limit=10)
stored_image = next(image for image in images if image.id == image_id)
print(stored_image.content)



##Add a message with an image dictionary

thread = memory.create_thread(
    thread_id="image-support-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=False,
        extraction_mode=MemoryExtractionMode.BACKGROUND,
    ),
)

message_id = thread.add_messages(
    [
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Please remember what is in this image."},
                {
                    "type": "image",
                    "bytes": image_bytes,
                    "mime_type": "image/png",
                },
            ],
        }
    ]
)[0]



##Generate an attached image description

memory.wait_for_memory_extraction()

stored_message = thread.get_message(message_id)
attached_image = next(
    part for part in stored_message.content if isinstance(part, ImageContent)
)
print(attached_image.description)



##Add typed multimodal message content

typed_message_id = thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="This message uses the typed content API."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)[0]
print(typed_message_id)



##Search for images

standalone_matches = memory.search(
    "red sample image",
    user_id="customer_123",
    record_types=["image"],
    max_results=5,
)
for match in standalone_matches:
    print(match.record.id, match.content)

attached_matches = thread.search(
    "red square on a white background",
    record_types=["image"],
    max_results=5,
)
for match in attached_matches:
    print(match.record.message_id, match.record.id, match.content)



##Retrieve standalone image bytes

#Raw bytes are omitted by default. To load them, use the image ID returned by
#the search result and pass the same user, agent, or thread ID used when
#adding it.
standalone_match = next(
    match for match in standalone_matches if match.record.id == image_id
)
loaded_image = memory.list_images(
    image_id=standalone_match.record.id,
    user_id="customer_123",
    include_bytes=True,
)[0]



##Retrieve attached image bytes

#get_message() omits image bytes unless included_image_ids selects them.
attached_match = next(
    match
    for match in attached_matches
    if match.record.message_id == typed_message_id
)
message_with_bytes = thread.get_message(
    attached_match.record.message_id,
    included_image_ids=[attached_match.record.id],
)
hydrated_image = next(
    part for part in message_with_bytes.content if isinstance(part, ImageContent)
)



##Extract memories from an image description

caption_thread = memory.create_thread(
    thread_id="caption-extraction-thread",
    user_id="customer_123",
    memory_extraction_config=MemoryExtractionConfig(
        extract_memories=True,
        extraction_mode=MemoryExtractionMode.INLINE,
        memory_extraction_frequency=1,
        memory_extraction_image_context=MemoryExtractionImageContext.CAPTION,
    ),
)

caption_thread.add_messages(
    [
        Message(
            role="user",
            content=[
                TextContent(text="Remember the color and shape in this image."),
                ImageContent(
                    bytes=image_bytes,
                    mime_type=ImageMimeType.PNG,
                    description="A red square on a white background.",
                ),
            ],
        )
    ]
)



##Update and delete a standalone image

memory.update_image(
    image_id,
    description="A red square used by the image support guide.",
    metadata={"source": "reviewed-profile-photo"},
)

updated_image = memory.list_images(
    image_id=image_id,
    user_id="customer_123",
)[0]
print(updated_image.content)
print(updated_image.metadata)

deleted = memory.delete_image(image_id)
print(deleted)