LangGraph에서 에이전트 메모리 단기 API 사용
에이전트는 종종 모든 차례의 모델로 전체 대화를 다시 전달하지 않고 최근 작업 컨텍스트를 보존해야합니다. 최신 메시지만 상태로 유지하면 이전 태스크 세부정보, 중간 진행률 또는 활성 스레드 토픽을 쉽게 추적할 수 없습니다.
Oracle 에이전트 메모리는 이 문제에 대해 두 개의 서로 다른 단기 도우미를 노출합니다.
get_summary()는content가 스레드 저장 내용을 압축하는 요약 객체를 반환합니다. 압축 시 transcript 압축만 필요할 때 선호합니다.get_context_card()는content가 스레드 요약, 검색 항목, 관련 레코드 및 요청된 경우 관련 또는 최근 대화 메시지를 포함하는 프롬프트 준비 컨텍스트 블록인 컨텍스트 카드 객체를 반환합니다. 압축 시 현재 회전에 대한 검색 인식 컨텍스트를 유지해야 하는 경우 선호합니다.
응용 프로그램에서 최신 N 메시지를 원래 메시지로 유지하는 경우 다음과 같이 카드의 파생 콘텐츠에서 동일한 꼬리를 제외할 수 있습니다.
#External-tail pattern:
#LLM Prompt = [context card, last N raw messages]
thread.get_context_card(except_last_messages=N, max_recent_messages=0)
결과 프롬프트는 하나의 컨텍스트 카드 메시지와 원래의 테일 메시지를 포함하는 목록입니다.
prompt_messages = [
{"role": "user", "content": "<context_card>...</context_card>"},
{"role": "user", "content": "The latest user message"},
{"role": "assistant", "content": "The latest assistant reply"},
]
여러 대화 회전에 대해 동일한 컨텍스트 카드를 유지하고 LLM 프롬프트에 새 메시지를 추가하므로 프롬프트 캐싱을 활용하는 것이 좋습니다.
또는 카드를 자체 포함시키고 다음과 같이 원시 꼬리를 내부에 유지할 수 있습니다.
#Self-contained-card pattern:
#LLM Prompt = [context card only]
thread.get_context_card(
except_last_messages=N,
max_recent_messages=N,
)
결과 프롬프트는 하나의 메시지(<recent_messages> 섹션에 원시 테일이 포함된 컨텍스트 카드 메시지)만 포함된 목록입니다.
prompt_messages = [
{
"role": "user",
"content": (
"<context_card>..."
"<recent_messages>...</recent_messages>"
"</context_card>"
),
},
]
except_last_messages가 0이 아닌 경우 max_recent_messages은 0 또는 동일한 값이어야 합니다. 이렇게 하면 컨텍스트 카드 내부 및 외부에서 동일한 원시 테일이 제공되지 않습니다.
이 설명서에서는 미리 빌드된 에이전트 주위에 LangGraph 미들웨어를 사용하므로 실행 중인 프롬프트가 너무 커지면 Oracle 에이전트 메모리가 자동으로 회전을 지속하고 Oracle 컨텍스트 카드를 주입할 수 있습니다. 즉, 미들웨어는 구성된 임계값을 통과한 후 프롬프트를 압축합니다. 이 예제에서는 압축을 통해 검색 인식 컨텍스트를 보존해야 하므로 get_context_card()를 선택합니다.
현재 스레드에서 검색된 메시지를 포함하여 컨텍스트 카드의 관련 정보에 표시되는 레코드 유형을 조정하려면 컨텍스트 카드 콘텐츠 사용자정의를 참조하십시오.
중요: 요약, 컨텍스트 카드, 검색된 레코드 및 자동으로 추출된 메모리는 모델 파생 또는 검색된 텍스트이므로 신뢰할 수 없는 것으로 처리되어야 합니다. 자동 추출 또는 요약이 사용으로 설정된 경우 응용 프로그램이 특정 중간 값을 검토할 수 있기 전에 나중에 메모리 추출, 요약, 컨텍스트 카드 또는 에이전트 프롬프트와 같은 프롬프트에서 SDK에서 해당 텍스트를 재사용할 수도 있습니다. 애플리케이션이 소비하는 출력을 검토하고, 메모리에서 파생된 텍스트가 권한 있는 작업을 승인하지 않도록 하고, 파생된 텍스트가 나중에 추출 또는 컨텍스트 구성에 영향을 줄 수 있기 전에 워크플로우에 검토가 필요한 경우 memory_extraction_config=MemoryExtractionConfig(extract_memories=False) 또는 명시적 메모리 쓰기를 사용합니다.
이 자습서에서는 다음을 수행하는 방법에 대해 알아봅니다.
- Oracle 에이전트 메모리를
Embedder, Oracle 메모리 LLM으로 구성한 다음 LangGraphChatOpenAI모델과 쌍을 이룹니다. - 사전 구축된 LangGraph 에이전트를 새로운 회전을 지속하는 미들웨어로 래핑하고 토큰 압력 상승 및 신속한 압축 시작 시
get_context_card()출력을 주입합니다. - 답변은 나중에 전체 기록을 재전송하는 대신 Oracle 에이전트 메모리 스레드의 단기 컨텍스트로 바뀝니다.
힌트: 패키지 설정은 에이전트 메모리 시작하기를 참조하십시오. 이 예에 대한 로컬 Oracle AI Database가 필요한 경우 로컬에서 Oracle AI Database 실행을 따르십시오.
에이전트 메모리 및 LangGraph 모델 구성
Oracle DB 연결 또는 풀을 사용하여 Oracle 에이전트 메모리 클라이언트를 생성하고, 벡터 검색을 위해 Embedder를 구성하고, 컨텍스트 카드 분석을 위해 Oracle 메모리 LLM을 제공하고, LangGraph 에이전트에 ChatOpenAI를 사용합니다.
from typing import Any
import oracledb
from langchain.agents import create_agent
from langchain.agents.middleware import AgentMiddleware
from langchain_core.messages import AIMessage, BaseMessage, HumanMessage, RemoveMessage
from langchain_core.messages.utils import count_tokens_approximately
from langchain_openai import ChatOpenAI
from langgraph.graph.message import REMOVE_ALL_MESSAGES
from langgraph.runtime import Runtime
from oracleagentmemory.core import MemoryExtractionConfig, SchemaPolicy
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
from oracleagentmemory.core.oracleagentmemory import OracleAgentMemory
embedder = Embedder(
model="YOUR_EMBEDDING_MODEL",
api_base="YOUR_EMBEDDING_BASE_URL",
api_key="YOUR_EMBEDDING_API_KEY",
)
memory_llm = Llm(
model="YOUR_MEMORY_LLM_MODEL",
api_base="YOUR_MEMORY_LLM_BASE_URL",
api_key="YOUR_MEMORY_LLM_API_KEY",
temperature=0,
)
langgraph_llm = ChatOpenAI(
model="YOUR_CHAT_MODEL",
base_url="YOUR_CHAT_BASE_URL",
api_key="YOUR_CHAT_API_KEY",
temperature=0,
)
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_ST_MEMORY"
agent_memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=memory_llm,
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
)
thread_id = "langgraph_short_term_demo"
user_id = "user_123"
agent_id = "assistant_456"
| API 참조: OracleAgentMemory | 오라클 스레드 |
미들웨어 및 사전 구축된 에이전트 구성
미들웨어는 새 사용자를 유지하며 Assistant가 Oracle Agent Memory로 바뀝니다. 실행 중인 프롬프트가 토큰 임계값을 초과하면 전체 메시지 목록을 합성 oracle_context_card 메시지와 최신 원시 회전의 작은 테일로 바꾸어 상태를 압축합니다. 이렇게 하면 사전 구축된 에이전트 검색 인식 단기 컨텍스트를 제공하면서 LangGraph 상태가 압축 상태로 유지됩니다.
이 설명서에서는 토큰 기반 압축을 사용하지만 몇 번씩 압축하거나 다른 응용 프로그램 특정 트리거를 수행한 후 압축과 같은 다른 정책에 맞게 동일한 패턴을 조정할 수 있습니다.
def _message_text(message: BaseMessage | Any) -> str:
content = getattr(message, "content", "")
if isinstance(content, str):
return content
return str(content)
def _is_context_card_message(message: BaseMessage) -> bool:
return isinstance(message, HumanMessage) and (
getattr(message, "name", None) == "memory_context_card"
)
class OracleShortTermMemoryMiddleware(AgentMiddleware):
"""Persist LangGraph turns and compact prompts with an OracleAgentMemory context card.
Notes
-----
- ``before_model()`` receives the current LangGraph message state for this turn.
After compaction, that state already includes the synthetic ``memory_context_card``
message returned by a previous ``before_model()`` call.
- The middleware strips that synthetic message back out before persisting or
measuring token usage so OracleAgentMemory only stores real user/assistant turns
and the compaction threshold is based on the organic conversation.
- When compaction triggers, the middleware replaces the message history with one
context-card message plus the most recent raw turns. On the next turn, that
same injected message is seen again and filtered out before recomputing the
next compacted prompt.
"""
def __init__(
self,
memory: OracleAgentMemory,
thread_id: str,
user_id: str,
agent_id: str,
compaction_token_trigger: int,
kept_message_count: int,
) -> None:
self._thread = memory.create_thread(
thread_id=thread_id,
user_id=user_id,
agent_id=agent_id,
memory_extraction_config=MemoryExtractionConfig(
context_summary_update_frequency=4
),
)
self._compaction_token_trigger = int(compaction_token_trigger)
self._kept_message_count = int(kept_message_count)
self._persisted_message_ids: set[str] = set()
def before_model(
self,
state: dict[str, Any],
runtime: Runtime[Any],
) -> dict[str, Any] | None:
del runtime
messages = list(state["messages"])
#^ This will contain the context card message once the compaction occurs
raw_messages = [message for message in messages if not _is_context_card_message(message)]
self._persist_new_messages(raw_messages)
#we exclude the context card from the token counting
if count_tokens_approximately(raw_messages) < self._compaction_token_trigger:
return None
#External-tail pattern: exclude the raw tail from derived card
#content, then append the original messages separately.
context_card = self._thread.get_context_card(
except_last_messages=self._kept_message_count,
max_recent_messages=0,
).content
#Self-contained-card alternative: use the matching count for
#max_recent_messages and omit the raw-message tail from the return
#value.
#context_card = self._thread.get_context_card(
#except_last_messages=self._kept_message_count,
#max_recent_messages=self._kept_message_count,
#).content
if not context_card:
context_card = "<context_card>\n No relevant short-term context yet.\n</context_card>"
return {
"messages": [
RemoveMessage(id=REMOVE_ALL_MESSAGES), # Clear existing message state.
HumanMessage(content=context_card, name="memory_context_card"),
*raw_messages[-self._kept_message_count :],
]
}
def _persist_new_messages(self, messages: list[BaseMessage]) -> None:
persisted: list[dict[str, str]] = []
for message in messages:
#Persist only the conversational roles that map directly to short-
#term memory turns. Tool/system/synthetic messages are skipped here.
role = (
"user"
if isinstance(message, HumanMessage)
else "assistant" if isinstance(message, AIMessage) else None
)
if role is None:
continue
content = _message_text(message).strip()
if not content:
continue
#LangGraph messages usually have stable IDs. When they do not, fall back
#to a content-derived key so the same turn is not persisted repeatedly if
#the caller reuses the returned message list across later invocations.
message_id = str(getattr(message, "id", "") or f"{role}:{hash(content)}")
if message_id in self._persisted_message_ids:
continue
#Track what this middleware instance has already written so each real turn
#is added to Oracle once even though later turns may still carry the same
#messages in the LangGraph state.
self._persisted_message_ids.add(message_id)
persisted.append({"role": role, "content": content})
if persisted:
self._thread.add_messages(persisted)
short_term_middleware = OracleShortTermMemoryMiddleware(
memory=agent_memory,
thread_id=thread_id,
user_id=user_id,
agent_id=agent_id,
compaction_token_trigger=6000,
kept_message_count=3,
)
agent = create_agent(
model=langgraph_llm,
tools=[],
middleware=[short_term_middleware],
)
미들웨어 주입 컨텍스트를 사용하여 나중에 정답 제시
사용자 추가는 사전 구축된 에이전트의 실행 중인 메시지 목록으로 전환하고 미들웨어가 컨텍스트 카드를 삽입할 시기를 결정하도록 합니다. 나중에 차례가 오면 에이전트는 Oracle Agent Memory 단기 컨텍스트를 포함하는 Compact 상태로 응답할 수 있습니다. 예제에서는 삽입된 컨텍스트 카드를 인쇄하고 잘린 샘플을 포함하므로 전체 블록 인라인을 덤프하지 않고 프롬프트에 삽입된 압축을 검사할 수 있습니다.
messages: list[BaseMessage] = []
def print_current_context_card(messages: list[BaseMessage]) -> None:
for message in messages:
if _is_context_card_message(message):
print(_message_text(message))
return
print("<context_card>\n No injected context card yet.\n</context_card>")
def run_turn(user_text: str) -> str:
messages.append(HumanMessage(content=user_text))
result = agent.invoke({"messages": messages})
messages[:] = list(result["messages"])
assistant_message = next(
message for message in reversed(messages) if isinstance(message, AIMessage)
)
return _message_text(assistant_message)
run_turn(
"I'm Maya. I'm migrating our nightly invoice reconciliation workflow "
"from cron jobs to LangGraph."
)
run_turn("The failing step right now is ledger enrichment after reconciliation.")
final_answer = run_turn(
"What workflow am I migrating, which step is failing, and who am I?"
)
print_current_context_card(messages)
#<context_card>
#<topics>
#<topic>invoice reconciliation migration</topic>
#<topic>ledger enrichment failure</topic>
#...
#</topics>
#<summary>
#Maya is migrating the nightly invoice reconciliation workflow from cron jobs
#to LangGraph. The failing step is ledger enrichment after reconciliation.
#</summary>
#...
#</context_card>
print(final_answer)
#You're Maya, migrating your nightly invoice reconciliation workflow from cron jobs
#to LangGraph, and the ledger-enrichment step after reconciliation is currently failing.
결론
이 가이드에서는 get_summary()와 get_context_card()를 구분하고, 사전 구축된 LangGraph 에이전트를 중심으로 Oracle Agent Memory 단기 컨텍스트를 구성하고, 대화가 너무 커서 동사를 유지할 수 없을 때 미들웨어에서 컨텍스트 카드로 프롬프트를 압축하는 방법을 배웠습니다.
→ LangGraph 플로우에 단기 스레드 컨텍스트를 추가하는 방법을 배웠을 때 이제 Use Oracle Agent Memory with LangGraph로 이동할 수 있습니다.
전체 코드
뒤에 나오는 전체 코드를 복사합니다.
#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.
#Oracle Agent Memory Code Example - LangGraph Short-Term Memory
#--------------------------------------------------------------
##Configure Oracle Agent Memory and LangGraph models for short term context
from typing import Any
import oracledb
from langchain.agents import create_agent
from langchain.agents.middleware import AgentMiddleware
from langchain_core.messages import AIMessage, BaseMessage, HumanMessage, RemoveMessage
from langchain_core.messages.utils import count_tokens_approximately
from langchain_openai import ChatOpenAI
from langgraph.graph.message import REMOVE_ALL_MESSAGES
from langgraph.runtime import Runtime
from oracleagentmemory.core import MemoryExtractionConfig, SchemaPolicy
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
from oracleagentmemory.core.oracleagentmemory import OracleAgentMemory
embedder = Embedder(
model="YOUR_EMBEDDING_MODEL",
api_base="YOUR_EMBEDDING_BASE_URL",
api_key="YOUR_EMBEDDING_API_KEY",
)
memory_llm = Llm(
model="YOUR_MEMORY_LLM_MODEL",
api_base="YOUR_MEMORY_LLM_BASE_URL",
api_key="YOUR_MEMORY_LLM_API_KEY",
temperature=0,
)
langgraph_llm = ChatOpenAI(
model="YOUR_CHAT_MODEL",
base_url="YOUR_CHAT_BASE_URL",
api_key="YOUR_CHAT_API_KEY",
temperature=0,
)
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_ST_MEMORY"
agent_memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=memory_llm,
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
)
thread_id = "langgraph_short_term_demo"
user_id = "user_123"
agent_id = "assistant_456"
##Configure short term memory middleware and a prebuilt LangGraph agent
def _message_text(message: BaseMessage | Any) -> str:
content = getattr(message, "content", "")
if isinstance(content, str):
return content
return str(content)
def _is_context_card_message(message: BaseMessage) -> bool:
return isinstance(message, HumanMessage) and (
getattr(message, "name", None) == "memory_context_card"
)
class OracleShortTermMemoryMiddleware(AgentMiddleware):
"""Persist LangGraph turns and compact prompts with an OracleAgentMemory context card.
Notes
-----
- ``before_model()`` receives the current LangGraph message state for this turn.
After compaction, that state already includes the synthetic ``memory_context_card``
message returned by a previous ``before_model()`` call.
- The middleware strips that synthetic message back out before persisting or
measuring token usage so OracleAgentMemory only stores real user/assistant turns
and the compaction threshold is based on the organic conversation.
- When compaction triggers, the middleware replaces the message history with one
context-card message plus the most recent raw turns. On the next turn, that
same injected message is seen again and filtered out before recomputing the
next compacted prompt.
"""
def __init__(
self,
memory: OracleAgentMemory,
thread_id: str,
user_id: str,
agent_id: str,
compaction_token_trigger: int,
kept_message_count: int,
) -> None:
self._thread = memory.create_thread(
thread_id=thread_id,
user_id=user_id,
agent_id=agent_id,
memory_extraction_config=MemoryExtractionConfig(
context_summary_update_frequency=4
),
)
self._compaction_token_trigger = int(compaction_token_trigger)
self._kept_message_count = int(kept_message_count)
self._persisted_message_ids: set[str] = set()
def before_model(
self,
state: dict[str, Any],
runtime: Runtime[Any],
) -> dict[str, Any] | None:
del runtime
messages = list(state["messages"])
#^ This will contain the context card message once the compaction occurs
raw_messages = [message for message in messages if not _is_context_card_message(message)]
self._persist_new_messages(raw_messages)
#we exclude the context card from the token counting
if count_tokens_approximately(raw_messages) < self._compaction_token_trigger:
return None
#External-tail pattern: exclude the raw tail from derived card
#content, then append the original messages separately.
context_card = self._thread.get_context_card(
except_last_messages=self._kept_message_count,
max_recent_messages=0,
).content
#Self-contained-card alternative: use the matching count for
#max_recent_messages and omit the raw-message tail from the return
#value.
#context_card = self._thread.get_context_card(
#except_last_messages=self._kept_message_count,
#max_recent_messages=self._kept_message_count,
#).content
if not context_card:
context_card = "<context_card>\n No relevant short-term context yet.\n</context_card>"
return {
"messages": [
RemoveMessage(id=REMOVE_ALL_MESSAGES), # Clear existing message state.
HumanMessage(content=context_card, name="memory_context_card"),
*raw_messages[-self._kept_message_count :],
]
}
def _persist_new_messages(self, messages: list[BaseMessage]) -> None:
persisted: list[dict[str, str]] = []
for message in messages:
#Persist only the conversational roles that map directly to short-
#term memory turns. Tool/system/synthetic messages are skipped here.
role = (
"user"
if isinstance(message, HumanMessage)
else "assistant" if isinstance(message, AIMessage) else None
)
if role is None:
continue
content = _message_text(message).strip()
if not content:
continue
#LangGraph messages usually have stable IDs. When they do not, fall back
#to a content-derived key so the same turn is not persisted repeatedly if
#the caller reuses the returned message list across later invocations.
message_id = str(getattr(message, "id", "") or f"{role}:{hash(content)}")
if message_id in self._persisted_message_ids:
continue
#Track what this middleware instance has already written so each real turn
#is added to Oracle once even though later turns may still carry the same
#messages in the LangGraph state.
self._persisted_message_ids.add(message_id)
persisted.append({"role": role, "content": content})
if persisted:
self._thread.add_messages(persisted)
short_term_middleware = OracleShortTermMemoryMiddleware(
memory=agent_memory,
thread_id=thread_id,
user_id=user_id,
agent_id=agent_id,
compaction_token_trigger=6000,
kept_message_count=3,
)
agent = create_agent(
model=langgraph_llm,
tools=[],
middleware=[short_term_middleware],
)
##Answer later turns with the middleware backed agent
messages: list[BaseMessage] = []
def print_current_context_card(messages: list[BaseMessage]) -> None:
for message in messages:
if _is_context_card_message(message):
print(_message_text(message))
return
print("<context_card>\n No injected context card yet.\n</context_card>")
def run_turn(user_text: str) -> str:
messages.append(HumanMessage(content=user_text))
result = agent.invoke({"messages": messages})
messages[:] = list(result["messages"])
assistant_message = next(
message for message in reversed(messages) if isinstance(message, AIMessage)
)
return _message_text(assistant_message)
run_turn(
"I'm Maya. I'm migrating our nightly invoice reconciliation workflow "
"from cron jobs to LangGraph."
)
run_turn("The failing step right now is ledger enrichment after reconciliation.")
final_answer = run_turn(
"What workflow am I migrating, which step is failing, and who am I?"
)
print_current_context_card(messages)
#<context_card>
#<topics>
#<topic>invoice reconciliation migration</topic>
#<topic>ledger enrichment failure</topic>
#...
#</topics>
#<summary>
#Maya is migrating the nightly invoice reconciliation workflow from cron jobs
#to LangGraph. The failing step is ledger enrichment after reconciliation.
#</summary>
#...
#</context_card>
print(final_answer)
#You're Maya, migrating your nightly invoice reconciliation workflow from cron jobs
#to LangGraph, and the ledger-enrichment step after reconciliation is currently failing.