How To Improve Search Result Relevance
Search can retrieve useful records for a query, but a broad search can still return more direct results than an application can use, including results that are only loosely related to the question.
Search post-processing lets an application control the size and focus of that result set after retrieval. This is especially useful when search results are inserted into an agent prompt or a context card with limited space.
This guide explains how to configure post-search behavior for regular searches
and context-card retrieval. Use TopKMemorySearchConfig to return a fixed
maximum number of direct search results and optionally rerank them. Use
PruningMemorySearchConfig to additionally remove less relevant results
with an LLM.
Note: When building a context card, Oracle Agent Memory searches memory-like
records separately from messages in the current thread. When reranking or
pruning is configured, it is applied separately to each set of results
before the sets are combined. The combined results are limited by
max_relevant_results and the token budget.
When min_relevant_results_by_type is used, each requested record type
is searched separately. Reranking or pruning, when configured, is also
applied separately to the results of each search. Another search across
all memory-like record types can then use any result slots that remain. By
comparison, if reranking or pruning is configured for a regular search that
requests multiple record_types, it is applied once to the combined
search results.
In this guide, you will:
- configure a fixed maximum number of direct search results
- use reranking
- use LLM-based pruning
- use configured and per-call token budgets
Hint: For package setup, see the Get Started with Agent Memory. If you need a local Oracle AI Database for this example, follow Run Oracle AI Database locally.
Configure Search Components
Create the model components used by the search configuration. The reranker receives the query and raw document text, while the pruner uses the LLM to remove less relevant direct results.
import oracledb
from oracleagentmemory.core import (
OracleAgentMemory,
PruningEvaluationMode,
PruningMemorySearchConfig,
Reranker,
SchemaPolicy,
TopKMemorySearchConfig,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
search_llm = Llm(model="provider/search-llm-model")
reranker = Reranker(model="provider/reranker-model")
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_SEARCH_RESULTS"
| API Reference: Reranker | Llm |
Use a Fixed Direct-Result Limit and Reranking
Pass TopKMemorySearchConfig when creating the client or thread. The
search first retrieves at most max_results direct results. The optional
reranker then receives only those results and changes their order without
changing their identity.
top_k_config = TopKMemorySearchConfig(
max_results=5,
reranker=reranker
)
memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=search_llm,
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
search_config=top_k_config,
)
user_id = "user_123"
| API Reference: TopKMemorySearchConfig | OracleAgentMemory |
Use LLM-Based Pruning
Set pruner_llm on OracleAgentMemory to enable client-wide pruning with
the default FAST evaluation mode. New threads inherit this setting, while
threads that already have a stored search configuration keep it when reopened.
Use PruningMemorySearchConfig when the initial search contains too many
less relevant results and you need to tune the pruning behavior. It keeps the
ranked direct results and uses the configured LLM to identify additional
results that should be removed.
The direct-result limit is applied before pruning, so the pruner evaluates
only the results selected by that limit and may reduce the result count
further. Supply max_results on a regular search call. For context-card
retrieval, max_relevant_results provides the corresponding initial limit.
pruning_memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=search_llm,
memory_store_id=memory_store_id,
pruner_llm=search_llm,
)
thread = pruning_memory.create_thread(
thread_id="search_result_pruning_demo",
user_id=user_id,
)
pruned_results = thread.search(
query="Which database does the service use?",
max_results=50,
)
#Use PruningMemorySearchConfig for advanced modes and tuning.
extended_pruning_config = PruningMemorySearchConfig(
pruner=search_llm,
evaluation_mode=PruningEvaluationMode.EXTENDED,
num_probe_points=3,
protected_fraction=0.1,
)
How Pruning Works in a Nutshell
The pruner evaluates ranked results to determine how many documents should be
retained. protected_fraction ensures that the most relevant documents are
always included. For example, with protected_fraction=0.1, the top 10% of
ranked documents are protected from pruning and remain in the search results.
Set evaluation_mode with PruningEvaluationMode to control how much
evaluation work pruning performs:
FASTis the default. It prioritizes low latency and may stop evaluation early.EXTENDEDevaluates a broader portion of the results, with additional latency and LLM usage.EXHAUSTIVEevaluates every candidate result individually and has the highest latency and LLM usage.
num_probe_points controls the number of ranked-result regions considered
in FAST and EXTENDED modes. Omit it when using EXHAUSTIVE.
Note: Pruning uses an LLM, so it can make searches slower. Use it when getting a more focused result set is more important than minimizing search latency.
| API Reference: PruningMemorySearchConfig | PruningEvaluationMode |
Apply Token Budgets
Use soft_token_budget to target the estimated size of the formatted output.
Complete results are considered in rank order, and the result that reaches or
crosses the target is retained. This keeps at least the first result when the
search finds a match.
Use token_budget as a hard limit. A complete result that would exceed this
limit is omitted, so a hard limit can produce no output when the first result
is too large. Individual results are never split. We recommend setting both
budgets, with the hard limit larger than the soft target. A hard limit about
three times the soft target is a useful starting point.
For example, with soft_token_budget=800 and results estimated at 420, 300,
and 250 tokens, all three are returned. The third result takes the cumulative
estimate from 720 to 970 tokens and completes the soft-budget selection. If
token_budget=900 is also set, only the first two are returned because the
third would exceed the hard limit.
Configured budgets apply by default. Positive per-call values override their corresponding configured budgets, while non-positive values disable them.
results = memory.search(
query="Which database does the service use?",
user_id=user_id,
soft_token_budget=1_000,
token_budget=3_000,
)
#Per-call values override the configured budgets.
larger_results = memory.search(
query="Which database does the service use?",
user_id=user_id,
soft_token_budget=4_000,
token_budget=12_000,
)
| API Reference: MemorySearchConfig | TopKMemorySearchConfig |
Conclusion
In this guide we learned how to configure direct-result limits, reranking, LLM-based pruning, and soft and hard formatted-result token budgets.
→ Having learned how to customize search result selection, you may now proceed to Customize Context Card Content.
Full Code
The complete example is included in this guide for you to copy and run.
#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.
#Oracle Agent Memory Code Example - Improve Search Result Relevance
#------------------------------------------------------------------
##Configure search components
import oracledb
from oracleagentmemory.core import (
OracleAgentMemory,
PruningEvaluationMode,
PruningMemorySearchConfig,
Reranker,
SchemaPolicy,
TopKMemorySearchConfig,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
search_llm = Llm(model="provider/search-llm-model")
reranker = Reranker(model="provider/reranker-model")
db_pool = oracledb.SessionPool(
user="YOUR DB USER",
password="YOUR DB PASSWORD",
dsn="localhost:1521/...",
)
memory_store_id = "T_SEARCH_RESULTS"
##Configure top k search
top_k_config = TopKMemorySearchConfig(
max_results=5,
reranker=reranker
)
memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=search_llm,
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
memory_store_id=memory_store_id,
search_config=top_k_config,
)
user_id = "user_123"
##Configure pruned search
pruning_memory = OracleAgentMemory(
connection=db_pool,
embedder=embedder,
llm=search_llm,
memory_store_id=memory_store_id,
pruner_llm=search_llm,
)
thread = pruning_memory.create_thread(
thread_id="search_result_pruning_demo",
user_id=user_id,
)
pruned_results = thread.search(
query="Which database does the service use?",
max_results=50,
)
#Use PruningMemorySearchConfig for advanced modes and tuning.
extended_pruning_config = PruningMemorySearchConfig(
pruner=search_llm,
evaluation_mode=PruningEvaluationMode.EXTENDED,
num_probe_points=3,
protected_fraction=0.1,
)
##Apply search token budgets
results = memory.search(
query="Which database does the service use?",
user_id=user_id,
soft_token_budget=1_000,
token_budget=3_000,
)
#Per-call values override the configured budgets.
larger_results = memory.search(
query="Which database does the service use?",
user_id=user_id,
soft_token_budget=4_000,
token_budget=12_000,
)