How To Improve Search Result Relevance

Search can retrieve useful records for a query, but a broad search can still return more direct results than an application can use, including results that are only loosely related to the question.

Search post-processing lets an application control the size and focus of that result set after retrieval. This is especially useful when search results are inserted into an agent prompt or a context card with limited space.

This guide explains how to configure post-search behavior for regular searches and context-card retrieval. Use TopKMemorySearchConfig to return a fixed maximum number of direct search results and optionally rerank them. Use PruningMemorySearchConfig to additionally remove less relevant results with an LLM.

Note: When building a context card, Oracle Agent Memory searches memory-like records separately from messages in the current thread. When reranking or pruning is configured, it is applied separately to each set of results before the sets are combined. The combined results are limited by max_relevant_results and the token budget.

When min_relevant_results_by_type is used, each requested record type is searched separately. Reranking or pruning, when configured, is also applied separately to the results of each search. Another search across all memory-like record types can then use any result slots that remain. By comparison, if reranking or pruning is configured for a regular search that requests multiple record_types, it is applied once to the combined search results.

In this guide, you will:

Hint: For package setup, see the Get Started with Agent Memory. If you need a local Oracle AI Database for this example, follow Run Oracle AI Database locally.

Configure Search Components

Create the model components used by the search configuration. The reranker receives the query and raw document text, while the pruner uses the LLM to remove less relevant direct results.


import oracledb

from oracleagentmemory.core import (
    OracleAgentMemory,
    PruningEvaluationMode,
    PruningMemorySearchConfig,
    Reranker,
    SchemaPolicy,
    TopKMemorySearchConfig,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm

embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
search_llm = Llm(model="provider/search-llm-model")
reranker = Reranker(model="provider/reranker-model")
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_SEARCH_RESULTS"
API Reference: Reranker Llm

Use a Fixed Direct-Result Limit and Reranking

Pass TopKMemorySearchConfig when creating the client or thread. The search first retrieves at most max_results direct results. The optional reranker then receives only those results and changes their order without changing their identity.

top_k_config = TopKMemorySearchConfig(
    max_results=5,
    reranker=reranker
)
memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=search_llm,
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
    search_config=top_k_config,
)
user_id = "user_123"
API Reference: TopKMemorySearchConfig OracleAgentMemory

Use LLM-Based Pruning

Set pruner_llm on OracleAgentMemory to enable client-wide pruning with the default FAST evaluation mode. New threads inherit this setting, while threads that already have a stored search configuration keep it when reopened. Use PruningMemorySearchConfig when the initial search contains too many less relevant results and you need to tune the pruning behavior. It keeps the ranked direct results and uses the configured LLM to identify additional results that should be removed.

The direct-result limit is applied before pruning, so the pruner evaluates only the results selected by that limit and may reduce the result count further. Supply max_results on a regular search call. For context-card retrieval, max_relevant_results provides the corresponding initial limit.

pruning_memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=search_llm,
    memory_store_id=memory_store_id,
    pruner_llm=search_llm,
)
thread = pruning_memory.create_thread(
    thread_id="search_result_pruning_demo",
    user_id=user_id,
)
pruned_results = thread.search(
    query="Which database does the service use?",
    max_results=50,
)
#Use PruningMemorySearchConfig for advanced modes and tuning.
extended_pruning_config = PruningMemorySearchConfig(
    pruner=search_llm,
    evaluation_mode=PruningEvaluationMode.EXTENDED,
    num_probe_points=3,
    protected_fraction=0.1,
)

How Pruning Works in a Nutshell

The pruner evaluates ranked results to determine how many documents should be retained. protected_fraction ensures that the most relevant documents are always included. For example, with protected_fraction=0.1, the top 10% of ranked documents are protected from pruning and remain in the search results.

Set evaluation_mode with PruningEvaluationMode to control how much evaluation work pruning performs:

num_probe_points controls the number of ranked-result regions considered in FAST and EXTENDED modes. Omit it when using EXHAUSTIVE.

Note: Pruning uses an LLM, so it can make searches slower. Use it when getting a more focused result set is more important than minimizing search latency.

API Reference: PruningMemorySearchConfig PruningEvaluationMode

Apply Token Budgets

Use soft_token_budget to target the estimated size of the formatted output. Complete results are considered in rank order, and the result that reaches or crosses the target is retained. This keeps at least the first result when the search finds a match.

Use token_budget as a hard limit. A complete result that would exceed this limit is omitted, so a hard limit can produce no output when the first result is too large. Individual results are never split. We recommend setting both budgets, with the hard limit larger than the soft target. A hard limit about three times the soft target is a useful starting point.

For example, with soft_token_budget=800 and results estimated at 420, 300, and 250 tokens, all three are returned. The third result takes the cumulative estimate from 720 to 970 tokens and completes the soft-budget selection. If token_budget=900 is also set, only the first two are returned because the third would exceed the hard limit.

Configured budgets apply by default. Positive per-call values override their corresponding configured budgets, while non-positive values disable them.

results = memory.search(
    query="Which database does the service use?",
    user_id=user_id,
    soft_token_budget=1_000,
    token_budget=3_000,
)

#Per-call values override the configured budgets.
larger_results = memory.search(
    query="Which database does the service use?",
    user_id=user_id,
    soft_token_budget=4_000,
    token_budget=12_000,
)
API Reference: MemorySearchConfig TopKMemorySearchConfig

Conclusion

In this guide we learned how to configure direct-result limits, reranking, LLM-based pruning, and soft and hard formatted-result token budgets.

→ Having learned how to customize search result selection, you may now proceed to Customize Context Card Content.

Full Code

The complete example is included in this guide for you to copy and run.

#Copyright © 2026 Oracle and/or its affiliates.
#This software is under the Apache License 2.0
#(LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0) or Universal Permissive License
#(UPL) 1.0 (LICENSE-UPL or https://oss.oracle.com/licenses/upl), at your option.

#Oracle Agent Memory Code Example - Improve Search Result Relevance
#------------------------------------------------------------------

##Configure search components


import oracledb

from oracleagentmemory.core import (
    OracleAgentMemory,
    PruningEvaluationMode,
    PruningMemorySearchConfig,
    Reranker,
    SchemaPolicy,
    TopKMemorySearchConfig,
)
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm

embedder = Embedder(model="YOUR_EMBEDDING_MODEL")
search_llm = Llm(model="provider/search-llm-model")
reranker = Reranker(model="provider/reranker-model")
db_pool = oracledb.SessionPool(
    user="YOUR DB USER",
    password="YOUR DB PASSWORD",
    dsn="localhost:1521/...",
)
memory_store_id = "T_SEARCH_RESULTS"



##Configure top k search

top_k_config = TopKMemorySearchConfig(
    max_results=5,
    reranker=reranker
)
memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=search_llm,
    schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
    memory_store_id=memory_store_id,
    search_config=top_k_config,
)
user_id = "user_123"



##Configure pruned search

pruning_memory = OracleAgentMemory(
    connection=db_pool,
    embedder=embedder,
    llm=search_llm,
    memory_store_id=memory_store_id,
    pruner_llm=search_llm,
)
thread = pruning_memory.create_thread(
    thread_id="search_result_pruning_demo",
    user_id=user_id,
)
pruned_results = thread.search(
    query="Which database does the service use?",
    max_results=50,
)
#Use PruningMemorySearchConfig for advanced modes and tuning.
extended_pruning_config = PruningMemorySearchConfig(
    pruner=search_llm,
    evaluation_mode=PruningEvaluationMode.EXTENDED,
    num_probe_points=3,
    protected_fraction=0.1,
)



##Apply search token budgets

results = memory.search(
    query="Which database does the service use?",
    user_id=user_id,
    soft_token_budget=1_000,
    token_budget=3_000,
)

#Per-call values override the configured budgets.
larger_results = memory.search(
    query="Which database does the service use?",
    user_id=user_id,
    soft_token_budget=4_000,
    token_budget=12_000,
)