Understand Hybrid Search

With hybrid search, you can search through your documents by performing a combination of full-text queries and vector-based similarity queries, using out-of-the-box or custom scoring techniques.

Here are the key points to note when using hybrid search:

For the sake of the following explanations, we are using the same use case as seen in Understand Hybrid Vector Indexes:

Description of image follows

Description of the illustration overview_hybrid_vector_index.png

Pure Semantic in Document Mode

The pure semantic in document mode performs vector-only search to fetch document-level results.

The following SQL statement queries a hybrid vector index in this mode:

select json_Serialize(
  DBMS_HYBRID_VECTOR.SEARCH(
    json(
      '{ "hybrid_index_name" : "my_hybrid_idx",
         "vector":
          {
             "search_text"   : "galaxies formation and massive black holes",
             "search_mode"   : "DOCUMENT",
             "aggregator"    : "AVG"
          },
         "return":
          {
             "values"        : [ "rowid", "score", "vector_score" ],
             "topN"          : 10
          }
      }'
    )
  ) RETURNING CLOB pretty);

The result of this query may look like the following, where you see the ROWIDs of the DOCS table rows corresponding to documents as well as its vector score (vector_score). Here, the final score (score) is the same as the vector score because there is no keyword score in this case.

Here is an excerpt from the results:

[
{
    "rowid"        : "AAASBEAAEAAAUpqAAB",
    "score"        : 71.04,
    "vector_score" : 71.04
  },
  {
    "rowid"        : "AAASBEAAEAAAWBKAAE",
    "score"        : 67.82,
    "vector_score" : 67.82
  },
]

Here is how to interpret this statement:

  1. The system runs a similarity search on all the vectors and extracts the top k ones at most. The value k is internally calculated. Each is given a vector score.

  2. These k vectors (at most) are then grouped by document IDs and for each identified document, the semantic score of each associated vector found for that document is used to compute the average (in this example) of these scores for that particular document.

  3. The top 10 documents (at most) with the highest averages are returned.

This is illustrated here:

Description of image follows

Description of the illustration pure_semantic_document_mode_query.png

Pure Semantic in Chunk Mode

The pure semantic in chunk mode is the SQL semantic search equivalent. This mode performs a vector-only search to fetch chunk results.

The following SQL statement queries a hybrid vector index in this mode:

select json_Serialize(
  dbms_hybrid_vector.search(
    json(
      '{ "hybrid_index_name" : "my_hybrid_idx",
         "vector":
          {
             "search_text"   : "galaxies formation and massive black holes",
             "search_mode"   : "CHUNK"
          },
         "return":
          {
             "values"        : [ "score", "chunk_text", "chunk_id" ],
             "topN"          : 3
          }
      }'
    )
  ) RETURNING CLOB pretty);

Here is how to interpret this statement:

  1. The system runs a similarity search on all the vectors and extracts the top k ones at most. The value k is internally calculated. Each is given a vector score.

  2. The top 3 chunks (at most) with the highest scores are returned.

The result of this query may look like the following, where you can see the chunks corresponding to their semantic scores:

[
  {
    "score"      : 61,
    "chunk_text" : "Galaxies form through a complex process that begins with small fluctuations in the density of matter in the early
universe. Massive black holes, typically found at the centers of galaxies, are believed to play a crucial role in their formation and evolution.",
    "chunk_id"   : "1"
  },
  {
    "score"      : 56.64,
    "chunk_text" : "The presence of massive black holes in galaxies is closely linked to their morphological characteristics and star formation rates.
Observations suggest that as galaxies evolve, their central black holes grow in tandem with their host galaxy's mass.",
    "chunk_id"   : "3"
  },
  {
    "score"      : 55.75,
    "chunk_text" : "Black holes grow by accreting gas and merging with other black holes. Their gravitational influence can regulate star
formation and drive powerful jets of energy, which can impact the surrounding galaxy.",
    "chunk_id"   : "2"
  }
]

Pure Keyword in Document Mode

The pure keyword search in document mode is equivalent to the traditional CONTAINS query using Oracle Text. This mode performs a text-only search to fetch document-level results.

The following SQL statement queries a hybrid vector index in this mode:

select json_Serialize(
  dbms_hybrid_vector.search(
    json(
      '{ "hybrid_index_name" : "my_hybrid_idx",
         "text":
          {
             "contains"      : "galaxies, black holes"
          },
         "return":
          {
             "values"        : [ "rowid", "score" ],
             "topN"          : 3
          }
      }'
    )
  ) RETURNING CLOB pretty);

Here is how to interpret this statement:

  1. The system runs a CONTAINS query that returns a maximum number of documents. This maximum number is internally calculated. Each document is given a keyword score.

  2. The top 3 documents (at most) with the highest scores are returned.

The result of this query may look like the following, where you can see the ROWIDs of the DOCS table rows corresponding to documents as well as their keyword scores:

[
  {
    "rowid" : "AAAR9jAABAAAQeaAAB",
    "score" : 68
  },
  {
    "rowid" : "AAAR9jAABAAAQeaAAA",
    "score" : 35
  },
  {
    "rowid" : "AAAR9jAABAAAQeaAAD",
    "score" : 2
  }
]

Keyword and Semantic in Document Mode

Let us examine a non-pure case of hybrid search where keyword scores and semantic scores are combined.

The following SQL statement performs a keyword and semantic search to fetch document-level results:

select json_Serialize(
  DBMS_HYBRID_VECTOR.SEARCH(
    json(
      '{
         "hybrid_index_name" : "my_hybrid_idx",
         "search_scorer"     : "rsf",
         "search_fusion"     : "UNION",
         "vector":
          {
             "search_text"   : "How can I search with hybrid vector indexes?",
             "search_mode"   : "DOCUMENT",
             "aggregator"    : "MAX",
             "score_weight"  : 1,
             "rank_penalty"  : 5
          },
         "text":
          {
             "contains"      : "hybrid AND vector AND index"
             "score_weight"  : 10,
             "rank_penalty"  : 1
          },
         "return":
          {
             "values"        : [ "rowid", "score", "vector_score", "text_score" ],
             "topN"          : 10
          }
      }'
    )
  ) RETURNING CLOB pretty);

Here is how to interpret this statement:

In document mode, the result of your search is a list of ROWIDs from your base table corresponding to the list of best files identified.

To get to this list, two searches are conducted:

After the searches complete, the system needs to merge the results and score them as illustrated here:

Scoring for Keyword and Semantic Search in Document Mode

Description of image follows

Description of the illustration keyword_semantic_document_mode_query.png

As outlined in the preceding diagram,

The possible fuse operators are illustrated here:

Fusion Operation in Document Search Mode

Description of image follows

Description of the illustration fuse_operators_doc_mode.png

Keyword and Semantic in Chunk Mode

Let us examine another non-pure case of hybrid search where keyword scores and semantic scores are combined to fetch chunk results.

The following SQL statement performs a keyword and semantic search in chunk mode:

select json_Serialize(
  DBMS_HYBRID_VECTOR.SEARCH(
    json(
         '{
            "hybrid_index_name" : "my_hybrid_vector_idx",
            "search_scorer"     : "rsf",
            "search_fusion"     : "UNION",
            "vector":
                      {
                        "search_text"   : "How can I search with hybrid vector indexes?",
                        "search_mode"   : "CHUNK",
                        "score_weight"  : 1
                      },
            "text":
                      {
                       "contains"       : "hybrid AND vector AND index",
                       "score_weight"   : 1
                      },
            "return":
                      {
                        "values"        : [ "chunk_id", "score", "vector_score", "text_score" ],
                        "topN"          : 10
                      }
          }'
    )
  ) RETURNING CLOB pretty);

Here is how to interpret this statement:

In chunk mode, the result of your search is a list of best chunk identifiers from the files stored using your base table.

To get to this list, two searches are conducted:

After the searches complete, the system needs to merge the results and score them as illustrated here:

Scoring for Keyword and Semantic Search in Chunk Mode

Description of image follows

Description of the illustration keyword_semantic_chunk_mode_query.png

As outlined in the preceding diagram,

The possible fuse operators are illustrated here:

Fusion Operation in Chunk Search Mode

Description of image follows

Description of the illustration fuse_operators_chunk_mode.png

MINUS_VECTOR, UNION, and TEXT_ONLY are ignored. MINUS_VECTOR would eliminate all results. TEXT_ONLY and UNION are not possible because the right outer join excludes the non-overlapping text results.

See Also: