Understand Hybrid Vector Indexes

A hybrid vector index inherits all the information retrieval capabilities of Oracle Text search indexes and leverages the semantic search capabilities of Oracle AI Vector Search vector indexes.

Hybrid vector indexes allow you to index and query documents using a combination of full-text search and semantic vector search. A hybrid vector index is a class of specialized Domain Index that combines the existing Oracle Text indexing data structures and vector indexing data structures into one unified structure. A single index contains both textual and vector fields for a document, enabling you to perform a combination of keyword search and vector search simultaneously.

The purpose of a hybrid vector index is to enhance search relevance of an Oracle Text index by allowing users to search by both vectors and keywords in various combinations, using out-of-the-box and custom scoring techniques. By integrating traditional keyword-based text search with vector-based similarity search, you can improve the overall search experience and provide users with more accurate information.

When to Use a Hybrid Vector Index

Consider using a hybrid vector index for hybrid search scenarios where your query requires information that is semantically similar but pertains to a specific focus area, that is, involves a particular organization, user name, product code, technical term, date, or time. For example, a typical hybrid search query can be to find “top 10 instances of stock fraud for ABC Corporation”.

Such a query involves two separate components:

Pure keyword search may return results that specifically contain the query words like “stock”, “fraud”, “ABC”, or “Corporation” because it focuses on matching the exact keywords or surface-level representation of words or phrases with tokenized terms in a text index. Therefore, keyword search alone may not be suitable here because it can overlook the semantic meaning behind the words in our query, especially if the exact terms are not present in the content.

Pure vector search focuses on understanding the meaning and context of words or phrases rather than just matching keywords. Vector search considers semantic relationship between the query words, so it may include more contextually-relevant results like “corporate fraud”, “stock market manipulation”, “stock misconduct”, “financial irregularities”, or “lawsuits in the financial sector”. Vector search also may not be suitable here because it can include results about the broader topic of stock fraud involving ABC or similar organizations, especially if the exact phrase “stock fraud for ABC Corporation” is not present in the content.

Hybrid search can address both components of such a query by running keyword search and vector search on the same data and then combining the two search results into a single result set. In this way, you can utilize the strengths of both text indexes and vector indexes to retrieve the most relevant results.

Why Choose a Hybrid Vector Index?

Let us summarize the advantages of a hybrid vector index.

Use Case Examples of a Hybrid Vector Index

Here are some use case scenarios to understand how you can implement a hybrid vector index.

Hybrid Vector Index Creation Overview

You create a hybrid vector index by simply specifying on which table and column to create it along with some details, such as the local or remote location where all source documents are stored (datastore), the ONNX in-database embedding model to use for generating embeddings, and the type of vector index to create. You can specify additional parameters that are discussed later in this chapter.

As illustrated in the following diagram, the hybrid vector index DDL creates a single index that contains both textual fields (with derived text tokens) and vector fields (with extracted chunks and corresponding embeddings) for each indexed document.

Description of image follows

Description of the illustration overview_hybrid_vector_index.png

As you can see, the implementation of a hybrid vector index leverages the existing capabilities of the Oracle Text search index and the Oracle AI Vector Search vector index. You can define PL/SQL preferences to customize all these indexing pipeline stages for both the index types.

A document table DOCS contains IDs and corresponding document names or file names stored in a location called MY_DS datastore.

The indexing pipeline starts with reading the documents from MY_DS (datastore), and then passes the documents through a series of processing stages:

  1. Filter (conversion of binary documents such as PDF, Word, or Excel to plain text)

  2. Tokenizer (tokenization of data for keyword search) and Vectorizer (chunking and embedding generation for vector search)

  3. Indexing Engine (creation of secondary tables)

    As the system passes documents through this indexing engine, it populates and indexes a set of secondary tables that are collectively part of a hybrid vector index.

    The two main secondary tables created are:

    • $I is the same structure as the existing Oracle Text index, which contains the inverted indexed data with tokenized terms.

      An Oracle Text index is created over a textual column based on your specifications. For a deeper understanding of the Oracle Text indexing process, see Oracle Text Application Developer’s Guide.

    • $VR contains the indexed data with generated chunks and corresponding embeddings.

      A vector index is also created over the VECTOR column based on your specifications. $VR also contains a ROWID column that maps back to the document table’s row IDs, and a DOCID column that links back to Oracle Text document-level IDs. This creates a link between chunks, tokens, and documents.

      You can directly examine the $VR table using the dictionary view <index name>$VECTORS, which lets you query all row ids, chunks, and embeddings. See <index name>$VECTORS.

    Temporarily, the $D secondary table (not shown in the diagram) is also created to save a copy of the document if the original document is not already in the database or needs filtering. This table gets truncated after chunks are obtained.

Hybrid Vector Index Maintenance Operations

A hybrid vector index supports mostly all the traditional operations of an Oracle Text index, such as MAINTENANCE AUTO , SYNC, or OPTIMIZE. For details, see Guidelines and Restrictions for Hybrid Vector Indexes.

See Also: