HNSW Distribution on Oracle RAC

Learn how Hierarchical Navigable Small World (HNSW) indexes can be distributed across multiple Oracle Real Application Clusters (Oracle RAC) instances.

A distributed HNSW index on Oracle RAC enables scalable similarity search across multiple instances in a RAC (Real Application Clusters) environment. In a distributed HNSW index:

Why use a distributed HNSW index?

The following section summarize the advantages of using a distributed HNSW index.

Usage Notes

Distributed HNSW Index: High-Level Workflow

HNSW RAC Duplication vs. HNSW Index Distribution

Duplicated vs. Distributed

HNSW RAC Duplication HNSW Index Distribution
Data access

Each instance reads the entire vector dataset to build the full HNSW graph.
Data access

The vector dataset is split into vector distribution units (by ROWID RANGE, by [SUB]PARTITION). Each instance only reads a subset of vector data assigned to it.
Index structure

Each RAC instance has an identical, full copy of the HNSW graph for the entire vector dataset.
Index structure

Each instance holds only a portion (HNSW slice graph) of the full HNSW index.
Memory usage

High: Every instance must hold the entire index in memory.
Memory usage

Efficient: Memory load is split across instances, and each instance holds only a portion (HNSW slice graph) of the full HNSW index.
Index build process

The coordinating RAC instance creates a disk checkpoint of the HNSW graph that it created and then all the other participating RAC instances use that disk checkpoint to load the graph. Every instance independently loads the entire HNSW graph using the full vector dataset.
Index build process

Each instance builds an HNSW slice graph for a subset of vector data assigned to the instance.
Query execution

Queries execute on a single instance (no parallelism across RAC instances).
Query execution

Queries are distributed and parallelized across all instances holding HNSW slice graphs.
Resource utilization

Poor: Redundant storage and build compute on all instances.
Resource utilization

Good: Storage is not redundant and workload is balanced across all instances.
Scalability

Limited by the memory and processing power of a single instance.
Scalability

Scales with cluster size; can handle much larger datasets.
Cluster reconfiguration

All instances must reload or rebuild full HNSW index if cluster changes.
Cluster reconfiguration

Only affected HNSW slice graphs are redistributed or rebuilt.
Best use case

Small datasets
Best use case

Large or growing datasets.