07 / STORAGE & INDEXES

Vector database: choose the index for the workload

An index choice should follow a measurable workload. Evaluate useful results, latency, filters, memory, and recovery together instead of choosing from a single speed claim.

HNSW versus IVFFlat: graph connections and grouped vectors on a bright-yellow typography card.
01 / KEY IDEA

Exact baseline

Establish nearest neighbors under the chosen metric before introducing index approximation.

02 / KEY IDEA

Approximate indexes

Measure the recall, latency, build, and memory tradeoffs of the configuration you actually deploy.

03 / KEY IDEA

Operational fit

Include filters, updates, deletion, backups, migrations, and recovery in the acceptance criteria.

Start with data and query contracts

Define what each row represents and which embedding configuration produced it. Record the metric, dimension, numerical type, and normalization policy. A database can store arrays correctly while the application compares incompatible representations. The token vector guide covers this boundary.

Describe the queries that matter: short questions, identifiers, version constraints, and tenant restrictions. Measure the number and size of eligible records rather than assuming every request searches the same collection.

Establish exact neighbors before approximation

For a manageable reference set, compare each query against all eligible vectors. This defines the exact neighbors under the chosen representation and metric. It does not establish that those neighbors are useful answers; task relevance needs a separate labeled evaluation.

The pgvector documentation describes exact retrieval and approximate indexes including HNSW and IVFFlat. Use its deployment-specific details as a reference, then measure your own dataset rather than adopting a universal tuning recipe.

Include filters in the experiment

A tenant restriction or product-version condition can change the amount of eligible data and the behavior of approximate retrieval. Test realistic selectivity and inspect the actual query plan. Compare equivalent queries, not an unfiltered benchmark against a filtered production request.

Verify both result count and authorization. Returning fewer useful neighbors may require a different search strategy or index configuration, but it is never a reason to weaken security boundaries. Small eligible sets can justify different approaches from a large shared corpus.

Calculate payload, then measure the system

Raw vector payload is record count multiplied by dimension multiplied by bytes per coordinate for a fixed-width representation. That calculation excludes text, metadata, index structures, database overhead, replicas, backups, and temporary migration space. Keep it labeled as a payload estimate.

Sample representative documents to estimate passage counts and measure actual stored sizes. The vector database cost article uses transparent hypothetical arithmetic and separates ingestion, queries, compression, and replacement workloads.

Treat replacement and recovery as requirements

Measure build time, peak resources, update behavior, deletion propagation, and restoration. A configuration that is comfortable at steady state may need substantially different capacity while old and new collections coexist.

Preserve a known-good configuration and a rollback plan. Define acceptance thresholds before tuning so the choice reflects application requirements rather than whichever test makes a favored technology look strongest. The token vector search overview connects the storage decision to candidate quality, ranking, and an inspectable user experience.

FOLLOW THE CONNECTIONToken Vector Search