Choosing between HNSW and IVFFlat is not a contest to name the universally best vector index. It is a workload decision involving candidate recall, latency, memory, build time, update behavior, and filtering. A useful evaluation holds the vectors and queries constant, measures the tradeoffs, and includes the operational conditions your application will actually encounter.

This guide uses pgvector as a concrete reference while keeping the decision process broader than one database extension. It does not provide universal tuning numbers or vendor benchmarks. The aim is to help you prepare an experiment that distinguishes index behavior from embedding quality and produces a defensible deployment choice.

Establish what approximation changes

An exact nearest-neighbor search identifies the closest eligible vectors under a defined distance function. An approximate index searches more selectively to reduce work, which means it may miss some exact neighbors. This is separate from whether those neighbors answer the user’s question.

The pgvector project documentation describes exact search and its HNSW and IVFFlat index options. HNSW uses a layered graph, while IVFFlat groups vectors into lists and searches selected lists. The documentation also describes configuration and filtering behavior that should be checked against the extension version you deploy.

Begin with an exact baseline on a representative subset. Save the top results for a fixed query set using the same metric and filters as the approximate experiment. Without that baseline, a missing result could come from the encoder, the data, or the index, and your tuning will mix these causes together.

Understand the two search structures

HNSW organizes relationships between vectors in a graph with multiple layers. A query navigates that structure to locate promising candidates. A practical evaluation should consider not only query speed but also the memory and time needed to build and maintain the graph under your workload.

IVFFlat partitions the collection into lists and searches a selected subset of them. Its behavior depends on how the lists are built and how broadly a query searches them. Representative data matters when preparing the partitioning structure; an early collection that differs substantially from later data may not reflect the final workload.

These descriptions suggest different experiments rather than an automatic winner. A mostly static archive and a frequently updated operational collection can put pressure on different parts of the system. The vector database guide provides a broader framework for connecting index choices to application requirements.

Match the metric and operator

Choose the metric expected by the representation and configure the index accordingly. If the application ranks by cosine distance, the query and index configuration need to support that operation. A query that computes a related expression in a different form may not use the intended index path.

Inspect the actual execution plan in your database rather than assuming an index exists and therefore is used. Record the query shape, ordering direction, limit, filters, and relevant settings alongside benchmark results. A performance comparison is invalid when one test quietly uses a different execution strategy.

Also verify the stored vectors. Mixed encoder versions, inconsistent normalization, or a serialization error can produce poor neighbors regardless of the index. The cosine similarity article explains why metric selection and numerical conventions belong in the representation contract.

Design a workload that resembles production

Sample documents and queries across the application’s important segments. Include short queries, longer questions, frequent terms, rare identifiers, and content from small as well as large tenants where applicable. A random sample can overlook the very cases that dominate support incidents.

Test realistic concurrency and corpus size. A single warm query in an otherwise idle environment answers a different question from a burst of simultaneous requests while ingestion is active. Keep those scenarios separate so that their results remain interpretable.

Describe the hardware, storage, database version, index configuration, and cache conditions. You do not need an elaborate benchmark platform to make a fair comparison, but you do need enough context to reproduce it. Report latency distributions and failures rather than only the fastest request.

Measure recall without changing the target

Approximate-neighbor recall at k can be defined as the fraction of the exact top-k identifiers recovered in the approximate top-k result. Use the same eligible collection, metric, and query for both. Decide how ties and collections with fewer than k eligible records are handled before computing the score.

For an invented example, recovering nine of ten reference neighbors yields a recall of 0.9 for that query. Averaging over a test set describes that test set, not a universal property of the index. Inspect weak segments even when the overall average appears acceptable.

Pair neighbor recall with application relevance. Losing an exact neighbor might be harmless if several passages answer the question, or serious if the missing passage contains the only correct procedure. The retrieval experiment should therefore retain human-labeled examples alongside the vector baseline.

Treat filtering as part of the benchmark

A vector index is often queried with restrictions such as tenant, product, language, or lifecycle status. The execution strategy and approximate-search behavior can interact with these filters. Test the implemented behavior under the selectivity you expect rather than borrowing an unfiltered result.

Create cases where most records are eligible and where very few are eligible. Check whether the query returns enough results, how its latency changes, and whether all returned records obey the restriction. For security boundaries, verify enforcement using the actual application role and trusted authorization context.

A small tenant embedded inside a large shared collection may behave differently from the whole corpus. Consider workload-specific alternatives, such as a suitable partitioning strategy or exact search over a small eligible set, when supported by your database design. Evaluate the extra operational complexity before adopting them.

Include lifecycle and recovery costs

Index construction is only the beginning. Test inserts, updates, deletions, maintenance, backups, and restoration. A configuration that meets query goals but cannot be rebuilt within your recovery requirements may be unsuitable for the application.

Measure peak resource use during replacement or migration. Keeping an old index available while building a new one may require more capacity than steady-state operation. The same applies when changing embedding models and temporarily retaining two collections.

Document how stale records become unavailable and how a failed build is handled. Do not release a half-populated index without a clear policy. Our vector database cost model includes these transition states because the monthly steady-state footprint is only part of the engineering requirement.

Make the decision with explicit thresholds

Define acceptance before tuning

Set application-specific bounds for latency, neighbor recall, relevance, memory, and recovery. Label them as your requirements rather than universal standards. This prevents the experiment from ending whenever a preferred option happens to look attractive.

Change one dimension at a time

Vary search breadth or another documented parameter while holding the rest of the setup constant. Preserve the baseline and record the resulting curve. A sequence of unrelated configuration changes makes it difficult to know which change improved or harmed the result.

Keep a rollback path

Save the known-good configuration, collection version, and evaluation report. Test the deployment change using the same read role and query shape as production. A rollback should not depend on reconstructing forgotten settings after an incident begins.

Conclusion: choose the measured tradeoff

HNSW and IVFFlat are tools for navigating a performance and recall tradeoff. The right evaluation includes exact neighbors, application relevance, realistic filters, concurrency, and lifecycle operations. Decide from the resulting measurements and your requirements, not from a single headline about speed.

Continue with the token vector search guide to connect index behavior with the rest of the retrieval pipeline. A well-chosen index supports the application’s task; it does not replace a suitable embedding model, good source content, or access control.