
Hybrid Search: Combine Lexical and Vector Rankings
Combine lexical and vector candidates with deliberate ranking. Explore reciprocal rank fusion, candidate windows, duplication, and segment-level tests.
Separate mathematical neighbors from useful answers.
Evaluation gives each retrieval change a concrete question to answer. A metric implementation can be numerically correct while the representation retrieves irrelevant passages. An approximate index can reproduce exact neighbors while the final ranking still misses the user’s task. These articles distinguish those layers.
Start with human-readable relevance examples and an exact-search baseline. Check cosine and distance conventions independently, then measure index recall under the same query and eligibility rules. For hybrid search, retain lexical-only and vector-only baselines and inspect candidate coverage before judging the fusion stage. Preserve configuration and corpus versions with every result.
Use separate development and held-out examples. Include identifiers, paraphrases, selective permissions, duplicates, and no-answer requests. Define denominators, result counts, and tie handling before publishing a measurement. Report weak query segments and inspect failed passages rather than relying on an average alone. The most useful evaluation artifact is one another engineer can reproduce and challenge.

Combine lexical and vector candidates with deliberate ranking. Explore reciprocal rank fusion, candidate windows, duplication, and segment-level tests.

Work through cosine similarity, normalization, dot products, distance, and edge cases. Learn why a similarity score is not a confidence probability.

Compare HNSW and IVFFlat with your own workload. Measure recall, filtering, latency, resource use, updates, and recovery against an exact baseline.

Build an inspectable semantic retrieval baseline. Evaluate compatible embeddings, exact search, relevance labels, permissions, and candidate coverage.