04 / APPLIED SIMILARITY

Vector AI: turn similarity into a testable application

Similarity becomes useful when you define the relationship you need. Search, duplicate detection, and knowledge discovery should start with examples of success—not a confident-looking score.

Cosine, not confidence: two directional vectors illustrate an angle on a black neon-accented card.
01 / KEY IDEA

Define the relationship

Separate matching a topic from answering a question or identifying a duplicate issue.

02 / KEY IDEA

Choose a metric

Keep scoring direction, normalization, zero-vector handling, and model compatibility explicit.

03 / KEY IDEA

Measure the decision

Use labeled examples and difficult counterexamples before selecting thresholds or adding complexity.

Begin with a decision, not an embedding

Write down what your application needs to decide. A documentation tool should retrieve an answer-bearing passage. A duplicate detector should identify records describing the same underlying issue. A related-content browser may only need thematic neighbors. These are different outcomes even when they use similar infrastructure.

Select positive and negative examples before evaluating a model. Include cases that share vocabulary but disagree in meaning, such as enabling and disabling the same feature. This prevents a visually convincing neighbor list from becoming its own definition of correctness.

Understand what a similarity score says

Cosine similarity compares vector direction. Dot product includes magnitude unless the relevant vectors are normalized. The Sentence Transformers similarity reference describes supported comparison methods and the normalized relationship between these two operations.

A score is not automatically a calibrated probability that content is correct. When displaying results, source context and provenance may be more useful than an unexplained decimal. Thresholds need task-specific labels, and their behavior should be revisited when the representation or corpus changes.

Keep learned representations compatible

Two encoders can output vectors of the same length without producing interchangeable coordinates. Keep the model revision, input procedure, pooling, and normalization in a documented contract. Test both query and document paths. A simple shape check cannot establish semantic compatibility.

The token vector page provides the contract checklist. When planning a migration, preserve a known-good collection and compare the new version on the same labeled questions before switching the application.

Combine signals deliberately

An exact identifier can matter more than broad semantic proximity. A paraphrased question can be difficult to solve through shared terms alone. These are reasons to investigate hybrid retrieval rather than assume one signal should replace the other.

Keep lexical-only and vector-only baselines. Evaluate candidate coverage before final ranking, then inspect whether combining the signals helps the difficult query segments. The hybrid search guide introduces rank-based fusion without treating unrelated raw scores as directly comparable.

Make evaluation reproducible

Preserve queries, labels, corpus revision, and model configuration with each run. Separate development examples from a held-out release set. Record no-answer behavior, duplicates, and segment-level failures instead of reporting only a broad average.

A reliable experiment explains both what improved and what became worse. That makes it easier to decide whether another model, a better passage boundary, a ranking adjustment, or a simpler interface is the appropriate next change. Vector AI is a design space to test, not a shortcut around defining useful behavior.

FOLLOW THE CONNECTIONVector LLM