02 / REPRESENTATION

Token vectors: give every embedding a contract

A vector is more than an array with the right length. Know what it represents, which model produced it, and how it is intended to be compared before putting it into an index.

Context changes everything: a cyan diagram connects a token with contextual meaning.
01 / KEY IDEA

Identify the unit

Distinguish a vocabulary lookup, a contextual token occurrence, and an aggregated passage representation.

02 / KEY IDEA

Record the method

Keep the encoder revision, input role, pooling, dimension, numerical type, and normalization together.

03 / KEY IDEA

Test compatibility

Compare like with like. Matching vector dimensions does not establish a shared representation space.

From lookup to contextual representation

An initial embedding lookup selects a numerical row for a token identifier. A model can then transform representations using the surrounding input. A final hidden state and an initial lookup are therefore different objects, even when their dimensions happen to match. The Hugging Face model-output reference is useful when identifying which tensor a model returns.

Treat a representation as belonging to a particular stage. Record the layer or output field instead of describing everything as an embedding. During debugging, compare outputs produced under the same conditions before interpreting differences as a problem.

Decide whether you need one vector or many

Some workflows work with one representation per token position. Passage search often needs one representation per retrieval unit. An aggregation step, sometimes called pooling, connects those designs, but its suitability depends on the model. An arbitrary average is not a guarantee of useful semantic behavior.

When pooling variable-length sequences, review how padding is handled. Test the same text alone and in a mixed-length batch. A change caused only by unrelated padding can reveal an implementation error that a basic dimension check would miss.

Write a representation manifest

A useful manifest names the encoder, tokenizer, preprocessing rules, input limits, pooling method, output dimension, numerical type, and normalization policy. Include query-versus-document conventions where they apply. Give this configuration a stable version identifier so that every stored record can point to it.

Keep incompatible versions separate unless you have validated a compatibility or alignment method. Changing a model without changing a database column definition can still change the representation space. The manifest helps make this otherwise invisible change explicit.

Match similarity to the task

Define useful similarity in application language. Does a candidate answer a question, describe the same issue, or merely share a topic? These relationships can require different evaluations. A mathematically valid score does not establish that the representation captures the relationship you care about.

The vector AI overview connects these distinctions to application design. The cosine similarity article works through normalization and scoring with simple, inspectable vectors.

Validate before scaling

Test repeated encoding, empty input handling, length limits, batching, and a storage round trip. Use justified numerical tolerances rather than assuming byte-for-byte equality is always required. Check relevance separately with labeled queries and passages. Numerical stability and useful ranking answer different questions.

Preserve the test inputs and configuration with the results. A future model migration is easier to evaluate when the old system has a reproducible baseline. Start with a collection small enough to inspect and expand once the represented object and its intended comparison are clear.

FOLLOW THE CONNECTIONVector Tokenization