A token vector can mean an embedding looked up from a vocabulary table, a hidden state produced after contextual processing, or a representation retained by a retrieval system. These meanings are close enough to sound interchangeable in conversation and different enough to cause substantial implementation mistakes. Before comparing arrays, define which representation you have and which task it is supposed to support.
This article follows one conceptual path from token IDs to contextual representations and then to a passage-level search vector. It does not prescribe a universal architecture. Different models organize these steps differently. The useful habit is to examine the contract of the actual model rather than infer behavior from a field named embedding.
Begin with an indexed lookup
A basic embedding layer maps an integer identifier to a row in a learned table. If a vocabulary contains V entries and each row has D values, the table has V by D entries. A sequence of T token IDs selects T rows. The immediate result has one D-dimensional vector for each sequence position, before additional model-specific processing.
A repeated ID can select the same initial row even when it appears in different sentences. That does not imply that the final representation remains identical. Position information and contextual computation can change what is associated with that occurrence. The Hugging Face model-output documentation distinguishes hidden states and other returned tensors, which is essential when deciding what an application is actually reading.
As a review exercise, write the dimensions next to each variable in a notebook. A tensor shaped like batch by sequence by hidden size is not the same object as a matrix containing one vector per document. Explicit shape comments often expose a mistaken reduction before it reaches storage.
Context belongs to an occurrence
Take the invented sentences “the crane lifted the beam” and “the crane stood in the marsh.” A vocabulary entry alone does not identify which sense is intended. A context-dependent representation can incorporate surrounding information. The unit you are observing is now a token occurrence in a particular input, not simply a reusable dictionary entry.
This distinction matters for debugging. A developer may compare an input embedding table against a model’s final hidden states and interpret differences as corruption. They are different stages. A meaningful comparison holds the stage, layer, tokenizer, model revision, and input conditions constant.
It also matters for explanation. A colorful projection of a few vectors is a visualization, not proof that a model has formed a particular human concept. When presenting such a plot, label the projection method and the represented object. Avoid turning exploratory geometry into an unsupported claim about understanding.
Pooling is a design choice, not an automatic guarantee
A search application often needs one representation for a passage rather than one for every token. Pooling is one way to aggregate sequence information. Depending on the model, a designated position, an average over selected positions, or another learned mechanism may be appropriate. The correct choice depends on how the model was trained and how its outputs are intended to be used.
Suppose you average hidden states without excluding padding positions. The arithmetic still produces a vector with the expected dimension, so a shape check passes. Yet the result can depend on the unrelated lengths of other sequences in the batch. This is a reason to test batching invariance, not a reason to assume all averaging methods are wrong.
Our token vector topic page treats pooling and normalization as part of the representation contract. Record those choices with the model revision. A future maintainer should not have to reverse engineer the settings from an old notebook.
Compare the right kind of similarity
A representation trained for one purpose may not organize its space in the way your application needs. Predicting the next token, comparing short sentences, retrieving passages for questions, and identifying duplicate support requests are different tasks. The fact that a model returns numbers does not establish that distance between those numbers measures your chosen relationship.
Write down the intended relationship in ordinary language. For a documentation search tool, it might be “this passage contains enough information to answer the question.” For deduplication, it might be “these requests describe the same underlying issue.” Those statements lead to different labels and different counterexamples.
Then create a test set that includes both semantic neighbors and tempting mistakes. A passage discussing account cancellation may be related to a question about pausing an account without answering it. Similarity should be assessed against the application’s definition of usefulness, not merely thematic overlap.
Build a representation manifest
Treat the encoder as a versioned transformation. A manifest can identify the model, tokenizer revision, preprocessing steps, maximum input length, pooling policy, normalization method, dimension, numerical type, and query-versus-document conventions. These fields describe how the representation was made; they do not need to be repeated inside every user-visible result.
Give the manifest a stable identifier and store that identifier with each record. A batch that uses a different normalization policy should not silently enter the same collection. If the change is intentional, create a separate version and evaluate the migration. Matching dimensions alone are not sufficient evidence of compatibility.
An operationally useful manifest also records the source revision and the software release responsible for ingestion. When a retrieval regression appears, you can distinguish a changed document from a changed encoder. This is often more actionable than comparing two unexplained arrays.
Test properties before testing scale
Start with deterministic or tolerance-based checks suitable for your chosen model. Verify that repeated encoding under the same inference conditions is acceptably stable. Compare encoding an item alone with encoding it in a mixed-length batch. Inspect empty input, whitespace-only input, long input, and text containing identifiers or non-English characters.
Some environments introduce small numerical differences, so choose tolerances rather than demanding byte equality without justification. The important question is whether these differences alter decisions your application cares about. A stable representation may still produce poor relevance, and a tiny numerical change may be harmless if the ordering remains acceptable.
Add a round-trip test through storage. Write a known vector, retrieve it, and check dimension, numerical type, and values within the expected tolerance. This isolates serialization and database behavior from model behavior. A failed search should not require debugging both layers at once.
Avoid three misleading shortcuts
Averaging token IDs
Vocabulary identifiers are labels. Their arithmetic average does not provide a principled semantic representation. Two unrelated sequences can easily have similar averages. When you need a deliberately simple baseline, choose one whose meaning you can explain and evaluate rather than disguising ID arithmetic as an embedding model.
Mixing output layers
A model may expose several hidden-state tensors. Selecting a different layer changes the representation. Do not mix outputs from arbitrary layers in one index because their shapes happen to match. Treat layer selection as a configuration change requiring evaluation and a documented migration path.
Treating every score as confidence
A vector similarity score is not automatically a calibrated probability that a passage is correct. Its interpretation depends on the model, data, metric, and retrieval pipeline. The cosine similarity guide explains why a high score and a trustworthy answer are different claims.
Conclusion: keep the contract with the vector
The most useful question is not “does this model produce embeddings?” It is “which representation does it produce, for what task, under which conditions?” Distinguish initial token lookups from contextual states and passage representations. Document aggregation, normalization, and model revisions. Test simple invariants before adding scale or approximation.
To connect these ideas to a working retrieval flow, continue with the token vector search guide and the vector AI overview. Both start from a defined representation and work toward measurable application behavior rather than relying on the presence of an array alone.



