VECTORTOKEN LAB / CATEGORY

Foundations

Get the language right before you build.

Tokens, token IDs, contextual representations, and passage embeddings answer different questions. This collection starts at those boundaries and follows them into the geometry used by retrieval systems. It is designed for readers who need to review an architecture, understand a model output, or explain why two arrays should not be compared just because their dimensions match.

Begin with the vector token introduction, then follow the contextual embedding walkthrough. Use the cosine similarity article when you need to inspect normalization, sorting direction, or a threshold. Try its small numerical examples by hand before diagnosing a large search system. The suggested reading order moves from vocabulary to representation to comparison, keeping the represented unit explicit at every stage.

A useful outcome is a written representation contract: the model and tokenizer revisions, the input role, the output unit, pooling, normalization, and the intended metric. Carry that contract into ingestion and search instead of relying on an undocumented field named embedding.

THE READING PATH

03 articles to explore.

View all articles

Follow a subject