
Vector Database Costs: Budget Beyond the Embedding
Estimate vector payloads without confusing them with total costs. Include passage counts, indexing, metadata, copies, compression, and migration.
Know what the vector represents.
An embedding is useful in relation to a model, an input role, a represented unit, and a comparison method. This reading path follows that contract from token-level representations to passage retrieval, then connects output dimensions and numerical types with storage planning.
Start by distinguishing token IDs from learned representations. Continue with contextual states and pooling, then work through cosine similarity and normalization. Read the cost model when considering dimension, precision, or multiple retained versions. The same-length output of two encoders does not by itself establish compatibility; the manifest should explain why query and document representations belong in the same space.
Create a small test set containing paraphrases, related non-answers, and exact identifiers. Check batching behavior and a storage round trip separately from relevance. Keep a full-precision reference when experimenting with compressed representations. The goal is not merely to produce an array, but to preserve an interpretable transformation whose usefulness can be evaluated.

Estimate vector payloads without confusing them with total costs. Include passage counts, indexing, metadata, copies, compression, and migration.

Work through cosine similarity, normalization, dot products, distance, and edge cases. Learn why a similarity score is not a confidence probability.

Follow token IDs into contextual representations. Understand pooling, representation contracts, and the checks that make embeddings reproducible.

Separate tokens, vocabulary IDs, and embeddings. Build a precise vocabulary for vector AI, model inputs, storage, and semantic retrieval.