03 Retrieval Systems 3 min read 623 words

Embedding Models (Dec 2025)

Embedding models convert text into high-dimensional vectors. In late 2025, we have moved beyond "static vectors" to Multi-Resolution & Late-Interaction representations.

embeddingsretrievalmodel-selectionquantizationcore
01matryoshka embeddings

The Embedding Frontier: Matryoshka Embeddings

Traditionally, if you embedded text into 1,536 dimensions, you were stuck using all 1,536 dimensions for search.

2025 Innovation: Matryoshka Represenation Learning (MRL)

  • Models are trained to "store" the most important info in the first few dimensions.
  • The Win: You can embed at 1,536 dims, but index only the first 64 dims for a "fast search" pass, then refine the top results with the full 1,536 dims.
  • Efficiency: 20x reduction in memory/index size with <2% drop in accuracy.
02colbert v2

Late Interaction: ColBERT v2

Standard embeddings are "Bi-Encoders" (one vector per chunk). ColBERT (Contextualized Late Interaction over BERT) uses a "token-level" approach.

  • How: Instead of 1 vector per chunk, ColBERT stores 1 vector per token.
  • Interaction: At query time, the model compares every token in your query to every token in the documents (the "MaxSim" operation).
  • 2025 Status: ColBERT v2 is drastically compressed (PLAID indexing), making it feasible for production. It achieves much higher precision for "needle in a haystack" technical queries.
03int8 quantization

Binary and Int8 Quantization

Storing float32 vectors is expensive. In 2025, we use In-Model Quantization.

  • Binary Embeddings: Convert vectors to 1s and 0s.

    • Memory: 32x reduction.
    • Speed: Hamming distance (XOR operations) is 10x faster than Cosine similarity on modern CPUs.
  • Int8/Int4: Supported natively by models like text-embedding-3-small.
04criteria dec

Model Selection Criteria (Dec 2025)

ModelProviderFeaturesContext
Text-Embedding-4OpenAIMatryoshka, Native Int832k
Cohere Embed v3.5CohereBinary quantization, "Compressible"1M
BGE-M3Open SourceMultilingual, Multi-granularity8k
Jina-Embeddings-v3Jina AILate-interaction support128k
05embeddings

Multimodal Embeddings

In late 2025, "Text-only RAG" is dying.

  • CLIP (2025 version): Embeds images and text into the same space.
  • Architecture: You can search a library of schematics (images) using a natural language query ("Where is the emergency shutoff valve?").
06questions

Interview Questions

Q: What is the "Vocabulary Mismatch" problem in embeddings?

Strong answer: Embeddings rely on the semantic space learned during training. If a user query uses a newer term (e.g., "DeepSeek-V3") that wasn't in the embedding model's training set, the model might assign it a generic "AI" vector, missing the specific nuances. In 2025, we solve this with Hybrid Search (using BM25 to catch the specific keyword) or Cross-Encoder Reranking, which is better at handling out-of-distribution vocabulary by looking at the query and document tokens simultaneously.

Q: Why would you choose a Matryoshka model for a 1-billion-vector index?

Strong answer: Scaling to 1 billion vectors with standard float32 1536-dim embeddings requires ~6TB of high-speed RAM for an HNSW index, which is prohibitively expensive. With a Matryoshka model, I can use the first 128 dimensions (Binary quantized) for the initial retrieval. This reduces the memory footprint by over 90%, allowing the "Top 1,000" candidates to be found on significantly cheaper hardware. I can then fetch the full-resolution vectors for just those 1,000 candidates to perform the final reranking.

07references

References

  • Kusupati et al. "Matryoshka Representation Learning" (2022/2024 update)
  • Khattab et al. "ColBERT v1 & v2: Efficient Late Interaction" (2021/2023)
  • OpenAI. "Introducing New Embedding Models with Matryoshka Support" (2024)

Next: Vector Databases Comparison

summary · added by this rebuild

Key takeaways

01

Matryoshka embeddings can be truncated

Training packs the important signal into the leading dimensions, so indexing the first 64 of 1,536 gives roughly a 20x smaller index for under 2% accuracy loss.

02

ColBERT stores one vector per token

Late interaction compares every query token against every document token via MaxSim, lifting precision on needle-in-a-haystack technical queries; PLAID compression made it production-viable.

03

Binary vectors trade precision for throughput

One-bit embeddings cut memory 32x, and Hamming distance over XOR runs about 10x faster than cosine similarity on modern CPUs.

04

Two-stage search beats one huge index

A billion float32 1,536-dim vectors need roughly 6TB of RAM for HNSW; retrieving on 128 binary dimensions first cuts the footprint by over 90%.

05

Embeddings miss vocabulary they never saw

A term absent from training collapses into a generic vector, so hybrid BM25 search or cross-encoder reranking is needed to catch out-of-distribution keywords.