Interview Q&A · 4 · Retrieval & Long Context
Give me three cases where semantic search over embeddings retrieves the wrong document — and explain why.
Reveal the answer
Embeddings encode topic and paraphrase well; they encode negation, entities,
and numbers badly. Three concrete failures:
1. **Negation collapse.** "The API is safe from prompt injection" and
"The API is not safe from prompt injection" produce nearly identical
embeddings. Cosine similarity says match; the meaning is opposite.
2. **Entity swap.** "Anthropic released Claude 5" vs. "OpenAI released
GPT-6" — same shape, same topic — embed close. A user asking about
Claude gets an OpenAI passage back.
3. **Date and number blindness.** "revenue was $50M in Q2 2024" and
"revenue was $500M in Q2 2025" embed close because the sentence
structure dominates the vector.
The fix is almost never "swap embedding models." It's hybrid retrieval
(BM25 + vectors) so exact tokens like "not" and "$500M" get keyword
weight, plus a reranker that reads the query and the passage together.
Common variants
- Would fine-tuning the embedding model fix these failures?
- How do you evaluate whether retrieval or the LLM's answer is the bug?
- What does a reranker actually do, mechanically?
Track this card
Stored in this browser only — no signup, no sync. Clearing site data clears it.