Interview Q&A · 6 · Interview prep
Design a chatbot that answers questions about our internal documentation. Walk me through the components and where the failure modes are.
Reveal the answer
Sketch the pipeline out loud in stages; the interviewer is listening for
which trade-offs you name.
**Ingest.** Documents → chunker (structural first: headings and
paragraphs; fixed-size ~512 tokens with 64 overlap as fallback) →
embeddings (a single model, versioned) → vector store (pgvector for
<10M chunks, purpose-built for more) plus a lexical index (BM25)
because embeddings miss entities and numbers.
**Serve.** Query → hybrid retrieval (top 50 from each of vectors + BM25)
→ reranker (a cross-encoder or a small LLM call) → keep top ~5 →
passed as context to the answering LLM with a strict "cite your
passages" system prompt → response.
**Guardrail.** Refuse politely when the retrieved passages don't
contain the answer; never let the model fill from parametric memory
for company facts (that's how you get confident, wrong answers about
your own product).
Failure modes to name unprompted:
- Retrieval getting the wrong doc (entity swap, negation collapse).
- Chunk boundary cutting an answer in half.
- The model ignoring the "only from context" instruction.
- The moment someone renames a doc and the vector index goes stale.
- Prompt injection via an ingested document itself.
- Cost drift as chunk count grows.
Evals to name: retrieval@k, answer faithfulness (LLM-as-judge with
passages as ground truth), a small "unanswerable" set to check refusal
rate. Deploy behind a gateway with per-tenant rate limits.
Common variants
- The corpus is 100M chunks now — what changes?
- How do you keep the index fresh as docs are edited?
- How would you handle a doc that itself contains prompt-injection?
Track this card
Stored in this browser only — no signup, no sync. Clearing site data clears it.