System Design in the RAG Era: Vector DBs, Embeddings & the New Round
The round that didn't exist 18 months ago
Retrieval-Augmented Generation has gone from "cool research technique" to "the thing every ML team is building" in roughly eighteen months. If you're interviewing for any role that touches AI/ML, LLMs, or backend engineering at an AI company, you will get RAG questions — and they won't be theoretical. They'll be "have you actually built this" questions.
The shift matters because classical system design and AI-era system design test fundamentally different things. The classical round rewards you for knowing when to reach for Redis. The RAG round rewards you for knowing when not to reach for a vector database at all.
Classical system design vs RAG-era system design
Dimension | Classical round (2020-era) | RAG-era round (2026) |
|---|---|---|
Canonical problem | Design Twitter / URL shortener / rate limiter | Design a customer-support RAG agent / doc-search copilot |
Core data structure | Hash map, B-tree, log | Vector index (HNSW, IVF), embedding |
Scaling lever | Sharding, caching, replication | Chunking strategy, embedding model choice, re-ranking |
Failure mode | Hot shard, cache stampede | Hallucination, stale embeddings, retrieval misses |
Cost driver | Compute + storage | Token cost + embedding API calls + storage |
"Senior" signal | Knows when to denormalize | Knows when RAG is the wrong tool |
"RAG isn't just about memory; it's about traceability, real-time updates, data security, and source verification."
— TechInterview, RAG Interview Guide
The vector DB decision tree (2026 edition)
In 2026, "add a vector database" has become as standard in AI application architecture as "add Redis for caching" was a decade ago. The choice between Pinecone, Weaviate, pgvector, and Qdrant now involves real engineering trade-offs.
Vector DB | Best for | Watch out for | Cost posture |
|---|---|---|---|
Pinecone | Zero-ops managed search at any scale | Vendor lock-in, limited hybrid search | Premium managed |
Weaviate | Hybrid search (vector + BM25 + metadata) | Self-hosted ops burden if not using cloud | Mid |
pgvector + pgvectorscale | Postgres shops, <100M vectors | Performance ceiling at scale | Cheap (reuse infra) |
Qdrant | Budget-conscious teams, <50M vectors, best free tier | Smaller ecosystem | Low |
Milvus / Zilliz | Billions of vectors at lower cost | Requires engineering resources | Mid-low at scale |
ChromaDB | Prototyping and MVPs only | Not production-grade | Free |
Sources: Firecrawl 2026 vector DB comparison, Dev Note production comparison, Kunal Ganglani's Pinecone vs Weaviate deep-dive.
The decision isn't "which is best." It's "which fits your scale, sovereignty, and hybrid-search needs." Pinecone wins on zero-ops speed; Weaviate wins on hybrid search and data sovereignty.
The embedding tradeoff nobody mentions
Most candidates pick a vector DB and stop. Senior candidates know the embedding model choice has more impact on retrieval quality than the DB choice. The questions that separate levels:
Dimensionality — 1536 (OpenAI ) vs 768 (open-source) vs 3072 (large). Higher dims = better recall, more storage, slower queries.
Domain fit — general-purpose embeddings underperform on medical, legal, or code corpora. Fine-tuning or domain-specific models matter.
Chunking strategy — fixed-size, sentence-aware, semantic, or hierarchical. This is where most RAG systems silently fail.
Re-ranking — a cross-encoder re-ranker (Cohere, BGE) on top of vector retrieval routinely lifts precision by 15–25%.
"The answer to 'do we even need RAG anymore?' as context windows expand remains a resounding yes. RAG isn't just about memory — it's about traceability, real-time updates, data security, and source verification."
— TechInterview RAG Guide
The 7 questions that decide the round
Based on reported questions at OpenAI, Anthropic, and Google, here's the 2026 RAG-round question bank — annotated with what each is actually testing:
"Walk me through how you'd chunk a 10,000-page legal corpus for RAG." — Tests chunking strategy + domain awareness.
"Your retrieval precision is 60%. How do you debug?" — Tests eval mindset. Junior says "tune the model." Senior says "instrument retrieval first, then re-rank, then re-embed."
"When would you not use RAG?" — Tests architectural judgment. Long-context models, structured lookups, and classical search all have their place.
"Design the eval pipeline." — The killer question. Most candidates have no answer. Have one: retrieval metrics (recall@k, MRR), generation metrics (faithfulness, answer relevance), and human-in-the-loop spot checks.
"How do you handle stale embeddings when the source docs update?" — Tests production thinking. Incremental re-indexing, TTL on embeddings, versioned collections.
"Cost is blowing up. Where do you cut?" — Tests cost literacy. Embedding API calls, token usage, storage — in that order.
"How do you prevent hallucination?" — The trap. The honest answer is "you can't fully — you constrain it" via grounded retrieval, citation enforcement, and refusal-on-low-confidence.
How to prepare (the 14-day plan)
Days 1–3: Read the GitGood 2026 RAG interview guide end-to-end and the TechInterview advanced guide.
Days 4–6: Build one RAG pipeline end-to-end. Use pgvector if you already know Postgres, Qdrant if you don't. Chunk a real corpus — your own notes, a textbook, anything.
Days 7–9: Add a re-ranker. Measure recall@5 before and after. The number will surprise you.
Days 10–12: Build the eval pipeline. This is what 80% of candidates skip and 100% of senior roles require.
Days 13–14: Mock the 7 questions above, out loud, with a timer. Narrate every tradeoff.
The cultural shift underneath
The RAG round isn't just a new topic. It reflects a deeper 2026 reality: AI infrastructure is now core engineering infrastructure. The companies hiring for it have explicitly moved away from LeetCode-heavy loops toward "have you actually shipped this" questions. If you can design, build, and reason about a RAG pipeline, you're not just interview-ready — you're useful on day one.
Further reading:
GitGood — RAG Interview Questions 2026: The Complete Guide
TechInterview — RAG (Retrieval-Augmented Generation) Interview Guide
Medium — How to Prepare for System Design Interviews in the Age of AI (2026 Update)
Firecrawl — Best Vector Databases in 2026
Dev Note — Vector Databases in 2026: Comparing Pinecone, Weaviate, pgvector, Qdrant
#SystemDesign #RAG #VectorDatabase #AIEngineer #LLM #Pinecone #Weaviate #InterviewPrep #2026 #AICareer
official@dsaquest.com • Contributor
Discussion (0)
No comments yet. Be the first to start the discussion!