Question bank

4,310 interview questions, answered.

Reference answers, what the interviewer is really testing, how it is graded, and the follow-ups that come next.

Easy 190Medium 2,498Hard 982
FiltersDifficulty, topic, company

4,310 questions

  1. 2209Paged attention (vLLM)HardAI ML
  2. 2210Planning and decomposition in multi-step agent tasksHardAI ML
  3. 2211Precision vs recall vs F1 vs AUCEasyAI ML
  4. 2212Prompt caching and KV cache prefix sharingHardAI ML
  5. 2213QLoRA vs LoRAHardAI ML
  6. 2214Quantization: INT8, INT4, FP8, AWQ, GPTQHardAI ML
  7. 2215RAGAS: why it's popular and where it breaksMediumAI ML
  8. 2216Regression testing before model/prompt changes shipHardAI ML
  9. 2217Retrieval over structured data: tables, code, JSONHardAI ML
  10. 2218RLHF in plain terms: what does it solve that SFT cannot?HardAI ML
  11. 2219Scaling LLM inference for traffic spikesHardAI ML
  12. 2220Sinusoidal positional encodings vs RoPEHardAI ML
  13. 2221Speculative decoding: when it helps, when it failsHardAI ML
  14. 2222The lineage from PPO to DPO to GRPOHardAI ML
  15. 2223The "lost in the middle" phenomenonHardAI ML
  16. 2224The main parts of a RAG systemEasyAI ML
  17. 2225The ReAct frameworkMediumAI ML
  18. 2226Tools in agentic AI and how agents pick themEasyAI ML
  19. 2227Top-k, top-p (nucleus) and temperature samplingMediumAI ML
  20. 2228What are embeddings, beyond similarity search?EasyAI ML
  21. 2229What is an AI agent vs a prompt vs a chain?EasyAI ML
  22. 2230What is instruction tuning and why does it matter?MediumAI ML
  23. 2231What is KV cache and how does it speed up inference?MediumAI ML
  24. 2232Why do transformers need positional embeddings at all?MediumAI ML