Question bank

4,310 interview questions, answered.

Reference answers, what the interviewer is really testing, how it is graded, and the follow-ups that come next.

Easy 190Medium 2,498Hard 982
Filters · onDifficulty, topic, company

982 questions

  1. 649Catastrophic forgetting during fine-tuningHardAI ML
  2. 650Common agent failure modes in productionHardAI ML
  3. 651Designing LLM benchmarks: MMLU, GSM8K, HELMHardAI ML
  4. 652Detecting distribution shift in production LLM trafficHardAI ML
  5. 653DPO vs RLHF/PPO: when would you pick each?HardAI ML
  6. 654Evaluating reasoning models vs chat modelsHardAI ML
  7. 655Fine-tuning vs RAG vs prompt engineeringHardAI ML
  8. 656FlashAttention and why it mattersHardAI ML
  9. 657Guardrails for agents that call external APIsHardAI ML
  10. 658How do you evaluate a RAG system?HardAI ML
  11. 659How do you evaluate an agent?HardAI ML
  12. 660Knowledge distillation for LLMsHardAI ML
  13. 661Latency vs throughput vs cost: choosing batch sizesHardAI ML
  14. 662LLM-as-judge: when it works, when it failsHardAI ML
  15. 663LLM observability vs classical ML observabilityHardAI ML
  16. 664Measuring hallucination rate without expensive human evalHardAI ML
  17. 665Mitigating hallucinations in RAGHardAI ML
  18. 666Monitoring RAG in productionHardAI ML
  19. 667Multi-agent orchestration: when does it help?HardAI ML
  20. 668Paged attention (vLLM)HardAI ML
  21. 669Planning and decomposition in multi-step agent tasksHardAI ML
  22. 670Prompt caching and KV cache prefix sharingHardAI ML
  23. 671QLoRA vs LoRAHardAI ML
  24. 672Quantization: INT8, INT4, FP8, AWQ, GPTQHardAI ML