← All questions
HardAI MLTechnical screen20 min to answer
How would you evaluate a RAG pipeline end to end?
1Give yourself 20 minutes
2Answer out loud, not in your head
3Then compare with the answer below
Stuck? Show a way to structure it+
- 01Clarify the objective, user, constraints, and acceptable failure modes.
- 02Structure the response around retrieval recall, ranking, faithfulness, task outcome.
- 03Compare credible alternatives and state the decision criteria.
- 04Finish with validation, monitoring, and what would change the decision.
Reference answer
Then expect these follow-ups
What would your evaluation dataset contain?
Tests: depth
What if offline and online results disagree?
Tests: adaptability
How would you set the launch threshold?
Tests: adaptability
Free to read · better with Enzo
Practice this out loud with Enzo
Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.
Next question
- How would you measure whether an LLM correctly executes a user’s requested actions?
- How would you launch an AI feature whose output cannot always be objectively graded?
- Offline evaluations improved but production metrics fell. What happened?
- How would you build an evaluation dataset for an AI support agent?
- What is data leakage, and how can it invalidate an evaluation?
- An agent revision improves 180/200 task passes to 188/200 but creates three unauthorized writes. The old version had none in this set. Give a release decision and the smallest useful investigation.