← All questions
HardDataAI ML
An agent revision improves 180/200 task passes to 188/200 but creates three unauthorized writes. The old version had none in this set. Give a release decision and the smallest useful investigation.
1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0
Reference answer
Then expect these follow-ups
What if all three failures share one tool version?
Free to read · better with Enzo
Practice this out loud with Enzo
Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.
Next question
- How would you measure whether an LLM correctly executes a user’s requested actions?
- How would you launch an AI feature whose output cannot always be objectively graded?
- How would you evaluate a RAG pipeline end to end?
- Offline evaluations improved but production metrics fell. What happened?
- How would you build an evaluation dataset for an AI support agent?
- What is data leakage, and how can it invalidate an evaluation?