← All questions
HardBehavioralSystem design

You are the founding engineer at the end of your first week on a fictional report product. The agreed pilot outcome is that valid accepted exports produce correct, account-authorized results within five minutes. The following complete observation window covers four pilot accounts and 12 distinct accepted exports: eight completed correctly within five minutes, two completed correctly after five minutes, and two failed. Both failed jobs used the documented, supported format B; you reproduced a parser mismatch with a synthetic B fixture. No other failure cause has been established.

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below

The problem

You are the founding engineer at the end of your first week on a fictional report product. The agreed pilot outcome is that valid accepted exports produce correct, account-authorized results within five minutes. The following complete observation window covers four pilot accounts and 12 distinct accepted exports: eight completed correctly within five minutes, two completed correctly after five minutes, and two failed. Both failed jobs used the documented, supported format B; you reproduced a parser mismatch with a synthetic B fixture. No other failure cause has been established.

You traced the request/job records and built the reproducing fixture. A teammate implemented and deployed a separate retry-cap change. A colleague checked the ten completed outputs and their account access. The provider also increased its quota that week. An earlier local run used a smaller dataset, so no comparable before/after speed result exists.

Three engineer-days are available for the next milestone. Estimates are: two days to repair format B and add focused regression checks, including review; half a day to run an authorized, bounded pilot verification after that repair; half a day to review outcomes and write the next decision; or eight days to replace the queue architecture. The format repair does not change stored-data representation; the previous parser path can remain available for rollback. Real workload volume and longer-term reliability are still unmeasured.

Write a short ownership update and choose the next milestone. Separate your contribution from others', calculate the observed timely-completion measure, identify what is known versus unverified, and specify success and stop conditions for a plan that fits three days.

Reference answer

Then expect these follow-ups

  • If the B fixture passes but a pilot output differs, what should stop and which evidence should be preserved?

  • If the repair estimate expands beyond two days, how would you revise the commitment without hiding verification time?

  • Which additional observation would justify investigating queue architecture?

  • Optionally, compare this decision with a real experience, clearly separating that experience from the fictional case.

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.