Lesson 3 of 4 · 25 min

Build a supported path with an escape route

Design a developer platform contract around recurring user work.

Mechanism and reasoning

A developer platform should remove repeated work for application teams. Begin with the task users need to complete, such as creating a service with ownership, deployment, observability and access controls. A portal page is one interface to that task. It is not evidence that the underlying workflow is useful.
A supported template can encode defaults and produce a working starting point. Its output needs an owner and an update path. If the platform generates code once and never helps teams adopt later fixes, the organization accumulates many slightly different copies. Decide which concerns belong in a shared component and which belong in generated project code.
Treat the platform contract as an API. Name required inputs, generated outputs, supported customization and failure behavior. A team should know what it can change without leaving support. An escape route is useful when a workload differs, but exceptions need an owner and a route back when the platform gains the missing capability.
Measure completion, elapsed effort and support demand. Counting portal visits or template executions can reward repeated failures. A service created five times because the first four attempts failed should not look like five successful adoptions. Use task identities and inspect outcomes.
In an interview, select one recurring workflow and design it end to end. Explain how a new engineer learns the path, how a failed run recovers and how the generated service receives security updates. Avoid proposing a universal platform before showing that one path works for its intended users.

Service-creation contract

Teaching user need: a team wants a small internal HTTP service. The supported path accepts a service name, owning team, repository location and approved runtime. It creates a repository from a reviewed template, registers ownership, configures a deployment pipeline and adds a health/metrics contract. The result is ready for the team to add business code. It is not automatically authorized to expose sensitive data or create production spending without the organization's normal controls.
Write inputs and outputs as a decision artifact:
Input or outputContractFailure behavior
Service nameUnique within the catalog namespaceReturn a conflict before creating duplicates
OwnerExisting accountable teamBlock creation if ownership cannot be resolved
RuntimeOne supported version familyExplain unsupported choice and exception route
RepositoryCreated once under task identityRetry reuses the same repository
Catalog entryLinks owner, repo and serviceReconcile partial registration
DeploymentUses shared reviewed workflowReport exact failed step and recovery path
Partial failure is normal. Suppose repository creation succeeds but catalog registration fails. A retry should not create a second repository. Keep a task identity and durable step outcomes, or make each action idempotent through its provider's contract. Show the user the existing repository and a clear retry action for registration. A spinner that ends in 'something went wrong' transfers the platform's state problem to the developer.
Now consider template maintenance. The original template pins runtime version R1. Six months later R1 needs a security update. If every generated repository contains a copied deployment workflow, the platform needs a migration mechanism or automated reviewable update. If the workflow is a shared versioned component, users still need controlled version adoption and compatibility tests. Shared code reduces duplication but can also create a wide failure domain if changed without versioning.

User-study artifact

An invented trial has five developers attempt the workflow. Their completion times are 12, 15, 14, 60 and 18 minutes. The sixty-minute case initially fails during registration, then completes with manual support; sixty minutes is its eventual completion time. Sort the values: 12, 14, 15, 18, 60. The median is fifteen minutes. That median hides the failed or assisted experience, so also report four unassisted completions out of five and one support incident.
After a retry repair, a second trial has times 11, 13, 14, 15 and 16, all unassisted. The median is fourteen and unassisted completion reaches five of five. This small trial suggests the failure path improved. It does not establish organization-wide productivity gains. Preserve the task definition and participant context before making a larger claim.
The interview design should include an exception. A GPU batch service may not fit the HTTP template. The platform can offer an explicit unsupported-workload route with a short design review, rather than force the service into unsuitable probes and scaling. Record the recurring unmet need. If many teams need it, build another supported path. If only one team does, a documented exception may be cheaper and clearer.
The final product is a maintained workflow with useful errors, ownership and upgrades. The portal can make that workflow easy to find. It cannot compensate for an unreliable sequence of provider calls or a template nobody owns.
For the next trial, ask a participant who did not help build the template to complete the task from a clean checkout. Record where they stop to ask for help. This reveals missing instructions and hidden local dependencies that an expert author may skip unconsciously. Use the same supported runtime and task definition so the comparison remains meaningful. A successful demonstration by the platform team is weaker evidence than an unassisted completion by the intended user. Keep the sample small enough to observe closely, then repeat after the identified failure is repaired. Do not convert five participants into a claim about every engineering team.

Worked example

Teaching failure trace: task T9 creates repository service-a, records its URL, then catalog registration times out. On retry, T9 checks the recorded repository and attempts registration again using service-a's stable identity. It does not create service-a-2. The user sees 'Repository ready; catalog registration pending' with the existing link. A later reconciliation confirms one repository and one catalog entry. The task is complete only when the declared outputs exist and agree.

Exercise

Six template runs occur for one intended service. Five fail after repository creation and one succeeds. A dashboard reports six adoptions. Define better measures and the required retry behavior.

Model solution and rubric

Count one intended task, its eventual completion, the number of attempts and whether manual support was required. Report time to a usable service and failure rate by step. Reuse the existing repository under the task identity rather than create a new one on each attempt. Verify exactly one service and one catalog record. Template executions are activity, not adoption, and repeated attempts can indicate platform friction.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: more portal use proves more developer value. Repeated failures can increase clicks and runs while wasting time. Misconception 2: a template ends at generation. Its output still needs an upgrade, support and ownership contract or copies will drift.

Interview probe

Evidence class: recommended. Original practice.
How would you prove that your internal platform helps developers?
Strong answer: I would choose a defined task and measure unassisted completion, time to a usable result, failure points and ongoing support. I would also verify that generated services stay maintainable through upgrades.
Follow-up: What belongs in a shared component versus generated project code?
Weak answer indicators: Measuring visits alone; ignoring partial failure; forcing every workload into one template; no owner for generated output.

Sources

Technical references: Backstage software templates; Kubernetes Deployments; Kubernetes probes. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsBackstage software templatesbackstage.iodocsKubernetes Deploymentskubernetes.iodocsKubernetes probeskubernetes.io

Checkpoint

A repository is created but registration fails. A retry should?

ACreate a new repository nameBReuse the task identity and existing outputCMark the whole task completeDDelete all repositories with the prefix
Sign up free to answer and see why

Checkpoint

Five times are 10, 12, 13, 14 and 90 minutes. Median is?

A13B27.8C90D12
Sign up free to answer and see why

Checkpoint

A template's median completion time drops, but one in five users still needs manual rescue. Which report best represents the trial?

AReport only the faster medianBReport median plus unassisted completion and failure-step evidenceCCount manual rescues as new successful adoptionsDRemove the slowest trial because it is an outlier
Sign up free to answer and see why

Checkpoint

A security fix affects generated workflows in 200 repositories. Which design addresses long-term maintenance?

AUpdate only the template for future servicesBAsk teams to remember the fix without tracking adoptionCFreeze old services permanentlyDProvide versioned shared components or reviewable migration updates with owners
Sign up free to answer and see why

Checkpoint

One team needs a GPU batch job outside the HTTP template contract. Which response best preserves platform clarity?

ADefine an owned exception and assess demand for a separate supported pathBAdd hidden batch behavior to the existing template without documentationCForce the HTTP autoscaling policy on the workloadDRemove the service from ownership tracking
Sign up free to answer and see why

Explain how you would design a developer platform contract around recurring user work without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Design one useful workflow, including partial failure and upgrades. Measure completed user work instead of interface activity.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.