Lesson 3 of 4 · 25 min

Design a safe self-service scaling API

Specify authorization and concurrency for a platform mutation.

Mechanism and reasoning

A platform API that changes replica counts is a small interface with large consequences. The request needs an authenticated caller, an authorized target, a bounded desired value and a concurrency rule. A friendly portal does not remove any of those requirements.
Separate the API's identity from the end user's identity. A backend service account may have permission to modify many workloads, while the caller should control only their team's service. The API must enforce the caller's scope before using its stronger backend authority. This is a common confused-deputy risk.
A read-modify-write sequence can lose concurrent changes. If two clients read the same state and submit different desired values, the later write may overwrite the earlier intent. Use an explicit version precondition or a clearly documented last-write policy where appropriate. A scaling interface should also define interaction with autoscalers, which may own the same field.
Validate bounds based on quota and service policy. Negative counts, extreme scale-out and scaling a protected service to zero should not reach the cluster accidentally. A syntactically valid integer can still violate operational policy. Return a useful rejection reason without exposing unauthorized workload details.
The interview deliverable is a request contract, decision table and a failure trace. Explain idempotency for retries and audit records for accepted changes. An HTTP success should say whether desired state was accepted or the requested capacity became ready. Do not make those meanings ambiguous.

API contract artifact

Teaching endpoint changes desired replicas for a service owned by the caller's team. It accepts a target service ID, requested count, expected configuration version and an idempotency key. The server resolves the target from an internal registry rather than accepting an arbitrary Kubernetes namespace and object path.
FieldExampleRule
service_idinvoice-apiMust resolve to a caller-authorized service
replicas6Integer within service range 2–10
expected_version18Must match the current version for this mutation
idempotency_keychange-71Same key and same payload reuse the recorded outcome
actorDerived from authenticationNever trusted from a caller-supplied display field
The response can say accepted desired replicas six at version nineteen, with an operation ID for observing readiness. It should not claim six ready replicas until that condition is measured. A separate status endpoint can show desired, ready and failure reason without granting mutation rights.

Concurrent trace

Clients A and B both read version eighteen with four desired replicas. A submits six with expected version eighteen and succeeds, producing version nineteen. B submits two with expected version eighteen. The server rejects the stale precondition rather than silently replacing A's change. B must reread and decide whether its new intent is still appropriate.
Now A retries because its response was lost. With the same idempotency key and payload, the API returns the recorded accepted result. It should not apply the mutation again under a new version merely to answer the retry. If the same key arrives with a different requested count, return a conflict. This prevents a client bug from treating one identity as several different operations.
The platform must also consider the autoscaler. If an HPA owns replica count, the self-service endpoint might change an approved minimum/maximum policy rather than directly fight the current replica field. Alternatively, it can reject manual scaling while that mode is active. The correct choice is a product decision. The API must expose the ownership rule so the user understands why their count reverts.

Authorization and audit

An audit record can include actor ID, service ID, old and requested values, policy decision, request identity, accepted version and timestamp. Do not store bearer tokens or full credentials. Failed authorization should record enough to investigate abuse without revealing another team's sensitive configuration to the caller.
Test a member of team red against a team-blue service. The result should deny the action even though the backend service account could technically perform it. Then test a permitted team-red service with a count above quota. Authorization passes, but policy validation fails. These are separate checks and should produce separate internal reasons.
A rate limit protects the control plane from repeated mutations, but it does not replace authorization or idempotency. Ten unauthorized requests per minute are still unauthorized. A single allowed scale-to-ten request can still exceed cost policy if the range is wrong. Defense comes from a clear contract across identity, target scope, value bounds and state ownership.
This case is original practice. Teleport publishes a related SRE challenge involving a Go service that reads and changes Kubernetes Deployment replicas. That first-party sample supports the relevance of this practical interface, but the request shape, version rules and scenario here are our own. Do not label these exact questions as Teleport's wording or claim that the sample is currently used for every candidate.

Worked example

The application owns a durable desired-configuration record. Acceptance changes that record; it does not mean the Kubernetes update has finished. The operation record, desired configuration and reconcile intention commit in one database transaction. All writers use this boundary. The version test must be part of the protected update, not a separate unprotected read.
code
1actor = authenticate(request)2target = resolve_service(request.service_id)3authorize_current_caller(actor, target, "scale")4scope = (actor.tenant, actor.id, target.id, "scale", request.key)5transaction:6    lock target desired-configuration record7    prior = find_operation(scope)8    if prior exists:9        if prior.payload != request.payload: return key_conflict10        return prior.recorded_outcome11    validate_range_and_field_ownership(target, request.replicas)12    require target.version == request.expected_version13    update desired replicas and increment target.version14    record operation with unique scope and accepted version15    record reconcile intention for that version16commit17return accepted operation, desired version and desired replicas
This is illustrative pseudocode, not an implementation. Recheck current authorization before returning a prior operation. Its stored payload includes the target, desired count and original expected version. A retry does not need to pass that now-stale version test again after it matches the authorized stored operation. A new intent needs a new key and the current version. The unique scoped key also prevents two simultaneous attempts from recording separate operations.
An asynchronous reconciler reads current desired state and the current Kubernetes object. It uses a conditional update with the observed Kubernetes resourceVersion; a conflict requires a fresh read and a new decision. Application version nineteen and Kubernetes resourceVersion are different fields. Store the applied application generation with the replica update on the same Kubernetes object. Never replace a higher applied generation with a lower one, and never retry a rejected old object without reading its current generation. These rules prevent a stale reconciler from undoing a newer applied request when all writers follow the contract. External writers and an autoscaler need the ownership rule described above.
A lost Kubernetes response leaves the outcome unknown. Read the current object and reconcile the desired generation before trying another mutation. Record accepted, observed-applied and ready states separately. No database transaction makes a remote cluster update atomic with local acceptance. The durable reconcile intention makes unfinished work recoverable, while conditional cluster updates and generation checks prevent stale writes. Test crashes before local commit, after commit, after the remote update, and before the completion record.

Exercise

Caller belongs to team green. They submit replicas 8 for a team-purple service. The backend account can update both namespaces. What should happen? Then describe a retry after a permitted request succeeds but its response is lost.

Model solution and rubric

Deny the cross-team mutation before using backend authority. Backend capability is not caller permission. For a permitted request with a lost response, the same idempotency key and identical payload should return the existing outcome or an explicit still-pending state. It should not create a new logical change. Test a reused key with a different payload and a stale expected version. Audit the actor and target without storing credentials.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: backend RBAC alone enforces each portal user's scope. The backend may be more privileged than the caller, so application authorization is necessary. Misconception 2: idempotency solves concurrent intent. It handles retries of one operation; version checks or ownership rules handle different competing operations.

Interview probe

Evidence class: recommended. Original practice.
How do you prevent a self-service scaling endpoint from becoming a cluster-admin proxy?
Strong answer: Resolve approved targets, authorize the caller's team and action, bound values, use a least-privileged backend identity, and record the mutation. I also define concurrency and autoscaler ownership instead of forwarding arbitrary cluster requests.
Follow-up: How do you recover when the cluster accepted the update but your idempotency record is uncertain?
Weak answer indicators: Trusting a caller-supplied team field; relying only on backend authority; treating a response timeout as proof no mutation occurred.

Sources

Technical references: Kubernetes API update preconditions; AWS transactional outbox pattern; Teleport public SRE challenge; OWASP authorization; Kubernetes Deployments. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes API update preconditionskubernetes.iodocsAWS transactional outbox patterndocs.aws.amazon.comdocsTeleport public SRE challengegithub.comdocsOWASP authorizationcheatsheetseries.owasp.orgdocsKubernetes Deploymentskubernetes.io

Checkpoint

The API backend can update both team-red and team-blue namespaces. A signed-in red user requests a blue service. Which control must reject this request?

AA numeric replica bound enforced before cluster accessBApplication authorization for the trusted caller on the resolved service and actionCAn idempotency key unique to this browser sessionDThe broad backend Kubernetes role used for both teams
Sign up free to answer and see why

Checkpoint

A and B read application version 10. A commits version 11. B submits a distinct operation with expected_version 10. Under the stated optimistic contract, what should the server do?

AReapply B using version 11 without asking its callerBReturn A's operation because the target is the sameCQueue B unchanged until A becomes ready, then ignore the versionDReject B's stale precondition so the caller can reread and reconsider
Sign up free to answer and see why

Checkpoint

An authorized caller reuses a stored scoped idempotency key but changes replicas from 4 to 7. What should the service return under this course contract?

AThe original success without disclosing the payload mismatchBA new accepted operation because the count remains in rangeCA key/payload conflictDThe latest target status as proof the new request succeeded
Sign up free to answer and see why

Checkpoint

The API returns operation op-9, accepted application version 19 and desired_replicas 6. Its contract records these in the local transaction before cluster reconciliation. What is still unproven?

ASix replicas are ready and serving in the clusterBA local operation record existsCThe desired configuration was accepted at version 19DThe response identifies the accepted operation
Sign up free to answer and see why

Checkpoint

An HPA owns the current replica count. Users report that direct manual scale changes soon revert. What should the platform contract decide first?

AUse a fresh idempotency key on every retryBPermit longer request timeouts for manual changesCRepeat manual writes until the HPA stops changing the fieldDWhether manual requests change approved HPA bounds or are rejected in this ownership mode
Sign up free to answer and see why

Explain how you would specify authorization and concurrency for a platform mutation without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • A platform mutation needs caller scope, bounds, retry identity and concurrency. Accepted intent is separate from ready capacity.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.