Treat infrastructure state as shared operational data
Explain state locking and recover from an uncertain infrastructure run.
Mechanism and reasoning
Infrastructure code describes intent, while state records the tool's mapping between configuration and remote resources. Losing or corrupting that mapping can make the next plan unsafe or confusing. State deserves access control, versioning, backup and a clear ownership boundary.
A lock prevents concurrent state writers when the selected backend supports it. It does not make every remote API call transactional. A run can create a resource and fail before the final state update completes. Recovery needs to inspect both remote reality and the recorded state. Repeating an apply blindly can create duplication or a destructive plan.
A stale-looking lock may still belong to an active run. Check the owner, process and backend evidence before force-unlocking. Removing an active lock can permit two writers and compound the incident. The correct response is not determined by how long the user has waited.
Plans also have a time boundary. A plan describes a comparison at a particular point using particular inputs and provider behavior. Changes to configuration, state or remote resources can make an earlier plan stale. Use the tool's supported saved-plan behavior and review process rather than treating a screenshot as an executable guarantee.
Keep secrets out of broad state access. Marking a value sensitive in output often changes display behavior, not whether it exists in state. The exact backend and provider determine storage details. In an interview, name who can read and write state, how runs serialize, and how the team verifies recovery after an interrupted operation.
Concurrent-run timeline
Teaching environment has one state file for a shared network. Run A acquires the backend lock and starts creating a subnet. Run B starts a minute later and cannot acquire the lock. An impatient operator disables locking for B. Both runs now work from state that may not include the other's actions.
At 10:02, A creates subnet S1. At 10:03, B sees no recorded subnet in its stale view and attempts a competing change. Depending on remote API constraints, B can create a duplicate, fail or modify something A expects to own. The lock was a coordination control. Bypassing it did not remove work; it removed serialization around that work.
Now use a different failure. A creates S1, then the process loses connectivity before it can record a completed result. The lock remains. The safe investigation checks whether A is still alive, whether the remote subnet exists, what the backend state version contains and whether the operation has an outstanding provider response. Only after establishing that no writer remains should the owner follow the backend's supported recovery procedure.
Evidence
What it establishes
What it does not establish
Lock owner and ID
Which run claims write ownership
That the process is definitely alive
Process status
Whether the known runner is active
That every remote operation finished
Remote resource lookup
Whether S1 currently exists
Whether state records it correctly
State version history
Recorded mappings over time
Complete remote transaction atomicity
The next plan is part of verification. If it proposes creating a subnet that already exists, reconcile the mapping through supported import or recovery mechanisms under review. If it proposes replacing a critical resource unexpectedly, stop and inspect the mismatch. Do not approve a destructive plan merely because the previous run was labelled failed.
State ownership decision
Partition state by operational ownership and failure scope. One enormous state can serialize unrelated teams and enlarge the impact of a mistake. Many tiny states can create dependency coordination and output-sharing problems. Choose boundaries where resources change together and have a clear owner. A network shared by several services may need a separate controlled state from each application's compute.
Define a recovery drill using disposable resources. Interrupt a run after a known remote create, inspect recorded and actual state, then restore a clean plan without creating duplicates. Record the tool and provider versions because recovery behavior can differ. This drill teaches the difference between code, state and remote reality better than a diagram alone.
Access review is equally practical. A read-only dashboard user may not need state-file access, because state can contain sensitive values. A CI runner that applies changes needs write authority to its specific backend and the relevant remote resources, not every environment. Log run identity and approvals without printing secrets. The lesson's point is operational integrity: the mapping, writer and observed infrastructure must agree before the next change.
A useful recovery record names the last confirmed state version, the uncertain remote operation and the exact next verification. For the subnet example, record that S1 exists, the runner is stopped and the state mapping is absent or present. The next action then follows from evidence rather than a generic failed-run label. Have another authorized reviewer inspect a proposed replacement or deletion when the plan affects shared infrastructure. Keep the original state version available under the backend's retention policy. Restoring old state blindly can also be wrong because the remote resource may have changed since that version. Recovery must reconcile both sides.
Worked example
Teaching timeline:
Time
State
Remote reality
Run status
10:00
No subnet
No subnet
A holds lock
10:02
No confirmed subnet mapping
S1 created
Response uncertain
10:03
Last version unchanged
S1 exists
A disconnected
A failed run does not imply no change occurred. Verify A's status and S1 before recovering the lock or retrying. The expected final state is one mapping to S1 and a subsequent reviewed plan with no unintended duplicate create or replacement.
Exercise
A lock is fifteen minutes old. The runner still shows an active apply, and the cloud API shows an operation in progress. Should another operator force-unlock? State the evidence needed before any recovery.
Model solution and rubric
No. There is evidence of an active writer and remote operation. Coordinate with the run owner and observe or stop the run through the approved mechanism. Before recovering a failed unlock, establish that the writer is no longer active, identify the exact lock and inspect remote/state outcomes. Then use the supported backend procedure and review the next plan. Age alone is not proof that a lock is abandoned.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: state locking makes the entire cloud change atomic. It serializes state writers but remote operations can partially complete. Misconception 2: sensitive output means secret-free state. Display redaction does not necessarily remove stored values, so backend access still needs control.
Interview probe
Evidence class: recommended. Original practice.
An apply failed after creating a resource. What do you do before retrying?
Strong answer: Inspect the remote resource, recorded state and run ownership. I would resolve any mapping discrepancy through supported recovery, then review a fresh plan. I would not infer that failure means nothing changed.
Follow-up: When would you split a shared state into separate ownership boundaries?
Weak answer indicators: Force-unlock based only on age; disabling locks to avoid waiting; assuming failed runs have no side effects.
Sources
Technical references: Terraform state locking; Terraform state; OWASP secrets management. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
A resource exists remotely after an apply timeout, but the last state version has no mapping. What is the next defensible step?
ARetry the create because the exit code failedBInspect run status and reconcile remote/state mapping through supported recoveryCDelete state and let the tool rediscover all resourcesDAssume the remote resource belongs to a different team
A secret output is marked sensitive and hidden by the CLI. Which access decision is justified?
AGrant all developers state read access because output is redactedBPublish the state after removing its output blockCKeep backend access restricted until actual stored content is assessedDTreat state encryption as permission for unrestricted reads
Explain how you would explain state locking and recover from an uncertain infrastructure run without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Keep code, state and remote reality distinct. Serialize writers and verify partial outcomes before retrying.