Calculate one logistic-regression update and identify a sign error.
When an interview includes model code, a small hand calculation is often more useful than a broad architecture sketch. Choose one example whose expected direction is obvious. A positive example with a positive feature should push its predicted probability upward under a correct loss-minimizing update, assuming no competing terms. If the update does the opposite, inspect the sign and label convention before tuning hyperparameters.
For binary logistic regression, the score is z = wx + b and probability is sigmoid of z. Under binary cross-entropy, the derivative with respect to w for one example is the predicted probability minus the label, multiplied by x. The derivative for b is the predicted probability minus the label. Gradient descent subtracts the learning rate times the derivative.
Keep the loss definition explicit. Some libraries expect logits and apply the sigmoid internally. Applying sigmoid first to a logits-based loss can change the intended calculation. Other errors arise from averaging twice, forgetting to reset gradients, or including padded examples in the denominator. A tiny fixture makes these mistakes visible without requiring a large training run.
A correct local update does not establish a useful model. It is a test of the implementation. Next test whether the model can overfit a tiny clean dataset, whether example shuffling preserves feature-label pairs, and whether evaluation uses the intended mode and preprocessing. These checks isolate mechanics before expensive searches.
Worked example
Use this invented single example: x = 2, y = 1, w = 0, b = 0, and learning rate 0.1. Initially z = 0 and p = 0.5. The weight gradient is negative one and the bias gradient is negative 0.5. After gradient descent, w becomes 0.1 and b becomes 0.05. The new logit is 0.25 and probability is about 0.562.
A sign error that adds the gradient instead produces w = negative 0.1 and b = negative 0.05. The probability falls to about 0.438. That is the wrong direction for this positive example. A large learning rate could also produce instability, but the wrong direction in this controlled first step identifies a more basic defect.
Exercise and solution
Repeat the update for x = 2 and y = 0, with the same initial parameters. State the expected direction and new probability.
The weight gradient is 1 and the bias gradient is 0.5. The updated parameters are w = negative 0.1 and b = negative 0.05. The new logit is negative 0.25, so probability is about 0.438. Award one point each for the two gradients, updated parameters, and direction. Explain why the positive and negative cases form a useful pair of tests. They check label meaning as well as the update sign.
Lab artifact: inspect the update rather than the final loss
The following scalar implementation deliberately avoids an automatic differentiation framework. It is a teaching fixture for modest values, not a numerically stable production loss. Its purpose is to give a reference against which a library calculation can be checked.
python
1from math import exp, log23def loss(w, b, x, y):4 p = 1 / (1 + exp(-(w * x + b)))5 return -y * log(p) - (1 - y) * log(1 - p)67x, y, w, b, rate = 2.0, 1.0, 0.0, 0.0, 0.18p = 0.59dw, db = (p - y) * x, p - y10next_w, next_b = w - rate * dw, b - rate * db11# Reference: dw=-1, db=-0.5, next_w=0.1, next_b=0.05
For large magnitudes, directly computing logarithms of sigmoid probabilities can overflow or reach log zero. The framework's logits-based loss combines the operations in a more stable formulation. Do not use the simplicity of this fixture as a reason to replace a stable library implementation in the product. Compare the same defined function and reduction, with matching data types, before interpreting a numeric difference as a bug.
A finite-difference check adds an independent path. Approximate the weight derivative by [L(w + epsilon) − L(w − epsilon)] / (2 epsilon), holding b and the example fixed. With epsilon equal to 0.00001 and the initial parameters above, the result should be close to negative one. Try a small range of epsilon values. Too large a step introduces approximation error; too small a step can amplify floating-point cancellation. Agreement supports this local derivative, not every batch or every branch of the implementation.
A second failure case: cancellation hides the defect
A batch with x values [2, 2], labels [1, 0], and zero parameters has weight gradients [-1, 1]. Their mean is zero. A broken implementation that accidentally returns zero can pass that single symmetric fixture. Use asymmetric cases and inspect individual contributions.
Example
x
y
p
Weight gradient
A
2
1
0.5
-1
B
1
0
0.5
0.5
Batch mean
-0.25
The bias gradients are negative 0.5 and positive 0.5, so the mean bias gradient remains zero. Under rate 0.1, the mean-reduced batch updates w to 0.025 and leaves b at zero. A sum-reduced loss updates w to 0.05 instead. Neither reduction is inherently wrong, but silently comparing them can look like a learning-rate defect.
Exercise: distinguish three implementations
On that asymmetric batch, implementation P gives weight gradient negative 0.25, Q gives negative 0.5, and R gives positive 0.25. All claim to use mean binary cross-entropy without regularization. Explain the first targeted check for each.
P matches the hand calculation. For Q, inspect whether the loss sums rather than averages, or whether gradients accumulated across two steps. For R, inspect a reversed derivative residual or reversed label meaning. A wrong descent sign changes the parameter update, not the derivative value itself; inspect that boundary separately. These are hypotheses to test, not proven causes from one number. Award one point for the expected gradient, one for each plausible targeted check, and one for the distinction between local correctness and generalization.
Misconceptions to correct
“A decreasing training loss proves labels have the intended meaning” fails if a swapped label convention is applied consistently: the code can optimize the wrong task perfectly. Check known semantic examples. “A tiny overfit test estimates production quality” fails because memorizing a few selected rows tests capacity and mechanics, not population performance.
When presenting the debugging result, report the loss definition, reduction, initial parameters, feature values, labels, and first update. Another engineer should be able to reproduce the result without your full dataset. This artifact is stronger than a screenshot of a falling curve because it identifies the exact calculation being checked.
Interview probe
Original practice: Training loss rises immediately. What do you check before adding a scheduler? A strong answer uses a tiny known example to inspect loss inputs, gradients, sign, and update magnitude. Follow up with a loss that expects logits. A weak answer starts a wide hyperparameter sweep without a mechanical test.
Why can a symmetric batch with gradients -1 and 1 miss a bug that returns zero?
AThe correct batch mean is also zero.BMean reduction discards every negative contribution.CBalanced labels guarantee that every feature produces a zero gradient.DAgreement on one aggregate proves the individual contributions are correct.
A local finite-difference derivative matches automatic differentiation. What has been established?
AThe model generalizes to production.BThe checked local derivative is consistent under the fixture.CAll loss branches and shapes are correct.DThe chosen labels express the desired business target.
Can you calculate the update, distinguish sum from mean reduction, and state what a local fixture does not prove? Rate confidence from 1 to 5 and identify a case that could still hide a bug.
Not yetGetting thereConfident
Wrap-up
A hand-checkable update can isolate a training bug. Establish mechanics before measuring generalization.