Lesson 4 of 4 · 35 min

Make training and serving use the same feature meaning

Specify a model package that preserves preprocessing and missing-value semantics.

A trained model expects a particular feature representation. Column names alone do not capture that representation. Units, category vocabularies, missing-value treatment, scaling parameters, and feature order all affect predictions. Training-serving skew occurs when the live path supplies a different meaning than the path used during model development.
Fit learned transformations on training data only. Then carry those fitted transformations with the estimator. If a scaler learns a mean from the full dataset before validation, the evaluation has used information from outside training. If serving refits the scaler on each batch, the same user can receive different predictions depending on who else is in the batch. Neither behavior matches a stable fitted model.
Define handling for unseen categories and absent fields. An unknown country code should not silently become a known category through positional indexing. A missing numeric feature should follow the documented imputation or rejection rule. Keep a schema version and a transformation version with every prediction. This makes rollback and incident comparison possible.
Parity tests should use small fixed examples that pass through both paths. Compare the transformed feature vector as well as the final score. A model score can coincidentally match while the underlying features differ. Include edge cases: missing input, unseen category, boundary timestamp, and different units. These tests do not prove statistical quality, but they establish that the evaluated function is the function being served.

Worked example

A fictional model expects account age in days and purchase amount in rupees. Training uses age 30 and amount 500. Serving accidentally supplies age in seconds, 2,592,000. A schema that only says "number" accepts it. A semantic contract with units, plausible ranges, and a parity fixture detects the mismatch.
For a simpler scale example, training mean is 100 and standard deviation is 20. An amount of 140 transforms to 2. Serving must reuse those stored values. If it refits on a live batch with mean 140, the same amount transforms to 0. The model has received a different input even though the raw amount is identical.

Exercise and solution

Training categories are basic, plus, and enterprise. Serving receives "student". The model package specifies an explicit unknown bucket and stored category order. Explain the correct representation and two checks.
The value maps to the unknown bucket without changing the learned category order. A fixture compares training and serving vectors for "student". A second fixture checks an absent plan value under its separate missing-value rule. Award one point each for stable vocabulary, explicit unknown handling, separate missing semantics, and vector-level parity. Refit-on-request is not a valid repair because it changes the feature space.

Inspect the feature contract

This teaching schema records meaning that a plain numeric type cannot express.
code
1feature_schema: account-risk-v42ordered_features:3  age_days:4    unit: days5    missing: reject6    valid_range: [0, 36500]7  purchase_rupees:8    unit: INR9    missing: impute_with_training_median10    scaler: mean=100, standard_deviation=2011  plan:12    categories: [basic, plus, enterprise, unknown]13    unseen: unknown14    missing: separate_missing_flag15transform_version: T416model_version: M9
The example's ranges and imputation choices are invented for teaching. A real contract needs domain review and data evidence. A range can detect seconds accidentally supplied as days, but it must not reject valid old accounts merely because a convenient threshold was guessed. The contract should distinguish a hard semantic impossibility from an unusual but valid value.
Feature order is explicit because some model formats consume arrays without names. Swapping two numeric columns can pass type checks and produce plausible scores. The package should either bind names to positions or validate the exact ordered schema. A model file and transform file also need compatible versions; rolling back only one can create skew.

A second worked case: category order changes

Training encodes plan as [basic, plus, enterprise, unknown]. The raw value plus becomes [0,1,0,0]. A serving process sorts categories alphabetically and uses [basic, enterprise, plus, unknown], producing [0,0,1,0] for the same value. Both vectors have length four and contain one active element. A shape check passes while the semantic input changes.
Raw inputExpected vectorBuggy vector
basic[1,0,0,0][1,0,0,0]
plus[0,1,0,0][0,0,1,0]
enterprise[0,0,1,0][0,1,0,0]
student[0,0,0,1][0,0,0,1]
Testing only basic and student would miss the swapped known categories. Choose fixtures that cover every meaningful mapping, not just two convenient examples. The expected vector must come from the declared training contract, not from calling the same potentially buggy encoder in both paths.

Test batch independence

A fixed fitted transform should produce the same feature vector for one raw request regardless of unrelated companion requests, unless the model explicitly includes batch-dependent semantics. Test the same request alone and beside a very large purchase value. If its standardized feature changes, serving may be refitting statistics on the batch.
This check is distinct from model mode. Some neural components use different behavior in training and evaluation mode. Verify that the serving path uses the intended mode and stored statistics. The broader lesson is that the deployed function includes preprocessing and mode state, not only parameter tensors.

Misconceptions to reject

"Matching tensor shape proves feature parity" ignores units, ordering, vocabulary, and missing-value semantics. A wrong vector can look structurally valid.
"Refitting preprocessing on live data keeps the model fresh" changes the feature space without retraining and validating the estimator under that change. Freshness needs a deliberate model-update procedure, not an accidental request-time fit.

Transfer exercise

A model expects temperature in Celsius, and its stored scaler uses mean 20 and standard deviation 10. A sensor sends 68 Fahrenheit but labels the field temperature. What should the parity test and adapter do?
The adapter must convert 68 Fahrenheit to 20 Celsius before the stored transform, yielding standardized value 0. Passing 68 directly yields 4.8 and changes the meaning. The fixture should include raw unit metadata, converted value, and final vector. Award one point for conversion, one for the correct standardized result, one for rejecting an unlabeled unit assumption, and one for checking the same path in training and serving.

Check imputation fit state

Suppose observed training values are 10, 20, and 30, so a training-median imputer stores 20. A held-out observed value is 1,000. Fitting the imputer on all four observed values would use the middle pair 20 and 30, giving median 25. That is a different fitted transformation informed by held-out data. The magnitude of the change is not the key issue; the evaluation boundary was crossed.
For a missing live value, the packaged training imputer should supply 20 under this contract. A request-time median based on companion rows may supply something else or fail when every value is missing. Add a fixture with an all-missing request batch and define whether the stored imputer can handle it. This checks a realistic serving condition without pretending that imputation itself guarantees good predictions for missing-heavy traffic. Report missingness slices in model evaluation because preserving the intended transform does not prove that its imputed predictions are accurate.

Interview probe

Original practice: Why package preprocessing with model weights? A strong answer explains that transformations are part of the learned function and must preserve fit state. Follow up with batch-wise rescaling. A weak answer checks only the estimator file hash.

Sources

docsscikit-learn: common pitfalls and data leakagescikit-learn.orgdocsGoogle: rules of machine learningdevelopers.google.com

Checkpoint

A fitted training scaler is needed at serving. What should the service do?

ARefit it on each request batch.BReuse its stored fit state to transform inputs.CFit it on validation plus live rows.DSkip it whenever inputs look unusual.
Sign up free to answer and see why

Checkpoint

Equal-length encoders swap plus and enterprise positions. Which check detects the semantic mismatch?

ACompare only total active entries per vector.BCompare only unknown-category handling.CCompare only shape and dtype.DTest both raw categories against their expected ordered vectors.
Sign up free to answer and see why

Checkpoint

Celsius scaler mean 20, std 10 receives 68 Fahrenheit. Correct normalized value after conversion?

A0B2C4.8D6.8
Sign up free to answer and see why

Checkpoint

Training values 10, 20, 30 give median 20. Adding held-out 1000 makes joint median 25. Which imputer should final validation use?

AA new median per validation row.BThe joint median 25 because more data improves estimates.CThe training-fitted median 20.DThe validation median 1000.
Sign up free to answer and see why

Checkpoint

A fixed request's transformed vector changes when unrelated companion rows change. Which cause should you inspect first?

AA fixed training-fitted scaler applied identically to each row.BRequest-time refitting or unintended batch-dependent preprocessing.CA changed model threshold after the feature vector is produced.DThe label horizon used only for later evaluation.
Sign up free to answer and see why

Explain how you would prove raw-input-to-feature parity across training and serving. Use one supplied fixture and identify a condition that would invalidate your conclusion. Rate confidence from 1 to 5.

Not yetGetting thereConfident

Wrap-up

  • Ship the fitted input function with the model. Test feature vectors at the training-serving boundary.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.