Lesson 1 of 4 · 35 min

A threshold is an operating decision

Calculate the cost and workload of two candidate thresholds.

A binary classifier usually produces a score before an application turns it into an action. The threshold determines which examples trigger that action. A default threshold of 0.5 has no special claim to business correctness. Its usefulness depends on probability calibration, error costs, and operational constraints.
Write the decision before choosing the metric. A fraud score might send a transaction to manual review. A false positive then consumes reviewer time and delays a legitimate customer. A false negative leaves fraud unreviewed. If the action instead blocks a transaction, the costs differ. The same model can need different thresholds for different actions.
A confusion matrix turns a proposed threshold into counts. From true positives and false positives, calculate precision. From true positives and false negatives, calculate recall. Then apply a stated cost model. If manual review has limited capacity, inspect the number of flagged cases as well. A threshold with lower expected error cost may be infeasible if it exceeds the available queue.
These calculations assume the evaluation population represents deployment and the stated costs are appropriate. If the class prevalence changes, precision can change even when conditional error rates remain similar. If review quality varies under overload, a fixed per-case cost is too simple. Explain those assumptions instead of presenting the arithmetic as a complete policy.

Worked example

The following invented holdout contains 1,000 transactions and 50 fraud cases.
ThresholdTPFPFNTNReviews
High30202093050
Low408010870120
At the high threshold, precision is 30/50, or 60%, and recall is 30/50, or 60%. At the low threshold, precision is 40/120, about 33.3%, and recall is 80%. Let false negatives cost 100 units and false positives cost 2. High costs 2,040 units; low costs 1,160. If capacity is 80 reviews, however, the low threshold cannot run unchanged. The next step is to evaluate a feasible threshold or a top-80 policy on validation data.

Exercise and solution

A third threshold has TP 36, FP 44, FN 14, and TN 906. Calculate reviews, precision, recall, and cost under the same assumptions.
It produces 80 reviews, precision 45%, recall 72%, and cost 1,488 units. It fits the 80-review capacity and improves estimated cost over the high threshold. It does not establish that it is globally optimal among all thresholds. Award one point each for the four calculations and one for the limited conclusion. Select the policy on validation data, then estimate its performance on untouched test data.

Lab artifact: the queue contract

A threshold is incomplete without the unit of capacity. Eighty reviews per minute is not the same constraint as eighty reviews per day. Nor is a daily average a guarantee that an hourly queue will stay within its service limit. Record the arrival window, reviewer throughput, and maximum acceptable age. These determine whether backlog can recover after a burst.
code
1Policy candidate: score >= 0.632Scoring window: one day3Validation reviews: 80 of 1,000 transactions4Review capacity: 80 completed reviews per day5Initial backlog: 156Max acceptable backlog at day end: 107Expected end backlog: 15 + 80 - 80 = 158Capacity conclusion: daily inflow fits; backlog target fails
The earlier third threshold fits the simple arrival constraint, but it fails this stricter backlog contract. To reduce the backlog from fifteen to ten, today's admitted work must be at most seventy-five if eighty reviews are completed. This is an original queue accounting example, not a statement about a particular company's review system. The calculation assumes completed reviews are genuinely available throughput, not scheduled staff time. Breaks, investigation complexity, and rework can reduce effective capacity.
If the business wants a top-seventy-five policy, select and evaluate that policy explicitly. Do not calculate threshold metrics for eighty reviews and silently drop five later. A top-k policy ranks within a specified window. It can have a different score cutoff every day and needs a tie rule. When urgent items arrive after a batch has consumed the allowance, a daily batch policy may be unsuitable. A streaming priority queue, an emergency reserve, or a second action such as temporary restriction would be a new policy requiring its own evaluation.

A second failure case: counting the wrong cost

The first cost model assigns two units to each false positive. Suppose reviewers cost one unit per reviewed transaction, and a legitimate customer delay costs two additional units. The cost function becomes 100 × FN + 2 × FP + 1 × (TP + FP). Review labor now includes true positives. For the third threshold, cost is 1,400 + 88 + 80 = 1,568 units. Simply changing FP cost from two to three would omit the thirty-six units spent reviewing true positives.
Use a small executable specification to make the accounting inspectable:
python
1def policy_report(tp, fp, fn, tn, initial_backlog, throughput):2    reviews = tp + fp3    return {4        "precision": tp / reviews if reviews else None,5        "recall": tp / (tp + fn) if tp + fn else None,6        "cost": 100 * fn + 2 * fp + reviews,7        "end_backlog": max(0, initial_backlog + reviews - throughput),8    }9# Counts are evaluation outcomes, not information available at decision time.
Undefined precision for zero reviews should remain undefined. Replacing it with one can make a useless no-action policy look attractive in a dashboard. The cost model might still favor no review in a different setting, but that conclusion must come from costs, not a fabricated precision value.

Exercise: defend a constrained comparison

Candidate P has TP 28, FP 32, FN 22, TN 918. Candidate Q has TP 35, FP 45, FN 15, TN 905. Use the original error-only cost model, a seventy-five-review admission limit, and an initial backlog of fifteen with throughput eighty. Calculate both costs and explain the decision.
P flags sixty, costs 2,264, and ends with backlog zero under this daily approximation. Q flags eighty, costs 1,590, and ends with backlog fifteen. Q has lower estimated error cost but violates both the admission and end-backlog limits. P is a feasible reference. The evidence does not prove that P is the best feasible policy; evaluate a revised Q cutoff or a defined top-seventy-five policy on development data. Award one point for each cost, one for queue arithmetic, and two for separating feasibility from global optimality.

Misconceptions to correct

“Higher recall means better review decisions” fails when added false positives consume the queue or delay more valuable cases. “Top-k automatically solves capacity” fails when the chosen window, initial backlog, or variable service time makes k a poor representation of actual throughput. State which constraints the evaluation really checks before recommending release.

Interview probe

Original practice: Would you choose the model with higher F1? A strong answer first identifies the action, costs, prevalence, and capacity, then explains whether F1 matches them. Follow up with a review queue limit. A weak answer chooses the largest single metric without translating it into decisions.

Sources

docsscikit-learn: model evaluation metricsscikit-learn.orgdocsscikit-learn: probability calibrationscikit-learn.org

Checkpoint

A lower error-cost threshold sends 120 items per day to capacity 80. Which next step supports a release decision?

AEvaluate an explicit policy whose admitted workload fits capacity.BDeploy the threshold and exclude late reviews from metrics.CCap the queue at 80 but retain the 120-item policy's metrics.DUse the final test labels to select which 40 items to omit.
Sign up free to answer and see why

Checkpoint

There are 15 pending reviews, throughput 80 per day, and a desired end backlog at most 10. Maximum new admissions?

A80B90C75D65
Sign up free to answer and see why

Checkpoint

Labor costs one per review, false-positive delay costs two, and missed positives cost 100. Which formula matches?

A100 FN + 3 FPB100 FN + 2 FP + TP + FPC100 FN + 2 FP + FND100 FN + 2 FP + TP
Sign up free to answer and see why

Checkpoint

Two scores tie at the last top-k slot. What belongs in the policy before evaluation?

ABreak ties using the eventual label.BCount both as half a review.CInclude both even if k is a hard maximum.DSpecify a reproducible tie rule and evaluate its decisions.
Sign up free to answer and see why

Checkpoint

A policy sends no cases to review. What is its precision?

AOne, because there are no false positives.BUndefined; report workload and other costs separately.CZero recall, therefore precision must also be zero.DThe prevalence of the whole evaluation set.
Sign up free to answer and see why

Can you calculate workload, cost, and end backlog, then state exactly which constraints a policy satisfies? Rate confidence from 1 to 5 and identify the calculation you would recheck.

Not yetGetting thereConfident

Wrap-up

  • Choose a threshold as an operating policy. Show its error counts, costs, and workload.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.