Design the review queue alongside the screening model
Translate screening metrics into alert volumes and review work. At low prevalence, false positives can dominate the queue, so acceptance must include missed cases, review capacity and the consequences for people flagged.
Keep development and test groups separate before comparing performance. No measured results are shown.
A screening model’s sensitivity and specificity need an operating interpretation. Multiply them by the expected population and prevalence to estimate alerts, missed cases and review effort. Include the people who will work that queue in the acceptance decision.
Review capacity can become a constraint even when the classifier meets its stated targets. Record the assumptions about volume, handling time and acceptable delay. Test them against the workflow before accepting the operating model.
A rare target can produce more false alerts than true alerts unless the false-positive rate is sufficiently low. The result depends on prevalence, sensitivity and specificity. It is not inevitable for every rare-event classifier, but the calculation belongs in the service design.
The arithmetic, once
Take a hundred thousand cases a month, and suppose one in a thousand is the thing you are looking for. That is a hundred genuine cases. Now take a model that catches ninety-nine of every hundred genuine cases, and that wrongly flags one in every hundred ordinary ones — figures most teams would be pleased to report.
In this constructed example, the system produces 1,098 alerts: 99 true positives and 999 false positives. About 9% of flagged cases are genuine. That ratio describes the queue produced by these assumptions. It does not establish that staff will ignore alerts or that the service will perform worse than no screening.
The cybersecurity review of the base-rate problem discusses why rarity matters when interpreting detection results. Apply the calculation to the actual population and operating point. Do not use the worked example as a measured estimate of a particular screening service.
Why this is a design decision, not a tuning exercise
Raising a threshold can reduce alert volume while missing more genuine cases. Evaluate the whole trade-off at candidate operating points. Its acceptability depends on the harm from a missed case, the cost of a wrong flag and the capacity to review.
| The move | What it spends |
|---|---|
| Raise the threshold | Missed genuine cases. In a screening system, this is the harm the system exists to prevent. |
| Add reviewers | Additional staffing cost, with handling time and review quality to measure. |
| Narrow the population screened | Changes coverage and possibly prevalence, with distributional effects to assess. |
| Stage the screening | A further filtering stage with end-to-end recall, delay and cost to evaluate. |
Staged screening can reduce expensive review by applying a further check to selected cases. It can also lose genuine cases at the first stage or introduce correlated errors. Measure end-to-end recall, precision, delay and cost. It is one design option, rather than the only scalable approach or a way to avoid every trade-off.
Narrowing the screened population changes coverage and may change prevalence. Assess both instead of assuming precision will improve. The selection rule can affect groups differently and leave some risks outside screening. That decision needs business and governance review as well as statistical evaluation.
What to specify before the model is built
The last two are the ones that make this a governance artefact rather than a metrics table. A false positive is not a rounding error to the person it lands on: it is an account frozen, an application delayed, a claim investigated, a name on a list. Systems that record only the aggregate error rate have no way to see that cost, and no way to notice when it falls unevenly.
Give the threshold a named owner and a change process. Queue pressure is relevant evidence, but it should not silently redefine the tolerated balance of missed cases and false flags. Record the reason for a change, its expected effects and the measurements required afterwards.
What this does not tell you
The arithmetic above is a worked example with round numbers, not a claim about any real deployment. Prevalence varies enormously by domain, and the whole point of the exercise is that you have to compute it for your population rather than borrowing a figure from someone else’s.
This is a method for deciding whether and how screening fits the task. Acceptance concerns the classifier, reviewers, thresholds and resulting actions together. A model score alone cannot establish that the service improves outcomes or that its review burden is affordable.
The operating-model owner should ask for expected alerts per reviewer, missed cases and the consequences of each error. Require evidence for the handling process at the intended volume. Use those results to accept a bounded scope, revise the design or request further measurement.