Request a scoping call Contact
← Research

Design the review queue alongside the screening model

Translate screening metrics into alert volumes and review work. At low prevalence, false positives can dominate the queue, so acceptance must include missed cases, review capacity and the consequences for people flagged.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A screening model’s sensitivity and specificity need an operating interpretation. Multiply them by the expected population and prevalence to estimate alerts, missed cases and review effort. Include the people who will work that queue in the acceptance decision.

Review capacity can become a constraint even when the classifier meets its stated targets. Record the assumptions about volume, handling time and acceptable delay. Test them against the workflow before accepting the operating model.

A rare target can produce more false alerts than true alerts unless the false-positive rate is sufficiently low. The result depends on prevalence, sensitivity and specificity. It is not inevitable for every rare-event classifier, but the calculation belongs in the service design.

The arithmetic, once

Take a hundred thousand cases a month, and suppose one in a thousand is the thing you are looking for. That is a hundred genuine cases. Now take a model that catches ninety-nine of every hundred genuine cases, and that wrongly flags one in every hundred ordinary ones — figures most teams would be pleased to report.

What one month of alerts contains, at one genuine case in a thousand Fig. 01
  • Genuine cases found 99 alerts Of the 100 present in the population.
  • Ordinary cases wrongly flagged 999 alerts One per cent of the 99,900 that were not the thing.

Worked from a 100,000-case population at one-in-a-thousand prevalence

In this constructed example, the system produces 1,098 alerts: 99 true positives and 999 false positives. About 9% of flagged cases are genuine. That ratio describes the queue produced by these assumptions. It does not establish that staff will ignore alerts or that the service will perform worse than no screening.

The cybersecurity review of the base-rate problem discusses why rarity matters when interpreting detection results. Apply the calculation to the actual population and operating point. Do not use the worked example as a measured estimate of a particular screening service.

Why this is a design decision, not a tuning exercise

Raising a threshold can reduce alert volume while missing more genuine cases. Evaluate the whole trade-off at candidate operating points. Its acceptability depends on the harm from a missed case, the cost of a wrong flag and the capacity to review.

What each way out actually costs Fig. 02
The move What it spends
Raise the threshold Missed genuine cases. In a screening system, this is the harm the system exists to prevent.
Add reviewers Additional staffing cost, with handling time and review quality to measure.
Narrow the population screened Changes coverage and possibly prevalence, with distributional effects to assess.
Stage the screening A further filtering stage with end-to-end recall, delay and cost to evaluate.

Staged screening can reduce expensive review by applying a further check to selected cases. It can also lose genuine cases at the first stage or introduce correlated errors. Measure end-to-end recall, precision, delay and cost. It is one design option, rather than the only scalable approach or a way to avoid every trade-off.

Narrowing the screened population changes coverage and may change prevalence. Assess both instead of assuming precision will improve. The selection rule can affect groups differently and leave some risks outside screening. That decision needs business and governance review as well as statistical evaluation.

What to specify before the model is built

The last two are the ones that make this a governance artefact rather than a metrics table. A false positive is not a rounding error to the person it lands on: it is an account frozen, an application delayed, a claim investigated, a name on a list. Systems that record only the aggregate error rate have no way to see that cost, and no way to notice when it falls unevenly.

Give the threshold a named owner and a change process. Queue pressure is relevant evidence, but it should not silently redefine the tolerated balance of missed cases and false flags. Record the reason for a change, its expected effects and the measurements required afterwards.

What this does not tell you

The arithmetic above is a worked example with round numbers, not a claim about any real deployment. Prevalence varies enormously by domain, and the whole point of the exercise is that you have to compute it for your population rather than borrowing a figure from someone else’s.

This is a method for deciding whether and how screening fits the task. Acceptance concerns the classifier, reviewers, thresholds and resulting actions together. A model score alone cannot establish that the service improves outcomes or that its review burden is affordable.

The operating-model owner should ask for expected alerts per reviewer, missed cases and the consequences of each error. Require evidence for the handling process at the intended volume. Use those results to accept a bounded scope, revise the design or request further measurement.

Filed under · Method · Evaluation · Screening · Operating model Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.