Request a scoping call Contact
← Research

Do not ask a quality model for a score. Ask for an intervention.

A quality model can help engineers choose a process intervention. Its commercial value depends on a safe comparison that connects the change to saleable output, stability and cost.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A rise in scrap gives a plant manager a problem to investigate. A quality model may identify measurements associated with failed parts, but that association does not establish which adjustment will help. The next decision is a safe, testable process change with an agreed comparison and stopping rule.

A quality score can prioritise investigation. To support an operational decision, it needs to inform a specific intervention that an engineer can assess and production can test. The useful sequence connects the signal to an approved action, a controlled comparison and a decision about further use.

The mechanism

Plant quality, operations and engineering leaders already work with measurements, inspection results and knowledge of the equipment. The question for an AI assessment is therefore specific: which process change could improve the next controlled batch, and what evidence would justify trying it?

That boundary matters because quality data is rich in correlation and poor at proving what caused a defect. A model can rank a pressure, temperature or tool condition that appears alongside yield loss. An engineer still has to check whether the proposed setting is safe, available on the line and plausible in the process. The model should make the next experiment easier to choose and inspect.

Senoner, Netland and Feuerriegel tested this pattern in semiconductor manufacturing. Their explainable model used historical production data to select process improvement actions rather than only predicting yield. In a field experiment, the selected actions reduced yield loss by 21.7 per cent against the sample average. A later rollout on a different transistor product reported a 51.3 per cent reduction in yield loss. Senoner, Netland and Feuerriegel, Using Explainable AI to Improve Process Quality

The numbers are evidence for a workflow, not a promise for another plant. The useful idea is that a quality system can produce an action hypothesis with the signals behind it, then ask production to run a fair test.

What to do about it

Turn a quality signal into a production decision Fig. 01
  1. 01 Measure Define yield loss, scrap, rework and the process window from the current line.
  2. 02 Prioritise Rank an intervention and show the source measurements and affected machine.
  3. 03 Intervene Have an engineer approve a bounded change with a stop rule.
  4. 04 Compare Run a controlled batch or matched line and preserve the unchanged baseline.
  5. 05 Roll out Expand only when output, quality and stability clear the agreed threshold.

Start with one defect family and one line. Record the current yield loss, scrap cost, rework hours, downtime and throughput. Define the unit that the business will use for the decision, such as good parts per shift or contribution margin per wafer. A model that improves an intermediate score while the line makes fewer saleable units has not improved the operation.

Ask the system for an intervention record, not a probability in isolation. The record should name the proposed process or machine change, the measurements that support it, the expected direction of effect, the engineer who owns the test and the conditions under which the change must be reversed. Keep the recommendation separate from the policy that permits it. A plant quality lead should be able to reject a suggestion without changing the model’s history.

Run the test where the line can support a comparison. The semiconductor study used a new production batch of 24 wafers split into four groups of six, with 372 chips per group. The groups experienced the same conditions except for the selected process actions. That design made it possible to connect the action to the observed yield result rather than to a favourable week. The field-experiment design

The measurement plan should include the economics that made the problem worth solving. Track yield loss and scrap value first. Add rework time, line downtime, throughput, changeover time and any extra inspection or maintenance. Set a comparison group or matched baseline, a minimum effect worth keeping and a time window long enough to catch drift. Report overrides and stopped tests alongside successful interventions. Those are operating signals, not failures to hide.

Speed alone can also mislead. In a separate manufacturing field experiment, augmented-reality glasses cut completion time for a new task by 43.8 per cent, but after the glasses were removed those workers took 23 per cent longer than the paper-instruction group. The result is about augmented reality rather than AI, yet it gives a useful check: measure retained capability and normal-shift performance after the assistance ends. Raisch and colleagues, Seeing the Bigger Picture? Ramping up Production with the Use of Augmented Reality

The Inference Institute provides independent advice on that assessment. We can map the defect workflow, define the actions a model may recommend and specify a credible comparison. Architecture and risk advice can also clarify evidence records, review responsibilities and operating limits before a supplier is selected.

A useful assessment brief identifies the defect family, line, proposed intervention, owner, baseline, observation period, economic measures and stopping rule. Production teams can use it to evaluate the change under their own engineering and safety controls and decide whether further investment is justified.

What this does not tell you

The semiconductor evidence comes from one operating setting. Its field test was small, its model selected correlational actions and the later product rollout was not a randomised comparison. The augmented-reality study uses a different technology and measures a different task. Neither study establishes a general return on an AI quality programme.

A recommendation can be wrong when the process changes, a sensor drifts or an unmeasured constraint moves with the proposed action. Keep a human owner for the intervention, preserve the unchanged comparison and stop when quality or safety moves outside the agreed boundary. The plant leader should be able to explain which decision changed, what output improved and why the next rollout is justified.

A plant quality director needs evidence that a permitted intervention improves the operation. Model accuracy contributes to that assessment, but the funding decision should turn on good output, cost, process stability and the limits of the comparison.

Filed under · Method · Manufacturing · Quality · Process improvement Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.