Define the decision before the dataset

The operator sees more flags and assumes the new line produces more defects. An engineer sees a rise in uncertain scores and assumes the model needs retraining. Before either explanation wins, put a few original images beside the pilot images. The changed conditions may be visible without running another experiment.

Begin with what a person or machine will do after the prediction. Will it reject an item, prioritise a review queue, count objects or request another image? The consequence of a false positive differs from the consequence of a missed defect. Those differences should shape the data collection and the release criteria.

In an illustrative packaging inspection, flagging every uncertain image may protect quality but overwhelm the review team. Automatically accepting a weak prediction may reduce the queue while passing a costly defect. The operating decision needs both error rates and the capacity of the people reviewing the output.

Evaluate the conditions the camera actually sees

Collect examples across the relevant cameras, shifts, lighting, backgrounds and product variants. Keep a record of how labels were assigned and how disagreements were resolved. A dataset with neat images from one controlled setup may not represent the installation where the model will run.

WILDS collects real-world distribution shifts and benchmarks how models perform across them. Its datasets make a broad problem tangible: evaluation and deployment data can differ in consequential ways. The appropriate lesson for this installation is to test the variations it will actually encounter.

Split evaluation data to avoid near-duplicate images leaking between development and final assessment. Frames from the same object or sequence can make performance look stronger than it will be on unfamiliar material. Choose the split around the variation the production system must handle.

Report the important categories separately. A common easy class can dominate a headline score while a rare but costly defect remains weak. Include examples of misses and false alarms so the operations team can judge their practical meaning.

Give confidence scores an empirical meaning

A model score is not automatically a trustworthy probability of correctness. Guo and colleagues study calibration in modern neural networks and show why confidence and accuracy can diverge. A threshold that looks reassuring needs evaluation on data representative of the intended use.

Choose thresholds with the process owner using held-out examples. Measure how many cases enter review and which errors remain outside it. Review capacity is part of this choice: a queue that grows faster than staff can handle it is an operational failure even if the model is conservative.

Treat unfamiliar inputs explicitly. A new camera, a changed product or a damaged image can fall outside what was evaluated. Define when the system should request another image, hold the item or route the case to a person.

Close the feedback loop deliberately

Keep the original image beside the prediction and the region or evidence the reviewer needs. Record corrections with a clear label policy. Avoid automatically retraining a live system on every correction; review the new data and evaluate a candidate release before replacing the current model.

The monitoring concerns in The ML Test Score are useful here: readiness includes how a system is checked after it is deployed. In the imagined second line, the review queue and input conditions should be inspectable alongside model outputs.

Monitor changes in incoming material and the review workload after rollout. Investigate a rise in uncertain images before treating it as a model problem: a camera may have moved, a light may have failed or a supplier may have changed the packaging.

The operating envelope should be written plainly: supported inputs, evaluated conditions, thresholds, human review and fallback. Expanding that envelope is a new engineering decision with new evidence.

  • Evaluate by defect and operating condition.
  • Keep related images out of opposing dataset splits.
  • Choose thresholds with review capacity in mind.
  • Version data, labels, model and deployment together.

Back at the second line

Record the changed conditions and check whether the new line operates within those covered by the evaluation. Use the agreed fallback during the investigation and name the person who decides when normal operation can resume.

If the new setup becomes supported, preserve the cases that established it. Next time a camera moves or a supplier changes the packaging, the team should know which boundary to revisit. The system becomes more dependable when that knowledge survives the people who first built it.

Sources & further reading

  1. Guo et al. — On Calibration of Modern Neural Networks
  2. Koh et al. — WILDS: A Benchmark of in-the-Wild Distribution Shifts
  3. Breck et al. — The ML Test Score

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.