Can Foundry AI Learn from Visual Inspection? A Measurement-System Gate for Casting Quality Labels
- Castella

- 7 days ago
- 5 min read
Short answer: Foundry AI should learn from an inspection result only after the label-producing measurement process is shown to be repeatable, reproducible, stable and traceable for that specific decision. Correct shot-to-part linkage is necessary, but it can still connect perfect process data to an inconsistent “OK”, “scrap” or defect-class label.
The previous Castella guide established the shot-to-defect data link. The next engineering question is harder: can the linked quality result be trusted as a target for machine learning?
Why a linked label can still be wrong
A supervised model does not observe the physical truth directly. It learns the decision recorded by a visual inspector, leak tester, X-ray evaluation, dimensional gauge or laboratory method. Variation may enter through lighting, part temperature, surface condition, fixture, device resolution, calibration, operator interpretation, defect terminology or a revised acceptance rule.
NIST describes production measurement characterization in terms including repeatability, reproducibility, stability, bias, resolution and differences among gauges. ISO 22514-7 likewise frames measurement-process capability relative to the measurement task: uncertainty must be considered against the specification or the process variation. The practical implication is simple: “the gauge is calibrated” does not prove that the complete inspection process is capable of producing dependable AI labels.
Casting research shows why this matters. A foundry-focused study on defect classification explains that visual inspection, stochastic defect formation and overlapping acceptance classes can create data-space overlap that limits supervised learning. A 2025 magnesium HPDC study explicitly notes that later visual quality inspection is subjective and may contain errors. Another industrial HPDC study demonstrates strong part traceability through data-matrix codes, yet also shows how automatic early rejection can leave parts without downstream quality labels. Traceability, coverage and label reliability are separate controls.
Treat the inspection result as a measurement event
Do not store only a final defect code. Preserve the event that produced it:
part, shot, cavity and route identifiers;
characteristic or defect taxonomy and rule revision;
method, device, program and fixture identifiers;
inspector or automated-system version;
timestamp, shift and relevant environmental condition;
raw value or image reference, decision, confidence and review status;
reinspection, adjudication, rework and final disposition as separate events.
This extends the minimum data model for HPDC and LPDC. It prevents a late decision from silently overwriting the evidence needed to explain disagreement.
Eight steps to qualify casting quality labels for AI
Define the decision precisely. State the characteristic, unit, defect boundary, acceptance rule and intended model action. “Porosity” is not precise enough if the real decision concerns leak failure, machined-surface exposure or a location-specific X-ray class.
Build a representative study set. Include normal variation, known defect modes and borderline parts from relevant cavities, tools, alloys, shifts and process states. A study made only from obvious good and obvious scrap parts will hide the difficult decisions the model must later face.
Blind and repeat the inspections. Randomize part order and hide previous results. Have the same inspector or system repeat the decision, then repeat across inspectors, devices or programs. Keep handling and conditioning representative of production.
Use metrics that match the output. For continuous measurements examine resolution, bias, repeatability, reproducibility, linearity and stability. For categorical labels examine per-class repeat agreement, inter-inspector agreement, confusion patterns, unknown rates and class balance. A single overall accuracy can conceal failure on a rare, safety-relevant defect; kappa-like statistics also need context when classes are highly imbalanced.
Investigate disagreements physically. Do not decide every dispute by majority vote. Use an appropriate reference method—such as calibrated dimensional measurement, controlled leak testing, CT, sectioning, metallography or an expert panel—when the characteristic and risk justify it. Preserve “uncertain” when no defensible reference decision is available.
Correct the measurement process before the model. Improve defect definitions, lighting, cleaning, fixturing, calibration, reference images, operator training or automated-inspection thresholds. Re-run the study after material changes. Model tuning cannot repair a label process whose decision boundary moves by operator or shift.
Release labels through a training-data gate. Mark labels as verified, provisional, conflicting or excluded. Train and test on a frozen taxonomy and rule revision. Keep an independently reviewed reference set, and report performance by defect type, cavity, tool, shift and operating state—not only as one global score.
Monitor label drift in production. Track disagreement, unknown and override rates over time. A new inspector, camera, X-ray program, leak fixture, drawing revision or customer rule can change the target even when the casting process is stable. Route those changes through the model-drift and retraining gate.
Modelled example: do not optimize the classifier yet
Consider a modelled illustration, not a Castella customer result: 240 HPDC castings are selected across two cavities and include clear and borderline surface conditions. Three inspectors evaluate the randomized set twice using the current work instruction. Agreement is strong for obvious blisters but weak where cold-flow indications overlap cosmetic flow marks. A reference review finds that lighting angle, cleaning state and two competing defect examples are shifting the boundary.
The correct first action is not to compare more algorithms. It is to repair the inspection method, revise the examples, preserve uncertain cases and repeat the agreement study. Only then should the verified labels be joined to shot data and released for training.
What this evidence does—and does not—prove
Classical gauge R&R is directly suited to many continuous measurement tasks, but it should not be copied mechanically onto complex visual, radiographic or multi-feature decisions. ISO 22514-7 itself is primarily scoped to relatively simple one-dimensional measurement processes. Categorical inspection needs a designed agreement study, representative defect prevalence and, where possible, an independent reference method.
A capable label process does not guarantee a useful AI model. It only removes one major source of avoidable uncertainty. Feature coverage, causal relevance, time-based validation, class imbalance and deployment controls still matter. Likewise, disagreement is not always “operator error”; it may reveal an ambiguous specification or a physically continuous defect being forced into a binary class.
FAQ
Is 100% inspector agreement required before AI training?
Not automatically. The acceptance criterion should reflect the characteristic, decision risk and available reference method. Safety-critical or customer-conformance decisions justify stricter evidence than an advisory surface-screening model. The threshold and rationale should be documented before reviewing results.
Can model confidence compensate for unreliable labels?
No. Confidence describes the model under its learned data and assumptions; it does not convert inconsistent targets into physical truth. Label uncertainty should be recorded and used to exclude, quarantine or separately model ambiguous cases.
Will more training data overcome label noise?
More data can reduce sampling uncertainty, but repeated inconsistency can simply teach the model the inspection process’s bias. Improve the label process first, then collect more representative verified examples.
How should reworked parts be labelled?
Keep the original finding, rework action, reinspection result and final disposition as separate events. Choose the training target according to the model’s intended decision point; never overwrite the pre-rework evidence with the final status.
Sources
Castella’s position is deliberately conservative: verify the measurement process, preserve uncertainty and make the model earn authority on future production data.




Comments