A convincing synthetic image of a parts kit or a meal tray is no longer hard to produce. Ten thousand of them, each labelled accurately enough to train an inspection model, is a different problem. That second problem decides whether a computer-vision project reaches the line. It is also the part of a synthetic data pipeline that most buyers examine least.
Who this is for, and what “label” means here
This post is for engineering leads and quality owners working on assembly and kitting inspection. That means checking that a tray, kit or case holds the right items in the right places. In that setting, a label has two parts. The class label says what an object is and whether it is correct. The bounding box says where it is. A model learns from both, and an error in either becomes an error in the model.
Synthetic data here means training images generated by software rather than photographed on a line. The images are only useful if their labels are right. So the question this post asks is practical. When a pipeline hands over a labelled synthetic dataset, where did each label come from, and who or what checked it? The sections below take three questions in turn. Why does image generation alone no longer separate one pipeline from another? Why do the two parts of a label fail in different ways? And what should a team ask before trusting either?
Image generation is becoming a platform feature
The reason to focus on labels is that image generation is quickly becoming a standard platform capability. In 2026, the Taiwanese manufacturers Wistron and Inventec described using NVIDIA’s Defect Image Generation skill to create synthetic defect images for visual inspection. NVIDIA also publishes a ready-made defect-generation pipeline that produces labelled synthetic defect images, with masks and pseudo-labels for downstream detectors. Roboflow now packages the same NVIDIA tooling with labelling, training and evaluation.
This is good news for teams that need data. It also changes what a buyer should compare. When several tools can produce a realistic image of a scratched housing, realism alone stops being a useful test. Analysts tracking industrial vision already list annotation automation and pipeline compatibility alongside realism as the grounds on which vendors compete. The label has become the part worth examining, and it has two components that behave differently.
A label has two parts, and they fail in different ways
The class label and the bounding box come from different places, so they need different checks. Consider the class label first. If a pipeline specified the scene before generating it, the pipeline already knows what the scene contains. It knows how many units are in the frame, which are correct, and which defect each incorrect unit carries. In that design, the class label is a by-product of the specification, not a guess made by looking at pixels afterwards.
Synthetic Data Studio (SDS), by Aicadium, is built on that design. SDS produces each image and its annotations together, from the same specification. It supports two levels of detail. Whole-unit mode boxes each assembled unit and classes it correct or incorrect. Per-compartment mode boxes each compartment and assigns one of eight error types (complete, missing, substitution, arrangement, extra, portion, wrong_dish or occluded). The product calls this per-dish mode.
The bounding box is a different matter. Knowing that a tray holds a missing item does not tell you, to the pixel, where the compartment sits in the rendered image. SDS tightens boxes with general-purpose AI detectors, which find objects from a plain-text description of what to look for. Any pipeline that draws boxes this way shares the same strengths and weaknesses. These models are capable, and they are also fallible. In one 2026 study of synthetic surface defects, the researchers found that SAM 3, a segmentation model, could not reliably outline scratches. The researchers fell back on a simpler method, which the authors say can introduce minor errors in some labels.
That finding applies to the whole category, not to one vendor. Any pipeline that draws boxes with a general detector inherits the same kind of risk. The useful question is therefore how a pipeline checks the boxes it draws.
Why the checking step cannot be skipped
A detector that is usually right will still be wrong on some frames, and synthetic datasets are large enough for those frames to add up. The industry offers little help in measuring this. Mordor Intelligence lists the lack of standard quality metrics across vendors as a restraint on the synthetic data market. Without a shared metric, a vendor’s claim that its labels are accurate is difficult to compare with anyone else’s.
The problem is sharper for rare defects. Inventec notes that as manufacturing yields improve, real defect samples become harder to collect. Rare defects are exactly where synthetic data is most useful. They are also where there are fewest real examples against which to check a generated label.
So a credible pipeline needs checks at more than one level. SDS uses an automated assessment pass to verify labels against the rendered image. When Aicadium delivers a dataset, a person reviews every annotated frame before it ships. Teams running SDS themselves should keep the same step. Automated checks are an aid. The standard worth holding any pipeline to is that a person has looked at what the automation produced.
A useful review also leaves a record. For each dataset, a team should be able to see the plan each frame was generated from and which frames the automated pass flagged. It should also see what a reviewer changed or rejected. That record matters later. When a trained model makes an unexpected error on the line, the first question is whether the training labels were wrong. A team with a review record can answer that in an afternoon. A team without one has to relabel a sample and hope it finds the problem.
Four questions to put to any labelling pipeline
Those checks become concrete when they are framed as questions a team can ask before accepting a dataset.
Where do the class labels come from?
Ask whether classes are taken from the scene specification or worked out afterwards from the finished image. Specification-derived classes can be traced back to a written plan that someone approved.
What draws the boxes, and where does it struggle?
Ask which detector tightens the boxes and which object types it handles poorly. A supplier that knows its failure cases is easier to trust than one that claims to have none.
What is checked automatically, and what does a person check?
Ask for the split. An automated pass that flags frames is useful. A named human review step is what catches the errors automation misses.
Is performance measured on real images?
A dataset’s labels are ultimately judged by how the trained model performs on real frames from the line. That is the subject of the next post in this series.
Where this leaves teams building inspection models
Image generation will keep getting better and cheaper, and much of it will arrive as a feature inside larger platforms. For teams building assembly and kitting inspection, the lasting advantage sits one step further along. It lies in labels whose origin is written down and whose accuracy someone has checked. Ask where each label came from, what could have got it wrong, and who looked.
That leaves the other half of the standard, which is keeping real images for measuring what a model has learned. The next post in this series, Where real data earns its place, sets out how.

