Real production images are the most valuable data an inspection team has. They are also scarce, slow to collect and often restricted by customer or confidentiality terms. Synthetic data promises to relieve that pressure, and it can. As major platforms make synthetic images easy to produce, more teams will be tempted to train on them alone. The harder question is what the real images should then be used for. The evidence points to a clear answer. Keep them for the exam.
The terms, and the question this post answers
A computer-vision model is built with three kinds of data. Training data is what the model learns from. Validation data guides decisions during development. Test data is held back to measure how the finished model performs on images it has never seen. For an inspection model, the test that matters is performance on real frames from the production cameras.
Synthetic data is software-generated training imagery. It can supply volume, rare defects and controlled variation that a real line may take months to produce. This post asks how a team should divide its data when synthetic images are available. It looks at what independent research says about models trained only on synthetic data, and why the gap appears. It then sets out a division of labour that keeps results honest. The argument applies to any inspection task where correct and incorrect units can be defined, from electronics assembly to meal trays.
What happens when a model sees only synthetic data
The strongest evidence comes from studies that measure synthetic-trained models on real images. Researchers at Hitachi America trained a detector for an industrial inspection task on 550 synthetic images. According to their 2026 paper, the model was near-perfect in simulation. On real images it scored only 0.2516 mAP (mean average precision, a standard detector accuracy measure from 0 to 1). The researchers then used an adaptation method that combined 50 paired real images with 500 synthetic ones. That lifted real-image mAP to 0.8853.
Other work reaches a compatible conclusion by a different route. A 2026 study of synthetic surface defects found that synthetic data was most valuable when it strengthened scarce real datasets. It did not work as a substitute for them. A practitioner article published by the Edge AI and Vision Alliance makes the same point more bluntly, saying that teams still need real-world testing.
The finding is narrower than it may sound. A synthetic-only model can look excellent on synthetic images and still disappoint on the line. The only way to see that gap is to measure on real frames.
Why the gap appears
The gap between simulated and real performance has identifiable causes, which makes it manageable. Real cameras introduce lighting, reflections, sensor noise and wear that a generator may not reproduce. The same surface-defect study tested a general-purpose image model. It did not reliably reproduce what industrial defects look like unless it was first adjusted to industrial images. Rewording the instructions to the model was not enough on its own.
Three-dimensional rendering pipelines have a different trade-off. They give precise control and automatic labels, but the same study notes that they require detailed scene models and substantial manual effort. Every generation method, in other words, carries some distance from the real line. The practical response is to keep a reliable measure of that distance.
A division of labour between curriculum and exam
That measure is a held-back set of real images, and it is the most important use of real data a team has. The division of labour follows from the evidence. Synthetic data becomes the training material. It supplies volume, rare defects and variation in position, lighting and arrangement. A small, carefully chosen set of real images may join the training material too, as the Hitachi America study did. Real data becomes the exam. A separate set of real frames, captured at the production cameras and never used in training, measures what the model has actually learned.
This arrangement has a second benefit. Customer or confidentiality terms sometimes restrict how production imagery may be used. In those cases, a team may still be able to use a small real set for evaluation while synthetic data carries the training load. Whether that is permitted depends on the terms in each case, and it should be confirmed before any plan relies on it.
The division also changes how a team spends its effort. Real-data collection shifts from “gather as much as possible” to “gather a representative test set and keep it clean”. That is a smaller, more achievable target, and it protects the one measurement that matters.
How to run the comparison
Setting that measurement up is straightforward, and doing it early prevents surprises later. Build the real test set first, from the cameras and conditions the model will face. Then train comparable models on real data only, synthetic data only, and a mix. Report all three on the same real test set, never on synthetic images.
This comparison answers the questions a quality owner will be asked. Does synthetic data improve results on the line? By how much, and for which defect types? Where does the model still fail? The answers then point the next round of synthetic data at the kinds of errors the real test set exposed.
Synthetic Data Studio (SDS), by Aicadium, generates labelled training images and exports them in standard label formats (COCO JSON, YOLO and Pascal VOC). Those are formats existing training pipelines already read, so a team can run this comparison with the tools it already uses. SDS supplies the training data. Measuring the trained model on real images remains the team’s own test, and it should.
Common questions about real-data evaluation
The same practical questions come up whenever teams adopt this approach, so they are worth answering directly.
How many real images does the test set need? There is no universal number. It depends on how many product variants, defect types and camera conditions the model must handle. In the Hitachi America study, 50 carefully selected real images, used with an adaptation method, changed the picture sharply in one task. That is evidence that a small real set is valuable, not a target to copy.
What if we have no real images yet? Then the first job is to collect a small, representative set before training anything. Without it, there is no way to tell whether a synthetic-trained model will work on the line.
Should real images ever be used for training? Where terms allow, a small, carefully chosen set of real images can help training, as the Hitachi America result shows. The rule is to keep a separate real test set that never enters training, whatever else the team does.
How often should the test set be refreshed? Whenever the line changes in a way the model will see, such as a new product variant, a new camera or a change in lighting. A test set that no longer matches production measures the wrong thing.
Keeping the exam honest
A synthetic data strategy is only as credible as the real-data test behind it. Keep the real images separate, measure on them every time, and report results only from them. That discipline turns synthetic data from a promise into something a quality director can defend. It gives a clear answer when a customer or auditor asks how the model was validated.
The first post in this series, Generating the image is easy. Checking the label is the work., covers the other half of the standard, which is labels whose origin and checks are clear. The third, Why computer-vision pilots stall before the model does, looks at food-service inspection pilots. For how Synthetic Data Studio fits, see the SDS page.


