A catering operation decides to check trays with a camera. The model is chosen quickly, and a demonstration works. Then the pilot slows down. Someone has to photograph every dish, every tray layout and every kind of mistake, then label them all. The deadline arrives before the dataset does. The model was never the problem. The data plan was.
What a data plan is, and why food service is different
A data plan answers three questions. What images will the model learn from? Who labels them, and how? How will results be measured on the real line? A pilot that cannot answer all three has a demonstration, not a path to production.
Food service makes these questions harder than they first appear. A tray standard can change with every menu cycle. Plating varies from one shift to the next, and some of that variation is perfectly acceptable. Mistakes such as a missing dessert or a swapped main may be rare in any given week, which makes them slow to capture. Client terms can also limit how service imagery may be used. This post looks at why data plans stall and how starting from the tray specification helps. It ends with what to ask any supplier before a pilot is scaled.
The data plan is where pilots stall
Stalling on data is a common problem well beyond food service. In Gartner research, 63% of organisations were not sure they had the right data management practices for AI, or said they did not. In 2025, Gartner also predicted that organisations would abandon 60% of AI projects unsupported by AI-ready data through 2026.
For a tray inspection pilot, “AI-ready data” has a specific meaning. It is a labelled set of images that covers the dishes, layouts and errors the line will actually see. Collecting that set from live service takes time the pilot rarely has. Labelling it by hand adds more. The result is a pilot that proves the idea and then waits for data that never quite arrives.
Start from the specification you already have
Food-service operations already hold something that computer-vision projects often have to create from scratch. They have a written standard for what a correct unit looks like. Every tray, meal box or kit is meant to match a specification. That specification can become the starting point for the training data, instead of relying on live capture alone.
Synthetic Data Studio (SDS), by Aicadium, works this way. A project starts from a photograph of one correctly assembled unit, a written description, or both. An AI model describes what it sees, and the user confirms that reading before a dataset brief is written. That confirmation step is the moment to correct a misread item before it reaches every image in the dataset. A dessert read as a side dish, for example, takes seconds to fix at this stage. Left uncorrected, it would appear in every image generated from that brief.
SDS can then label each compartment individually, using eight error types (complete, missing, substitution, arrangement, extra, portion, wrong_dish and occluded). Those categories can be matched to the checks in your own tray standard. A portion issue is a different finding from a wrong dish, and a model trained on the distinction can report it.
When the menu changes
That link between specification and training data matters most when the specification changes. In food service, it changes often. A new menu cycle brings new dishes, new tray layouts and new ways for a tray to be wrong. A model trained on last season’s photographs has never seen any of them.
A dataset built from photographs of live service has to be rebuilt the slow way. The team captures the new trays, waits for enough mistakes and labels everything again. A dataset built from a specification can follow the specification. The team photographs one correct example of the new tray, confirms the reading and generates a fresh training set from a new brief.
This does not remove the need for real trays. The new menu still needs a held-back set of real trays to confirm that the updated model works. What changes is the starting point. The training data begins from the standard the quality team signs off, not from whatever service happened to produce.
For a quality director, that tie is the point. When a client asks how the inspection model knows what a correct tray looks like, the answer is the same document the operation already uses.
Measure what the line actually cares about
A tray pilot should be judged on the errors that matter to the operation, not on a single headline accuracy figure. That means reporting results per error type. A model that catches missing items reliably but misses substitutions has a specific gap, and an average hides it.
It also means measuring false alarms. Plating varies, and a tray with a slightly generous portion may still meet the standard. A model that flags acceptable variation as an error will quickly lose the trust of the people on the line. Count those false alarms as carefully as the misses.
Finally, the test set should reflect real operating conditions. Hold back real trays from different shifts, stations and menu cycles. A model that works on one shift’s trays under one set of lights has not yet been tested on the line.
Ask how results were measured
Starting from the specification solves the supply problem. It does not answer the question procurement will ask next, which is whether it works. That question deserves care, because published figures in this market often come from the vendors themselves.
Wistron, for example, reports that its synthetic defect pipeline improved detection accuracy by 5 to 20%. Roboflow says its pipeline compresses a multi-quarter inspection project into a few days. These may be accurate, but they are vendor claims, not independent measurements. Analysts also note the lack of standard quality metrics across vendors, which makes figures hard to compare.
A pilot sponsor can close that gap with three questions.
-
Was accuracy measured on real trays from our line, or on generated images?
-
Which error types were tested, and how did each perform?
-
Can we rerun the measurement ourselves on a set of real trays we hold back?
A supplier who welcomes the third question is one worth piloting with.
Keep people and real trays in the loop
Those questions point to the arrangement that holds up in production. Real service imagery is kept for evaluation, as a held-back set of real trays that measures whether the model works. Synthetic data carries the training load, including the rare mistakes and the acceptable variation that live service produces too slowly.
People stay in the loop as well. A quality director should be able to see how each label was produced and who reviewed it. When Aicadium delivers a dataset, a person reviews every annotated frame before it ships. Teams generating their own datasets in SDS should build in the same check. Automation carries the volume. Human review is what makes the output defensible to a client.
From pilot to production
A tray inspection pilot reaches production when its data plan is as solid as its model. Start from the specification the operation already has. Keep a set of real trays for honest measurement. Ask every supplier how its results were measured, and on what.
This is the third post in a series. Generating the image is easy. Checking the label is the work. looks at how synthetic labels are produced and checked. Where real data earns its place sets out how to divide real and synthetic data. For how Synthetic Data Studio supports food-service and assembly inspection, see the SDS page.

