Why human-captured data still matters in the age of synthetic data
Synthetic images are cheap and plentiful. Here is why real photos taken by people remain essential for models that must work in the real world.
By Dataset By Humans ·
Synthetic images are cheap, fast and infinitely scalable. So why would anyone pay people to walk around taking photos with their phones?
Because the real world is the test set.
Synthetic data learns from itself
Generative models produce images by sampling from what they have already learned. Research published in Nature in 2024 (Shumailov et al., "AI models collapse when trained on recursively generated data") showed that models trained repeatedly on generated data progressively lose the rare cases, the "tails" of the distribution, and drift away from reality.
The rare cases are exactly what breaks production systems: the motorbike carrying a refrigerator, the menu written by hand, the receipt folded in half under a fluorescent light.
Real photos carry real noise
A photo taken on a phone carries the fingerprints of the real capture pipeline:
- sensor noise that depends on ISO and temperature
- lens distortion and chromatic aberration
- the phone's own HDR, sharpening and noise reduction
- motion blur from a moving hand or a moving subject
- clutter, occlusion and messy backgrounds
Simulators and generators approximate some of this, but your model will meet the real thing in production. Training and evaluating on real captures closes that gap.
Coverage of underrepresented places
Large public datasets are dominated by a handful of regions. Traffic datasets rarely show dense motorbike traffic; OCR datasets rarely include Vietnamese diacritics; food datasets rarely include the dishes most of Southeast Asia eats every day. You cannot synthesise your way into a distribution you have never observed.
Provenance you can document
Regulation is catching up with training data. The EU AI Act requires providers of general-purpose AI models to publish a summary of the content used for training. Human-captured data with a clear chain of rights, privacy processing and a datasheet is far easier to document than data of unknown origin.
Synthetic and real work best together
None of this means synthetic data is useless. It is excellent for augmentation, for rare classes you can describe precisely and for pre-training at scale. The strongest pipelines combine both: synthetic data for volume, and real, human-captured data to anchor the model to reality and to evaluate it honestly.