How we collect and clean every photo
From the moment a photo is taken to the moment it lands in your training set: capture, de-duplication, blurring, GPS removal, review and documentation.
By Dataset By Humans ·
Every dataset we deliver goes through the same pipeline. This post explains each step, why it exists and what it means for the data you train on.
1. Capture: a person, a phone, a real place
Every image starts as an original photo taken with a smartphone by one of our collectors. We do not download images, scrape websites, buy stock photos or use generative AI to create or alter image content.
Phones are deliberate. They are the camera most of your users carry, so their noise, lens distortion, HDR processing and auto-exposure are exactly what production models need to handle.
Before shooting we check where photography is allowed. We do not photograph military areas, airports or government facilities, and we ask permission in shops, restaurants, homes and farms.
2. De-duplication
Collectors often take several shots of the same scene. Near-identical images inflate a dataset without adding information, and they can leak between your training and test splits. We remove exact duplicates with file hashes and near-duplicates with perceptual hashing.
3. Quality filtering
Out-of-focus, accidentally triggered or corrupted images are removed. Some datasets want blur or noise, such as our low-light topic. In those cases imperfection is labeled, not filtered.
4. Privacy: blur first, then look
Photos of public life inevitably include people and vehicles. Before anything else happens:
- Faces are detected and blurred.
- License plates are detected and blurred.
- Documents, screens and phone numbers are blurred where they appear.
- Children are excluded entirely.
Automatic detection misses things, so every image is then reviewed by a person. Anything that was missed is blurred by hand, or the image is dropped.
5. Metadata: keep what helps, remove what harms
EXIF data is useful for training: device model, ISO, shutter speed and focal length tell you how the image was formed. We keep it.
GPS coordinates can reveal where someone lives, so we remove them. Where location matters, we record it at city or province level only.
6. Labels and captions
Depending on the order, images get topic labels, human-written captions or bounding boxes. Captions are written by people. If AI tools are ever used to assist with drafting labels, the datasheet says so explicitly.
7. Packaging and documentation
Each delivery contains:
dataset-name-v1/
├── images/
├── metadata.jsonl
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256
The datasheet documents how the data was collected, by whom (in aggregate), where, how it was processed and what its known limitations are. For example, it states that images from a small team of collectors may share stylistic biases. If you are a provider of a general-purpose AI model under the EU AI Act, this documentation helps with your training-data summary.
8. Takedown
If someone recognises themselves in an image despite our processing, they can ask us to remove it through our takedown process. We remove the image, publish a new dataset version and notify customers who licensed it.