How we collect and clean every photo
- 1
Capture
Original smartphone photos taken by people in real places. No scraping, no stock, no generative AI. We only shoot where photography is permitted and ask permission in private spaces.
- 2
De-duplicate
Exact and near-duplicate images are removed so they do not inflate the dataset or leak across your train/test split.
- 3
Filter
Broken, corrupted and accidental shots are removed, unless imperfection is the point of the dataset, in which case it is labeled.
- 4
Blur
Faces, license plates, documents, screens and phone numbers are detected and blurred. Children are excluded.
- 5
Review
Every image is checked by a person. Missed details are blurred by hand or the image is dropped.
- 6
Clean metadata
GPS is removed; device and exposure EXIF is kept because it helps training. Location is recorded at city or province level only.
- 7
Label
Topic labels, human-written captions or bounding boxes, depending on your order. Any AI assistance in labeling is disclosed in the datasheet.
- 8
Document & deliver
Images, metadata.jsonl, datasheet, license and checksums, delivered privately. Each customer copy is uniquely fingerprinted to deter leaks.
Our commitments
- 100% human-captured images
- No web scraping or stock imagery
- No generative AI used to create or alter image content
- Faces and license plates blurred, every image human-reviewed
- GPS removed from every file
- A datasheet with every dataset, including known limitations
- A takedown process for anyone who recognises themselves
Honest about limitations
Images captured by a small team of collectors share devices, habits and locations. That can introduce bias in device characteristics, framing and geography. Every datasheet states these limitations so you can account for them, and custom collections can be designed to widen coverage.
Read the detailed write-up →