Custom Image Collection to Spec Beats Scraping the Web
TLDR: Scraped photos carry unknown licenses, unknown capture settings, and repeats you cannot measure. Specified stills give you known provenance plus coverage matched to deployment. This post shows how we collect custom images to spec.
If you already know your categories, Request a pilot and we will scope collection around them.
Why scraped stills fail in production
Scraped sets look large. Performance in the field tells a different story.
Licenses are murky. Capture conditions are missing. Duplicates pile up across sites that copy each other. A model trained that way learns the web, not your store aisle, your factory line, your clinic desk, your farm crate.
Specified collection flips the order. You state deployment first. We collect for it. Phones people carry, DSLRs on tripods, endoscopic probes and bench rigs where needed. Morning sun, tube light flicker, dim godowns, night flash. Close crops, mid range views, wide room context. Clean sweeps, busy shelves, oily benches, wet counters.
I get kinda excited about this part because the pipe matters as much as the water. Good pipes, clean flow.
How to define image categories that models learn from
Categories fail when they are vague. “Defect” means nothing by itself. “Crack under 2 cm on glazed tile in daylight” trains.
Use this shape for each category:
| Field | What to write | Example |
|---|---|---|
| Object plus state | Name the thing and its condition | Rust spot on mild steel coupon, unpainted |
| Setting | Where it sits in real use | Outdoor rack, factory yard, monsoon air |
| Capture spread | Devices, light, angle, distance | Phone and DSLR, sun plus shade, front and 45 degree, 30 cm and 1 m |
| Count plus split | How many, divided how | 2,000 frames, split across 4 sites and 6 phones |
| Exclusions | What does not count | Painted coupons, oiled surfaces, video grabs |
Four habits keep the list honest:
- Name states, not nouns. Bruised mango, overripe mango, green mango beat fruit.
- Tie each category to a decision. Accept, reject, reroute, escalate.
- Add buyer edge cases early. Rare SKUs, rare faults, rare papers.
- Freeze the wording before collection starts. Version any change after.
Maybe start with 8 to 12 categories. Enough to cover deployment, small enough to check by hand.
Negative and background coverage earns its keep
Positives get all the attention. Negatives decide precision.
We collect backgrounds and lookalikes on purpose. Empty shelves, clean steel, healthy leaves, blank forms, good cartons, intact seals. Plus near neighbors that often confuse models. Brown paper beside kraft parcels, shadows that look like cracks, printed tile veins that look like scratches, food garnish that looks like foreign material.
A short rule I use. For every positive setting, collect the same setting without the target, in the same light, with the same devices. That pair teaches the boundary.
Near duplicate control keeps counts honest
Burst mode and shared folders inflate datasets fast. Ten frames from one second are not ten samples.
Our controls are plain and strict:
- Cap frames per subject, per angle, per setup. Move the object or move the camera before more frames count.
- Spread subjects across places, days, hands, devices. Same mango shot by five people counts more than five shots by one person.
- Hash on ingest to catch reuploads, resized copies, crops passed off as new.
- Hold back a site or a device lot for testing. If scores drop there, coverage was thin.
Counts mean little without this. A thousand true variations beat ten thousand near copies every time.
When stills beat video frames
Video has its own place. Our sibling post covers it. For many tasks, stills win.
Stills win for high resolution inspection, where a 48 MP phone shot or a DSLR raw shows texture that compressed video smears. They win for documents, labels, meters, and packaging, where glare control and focus matter more than motion. They win for rare faults, where you can stage lighting and angles with care instead of hoping a frame catches it.
Video frames win for action, tracking, and dwell. If motion defines the label, use motion. If the label sits still and detail decides it, use a still camera and take the time to get it right.
What delivery looks like
No surprises at handoff. Every batch ships with immutable raws, a checksum manifest, and a spec sheet that maps each file to category, device, light, angle, background, site code, and consent record.
Faces appear only where consented and in scope. Otherwise delivery is codes only, no names, no addresses, no extra metadata riding along. Consent forms stay on file and out of the training folder.
We run paid pilots first. Small paid lot, full checks, then scale. You see quality before you commit to volume.
Related work lives in sibling posts on video collection, annotation, and provenance. This post stays on stills.
Faqs
How many images per category do we need to start?
Start with 500 to 2,000 true variations per category for a pilot, spread across devices, lights, sites. Scale after error review shows where misses cluster.
Can you match our exact phones and lights?
Yes, within reason. Send device models, OS camera defaults, light types, mounting heights. We add your mix to the collection plan and log it per file.
Do you clean and dedupe before delivery?
We hash for exact and near copies, cap bursts, and flag thin lots. You get the manifest plus the reject log, so counts stay auditable.
What about consent and faces?
Consent comes first. Where faces or private spaces fall in scope, we collect signed consent and deliver codes only. Where they fall out of scope, we exclude them at capture and again at review.
Send three categories, one deployment setting, and your device mix. We will return a collection plan with counts, sites, and pilot price. Talk to us or Request a pilot.
