How to Spec a Custom Video Dataset That Actually Trains
TLDR
- Resolution matters less than you think. Viewpoint, environment, and action coverage decide if your video trains or not.
- Posterior collects HD to 4K custom video with fixed cameras and first-person POV, indoors and outdoors, in controlled and natural settings.
- You define the viewpoints, backgrounds, lighting, and object interactions. We ship immutable raw files with per-file manifests and checksums.
- All collection is consent-first and delivered codes-only.
- Start paid pilots small, lock the spec, then scale. Request a pilot
Most video specs start with resolution. Ours start with the shot.
I have seen teams argue for a week about 1080p versus 4K, then leave viewpoint vague in one line. That order is backwards. A sharp clip from the wrong angle still teaches the model the wrong thing.
This is a practical guide to writing a video spec that holds up in training. If you want custom collection, Talk to us with this list filled in.
Resolution is the least interesting decision
HD to 4K covers almost every training need we see. Pick HD when you need volume and long durations. Pick 4K when you need small objects, distant text, or heavy cropping later.
What breaks models is coverage gaps. Same kitchen, same angle, same light, fifty times. Or all daytime clips and nothing at dusk. Resolution did not cause that. The shot list did.
So fix resolution early, then spend your time on viewpoint, environment, and action.
Fixed, POV, and mixed setups compared
Three setups cover most requests.
| Setup | Use it when | Watch out for |
|---|---|---|
| Fixed camera | You need repeatable angles for counting, tracking, or interaction with a zone | Tripod height and lens choice drift between sites unless you lock them |
| First-person POV | You need what a person saw while doing a task with their hands | Head motion blur and framing drift, so mounts and task instructions matter |
| Mixed indoor, outdoor, controlled, and natural | You need a model that works outside the lab | Backgrounds and light change fast, so you need explicit targets per environment |
Fixed gives you control. POV gives you realism. Mixed gives you range. Many buyers end up with a mix, maybe fixed for baseline clips plus POV for the hard cases.
We collect all three for Posterior buyers, and you can define viewpoints, backgrounds, lighting, and object interactions per batch.
What to put in the shot list
A good shot list fits on one page per scenario and leaves no guessing for the crew. Ours usually has these lines in plain language.
- Camera position and height plus lens angle, written so a second crew could redo it
- Who or what is in frame and what they do, step by step
- Background details that must stay, and ones that must change across takes
- Lighting condition for that clip, with time of day if natural
- Duration and number of repeats per variation
- What counts as a reject, stated before anyone shoots
That last line saves more money than any camera upgrade. If the crew knows a clip with a blocked hand or a missing object is a reject, you get fewer bad files and faster delivery.
A short example helps. Instead of writing outdoor pouring, write chest mount facing down at the cup, late afternoon sun from the left, pour from full kettle to empty cup three times, keep hands and spout in frame the whole time.
Lighting regimes worth specifying
Lighting is where natural video falls apart. Models learn the lab lighting, then fail in a dim hallway.
We ask buyers to pick regimes up front.
- Bright controlled light for clean baseline clips
- Low indoor light with practical lamps only
- Daylight, overcast, and direct sun as three separate targets
- Dusk or night where your use case needs it
You do not need all four every time. You need the ones your model will meet in use, with rough ratios. Something like 40 percent daylight, 30 percent indoor bright, 30 percent low light is enough to plan around.
A retakes policy that keeps data clean
Retakes feel wasteful until you train on a broken clip for a week.
Our rule is plain. If a clip misses the shot list, we reshoot it. If the manifest checksum fails, the file does not ship. If consent is incomplete for anyone visible, the clip never leaves storage.
That sounds strict. It keeps the dataset honest. Immutable raw files mean no silent fixes in post, so what you get is what the camera saw, plus a per-file manifest that proves it.
Paid pilots help here. Shoot 50 to 200 clips, review failures together, tighten the reject rules, then scale. It is cheaper to fix wording in a pilot than to reshoot 10,000 clips.
How Posterior ships video
- Custom HD to 4K video, fixed and POV, indoor and outdoor
- Buyer-defined viewpoints, backgrounds, lighting, and interactions
- Consent-first collection with codes-only delivery, no names attached to files
- Immutable raw files with per-file manifests and checksums
- Paid pilots to lock the spec before scale
Two related reads live alongside this one. One covers annotation choices for video, one covers privacy by framing, and one covers egocentric robotics. Each is a separate post so this guide can stay on collection.
FAQs
How much video do I need to start?
Enough to cover your variations once. For most pilots that is 50 to 200 clips across your main viewpoints and lights. Volume comes after the spec is stable.
Should I always shoot 4K?
Not always. Shoot 4K when detail is small or distant. Shoot HD when you need hours of action and motion matters more than pixels.
Can I mix fixed and POV in one dataset?
Yes, and many teams do. Keep the manifests clean so training can weight or split by viewpoint later.
What slows delivery most?
Vague reject rules and missing consent. Both are fixed in the pilot if you review early clips fast.
Write your shot list with camera position, action steps, and reject rules on one page. Send it over and we will tell you what will break before you pay for scale. Request a pilot
