Back to blog
Article

Why Variation Matters and How You Specify It

TLDR: Models fail on what training never varied. You fix that with a variation matrix you approve, recruitment planned to fill it, device coverage that matches deployment, and retakes logged as signal.

Your model meets accents, devices, lighting shifts, and cluttered backgrounds on day one. If training never saw that spread, you find out after deployment. That is a painful way to learn.

There is a simpler path. You state the variation you want up front, and collection is planned around it.

Models fail on what training never varied

A model that saw one microphone, one lighting setup, and one type of speaker will act surprised by everything else. That surprise shows up as missed words, missed detections, missed actions.

You cannot patch this with more of the same data. You need planned spread across the conditions your system will face in use.

The variation matrix you control

You define variation explicitly instead of hoping it appears in the pile.

A typical matrix looks like this:

DimensionWhat you specify
PeopleSpeaker or participant attributes relevant to your task
DevicesCamera, phone, microphone, and sensor types plus capture settings
LightControlled, low light, mixed, and natural regimes
CameraAngles, distances, placement, and handling
PlaceIndoor and outdoor environments in your scope

For image work, you cover all image categories in scope. For video, you specify HD to 4K and the viewpoints you need. For speech, you cover English, Hindi, Kannada, Telugu, Tamil, Malayalam, Marathi, and Sanskrit, plus more on request.

You approve the matrix before collection starts. That sheet becomes the plan everyone works from.

Recruitment to the matrix

Variation does not happen by chance. We recruit participants and schedule sessions to fill the combinations you approved.

If your matrix calls for particular speaker profiles, settings, or conditions, collection is planned around those cells. Maybe a session is all indoor Telugu utterances on two phone models. Another is outdoor Tamil at distance. Each one fills a gap on purpose.

You review coverage as collection proceeds. If a combination runs thin or runs heavy relative to the plan, you see it in tracking and can shift priority before delivery is locked.

Device diversity that reflects deployment

You ship on varied hardware, so your data should come from varied hardware.

We collect across cameras, phones, microphones, and sensor setups relevant to your product, including depth and LiDAR units and sensor and radar attachments where they apply. You state which devices are in scope and which settings matter, such as resolution, placement, handling, and background conditions.

Then your team can test a plain question. Does performance hold on the hardware your users carry or your robots mount.

Retakes logged as signal

Some takes fail the protocol. Framing drifts. Audio turns unclear. An action step goes missing. Light falls out of range.

When that happens we log the reason and capture again to spec. You get takes that meet the protocol, with metadata noting capture conditions for each file. You also get a record of what needed a retake and how it was resolved, with multimodal alignment kept intact across retakes where streams are grouped.

That log is useful. It shows which conditions were hard and where your model may need extra care.

What to specify first

Start small and concrete. Pick the people, places, light, camera positions, and devices that match your next eval. Write them into a one page matrix. We will recruit to it, track coverage against it, and log what was hard to capture.

Start with a paid pilot. Both sides check fit before scaling. Request a pilot or Talk to us.

A researcher studying an observatory at dusk

NEXT / YOUR SYSTEM

Find a clearer
way forward.

Whether you're an investor, a partner, or a builder — we'd love to hear from you.

Get in touch