Back to blog
Article

How We Collect Training Data to Your Spec

TLDR: Send us your brief. We turn it into a written protocol. You approve it before we record a single file. Then we collect, annotate, and ship data that matches your deployment, with consent records and per-file checksums to prove it.

Off the shelf sets rarely match product reality. Your cameras sit in different places. Your users speak in different ways. Your edge cases live in places no public set covers. That is why we collect to spec.

You approve the protocol before we record anything

You start with your task and your constraints. The brief is kinda messy at first, that is normal. We turn it into a written protocol that names framing and lighting expectations. It names audio conditions and file formats. It names consent terms and annotation targets.

Collectors work from that same document. Reviewers check against that same document. If your spec changes, we update the protocol first, then collection restarts from the new version. Nothing gets recorded on verbal instructions.

You defineWe lock in the protocol
Task and environmentsScene list, indoor and outdoor sites, controlled and natural conditions
Devices and viewpointsFixed setups and first-person views, HD to 4K, sensor pairing where needed
Speech and text needsScripts, speaking style, recording setting, domain and style guides
Output and useAnnotation schema, QA sampling, manifest and checksum format

Video and sensor coverage for where you will run

You tell us where your system has to work. We recruit and schedule to that plan so coverage reflects deployment, not convenience.

For video you can request HD to 4K across fixed camera setups and first-person views. You can cover indoor and outdoor settings. You define viewpoints and backgrounds. You define lighting regimes and object interactions.

When your task needs more than RGB, we pair video with depth and LiDAR. We add inertial sensors and radar where the device stack calls for it. Multimodal takes are time aligned at capture so you do not have to resync in post.

Speech in eight languages plus more

For speech you get coverage across English, Hindi, Kannada, Telugu, Tamil, Malayalam, Marathi, and Sanskrit, plus more on request. You supply scripts or prompts. You set speaking style with recording setting and microphone placement.

For text you set domain and style with task instructions. Contributors produce or transcribe to those instructions. Every file stays tied to its prompt and its recording conditions. You can trace what was said, how it was elicited, and where it was recorded.

You receive data collected under signed consent obtained before capture. Identities stay in a private registry. You receive codes only, never personal details.

Raw files are preserved unchanged after capture. If a file needs correction or replacement, you get a new version with a new checksum. We never silently overwrite. Your team can audit what was captured, when it was captured, and under what consent code.

What lands in your delivery

  • Video in HD to 4K from the viewpoints and environments you named
  • Speech across the eight languages above, plus more on request, with prompt level traceability
  • Depth, LiDAR, inertial, radar, and paired multimodal takes where your task needs them
  • Annotation to your schema with QA sampling you can inspect
  • Per-file manifests with checksums, consent codes, and capture conditions

You spend less time cleaning. You spend more time training on data that looks like the field.

Common questions

Can we change the spec midstream? Yes. We revise the protocol, you reapprove, and new captures follow the new version. Old files keep their old protocol tag so mixes never confuse training.

Do you replace public data? Sometimes we fill gaps around it. Most teams bring one public set that is close but thin on their scenes or speakers. We collect the missing slice to the same annotation schema.

Who owns consent liability? We do the paperwork before capture and store the signed forms. You get codes and manifests you can show auditors, without ever holding personal details.

Start with a paid pilot

We like pilots because both sides learn fast. Send two pages on your scenes, devices, speakers, and annotation target. We will return a protocol you can approve this week.

Request a pilot and Talk to us.

A researcher studying an observatory at dusk

NEXT / YOUR SYSTEM

Find a clearer
way forward.

Whether you're an investor, a partner, or a builder — we'd love to hear from you.

Get in touch