Back to blog
Article

Robot Action Data: How Teleoperation Demonstrations Train Manipulation Policies

TLDR: Policies learn to move from demonstrations recorded on the same embodiment they will run on. Each episode pairs multi-view video with timestamp synced joint states, end effector poses, gripper actions and proprioception, plus language labels for tasks and subtasks. Successes teach the motion. Failures and human takeovers teach the limits.

Want this on your hardware? Request a pilot

Why manipulation data takes most of the budget

Manipulation is 60 to 70% of robot data demand. That matches what buyers ask us for every week. Pick and place. Tool use. Assembly. Bimanual handling. Deformables.

Open sets helped the field agree on formats. DROID published 76K episodes. RoboMIND published 107K trajectories. HIW-500 published 500 hours of humanoid whole body motion. They are useful for pretraining.

They rarely match your arm, your gripper, your camera angles, your objects, your cycle time. A policy trained on a different embodiment needs retargeting and still misses contact details. That gap is where projects stall.

Most teams cannot stand up teleop stations plus operator training plus QA plus trajectory pipelines in house. It is a separate job from building the policy. Posterior runs collection as a service on the buyer target platform, from tabletop arms to mobile manipulators and humanoids. Our hardware partnerships are expanding, so tell us what you run and we scope around it.

What a usable teleop episode contains

A video clip alone does not train control. The policy needs to see what the human saw and feel what the robot felt, at the same clock.

SignalHow we record itWhat it trains
Multi-view videoTwo to four fixed views plus wrist view, all timestamp syncedVisual grounding for objects, occlusions, contact
Joint states and end effector posesStreamed at the controller rate with hardware timestampsTrajectory targets for imitation and diffusion policies
Gripper actions with proprioceptionWidth, force proxies, slip cues where availableGrasp timing, release timing, retry logic
Language labelsTask name plus subtask segments with start and end timesInstruction following and subtask scoring

Short burst. Then the longer run that matters. When those four streams share one clock, you can replay the exact moment the wrist rotated, the fingers closed, the object shifted, and the operator corrected. Without that sync, you have pictures of work instead of records of work.

Collection uses leader follower arms for contact heavy tasks and VR teleop where reach and speed matter. Operators train on your task list before we count an episode. Kinda like a driving test before we hand over the steering wheel.

Why failures and takeovers matter as much as successes

Successes show the path. Failures show the edge.

We deliver success episodes for behavior cloning and diffusion. We also deliver failure episodes as negative signal, with labels for what went wrong and when. Dropped grasp. Missed insertion. Object toppled during handoff. Those labels let teams filter, weight, or contrast during training instead of guessing why eval dipped.

Human takeover segments sit in between. The policy or script runs until it struggles, a human grabs control through teleop, completes the subtask, hands back control. Each segment keeps full state and video around the switch. Teams use these clips to patch specific subtasks without recollecting whole tasks.

Demo or die. If the data cannot show a recovery, the policy will not learn one.

Delivery formats your training code can read

We deliver in LeRobot or HDF5, with consistent schemas across tasks:

  • Video frames linked to state rows by timestamp, not by filename guesswork
  • Proprioception, actions, rewards and language labels in the same episode group
  • Train and eval splits separated by object set, scene layout and operator
  • QA reports with drop rates, sync drift checks and relabel notes

Need a different split or a different frame rate? Say so in the request. We cut the set to match your loader.

Sister posts cover sensor depth and egocentric video. This post stays on action. If you need those streams synced with teleop, we add them to the same episodes instead of building a second set.

FAQs

How many demonstrations does one task need?

It depends on variation. A fixed pick from a fixed bin can converge with a few hundred successes. Add new objects, new lighting, new clutter and the count rises into the thousands. We usually start with a pilot of 500 to 2,000 episodes per task family, measure eval, then scale the long tail.

Can you collect on our robot and our objects?

Yes. That is the point. Send the platform, the tooling, the objects and the task list. We bring operators, capture rigs and QA. If your hardware needs a new mount or driver tweak, we flag it during scoping.

Do you annotate language?

Every episode carries a task label. Longer tasks carry subtask labels with time ranges. Example labels read like “pick red mug” or “insert peg with left arm while right arm holds base.” Your instruction tuning reads those strings directly.

What about privacy and IP?

Collection happens under your task spec. Footage, trajectories and labels belong to you on delivery. We keep no separate copy for other buyers unless you permit it in writing.

How do we start?

Send three tasks your policy keeps failing plus a short video of the cell. We reply with a collection plan, episode counts and timeline. Request a pilot or Talk to us

A researcher studying an observatory at dusk

NEXT / YOUR SYSTEM

Find a clearer
way forward.

Whether you're an investor, a partner, or a builder — we'd love to hear from you.

Get in touch