Multimodal capture
Collection is designed for egocentric observation of physical work. Depending on the capture hardware used for a campaign, recordings can include first-person video, IMU, camera pose, depth, and optional audio. Not every recording includes every modality; availability is part of the dataset metadata.
MCAP
Raw synchronized sensor recordings can be delivered in MCAP so downstream robotics tools can read camera, inertial, pose and metadata streams from a common log. The aim is to preserve the original capture rather than forcing every buyer into one reduced training format.
Synchronization
Sensor streams are mapped to a common timebase. Episode boundaries, annotations and quality flags refer to that same timeline so a buyer can reconstruct what was observed at a given moment.
Episode segmentation
People often record longer stretches of real work, including failure-response takeovers in a robot’s operating environment. Those recordings are segmented into task-level episodes that represent meaningful physical interactions — the unit most robot-learning workflows actually train on.
Metadata
Structured fields describe the task, environment, collection context, sensor availability, hand visibility where annotated, and consent status. Metadata is what makes a corpus searchable instead of a pile of files.
Hand pose estimation
Nuuduu can derive 3D hand pose from egocentric video using WiLoR and reconstruct detailed hand meshes with MANO via PAD-Hand. This provides per-frame hand joint positions and orientations without requiring gloves or markers — useful for training manipulation policies from human demonstrations.
Quality control
Recordings and episodes can be evaluated for visual and sensor quality before delivery. The goal is not cosmetic polish; it is to let buyers reason about whether a demonstration meets the requirements of a given training run.