SynData: A Large-Scale Real-World Multimodal Dataset for Embodied Intelligence
- Type
- dataset
- Venue
- PsiBot / Hugging Face
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:35.173452+00:00
- Verified
- 2026-07-17T20:18:35.173452+00:00
Summary
A next-generation large-scale real-world multimodal dataset by PsiBot for embodied intelligence training, comprehensively covering vision, language, and action modalities. It includes 449k clips across four subsets: egocentric visual data (313,674 clips), original exoskeleton-glove manipulation data (95,383 clips), background-replaced glove data (3,526 clips), and glove data with tactile signals (36,780 clips). Data is collected using PsiBot's self-developed exoskeleton glove system achieving millimeter-level positioning accuracy and capturing full degrees of freedom of both hands and arms, combining high-precision structured capture with natural human interaction behavior for vision-action modeling and imitation learning.
Keywords
robotics embodied-ai egocentric manipulation imitation-learning vision-action exoskeleton-glove tactile bimanual embodied-intelligence
Topics
Robotics / Embodied AI
Research notes
- Modalities include head RGB, depth, camera intrinsics, head pose, head IMU, wrist pose, hand qpos, fingertip keypoints, and tactile signals. Includes both exoskeleton-based and bare-hand data. Cited in H-RDT paper (AAAI 2026) for human manipulation enhanced bimanual robotic manipulation.