Split design · head unit approx. 210g. Just wear it and capture imperceptibly in real environments — factories, homes, stores and more. Vision, depth, pose and tactile data: five synchronized multimodal streams, ready for model training.
Real human work is data production.
Moving data capture from “purpose-built setups” to “something that simply happens”.
Other ego datasets are video only. Every frame NOLO EgoCapture outputs carries millimeter-level 6-DoF pose ground truth. A worker assembling parts on the shop floor, a cook working in a kitchen, a clerk stocking shelves — while they do their normal jobs, vision, depth, hand motion and tactile information are recorded in sync and turned directly into high-quality robot training data.
Traditional capture depends on the robot itself and on purpose-built capture environments. Hardware and deployment cost a great deal, and the cost per usable sample stays high.
Calibrating and capturing person by person, with limited capture rooms and workstations, makes it hard to reach the tens of millions of samples training requires.
Capture environments are disconnected from real production, scenario diversity is insufficient, model generalization suffers, and the “last mile” never gets closed.
The EgoCapture headset and the OmniGlove smart glove form a lightweight capture kit: no robot body to depend on, no complex deployment — put it on and start collecting. Powered by our in-house PolarTraq® and StarTraq® positioning technologies, it holds sub-millimeter accuracy and millisecond latency even in complex real environments, outputting five synchronized multimodal streams at 30Hz, ready for model training.
The battery and storage move down into the handheld controller, keeping the head unit extremely light and comfortable over long sessions — capturing becomes truly imperceptible.

Only what capture requires: six cameras, a high-precision IMU and dual noise-cancelling microphones. With battery and storage moved elsewhere, the head unit stays extremely light and comfortable through long sessions.

Battery, storage and capture control are centralized in the handheld terminal. A 4.5-inch touchscreen handles live preview and task management, and 40W fast charging supports about 8 hours of continuous capture.
From ultra-wide field-of-view capture to ground-truth-grade data output: a complete multimodal data infrastructure for embodied AI model training.
Five global-shutter RGB cameras plus one ToF depth camera deliver a full-device image FOV of 208°×193°, giving all-around visual coverage from the head-worn viewpoint, with depth and color captured in sync.
Head, wrist and finger poses are output per frame with millimeter-level accuracy, strictly time-aligned with the video frames. The data is ready for motion retargeting and model training, removing tedious SLAM post-processing and failed re-captures.
Five data streams — RGB imagery, ToF depth, audio, gesture and tactile — are output at 30Hz with unified timestamps, keeping multi-device capture spatiotemporally consistent and fit for large-scale dataset construction.
Battery and storage move down to the handheld controller; the head unit keeps only the capture module. At roughly 210g it stays comfortable through long shifts for truly imperceptible capture.
The handheld controller houses a large 22000mAh battery and supports 40W PD fast charging, delivering around 8 hours of continuous capture per charge for full-day operation and long task capture.
Supports 1,000+ people capturing and uploading at the same time, with data solving capacity matched 1:1 to capture capacity. Timestamps across streams stay strictly aligned when multiple people and rigs capture together, scaling up training datasets.
The same EgoCapture headset, paired with the OmniGlove smart glove or with bare-hand tracking, switches capture granularity to fit the task.
Captures fingertip tactile feedback and dexterous manipulation with high precision: the tactile array outputs per-finger normal force values and tangential force directions in real time. Paired with the headset for fine-skill demonstration, it suits assembly, caregiving and surgery — anywhere dexterity matters.
No glove required: the headset's built-in algorithm identifies 21 hand keypoints in real time. Raise your hand and it captures, with natural unrestricted motion — ideal for batch data capture of high-frequency daily tasks such as retail restocking and home cleaning.
Capture end, sensing layer, algorithm layer and output layer are connected end to end, delivering multimodal datasets that go straight into model training.
Supports more than a thousand people capturing and uploading at once, meeting the throughput demands of a scaled data factory.
A single full capture session yields hundreds of gigabytes of raw data, covering an entire day of work.
Data solving capacity matches capture capacity 1:1 — solving completes as capture finishes, with no queueing.
For embodied AI robot training in industrial manufacturing, domestic services, retail and medical care — real people “work as they capture”.
Capturing fine operations in shop-floor assembly, machine tending and quality inspection
Capturing everyday household motions such as home cleaning and tidying up
Capturing store operations such as shelf restocking and item picking
Capturing clinical skills such as nursing procedures and instrument handling
Key parameters of the EgoCapture headset and the handheld controller. Final values are subject to the delivered version.
| Processor | RK3588, local or real-time upload; offline output of 6DoF and 21 hand keypoints |
| Memory / Storage | LPDDR5 12GB | UFS3.1 512G | TF card 512G |
| RGB Cameras | 5× global shutter, per camera 1920×1200 @60Hz, full-device image FOV 208°×193° |
| Depth Camera | 1× ToF, 640×480 @30Hz, FOV 121.5°×98° |
| Positioning | Full-device positioning FOV 180°×160°, 6DoF accuracy 10mm |
| IMU | Accelerometer + gyroscope, sampling at 1000Hz |
| IR Illumination | 5× 850nm LED, FOV 150°; visual enhancement for more robust positioning |
| Microphones | 2× with noise reduction |
| Wireless | WiFi 6 | 2.4G private link (IR positioning module expansion) |
| Ports | Type-C ×2 (handheld terminal link / USB3.1 wired real-time upload) |
| Weight | Head unit approx. 210g (soft head strap excluded) |
| Split Design | Battery, storage and capture control sit in the handheld controller; the head unit keeps only the capture module |
| System | Custom capture system based on Android 14 |
| Display | 4.5-inch touchscreen, resolution 450×845 |
| Battery | 22000mAh, 40W PD fast charging, approx. 8 hours of runtime |
| Ports | Type-C ×2 (headset link / charging) |
| Buttons | Power button ×1 |
| Companion App | DAS app (Android / iOS) for task management, data capture, preview and upload |
| Capture Control | Capture control button, capture status indicator |
* Specifications are taken from official NOLO product documentation and may change as versions are updated. Please refer to the delivered product.
Book a product demo, or discuss your embodied AI data capture plan with us. Based on your scenario, the NOLO team will advise on everything from hardware selection to closing the data loop.