You sent us a sample drop from your hardware, and we spent time actually running it rather than just reading the spec sheet. This is a straight accounting of what was in the drop, what we did with it, and what we'd want to look at next.
See it running on your footage →| SOURCE | SESSIONS / FILES | CONTENTS |
|---|---|---|
| Stereo cam | 1 shared calibration set + 3 recording sessions (“Sample 1”, “Sample 2”, “Sample 3”) | Each session is a synchronized left+right camera pair recording |
| Wrist cam | 4 activity sessions (“Cooking”, “Dishes”, “Folding Clothes”, “Time Sync Proof”) | Each session bundles three synchronized camera feeds — a head-mounted camera plus a left-wrist and a right-wrist camera — each with its own IMU data and timestamp sync files |
| UVC | 3 clips (“assembly”, “clothes”, “kitchen”) | Standalone video clips |
| SD card (Trinet) | 6 recordings | Short clips captured directly to the card |
That's a genuinely broad sample — four sensor families, thirteen distinct recording sessions, all from one shared calibration baseline.
We reconstructed 3D hand motion from three of your headcam sessions — Dishes, Cooking, and Folding Clothes — full articulated hand pose over time, recovered directly from the raw video. You can see it running against your own footage in the live demo linked above.
We did not touch the wrist cameras themselves, the “Time Sync Proof” session, any of the three stereo-cam sessions, the UVC clips, or the Trinet recordings for this pass — the demo is scoped to headcam-only reconstruction on those three sessions.
The sample drop has more in it than we've used so far, and the parts we haven't touched are exactly the parts that would let us go further:
Every one of the four wrist-cam sessions carries left- and right-wrist feeds with their own IMU streams, and none of it has been processed yet. That's a close-range, hand-proximate viewpoint sitting unused right next to the headcam data we already reconstructed from.
A calibrated left+right pair is a different kind of signal than a single headcam — it's the natural next thing to bring in once we've validated the monocular reconstruction end to end.
This one looks purpose-built for checking cross-sensor alignment, which is directly useful groundwork before combining headcam, wrist-cam, and stereo streams.
Everything processed so far — and everything in the sample, as far as we can tell — comes from a single location under a single lighting condition. Before drawing any general conclusions about how well this holds up, we'd want footage from a few more environments and lighting setups to see where it starts to bend.
None of this requires a new capture rig or a new data-sharing conversation — it's already sitting in what you sent us. A scoped pilot is the natural way to go pull on these threads systematically rather than one clip at a time.
What this sample confirms is that your capture stack — headcam, wrist cams, stereo pair, IMU sync, all in one rig — is a genuinely rich substrate for turning raw egocentric video into structured motion data, not just footage. The fact that a single sensor stream from your hardware was enough to get usable 3D hand motion out is the encouraging part; the multi-sensor, multi-condition material you've already provided is what would tell us how far that generalizes. We think there's a real pipeline here worth building out properly, and the fastest way to find its edges is to keep working through the data you've already handed us.