07 / EXAMPLE DATASET
A real pilot sample, delivered July 2026 to a leading VLA lab in their ingestion format. Every figure on this page is measured from the delivered dataset — nothing here is a mockup. This is step 4 of the deployment loop: the lab receives primitive-level annotated data, verified and provenanced.
Episode 000037 — "Wrap Takeaway Coffee Cups", captured in a working café. Egocentric, exocentric, and both wrists, on a single shared timeline.
Consecutive subtask spans, each individually checked against the footage. Below: the first spans of episode 000037, verbatim from the delivery.
// meta/custom_annotation.json — v1.4 { "episode_id": "000037", "spans": [ { "start": 3.2, "end": 5.2, "label": "pick up food package from drawer and hold in hand" }, { "start": 5.2, "end": 7.5, "label": "pick menu cover from drawer and place menu cover on drawer" }, { "start": 12.1, "end": 15.6, "label": "pick menu booklet from inside drawer and place menu booklet in hand" }, ... ]}
Pre-merged into the per-frame records — no joins required. Most vendors ship 2D video and captions; this is what deployment training actually needs.
| Camera pose | SLAM · position + rotation, per frame |
| Hand pose | 3D world + image space · 21 landmarks per hand |
| Body pose | Exocentric · 9 upper-body joints |
| Labels | Verified subtask spans · ~86% clip coverage |
| Format | LeRobot v3.0 — the lab's requested format |
| What was checked | Result |
|---|---|
| Human evaluation set | 240 spans, double-annotated |
| Checker vs human reviewers | 91% agreement on rejections |
| Labels screened | 9,384 / 9,384 — failures re-captioned or removed |
| Final random audit | 1,800 labels · 0 flagged |
| Stated accuracy | 90–95%, deliberately conservative |
| 3D track | Mean detection | Episodes ≥90% |
|---|---|---|
| Camera position (SLAM) | 96.2% | 552/600 |
| Upper body (exo) | 98.4% | 574/600 |
| Left hand, 3D world | 95.4% | 552/600 |
| Right hand, 3D world | 93.2% | 532/600 |
175 task cards across 141 environment×site scenes at 10 real venues — hospitality, retail, logistics, industrial, home, office — capped at ≤3.6 h per scene so no single environment dominates. The delivered mix landed within ±5% of the lab's pilot targets in every category.
| Category | Delivered share | Lab target ±5% |
|---|---|---|
| Others (office / conference / outdoor) | 22.4% | 20% |
| Home | 19.9% | 15% |
| Retail (retail floor, café) | 17.6% | 15% |
| Hospitality (hotel, reception, restaurant kitchen) | 17.2% | 15% |
| Logistics (packing, storage) | 12.4% | 15% |
| Industrial (recycling, workstations) | 10.5% | 10% |
This sample is the front end of a 1,000-hour pilot in flight with the same lab. The loop's final step — the model improves, the robot redeploys — is what the pilot exists to measure. Eval results land here when the lab reports them.