REC · ON-SITE CAPTURE SYNJUKU · 2026PROTOTYPE · INTERNAL DRAFT

07 / EXAMPLE DATASET

This is what deployment-grade data looks like

A real pilot sample, delivered July 2026 to a leading VLA lab in their ingestion format. Every figure on this page is measured from the delivered dataset — nothing here is a mockup. This is step 4 of the deployment loop: the lab receives primitive-level annotated data, verified and provenanced.

At a glance

10.0 hrs
600 episodes × 60 s
Four camera streams each
9,384
Verified action labels
0 contradictions in a 1,800-span audit
1.08M
Frames with merged 3D tracks
29.97 fps · camera + hands + body
±5%
Category mix vs the lab's targets
All categories within tolerance

One episode, four synchronized views

Episode 000037 — "Wrap Takeaway Coffee Cups", captured in a working café. Egocentric, exocentric, and both wrists, on a single shared timeline.

Egocentric view, episode 000037
E · EGOCENTRICHEAD-MOUNTED
Exocentric view, episode 000037
X · EXOCENTRICFIXED THIRD-PERSON
Left-wrist view, episode 000037
L · LEFT WRISTWRIST-MOUNTED
Right-wrist view, episode 000037
R · RIGHT WRISTWRIST-MOUNTED

The labels are verified, not asserted

Consecutive subtask spans, each individually checked against the footage. Below: the first spans of episode 000037, verbatim from the delivery.

// meta/custom_annotation.json — v1.4
{ "episode_id": "000037", "spans": [
  { "start": 3.2,  "end": 5.2,
    "label": "pick up food package from drawer and hold in hand" },
  { "start": 5.2,  "end": 7.5,
    "label": "pick menu cover from drawer and place menu cover on drawer" },
  { "start": 12.1, "end": 15.6,
    "label": "pick menu booklet from inside drawer and place menu booklet in hand" },
  ...
]}

Every frame carries 3D ground tracks

Pre-merged into the per-frame records — no joins required. Most vendors ship 2D video and captions; this is what deployment training actually needs.

Camera poseSLAM · position + rotation, per frame
Hand pose3D world + image space · 21 landmarks per hand
Body poseExocentric · 9 upper-body joints
LabelsVerified subtask spans · ~86% clip coverage
FormatLeRobot v3.0 — the lab's requested format

Quality is measured, then shipped

What was checkedResult
Human evaluation set240 spans, double-annotated
Checker vs human reviewers91% agreement on rejections
Labels screened9,384 / 9,384 — failures re-captioned or removed
Final random audit1,800 labels · 0 flagged
Stated accuracy90–95%, deliberately conservative
3D trackMean detectionEpisodes ≥90%
Camera position (SLAM)96.2%552/600
Upper body (exo)98.4%574/600
Left hand, 3D world95.4%552/600
Right hand, 3D world93.2%532/600
Labels that could not be confirmed against footage were removed, not delivered. Annotations are versioned v1.0 → v1.4 with every change logged, and each episode carries provenance — task card, capture site, recording date — retained for audit.

Long tail by design

175 task cards across 141 environment×site scenes at 10 real venues — hospitality, retail, logistics, industrial, home, office — capped at ≤3.6 h per scene so no single environment dominates. The delivered mix landed within ±5% of the lab's pilot targets in every category.

CategoryDelivered shareLab target ±5%
Others (office / conference / outdoor)22.4%20%
Home19.9%15%
Retail (retail floor, café)17.6%15%
Hospitality (hotel, reception, restaurant kitchen)17.2%15%
Logistics (packing, storage)12.4%15%
Industrial (recycling, workstations)10.5%10%

Where this goes next

This sample is the front end of a 1,000-hour pilot in flight with the same lab. The loop's final step — the model improves, the robot redeploys — is what the pilot exists to measure. Eval results land here when the lab reports them.

SYNJUKU · EXAMPLE DATASET FIGURES AS DELIVERED · 2026.07.02 CUSTOMER ANONYMIZED · PROTOTYPE FOR NARRATIVE REVIEW