Technology

Not people with cameras. A pipeline with people in it.

The hard part of real-world data is not recording it. It is knowing that what you recorded is usable — synchronised, complete, correctly labelled and traceable — before it reaches your training run. Every episode moves through the same instrumented path.

The pipeline

From capture device to your training bucket

Six stages. Nothing skips a stage, and every stage leaves a record.

01

Capture

Time-synced multi-sensor recording with on-device integrity checks. The operator sees a live status readout, so a failed sensor is caught during the session rather than a week later.

02

Secure Upload

Encrypted, resumable transfer from the field — designed for patchy rural connectivity. Custody is logged at handoff and the local copy is wiped once receipt is confirmed.

03

QA Engine

Automated checks run on every episode before a human ever opens it, so reviewers spend their time on judgement rather than on spotting broken files.

04

Metadata Generation

Task labels, environment tags, operator and device identifiers, calibration and timing — attached per episode, not in a spreadsheet someone maintains by hand.

05

Versioned Storage

Datasets are versioned, so a redelivery after a fix is a new version rather than an overwrite. Access is controlled and every read is logged.

06

Delivery

Pushed to your cloud in the structure and format your training pipeline already reads, with a manifest describing exactly what changed since the last drop.

Quality assurance

What gets rejected before you ever see it

Most unusable real-world data fails for boring, detectable reasons. We check for those automatically, then put a trained human on everything that survives.

Sync drift

Streams that have slipped against each other beyond the agreed tolerance.

Dropped frames

Gaps in a stream, truncated recordings, and files that end mid-episode.

Exposure & focus

Blown highlights, underexposed interiors, motion blur and out-of-focus segments.

Occlusion

Hands or the manipulated object out of frame during the part of the task that matters.

Calibration drift

Stereo or depth extrinsics that have moved since the rig was last calibrated.

Protocol compliance

Episodes that skipped a step, ran short, or deviated from the agreed task definition.

Duration outliers

Runs far shorter or longer than the distribution, usually a sign something went wrong.

Coverage gaps

Under-represented operators, sites, lighting conditions or object variants in the batch.

Consent & redaction

Missing consent record, or a redaction rule that has not been applied.

Rejected episodes are re-shot against the same protocol rather than quietly dropped from the count.

Metadata

What ships with every episode

A dataset is only as useful as its ability to be filtered. Every episode carries enough structure that you can slice by site, operator, lighting, device or task variant without re-watching anything.

Ask for a sample episode
  • TaskTask ID, variant, step boundaries, success or failure outcome, and the reason where it failed.
  • EnvironmentSite type, industry, indoor or outdoor, lighting condition, ambient noise, weather where relevant.
  • ParticipantPseudonymised operator ID, experience level, handedness — enough to study variation, not enough to identify.
  • DeviceRig configuration, sensor models, resolutions, frame rates, and the calibration used for that session.
  • TimingPer-stream timestamps, sync offsets, and the measured drift across the episode.
  • ProvenanceCapture date, QA verdict and reviewer, redaction profile applied, dataset version.

Delivery

In the format your pipeline already reads

We agree the schema at scoping and deliver against it. If you have an internal format, we target that instead of handing you a conversion job.

Formats we deliver in

Common robot-learning formats, or your own schema if you have one.

RLDSLeRobotHDF5WebDataset ParquetMP4 + JSONROS bagYour schema

How it reaches you

Delivered where you want it, on the cadence you want it.

  • Direct push to your S3, GCS or Azure bucket
  • Rolling weekly drops, or one batch at the end
  • Manifest describing every change since last drop
  • Checksums so you can verify what arrived

Roadmap

What exists today, and what is coming

We would rather be specific about the line between the two.

Live

Running today

  • Time-synced multi-sensor capture
  • Encrypted, resumable field upload
  • Automated QA checks on every episode
  • Per-episode metadata generation
  • Versioned, access-controlled storage
  • Delivery to customer cloud storage

In development

  • Customer dashboard for live batch progress
  • Programmatic APIs for dataset queries
  • In-house annotation platform
  • Company-owned capture hardware
  • Automated redaction at ingest

Want to see the output before you commit?

We can share a sample episode with its full metadata and QA record, so your team can judge the format against a real training run rather than a description of one.