Formats we deliver in
Common robot-learning formats, or your own schema if you have one.
Technology
The hard part of real-world data is not recording it. It is knowing that what you recorded is usable — synchronised, complete, correctly labelled and traceable — before it reaches your training run. Every episode moves through the same instrumented path.
The pipeline
Six stages. Nothing skips a stage, and every stage leaves a record.
Time-synced multi-sensor recording with on-device integrity checks. The operator sees a live status readout, so a failed sensor is caught during the session rather than a week later.
Encrypted, resumable transfer from the field — designed for patchy rural connectivity. Custody is logged at handoff and the local copy is wiped once receipt is confirmed.
Automated checks run on every episode before a human ever opens it, so reviewers spend their time on judgement rather than on spotting broken files.
Task labels, environment tags, operator and device identifiers, calibration and timing — attached per episode, not in a spreadsheet someone maintains by hand.
Datasets are versioned, so a redelivery after a fix is a new version rather than an overwrite. Access is controlled and every read is logged.
Pushed to your cloud in the structure and format your training pipeline already reads, with a manifest describing exactly what changed since the last drop.
Quality assurance
Most unusable real-world data fails for boring, detectable reasons. We check for those automatically, then put a trained human on everything that survives.
Streams that have slipped against each other beyond the agreed tolerance.
Gaps in a stream, truncated recordings, and files that end mid-episode.
Blown highlights, underexposed interiors, motion blur and out-of-focus segments.
Hands or the manipulated object out of frame during the part of the task that matters.
Stereo or depth extrinsics that have moved since the rig was last calibrated.
Episodes that skipped a step, ran short, or deviated from the agreed task definition.
Runs far shorter or longer than the distribution, usually a sign something went wrong.
Under-represented operators, sites, lighting conditions or object variants in the batch.
Missing consent record, or a redaction rule that has not been applied.
Rejected episodes are re-shot against the same protocol rather than quietly dropped from the count.
Metadata
A dataset is only as useful as its ability to be filtered. Every episode carries enough structure that you can slice by site, operator, lighting, device or task variant without re-watching anything.
Ask for a sample episodeDelivery
We agree the schema at scoping and deliver against it. If you have an internal format, we target that instead of handing you a conversion job.
Common robot-learning formats, or your own schema if you have one.
Delivered where you want it, on the cadence you want it.
Roadmap
We would rather be specific about the line between the two.
We can share a sample episode with its full metadata and QA record, so your team can judge the format against a real training run rather than a description of one.