The pipeline's first stage is a data-quality gate, so the data itself is a deliverable. This is what we published, the sources we used and cited, and the data challenges still open.
Open artefacts under the ubr-physical-ai organisation. Licences are split by content: annotations are ours to release; synthetic renders inherit the terms of the NVIDIA assets they were rendered from, which is the open licence question below.
| Dataset | What it is | Scale | Licence | |
|---|---|---|---|---|
| isaac-sdg-rescue-target | The hybrid search-and-rescue set: Isaac Replicator renders passed through Cosmos Transfer, with segmentation masks and merged detection annotations over 15 classes. main = v3, tag v3.0 | 133,760 images 37 webdataset shards 109 clips w/ masks 559,085 train boxes 139 provenance runs |
annotations CC BY 4.0; renders under NVIDIA asset terms | Public |
| cosmos3-i2v-survival-sdg | Image-to-video synthetic clips explored as an alternative augmentation route. Published as a negative result — the yield did not justify the route, and saying so is the point. | I2V clip set | CC BY 4.0 (ours) / NVIDIA terms on conditioning | Negative result |
| ubr-maze-nav | Navigation dataset from before the Codefest, kept public as prior work and reused for context, not for the SAR target. | pre-Codefest | CC BY 4.0 | Prior work |
| cosmos-i2v-ugv-v0 | Rover-domain conditioning clips. Kept private: derived from a workplace recording of identifiable people. Persons in the clips were anonymised before any run; raw segments never leave the bench. | private | not distributed | Private |
Not ours; used under their own terms and never redistributed. Where a source contains identifiable people, only derived frames are used and only aggregate numbers leave the bench.
| Source | How it is used | Terms we respect |
|---|---|---|
| PhysicalAI-Spatial-Intelligence-Warehouse NVIDIA, Hugging Face | The L0 spatial-reasoning benchmark: 1,942 validation items scored across distance, left/right and multiple-choice. Cited, gated, never redistributed. | CC BY 4.0; annotations also carry the Llama 3.1 Community License |
| PhysicalAI — SmartSpaces · Spatial-QA · NuRec NVIDIA | Subsets only: SmartSpaces negatives, Spatial-QA items, and NuRec reconstructed rooms turned into Isaac benchmark scenes. Attribution logged per subset. | NVIDIA dataset terms; the 210 GB Spatial-QA train image set is deliberately not pulled |
| Rover blackbox recording ours, private | The real sensor domain: ~108 h indoor, 791 person segments. Supplies the 474-frame real-mission slice that the L1 instrument scores. Never uploaded; derived frames only after face/plate review. | identifiable people — private, anonymised before use |
| SARD public | External reference for lying-person recall per size bucket — an outside yardstick for the rescue posture the detector must catch. | public research dataset, cited |