Ground truth is relative to the simulated scene
Simulation can produce reference outputs from known scene state. Those outputs describe the modeled scene and the configured renderer or sensor. They do not prove that the scene exactly matches a physical facility or that the simulated sensor matches real hardware.
Omniverse Replicator provides tooling for synthetic data generation. The useful question for a project is which outputs its configured pipeline actually produces and how those outputs are checked.
Specify the outputs your pipeline consumes
| Output | What to establish before production |
|---|---|
| RGB | Resolution, camera model, lighting and appearance requirements |
| Depth | Depth definition, units, invalid values and camera calibration |
| Semantic labels | Class vocabulary, unlabeled regions and mapping to training classes |
| Instance labels | Object identity policy and behavior across frames or resets |
| Normals | Coordinate space and encoding used by the consuming pipeline |
Do not assume that selecting a USD environment automatically supplies every modality. Dataset generation also requires camera setup, annotators or sensors, writers and quality control in your runtime.
Validate a small batch first
Produce a few representative frames before generating a large dataset. Overlay annotations on the corresponding RGB views and inspect object boundaries, occlusion and class assignments. Confirm depth with known scene distances using the chosen definition. Check file pairing and frame identifiers so that related outputs stay aligned.
For sequences, explicitly test identity continuity and timing. A visually plausible still frame does not establish that temporal annotations remain consistent.
Keep three kinds of evidence separate
Authored scene values describe the model. Validation checks test defined requirements. Real-world measurements establish how the model relates to a physical reference. Calling all three ground truth without qualification can hide important assumptions.
Record scene revision, camera configuration, taxonomy version, generator settings and any post-processing. Retain a held-out evaluation set with conditions relevant to the intended deployment.
Browse simulation environments as source content and use the semantic labeling guide to plan a taxonomy. For task-specific scenes and required output coverage, discuss your project before treating a library package as a complete data pipeline.