Photorealistic synthetic data for computer vision,
built from production 3D environments
Built by artists, ready for training. Evermotion prepares photorealistic 3D scenes for synthetic data for computer vision.
You render in your own pipeline - masks, classes, geometry and cameras come straight from the scene.
18 000+
ready-made 3D assets, plus thousands of complete scenes
Since 2005
photorealistic 3D content for commercial clients worldwide
461
Evermotion scenes used to build Apple's Hypersim dataset
Apple's Hypersim was built
on Evermotion scenes
Hypersim, Apple's photorealistic synthetic dataset for holistic indoor scene understanding, was generated from 461 Evermotion Archinteriors scenes: dense per-pixel semantic and instance labels, ground-truth geometry and camera parameters for every image. An independently documented example of Evermotion environments used for computer vision at scale, released by Apple as one of the public datasets researchers rely on for indoor scene understanding.
461
Evermotion scenes
77 400
photorealistic images rendered
74 619
images in the public release
Source: Hypersim, Apple Machine Learning Research · apple/ml-hypersim on GitHub. Public counts re-checked before release.
The render is only one layer.
Switch through the data underneath it.
Every layer below comes from the same camera, frame and scene state of one prepared Evermotion scene. The web images are display transforms; the delivered package keeps the numeric EXR data. Which layers a project receives is agreed before production.
Photorealistic RGB
The final beauty render: the appearance domain the model will see. Path-traced lighting, physically based materials, real-world scale.
rgb / beautyInstance identity
Every uniquely named object gets its own ID, so instance and semantic masks are derived from the scene rather than drawn by hand. 493 objects in this example, no ID collisions.
CMasking_Cryptomatte_NodeNameValidated masks
Object recognition grounded in the 3D scene: segmentation masks and bounding boxes come straight from the scene geometry and are validated against the manifest - ready-made supervision for detection and recognition models.
masks / 2D boxes / class labelsWireColor IDs
A second, independent identity channel. 493 unique colors, zero collisions, gamma-decoded before matching against the linear manifest.
CMasking_WireColorShading normals
Per-pixel surface orientation as shaded, including normal maps. Useful for surface and lighting-aware tasks.
CGeometry_NormalsShadingGeometry normals
Surface orientation from the mesh geometry only, without shading tricks. Pairs with depth and world position for 3D supervision.
CGeometry_NormalsGeometryDepth
ZDepth EXR is present for every view. Its range, units and numerical convention are documented and validated for the selected renderer before it is delivered as metric depth.
CGeometry_ZDepthSource color
Material color before lighting. A diagnostic layer for material-aware tasks and for auditing what the renderer was given.
CShading_SourceColorMaterial albedo
A diagnostic reflectance pass, not a base-color texture. Included when the task calls for it.
CShading_AlbedoWhere will your model run?
Start from a matching environment
These are examples, not claims that one dataset covers every domain. We start from more than 18 000 ready-made 3D assets and thousands of complete scenes, then adapt selected geometry, content distribution, lighting and annotations to the brief. This closes the domain gap between your training data and the target domain your CV model meets in production.

Retail
Product detection, planogram compliance and shelf monitoring in scenes with dense packaging, reflective glass and mixed color temperature.

Warehouse
Pallet detection and rack occupancy in environments with repeating structures, deep perspective and skylight glare.

Manufacturing
Assembly-line inspection and part recognition in workcells with specular metal, repeating structures and heavy occlusion.

Industrial vision
Inspection rigs and conveyors can be built from suitable client CAD, prepared for rendering and configured for sensor-placement studies.

Residential
Domestic interiors provide small-object instances, scale variation, clutter and partial visibility for task-specific capture plans.

Healthcare
Equipment tracking and room-state monitoring across clinical surfaces, flat lighting and large low-texture regions.

Public spaces
Signage recognition and left-object detection in deep scenes with overcast daylight and wet reflections.

Exterior
Damage assessment and debris segmentation from street-level or planned aerial viewpoints across unstructured geometry.
Start with the model objective.
Then configure the dataset around it.
-
Step 1 of 3Define
You and your computer vision engineers define the task: target classes, edge cases, viewpoints, output formats and success criteria - for object detection, segmentation or other computer vision applications, from autonomous vehicles to industrial inspection. We review the brief with you and propose a feasible data scope.
-
Step 2 of 3Configure
We map the scene: taxonomy, object structure, materials, cameras and variations are prepared to the brief. Missing scenarios, props or layouts are built or adapted.
-
Step 3 of 3Deliver
Selected passes, annotations, camera models, machine-readable manifests and the disclosed QA result ship together.
Only the outputs your task needs
We do not force every project into one annotation package. Evermotion scenes are editable, structured production assets: mapped to your taxonomy, rendered as multi-view synthetic data and exported with the ground truth the task actually needs.
Pixel-aligned ground truth
Depending on the task, one scene state can produce RGB together with instance identity, object IDs, shading and geometry normals, world position, alpha, source color and diagnostic albedo.
Annotations in your schema
Raster semantic masks, visible 2D boxes, class mappings, occlusion rules and material taxonomies are generated after the target task and acceptance tests are agreed. Systematic variation is defined as a project-specific domain randomization plan.
Cameras and 3D geometry
Cameras ship with derived 3×3 intrinsics and camera-to-world poses. Objects can carry stable IDs, semantic classes, oriented 3D bounding boxes and real-world dimensions.
Verification you can inspect
A delivery can include machine-readable manifests, semantic coverage, transform and scale checks, object separation checks, material coverage and disclosed advisories. Thresholds are agreed per project, so QA describes the delivered data rather than acting as a badge.
From production scene to structured data.
One inspectable example.
A documented example of preparing an existing Evermotion scene for synthetic data. It demonstrates the workflow; taxonomy, cameras, passes, formats and acceptance scope remain project-specific. Reflective glass, polished stone, foliage, thin geometry, repeated furniture and dense tableware create overlapping targets, narrow boundaries and partial visibility - relevant conditions for indoor perception data.
493
uniquely named scene objects
34
mapped semantic classes
11.87M
renderable triangles
229
tracked materials
14
views with exported camera models
27
render elements per view
Stable scene state, different poses, focal lengths and occlusion patterns. Each frame includes derived 3×3 intrinsics and a camera-to-world transform. They reproduce the render camera; they are not physical-target calibration.
Evidence from the manifests.
The limits are recorded too.
{
"scene_id": "AIXXX_001_Corona_test_SDR_CV100",
"renderer": "Corona 15 Hotfix 2",
"scene_complexity": {
"objects": 493,
"semantic_classes": 34,
"triangles": 11865783,
"materials_tracked": 229,
"camera_views": 14,
"render_elements_per_view": 27
},
"production_verification": {
"shots_completed": "14/14"
},
"validator_result": {
"grade": "A",
"readiness_checks": "8/8 (internal rubric)",
"semantic_rule_precision_percent": 99.8,
"material_full_classification_percent": 79.04,
"external_certification": false
}
}
| Check | Result and disclosed caveat |
|---|---|
| Semantic assignment | 493/493 objects assigned; 99.8% mapped through specific semantic rules |
| WireColor audit | 493 unique colors, 0 collisions; gamma 2.2 decoded before linear-manifest matching |
| Mesh readiness | 0 semantic merge blockers; 30 non-blocking topology and complexity advisories retained |
| Material metadata | 229/229 tracked; 79.04% fully classified |
| Camera model | Derived intrinsics plus camera-to-world extrinsics; no physical target calibration claimed |
| Depth status | ZDepth EXR present; range and convention require qualification before metric use |
This is what your training data
starts from
A short selection of Evermotion interiors and exteriors as they leave the studio: path-traced, physically based, real-world scale. Every environment here can be prepared for synthetic data capture. Click to view full size.
Photorealism is an input.
Real-world validation is the test.
Synthetic data narrows two gaps. Your benchmark decides whether it worked.
-
Appearance gap
Path-traced materials and lighting, tuned to the deployment domain.
We configure -
Content gap
Object frequency, clutter, occlusion, viewpoints and long-tail cases, per capture plan.
We configure -
Model performance
Accuracy, generalization and sim-to-real transfer are measured on your real-world benchmark.
You evaluate
Three ways to get your data
We did not buy the library - we built it. One counterparty for content, preparation, provenance and rights.
We build to your specification
Scenes adapted or built, captured with your cameras, annotations and QA record.
We supply the content
Selected Archinteriors and Archexteriors scenes under a dedicated ML agreement - for synthetic data platforms, simulation pipelines and research.
Ongoing content pipeline
Recurring scene production with review gates, delivered on an agreed cadence.
What this is
You license photorealistic 3D scenes, prepared for synthetic data capture and rendered in your own pipeline, or order ready-made image data generated from them, to train computer vision models.
What you get
The files, the labels and a written license that covers AI use - with the origin of the content confirmed in the contract, ready for your legal team.
How fast
Every brief gets an answer within 24 hours on business days: what is feasible, how it will be validated, what it will cost. A small validated sample comes before production volume.
How you pay
Per asset - a scene or a model - depending on license scope, order size and adaptation work.
Project-specific AI/ML rights,
not a shop license
Standard Evermotion licenses exclude AI and machine-learning use. The client keeps full rights to use what comes out of the training, and the detailed scope is agreed individually - we are flexible about formats, provenance records and how the data may be used.
| Use | Standard shop license | Project AI/ML agreement |
|---|---|---|
| AI / ML training and fine-tuning | No | Yes |
| Evaluation and benchmarking | No | Yes |
| Derived data and synthetic outputs | No | Yes |
| Source-scene delivery | No | If agreed |
| Redistribution of source 3D files | No | If agreed |
| Resale of datasets or derivatives | No | If agreed |
“The speed at which Evermotion worked was incredibly impressive as they moved from rough blocking based on STEP files generated by our CAD system and a few hand-drawn sketches to full detailed models in a timescale I wouldn’t have thought possible.”
What can be configured?
What still needs validation?
What ground-truth annotations can Evermotion deliver?
Depending on the project, a delivery can include instance identity, semantic masks, object IDs, visible 2D boxes, oriented 3D bounding boxes, surface normals, world position, alpha, material-related render passes, camera metadata and validated depth. The final set is defined by the target task and the tested export pipeline.
Can the dataset use our taxonomy and annotation schema?
Yes. Object naming and semantic classes can be mapped to a client taxonomy, including rules for class granularity, occlusion, visibility and fallback labels. The machine-readable schema and acceptance tests are agreed before production.
Which renderers and data formats are supported?
Production work is centered on 3ds Max with Corona or V-Ray, with Blender and Cycles or Unreal Engine 5 available when required. Scene and data formats are selected per project. A format such as COCO, YOLO, KITTI or OpenLABEL is only promised after the required fields and exporter have been implemented and validated for that delivery.
How is a scene verified before data generation?
Verification can cover physical units, coordinate conventions, transforms, semantic coverage, instance separation, material tracking, cameras, lighting and render-pass configuration. The checks, thresholds, advisories and output manifests are documented with the delivery.
Can Evermotion content be licensed for AI and machine-learning training?
AI/ML use is excluded from standard Evermotion product licenses. Approved training, fine-tuning, evaluation or dataset use requires a separate, project-specific B2B agreement defining the content, permitted uses, formats and downstream rights.
Who owns the trained model and its outputs?
You do. The license can grant the client full rights to use everything that results from training - the model, its weights and its outputs. Evermotion keeps ownership of the source 3D content itself.
Where does the content come from? Is anything scraped?
The library was built by our own studio since 2005 - modeled, textured and lit in-house for commercial rendering. Nothing is scraped and no third-party datasets are involved, which is why we can document provenance and put the license in front of your legal team.
Can the content be delivered white-label?
Yes. For synthetic data platforms and resellers, scenes or extracted data can be delivered without Evermotion attribution, under NDA, as part of the project agreement - your clients see your brand, not ours.
How is pricing structured?
Per asset - a model or a scene. The rate depends on three things: the scope of the license, the size of the order and the adaptation work needed to meet your specification. You get a production estimate with the feasibility response, and the terms are agreed individually.
Does photorealistic synthetic data guarantee sim-to-real performance?
No. Photorealistic assets and controlled scene variation can help address appearance and content gaps, but model accuracy and generalization must be evaluated on representative real-world data and your downstream benchmark.
Is synthetic data as good as real data for computer vision?
Synthetic data works best alongside real data, not as a full replacement. Combining synthetic renders with real images gives a model trained on large volumes of high quality training data, plus enough real examples to stay grounded in the deployment domain. Rendered frames also cover rare situations that are hard to capture in real life, so real and synthetic data together beat either source alone. Accuracy and sim-to-real transfer are measured on your real-world benchmark, not assumed.
How is synthetic data used to train computer vision models?
Synthetic data gives deep learning models annotated data to learn from without hand-labeling, so computer vision models train on exact labels from the first run. Rendered from virtual environments, it provides diverse datasets that cover real world scenarios and rare edge cases which are hard to collect from real footage. Because the images come from computer simulations of 3D scenes, you control the content and can match the statistical properties of your target domain.
Most teams use it for AI model training alongside real images, training models on this artificially generated data as a complement to real captures and data augmentation, not a replacement. Evermotion supplies the prepared scenes; the renders run in your own pipeline.
Does synthetic data remove manual labeling?
Yes. Labels are generated from the 3D scene, so masks, boxes, classes and depth come out pixel-perfect - perfectly labeled datasets straight from the render, with no manual labeling. Producing such data by hand is slow and adds label noise, so reading the visual data and its labels directly from the scene removes both the annotation cost and that noise.
How is synthetic data for computer vision generated?
Synthetic data generation for computer vision follows two broad approaches that create synthetic data in different ways: rendering from explicit 3D scenes, or synthesizing synthetic images from generative models. Evermotion scenes are built for 3D rendering, which reads exact, pixel-perfect ground truth straight from the scene and keeps full control over object positions, camera angles and lighting.
| Method | How it works | Ground-truth labels | Control |
|---|---|---|---|
| 3D rendering (Evermotion scenes) | Path-traced images from explicit 3D scenes | Exact, pixel-perfect, read from the scene | Full: objects, cameras, lighting |
| Game-engine or procedural | Rule-based scene assembly from an asset library, driven by mathematical models | Exact, from the engine | High, bounded by available assets |
| GANs | Synthesize images from a learned distribution | Approximate, hard to guarantee | Limited |
| Diffusion / generative AI | Denoise noise into images from a prompt | Approximate, needs extra labeling | Prompt-level only |
| Neural style transfer | Restyle existing real images towards a target domain | Inherited from the source, not new | Low |
How is this different from GAN or diffusion-generated images?
Evermotion renders from explicit 3D scenes, not from generative AI. Generative adversarial networks, diffusion models and neural style transfer synthesize pixels from learned distributions, which makes exact ground truth hard to guarantee. A rendered 3D scene knows every object, material and camera, so labels are exact and you keep full control over object positions, camera angles and lighting.
Can synthetic data reduce bias and privacy concerns?
Yes. Synthetic scenes contain no real people or captured locations, so there are no privacy concerns carried over from real footage. Because you set the content distribution, you can balance under-represented cases and reduce bias instead of inheriting whatever real data collection happened to capture.

































