The prediction that made synthetic data sound easy
In 2021, Gartner predicted that by 2024, 60% of the data used to develop AI and analytics projects would be synthetically generated. The prediction was repeated across the industry and reflected the enthusiasm around synthetic data at the time, including Gartner’s research note “Forget About Your Real Data — Synthetic Data Is the Future of AI” (Ramos & Subramanyam, 2021), which was later cited in academic work including Pashevich & Hansen, 2026.
The hype made synthetic data sound easier than it was, and not because synthetic data doesn’t work. It made generation sound like a push-button process. Press a button, get a dataset, skip the people. Anyone who has tried to produce good synthetic data knows how much human review it takes.
Most companies still aren’t using it
Perforce’s 2026 survey of 518 enterprise leaders found that 51% don’t use synthetic data in their AI and ML workflows. Of those, 23% have never used it, 15% experimented and stopped, and 13% evaluated it and decided it didn’t meet their needs. Only 28% use it extensively. (Perforce sells synthetic data and data masking products, so read it as a vendor survey.)
That isn’t surprising. Synthetic data is still relatively new, and many companies are still working on getting the data they already have AI ready. The 15% who tried it and stopped are the interesting group: it’s easy to generate data, and much harder to generate data you can trust.
We’ve seen this pattern before. AI-assisted labeling was supposed to dramatically reduce the need for human annotation. Instead, in our experience, demand for labeled data has kept growing. AI has gotten much better at prelabeling, but companies are also constantly pushing the boundaries of what labels they are attempting to create.
Synthetic data is data labeling with different tools
Think about what a person does in a labeling project. They look at an image, draw a bounding box, pick a class, and move on. Reviewers check a sample, and corrections feed back into the next batch.
Now think about what a person does in a synthetic data project. They write a prompt. They set the variables — the setting, the lighting, the weather, what has to be in the scene. They look at what comes back, reject what’s wrong, adjust the prompt, and try again.
The tools changed, but the loop didn’t: define, produce, review, correct, approve. In both cases, the output still needs validation before you can trust it. Prompts are the new bounding boxes.
How it works at MLtwist
We launched our synthetic data technology on the foundation of our data labeling platform, and customers are using it today. It works across multiple models and modalities: text prompts generate synthetic images, those images generate synthetic video, and LLMs review the results. A person needs to step in and ensure that the quality remains high every step of the way.
For video, the work starts in a spreadsheet. Each row sets the parameters of a scene: the key things that have to be in it, and the conditions around it. From there, the workflow looks a lot like labeling:
- AI prelabeling. Models generate a first pass at the scene, the way they prelabel images in an annotation project.
- Labelers. People refine the scenes, generate videos, and identify the changes that are needed.
- Reviewers. A second set of eyes checks the work against the spec.
- The customer. They see and approve what goes into their set.
- Training or validation. Only then does the data go to a model.
A person approves every finished video. The LLMs work in tandem with those people, not in place of them. We’ve found they have different strengths: people spot the obvious problems quickly, while LLMs can surface less obvious inconsistencies that people overlook. Together they catch more than either does alone, and the people make the final call.
Physics and simulation engines have people too
Prompts aren’t the only way to make synthetic data. Physics and simulation engines can generate it too, and people are still in the loop. Instead of writing prompts, artists and engineers build, configure, and adjust virtual scenes.
Our forecast is that these approaches will increasingly converge. Artists will use prompts to build and modify simulated worlds, combining the control of simulation with the speed of generative AI. The line between prompted and simulated synthetic data will increasingly blur.
Comparing today’s video models
The models behind synthetic video are changing quickly. Three of the five models below launched since February 2026. Pricing is generally trending down as new generations arrive. But generating the video is only one component of the cost of producing a usable synthetic dataset. Here is how recent video models from Google and ByteDance compare. We use 720p because it is one resolution supported across all of them.
| Model | Launched | Longest single clip | 720p API cost |
|---|---|---|---|
| Gemini Omni 1.1 Flash | Aug. 2026 | 10 seconds | ~$0.10/sec* |
| Veo 3.1 | Oct. 2025 | 8 seconds | $0.40/sec |
| Veo 2 | Dec. 2024 | 8 seconds | $0.50/sec |
| Seedance 2.5 | July 2026 | 30 seconds | $0.231/sec** |
| Seedance 2.0 | Feb. 2026 | 15 seconds | $0.15/sec |
*Omni is priced by output tokens. The figure shown is the effective cost of generating 720p video.
**Seedance 2.5 price is for 720p, 16:9 video with no video input. BytePlus prices other resolutions and video-reference inputs differently.
Clip lengths are for a single generation call. Most of these models can extend a clip across multiple calls to make longer videos. Veo 3.1 is the standard tier, and Google’s Fast and Lite tiers cost less. Google’s Omni and Veo 3.1 prices are from the Gemini API pricing page. The Veo 2 price is as reported by TechCrunch in February 2025.
Pricing and capabilities checked October 2026.
What this means for your team
If you’re planning to use synthetic data, plan for it the way you’d plan for labeling. In both cases, the obvious cost is often the inexpensive part. Labor rates are often the smaller part of the cost. Labeling becomes expensive when labelers are set up for failure: waiting for data, using the wrong tools, working without prelabels, repeating unnecessary work, or sitting behind bottlenecks in the pipeline.
Synthetic generation works the same way:
- Optimize before you generate. The expensive part is generating videos you don’t need. Build the right starting scene, test that the model can reliably work with it, reuse successful outputs, and clip or extend existing video rather than regenerating entire scenes.
- Build a generation pipeline, not a prompt box. At scale, you need infrastructure that can move images, video, and text seamlessly into and out of models, preserve continuity between clips, extend scenes from their last frame, retry only what failed, and assemble short generations into longer videos.
- Set people up to succeed. Just as a labeler should have the right tool, data, instructions, and prelabels before they start, people working with synthetic data should have good starting scenes, reusable prompts, reference material, and clear ways to review and correct outputs. Every unnecessary generation or manual workaround adds cost.
- Define what “good” means up front. Guidelines matter as much for a prompt as for a bounding box. Knowing what makes a generated example useful prevents wasted generations and inconsistent datasets.
- Keep track of what’s real and what’s generated. Real and synthetic data can sit in the same training set, but every file should record which is which and how it was made.
There is an entire infrastructure behind long-form synthetic video generation that goes far beyond providing a text prompt and hoping for the best. Just as efficient data labeling is about much more than finding the lowest labeling rates, efficient synthetic generation is about much more than finding the cheapest model.
Synthetic data doesn’t take people out of the loop. It moves them to a different part of it.
Learn more about how synthetic generation works at MLtwist.