Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 21 min read

Synthetic Data Generation - Complete Deep Dive

Stage 6 - Customization Lesson 28 of 36

Prerequisites: Evals, Prompt Engineering, Fine-Tuning Used in: Distillation and Small Models, LLM-as-Judge, AI Security


What is Synthetic Data?

Synthetic data is data a model generated rather than data a user produced. In practice that means using an LLM to manufacture the inputs, the labels, or both, for training, evaluation, or augmentation β€” because the real data does not exist yet, is too rare, is too sensitive to use directly, or is too expensive to label.

It is a legitimate and widely used technique. It is also the fastest way to build a system that looks measured and is not. Synthetic data carries a systematic bias that is invisible from inside: it is cleaner, politer, better-spelled, and more on-topic than anything a real user will ever send you. A model trained on it learns a world that does not exist. A model evaluated on it earns a score that describes a world that does not exist. The techniques in this page are mostly defences against that one problem.

Real-world analogy: A flight simulator. Pilots train thousands of hours in one, and it is unquestionably the right way to practise an engine failure you cannot safely stage in a real aircraft. But no regulator certifies an airframe on simulator hours alone, because the simulator only contains the physics somebody thought to model. Synthetic data is the simulator: excellent for training and for rehearsing the rare case, and never sufficient for the sign-off. Real traffic is the flight test.


The Legitimate Uses

1. Bootstrapping an eval set before you have traffic. The genuine chicken-and-egg problem. You cannot mine production failures on day one, and shipping without any eval is worse than shipping with an imperfect one. Generate plausible questions against your actual corpus, review them by hand, and treat the set as scaffolding that real traffic will replace β€” see Evals for what it is replacing it with.

2. Hard negatives for retrieval training. One of the strongest uses, because the task has a built-in correctness check. To train or fine-tune a reranker you need examples that look relevant and are not, and mining those from a corpus by hand is slow.
πŸ’‘ A reranker is a second-pass model that re-scores a short list of already-retrieved passages, pushing the genuinely relevant ones to the top. Generation plus a filter that verifies β€œthis passage does not answer this question” produces exactly the confusable pairs the model needs to learn a boundary.

3. Augmenting rare classes and edge cases. Your fraud category has a handful of real examples, your refund-outside-the-window case appears monthly, your one genuinely ambiguous intent is a fraction of a percent of traffic. Generating variations of the rare class is standard practice and usually beats leaving the class unlearnable β€” as long as you remember the model now sees a class whose distribution you invented.

4. Training data for distillation. Structurally the cleanest use, because the teacher’s output is precisely the target the student is supposed to imitate.
πŸ’‘ In distillation the teacher is the big expensive model and the student is the small cheap one you train to copy its answers. The mechanics are in Distillation and Small Models; the caveat there is the same one as here, which is that the student inherits the teacher’s errors along with its skills.

5. Privacy-safe stand-ins for sensitive records. Generating records that resemble real medical, financial, or personal data so engineers can build and test against realistic shapes without touching the real thing. Valuable, and more subtle than it appears β€” see the privacy section below, because β€œsynthetic” does not automatically mean β€œcarries no information about the source.”

6. Adversarial cases for red-teaming. Generating jailbreak attempts, prompt injections, and policy-boundary probes at a scale no human team can write by hand. The generator’s creativity is an asset here rather than a liability, and it is the one use where being unrealistic barely matters. Pair with AI Security.


Generation Techniques

Seed from real examples and vary them. The highest-quality approach available, and the one to default to. Take a small set of genuine inputs, then ask for paraphrases, harder variants, differently-phrased versions, and versions with the same intent but a different tone. Because the seeds carry the real distribution’s shape, the output stays anchored to how people actually write. Even 20 real seeds change the output character dramatically compared with generating from nothing.

Templating with controlled variation. Write a structural template and fill slots from real inventories β€” product names from your catalogue, cities from your service area, amounts from realistic ranges. This gives you exact control over coverage and guarantees a specified combination appears, which is exactly what you want for a systematic eval grid. Its weakness is the mirror of its strength: templated phrasing is uniform, so it must be mixed with free generation rather than used alone.

Persona or scenario conditioning. Instead of β€œgenerate 200 customer questions”, generate 20 questions each as five different people. Make them a first-time user in a hurry, a frustrated user on their third contact, an expert using internal jargon, a user typing on a phone with typos, and a user who has misunderstood how the product works. Conditioning on a scenario is the cheapest large gain in diversity available, because it forces the generator out of its default register.

Self-instruct style bootstrapping. Start with a few seed tasks, have the model generate new task instructions, generate responses to those, filter, add the survivors to the pool, and repeat. It scales enormously from a small seed set. It is also where collapse and self-reinforcement live, since each round is drawing from a distribution the previous round produced β€” so the filter between rounds is the load-bearing component, not the generator.

# The shape that matters: real seeds, varied conditioning, dedup, and provenance.
for seed in real_seeds:                      # anchored in actual traffic
    for scenario in SCENARIOS:               # forces register variety
        cand = generate(seed, scenario, temperature=HIGH)
        if not schema_ok(cand):              continue
        if near_duplicate(cand, pool):       continue   # the step most teams skip
        pool.add(cand, provenance={"seed": seed.id, "scenario": scenario, "gen": MODEL_ID})

Provenance on every generated row is not bookkeeping. It is what lets you delete a whole cohort later when you discover the generator or the prompt that produced it was flawed.


Diversity Is the Whole Game

Naive generation collapses. Ask a model for 500 customer support questions and you get 500 items that are superficially different and structurally identical: similar length, similar politeness, similar sentence shape, clustered on the same handful of intents, all grammatical, all on-topic. It looks varied and it is not, which is the dangerous part β€” the variety is visible at the level of wording and the sameness is at the level of distribution.

What to do about it, in rough order of effect: raise sampling temperature and generate in many small independent batches rather than one large request, because a single response drifts toward internal consistency.
πŸ’‘ Temperature controls how much randomness the model allows when picking each next word, so a higher setting gives you more varied wording. Condition on personas and scenarios. Seed from real examples so the anchor is real. Generate deliberately against an axis grid β€” intent by tone by length by formality by error type β€” so coverage is designed rather than hoped for. Use more than one generator model, since each has its own stylistic attractor.

And then measure it rather than eyeballing it. Embed the generated set and look at how tightly it clusters, and compare that spread against the spread of a real sample β€” the machinery is in Embeddings.
πŸ’‘ An embedding turns a piece of text into a list of numbers, positioned so that texts meaning similar things land close together. Check the distribution of length, vocabulary size, intent labels, and question type against real traffic. Near-duplicate rate within the set is a direct diversity signal: a high rate means you paid for tokens and bought one example many times.


The Failure Modes

This is the heart of the page. Every one of these has shipped a system that looked measured and was not.

Distribution mismatch

Synthetic data is cleaner, politer, and more on-topic than real traffic, in a consistent direction. Real users send fragments, typos, no punctuation, multiple questions in one message, questions about things you do not do, pasted logs, mixed languages, and hostility. A generator asked for β€œcustomer questions” writes well-formed ones.

Two distinct harms follow. Train on it and the model is tuned for inputs it will rarely see, so it degrades on the messy ones that dominate. Evaluate on it and the score describes performance on a world that does not exist β€” which is worse, because it is confidence rather than just weakness. A team with a 92 percent synthetic eval score and a stream of user complaints is reading a real number about the wrong population.

Model collapse and self-reinforcement

Train a model on its own outputs, repeatedly, and quality degrades in a characteristic way: the distribution narrows, unusual-but-valid patterns fall out first, and the output converges on the model’s most typical behaviour. The tails go first, and the tails are where most of your hard cases live.

The loop does not have to be that explicit to bite. Generating data with model A, training model B on it, then using B to generate more data for the next round is the same loop with extra steps. Defences: always mix real data in rather than replacing it, keep the number of self-generated rounds small and deliberate, and re-anchor each round on fresh real seeds instead of on the previous round’s output.

Inherited label noise

When the generator produces both input and label, every systematic mistake the generator makes becomes a training target. If it misclassifies a category boundary 8 percent of the time, your student learns that error as if it were ground truth β€” and learns it confidently, because it appears consistently.
πŸ’‘ Ground truth is the answer you have decided is definitely correct, and the thing every other answer gets scored against. Worse, the error is now invisible: it is in the labels, so nothing downstream disagrees with it. This is why any generated label used for training deserves a human-verified sample estimating its error rate. It is also why the generator should be your strongest available model, even when the student will be small.

Validating synthetic data with the model that produced it

The specific trap, and the most common one. You generate a set with a model, then ask the same model to judge whether each item is high quality and realistic. It says yes, because you are asking a distribution to rate how likely its own samples are. The filter is correlated with the generator, so it removes almost nothing β€” and what it does remove is often the unusual-but-valid items you most wanted to keep.

A judge that is worth anything is a different model, and one whose agreement with human labels you have actually measured. That measurement is the entire subject of LLM-as-Judge, and it is a prerequisite here rather than a nice-to-have.


The Quality Control Pipeline

flowchart LR
    SEED[Real seed examples] --> GEN[Generate with varied personas and scenarios]
    GEN --> FIL[Automated filters - schema and dedup and near duplicate check]
    FIL --> DIV[Diversity check against a real traffic sample]
    DIV --> HUM[Human review of a sample]
    HUM --> MIX[Mix with real data - never replace it]
    MIX --> TRAIN[Training or augmentation set]
    REAL[Held out real traffic slice] --> EVAL[Evaluation - the source of truth]
    TRAIN --> EVAL

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class SEED client
    class GEN async
    class FIL,DIV service
    class HUM client
    class MIX service
    class TRAIN data
    class REAL edge
    class EVAL data

The controls that turn raw generation into something usable, in the order they appear:

  1. Automated filters first, because they are free. Schema validity, length bounds, required fields, language detection, exact deduplication, and near-duplicate detection β€” the step teams skip and the one that most often reveals a set is a fraction of the size it appears to be.
  2. Diversity measured against real data, not against itself. Length, vocabulary, intent mix, and embedding spread compared with a real sample.
  3. Human review of a sample. Not every row β€” a sample large enough to estimate the error rate and to notice systematic weirdness. An hour of reading 50 generated rows reliably finds the artefact that would otherwise be trained into the model.
  4. Mix, never replace. Synthetic data supplements real data. A blend keeps the real distribution’s shape while filling the gaps you generated for, and the mixing ratio is itself a parameter to tune by eval rather than pick by feel.
  5. Hold a real-data evaluation slice as the source of truth, kept out of the training mix entirely. Note the diagram: the eval node’s authority comes from the real branch, not the synthetic one.

The Hard Rule

Synthetic data is acceptable for training and augmentation. A purely synthetic eval set will lie to you.

The asymmetry is not arbitrary. Training is robust to some distribution error, because you are shaping a model that will then be measured against reality β€” and if the synthetic training data hurt, the real eval catches it. Evaluation has no such backstop. It is the backstop. An eval set drawn from a generator measures your system’s performance on generated text, and reports it as performance, with no mechanism anywhere in the pipeline to notice the substitution.

So: bootstrap with synthetic cases when you have no traffic, absolutely. Then treat every one of them as provisional, and replace them with real cases as fast as traffic allows. The moment you have production requests, the eval set’s centre of gravity shifts to them permanently, and anything about your users’ behaviour that you learn from a synthetic case is a hypothesis rather than a finding.


Privacy Is Not Automatic

Synthetic data is often justified on privacy grounds, and the justification is frequently made too casually. Data generated from sensitive records can leak information about those records. A generator conditioned on real patient notes, prompted with real examples, or fine-tuned on a sensitive corpus can reproduce distinctive details from it. That means verbatim quotes, rare combinations of attributes, an unusual case that only one person matches. β€œThe model wrote it” is not a legal or ethical anonymisation argument.

Treat it accordingly. Understand whether the generator was trained, fine-tuned, or merely prompted on the sensitive data, because each creates a different exposure. Scan output for memorised spans from the source. Watch for rare-attribute combinations that re-identify an individual even with names removed. Keep the synthetic set inside the same access controls as its source until someone qualified has signed off that it can leave them. And involve whoever owns privacy review at your organisation rather than deciding this from an engineering seat β€” the question of whether a derived dataset is personal data is not an engineering question.


Use Case Scorecard

Use case Technique Main risk
Bootstrap an eval set with no traffic Generate against your real corpus, then review every case by hand The score measures a world that does not exist; must be replaced by real cases
Hard negatives for a reranker Generate confusable passages, filter with a relevance check Negatives that are too easy, teaching a boundary that is not the real one
Augment a rare class Seed from the few real examples and vary tone and phrasing You invented the class distribution, so the model over-trusts your version of it
Distillation training data Teacher generates on real or realistic inputs The student inherits the teacher’s systematic errors as ground truth
Privacy-safe test records Template from real schemas with synthesised values Leakage of distinctive details; not anonymous by default
Red-team and adversarial cases Adversarial prompting plus mutation of known attacks False confidence - you covered the attacks you imagined
Instruction tuning at volume Self-instruct bootstrapping with per-round filtering Collapse and narrowing across rounds; label noise at scale

Bad to Good to Great

Bad - generate a thousand examples and call it a dataset

One prompt, one large batch, no filtering, no deduplication, no human read. The set is far less diverse than its row count suggests, its distribution is politer and cleaner than reality, and if it is also serving as the eval set, every number produced from it is confidently wrong about your users.

Good - seed from real examples, filter, and mix with real data

Real seeds, schema and duplicate filters, a blend rather than a replacement. A reasonable working setup. What is still missing: no diversity measurement against real traffic, no human sample review, and no estimate of label error from the generator. There is no provenance either, so a bad cohort cannot be retracted wholesale. And the one that matters most - nothing guarantees the eval slice is real.

Great - generated, filtered, sampled by a human, mixed, and evaluated on real data only

  1. Seeded from real examples, with persona and scenario conditioning and several independent batches.
  2. Automated filters for schema, length, language, exact duplicates, and near-duplicates.
  3. Diversity measured against a real sample β€” embedding spread, length, vocabulary, intent mix.
  4. A human reads a sample and the estimated error rate is written down, not assumed.
  5. Judged by a different model than the generator, with measured agreement against human labels.
  6. Provenance on every row β€” generator, prompt version, seed, scenario β€” so a flawed cohort can be deleted wholesale.
  7. Mixed with real data at a ratio tuned by eval, never used as a replacement.
  8. Evaluated only on a held-out real slice, which stays the source of truth permanently.

When to Use

βœ… Generate synthetic data when:

❌ Do not use synthetic data when:


Common Interview Questions

Q1: When is synthetic data appropriate, and when is it not?

It is appropriate for training and augmentation β€” bootstrapping a first eval set before you have traffic, generating hard negatives for retrieval where a correctness filter exists, augmenting a rare class, producing distillation data, and creating red-team cases at scale. It is not appropriate as the sole basis for evaluating a live system. The asymmetry is structural: bad synthetic training data gets caught by a real evaluation, but bad synthetic evaluation data has no backstop, because it is the backstop. So a real held-out slice stays the source of truth, and synthetic eval cases are scaffolding I plan to replace with production traffic as soon as it exists.

Q2: What is distribution mismatch and why does it matter more than it sounds?

Generated data is cleaner, politer, better-spelled, and more on-topic than real traffic, consistently and in one direction. Real users send fragments, typos, several questions at once, requests about things you do not do, pasted logs, and hostility. Training on the clean version tunes the model for inputs it will rarely see, so it degrades on the messy ones that actually dominate. Evaluating on it is worse, because the score is real but it describes the wrong population β€” which is how a team ends up with a high eval number and a queue of complaints. The fix is to anchor generation on real seeds, condition on messy personas deliberately, and compare the synthetic distribution against a real sample rather than trusting that it looks varied.

Q3: What is model collapse and how do you avoid it?

Training a model repeatedly on its own output narrows its distribution. The unusual-but-valid patterns go first and the output converges on the model’s most typical behaviour, which is exactly backwards from what you want, since the tails are where the hard cases live. It also happens indirectly β€” generate with model A, train model B, generate the next round with B β€” so it does not require an obvious loop. I avoid it by always mixing real data in rather than replacing it, and by keeping self-generated rounds few and deliberate. Every round re-anchors on fresh real seeds rather than on the previous round’s output. And I watch diversity metrics across rounds, so narrowing shows up as a number instead of a surprise.

Q4: Can you use the same model to generate and to validate synthetic data?

No, and it is the most common mistake in this area. Asking a model whether its own samples are realistic is asking a distribution to rate the likelihood of its own output. It says yes, so the filter removes almost nothing - and what it does remove tends to be the unusual-but-valid items worth keeping most. A useful filter has to be uncorrelated with the generator: a different model as judge, with its agreement against human labels actually measured, plus cheap deterministic checks that do not involve a model at all, plus a human reading a sample. If a generated label is going into training, I also want a human-verified estimate of the generator’s error rate, because inherited label noise gets learned as ground truth and is invisible afterwards.

Q5: Is synthetic data a solution to privacy constraints?

Not by itself. Data generated from sensitive records can carry information about them β€” a generator that was fine-tuned on, prompted with, or conditioned on a sensitive corpus can reproduce distinctive spans and rare attribute combinations that re-identify an individual even with direct identifiers removed. β€œA model wrote it” is not an anonymisation argument. What I would actually do, in order. First, establish whether the generator was trained, tuned, or only prompted on the sensitive data, since each is a different exposure. Then scan output for memorised spans from the source, check for rare-attribute combinations that uniquely match someone, and keep the synthetic set under the source’s access controls until it is cleared. Last, get sign-off from whoever owns privacy review, because whether a derived dataset counts as personal data is not an engineering call.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access