Skip to content
Tim Frenzel

// Insight

Borrowed structure: training data from sources you cannot use

13 min read
synthetic-datatraining-datavalidationdata-engineering

Every team that tunes a model on internal data eventually hits the same wall. The records that would make the training set realistic are the records you cannot move. They sit behind a contract, a jurisdiction, a client agreement, or an internal policy that nobody is going to relitigate for your fine-tune. The three usual substitutes all fail in the same direction. Public corpora have the wrong shape. Tasks invented by a language model are plausible in tone and wrong in structure. A hand-built simulator has structure and no realism.

There is a fourth option. A national statistical agency, an audit-software firm and two research groups have converged on it independently. Reduce the real source to a bounded statistical summary. Then build whatever you need on top of that summary, freely, because the summary is the only thing that ever touched the protected data. The architecture is worth understanding even if privacy is not your problem, because the summary step is the only stage in the pipeline that comes with a published error budget.

What follows is how it works, where it silently fails, plus the acceptance tests that catch the failures before you spend the compute.

What is known at each stage
BOUNDED SUMMARYstatistics36 mechanismsprivacy budgetrho 0.2505GENERATIVE RULESgraph250 nodes, 1,000 edgeserror budgetnone existsARTEFACT6 binders90 graded taskserror barsstill none
Amber marks a budget appearing. One appears at the summary and is never spent again. The other never appears, so the artefact ships with no error bars.

Are you building an evaluation set or a training set?

This is the fork that decides everything downstream. Conflating the two is the most expensive mistake available here.

An evaluation set can construct its ground truth. FinancialAuditBench, published in September by a team from Illinois, Stanford and an audit firm, generates six synthetic companies and then has licensed accountants write the answer key. Because the truth is authored, the data only has to be plausible enough to pose the task. It does not have to match any real distribution. The payoff is a benchmark with real discriminating power: the strongest model completes every rubric check on 69.17% of tasks, and on only 44.44% when the same task is run eight times.

A training set cannot construct its ground truth. It has to carry the distribution. The part of that distribution you usually want is the tail. This is where the evidence turns hostile. Stadler and colleagues ran the first quantitative comparison of synthetic data publishing against traditional anonymisation and found that generative models with formal privacy guarantees “do not preserve the fine-grained statistical patterns needed for outlier analysis,” naming financial fraud as one of the two use cases that motivate synthetic data in the first place. Synthetic data is weakest at exactly the rare events most teams are trying to augment, which inverts the usual reason for reaching for it.

Two artefacts, two acceptance tests
Evaluation setTruth is authoredPlausibility is enoughRubric decides pass69.17% at one try, 44.44% at eightTraining setTruth is inheritedFidelity is the productTails decide valueWeakest on rare events
An eval set authors its answer key, so plausibility is enough. A training set inherits its distribution, and inherits whatever the summary failed to carry.

So decide first. If you are building an eval harness, the bar is internal consistency and a defensible answer key. If you are building training data, the bar is distributional coverage. The amber panel is your risk.

How do you build one?

The summary comes first. The reason it is cheap is structural. A query whose cells are mutually exclusive and exhaustive costs the same to measure whether it has two cells or two thousand. AIM states it plainly: “we can measure every cell of a marginal at the same privacy cost of measuring a single cell.” The 2020 Census exploited the same property to release roughly 2,600 cells per geographic node at an L2 sensitivity of at most the square root of 2. FinancialAuditBench spends a total zero-concentrated budget of about 0.2505 across 36 mechanisms, 24 of them discrete-Gaussian histograms at sensitivity 1, plus six quantile mechanisms returning 7 quantiles each.

Two design notes matter more than the budget. First, the privacy unit. FinancialAuditBench protects one entire audit binder, meaning every document for one company-year, which is a far coarser unit than one record and therefore a far stronger statement at the same numbers. A budget quoted without its unit is not a number at all. Never copy an epsilon out of a paper.

Second, the summary is not the artefact. What turns 36 statistics into usable data is a generator built from domain rules, the same move the synthetic-company piece traced through a full pipeline. The audit case uses a computational graph of roughly 250 nodes and 1,000 edges with three node types: source nodes holding chosen inputs, conditionally sampled nodes drawing from distributions set by their parents, and deterministic nodes applying accounting identities. Industry picks a payroll-to-revenue ratio. That ratio times revenue gives payroll. The ledger ties because arithmetic made it tie. Generated feasibility tests then reject infeasible states, such as negative month-end balances, and trigger a resample.

That division of labour is the transferable part. The statistics carry the realism while the deterministic rules carry the consistency, with neither able to substitute for the other. There is a third thing neither of them carries. The audit case accounts for six financial ratios, 16 features describing financial records, and summary statistics for two transaction types. Everything else about those six companies is invented. Your artefact is faithful to the quantities you chose to measure and fabricated everywhere else, which makes the choice of what to summarise the highest-leverage decision in the build. Sim-PE makes the point at the opposite extreme by replacing the whole generator with a non-neural simulator steered by a single noisy nearest-neighbour histogram. On MNIST the simulator alone scores 11.6%. The prior method reaches 27.9%. The simulator steered by one noisy histogram reaches 89.1%. Worth noting against the hype: the best dedicated baselines still beat it at 94.0%. The authors say so themselves. That ablation is also the evidence for the claim this piece opened with. A simulator on its own carries structure and no realism, which is what 11.6% looks like. The bounded summary is the part that supplies the realism.

Here is the part nobody puts in the README. Every one of these systems invokes the same theorem to justify the build step, which is that post-processing a differentially private release costs no further privacy. That theorem is true. It is also a statement about privacy only. The Census team, who have the only national-scale post-mortem in existence, are explicit that the damage in their system came from the step after the summary. Controlling the biases and data-dependent errors, they write, are “properties due entirely to post-processing and not the DP mechanisms” and “the hardest implementation challenge.” A footnote puts it beyond doubt: the problem “is a direct consequence of requiring TDA to produce microdata. It is not due to the use of DP mechanisms.”

The privacy ledger closes at the summary and the error ledger never closes at all.

The same paper concedes that its error distributions are “data-dependent in a fashion that cannot be expressed in closed form and depends on confidential data values that cannot be published.” You cannot compute your own accuracy. The reason is the same wall that let you publish in the first place.

How much real data has to stay in?

Most teams are not locked out entirely. They hold a small real set and want more of it, which makes this a mixing problem. Three results bound the answer.

The mixing ratio has a floor. The FCA’s expert group reported a controlled substitution experiment: against a real-data-only benchmark, a 50/50 mix cost 2.5% of accuracy, a 70/30 synthetic-to-real mix cost 5.0%, and training on synthetic data alone cost 32.0%. Their conclusion was that “a threshold of at least 30% real data” was needed to hold accuracy. Treat that as practitioner-reported, because the report names no firm, no model family and no definition of accuracy, and calls its own findings preliminary. The shape is still informative.

Accuracy against a real-only benchmark as real data is added back (%)
5-35real-only benchmark-32.00%-5.030%-2.550%0100%
The penalty is mild while any real anchor remains and collapses when the last of it goes. One use case, self-described preliminary, so read the shape and not the decimals.

Never discard the real data as you iterate. The recursive-training literature separates two regimes cleanly. Replacing each generation’s training data with the previous generation’s output produces catastrophic collapse. Accumulating instead, keeping the original real data and adding each generation to it, prevents variance divergence and is provably stable. The engineering instruction is to append, never overwrite.

The third result is the one that should change behaviour tomorrow, because it contradicts an almost universal habit. Filtering synthetic samples by how well they match your reference data makes things worse when the reference is a biased slice, which it always is when the slice is one institution. After 10 recursive iterations against a skewed local reference, two standard quality-selection methods were beaten by picking at random.

Quality filtering against a skewed reference, CIFAR-10 FID after 10 iterations, lower is better
CovMatch115CenterMatch116Random106K-means102
Selection against a biased local reference underperforms random selection. Recall falls from 0.48 to 0.35, because samples that deviate from the local reference include the valid rare modes.

The mechanism is a confirmation-bias loop. Samples that deviate from the local reference get discarded. Valid but underrepresented modes are the ones that deviate. The authors name the consequence directly: data scarcity is amplified into persistent tail pruning. Their own fix is instructive and cheap. Compute the reference as a Wasserstein barycenter across silos. One silo’s manifold is the thing causing the damage. The noise needed to compute a barycenter privately costs almost nothing in output quality.

How do you know it is good enough?

Validation here has three separate jobs. Teams routinely run only the third.

Does it tie internally? This is the cheapest and most neglected check. Deterministic accounting nodes and generated feasibility tests give you it for free at build time, which is why the audit case can ship 179 files per engagement whose balances agree across every document.

Does it match reality? This is where the standard metric fails. Train-on-synthetic-test-on-real is the near-universal utility measure. It is also close to blind to structure. Across 13 generators, TabStruct finds that this measure correlates with global conditional-independence structure at a Spearman rho of 0.14. A structural measure correlates at 0.84.

Two ways to rank the same five generators
TSTRCIReal data0.991.00SMOTE0.920.30CTGAN0.800.08TVAE0.780.40Bayesian net0.410.35
TSTR is train-on-synthetic-test-on-real, CI is global conditional-independence structure. Rank by TSTR and you pick SMOTE or CTGAN, which sit near the bottom on structure.

Read the two columns against each other. The two generators that look best on the conventional measure, at 0.92 and 0.80, sit at 0.30 and 0.08 on structure. Selecting a generator on train-on-synthetic-test-on-real actively steers you toward the ones that reproduce the target relationship and discard the relationships between features. The structural check costs 0.64 seconds per 1,000 rows and returns the same generator ranking as a far more expensive configuration, which makes it roughly half the price of running the conventional metric properly.

Does it help the model? Only a real holdout answers this. The mixing floor above tells you what to expect.

Then there is the fourth job, which cannot be done. You cannot price the error your build step introduced. AIM is the one mechanism in this corpus that partially escapes, deriving per-marginal error bars from the private measurements themselves and requiring no additional budget to do it, though the guarantee covers only aggregate error per marginal and says nothing about any individual cell. AIM’s own framing of the field is the honest summary: existing methods “fail to produce accurate synthetic data” and “provide no way for end-users to detect the inaccuracy,” so differentially private synthetic data generation “remains an unsolved problem.”

The audit team hit the same wall from the other side. Their admission is the most useful sentence in the paper. Their accountants “could not always confidently assess whether these modeled values were realistic without being able to consult reference documents containing privacy-sensitive data.” The boundary that made the artefact publishable is the boundary that stops anyone validating it.

Do not expect governance to catch this for you. The UK regulator’s expert group published its final report on synthetic data governance in August 2025. It mentions differential privacy exactly once, misspelled, and defines it as the act of “introducing noise into the dataset.” That is not what differential privacy is. The gap between that sentence and a sensitivity analysis is the whole engineering problem.

What to decide

Decide the artefact type before the architecture. An eval set authors its truth and needs plausibility. A training set inherits its truth and needs tails. The audit benchmark works because it chose the first.

Put the realism in statistics and the consistency in rules. Mutually exclusive, exhaustive marginal groups make cell count nearly free, which is how 36 statistics underwrite 90 graded tasks, and how roughly 2,600 cells per node cost sensitivity of at most the square root of 2.

Hold at least 30% real data, and append every generation to it. Synthetic-only training cost 32.0% of benchmark accuracy in the one controlled comparison available. Accumulating is provably stable. Replacing the pool collapses it.

Stop filtering against your own slice. Two standard selection methods lost to random choice after 10 iterations against a skewed reference, at a cost of 0.13 in recall. Use a cross-silo barycenter, or do not filter.

Gate on structure as well as utility. At a Spearman rho of 0.14 between them, a generator that passes the conventional check tells you almost nothing about whether inter-feature relationships survived. The structural check runs in 0.64 seconds per 1,000 rows.

The pattern this all serves is older than the tooling. Dwork and colleagues established in 2006 that you can answer a bounded set of questions about a sensitive dataset with calibrated noise and a provable limit on what leaks. Twenty years later the interesting consequence for an engineer has little to do with leaking. It is that a bounded summary is the only artefact in your pipeline that arrives with a number attached to its error.

The summary step is the only one anyone can price. Every transformation after it is free in privacy and uncosted in accuracy. The discipline is to keep that chain short and to measure what you still can.

Working on AI that needs to ship?

I help funds, fintechs, and data teams take AI from prototype to production.