Skip to content
Tim Frenzel

// Insight

The synthetic company: real data has a price, synthetic data has a method

64 min read
synthetic-datadata-engineeringevaluationllm-trainingquant-finance

In August, Google bid $10 million for the corporate records of Spirit Airlines, a carrier in bankruptcy. The lot holds about 100 million emails and 500 million Teams messages, plus spreadsheets, code and operational data. TIME read the bid as a lab buying private white-collar work to build reinforcement-learning environments for agents. The sale still needs a judge. A hearing is now slated for 30 September.

A year earlier, a security group built the same kind of data from nothing. Chimera, published at NDSS 2026, runs three simulated 20-person companies, in technology, finance and medicine, for a month each. Language-model agents play the employees and carry out 15 insider attacks drawn from real cases across the three. Five security experts rated its logs 4.20 for realism on a five-point scale. Logs recorded from real people in a controlled exercise scored 4.25.

Real data has a price. Synthetic data has a method. Epoch AI expects frontier training to use up the stock of public human text between 2026 and 2032. Expert data is priced by the hour, with Mercor paying industry experts up to $200 an hour. The method is spread across statistics, data engineering and evaluation.

The August opener drew the architecture of a synthetic company for a retail chain. A deterministic scaffold owns counts, amounts, keys and rules. A semantic layer owns the words. This special goes one level down. It treats each component as an estimation problem with a known failure mode, a decision the team must make and an acceptance test that code can run. Ten primary sources carry the argument, nine papers and one engineering case study, each read in full and checked against its tables. Five executed code blocks carry the tests.

The worked example is a systematic research house, for two reasons. The first is that finance stresses every component at once. A market has one history. AQR’s Can Machines Learn Finance? makes the point with one number: a five-year forecast of market returns over 1926 to 2018 rests on 19 distinct observations. That history is licensed. Nasdaq’s data policy says models trained on its information may be fee-liable. Its clock leaks, because a language model that writes about the past may know how the past ended. Its evaluation is adversarial, because researchers search thousands of variants on the same data. The second reason is experience. I have spent 18 years building and reviewing models on quant desks. The failure modes below are ones I have watched in production.

The rule of the series still holds. The model proposes. Code decides. The engineering question is which code, at which boundary and with what tolerance.

Seven stages: what a model proposes, what code decides, what the team decides
Numbersfitted generatorretraining screenstructure, sample sizeRare eventstail scenariosconsumer's own scorereplay or generateRecordsrules from 100 rowsfull-table checktolerance per ruleWordslabels, QA pairscode, solver, samplechecker per fieldTimedated documentsvintage + lexical gatewhich date countsEvaluationcandidate modelsledger, deflationwhat counts as a trialServingagent submissionshost-side verifiermetric aggregation
My synthesis of the sections below. First box: what a model or a fitted generator proposes. Second: the code that accepts or refuses it. Third: the decision the team owns.

What does a synthetic company need beyond its generators?

The August opener split every pipeline into two layers. The deterministic layer draws counts and amounts, holds keys and types and runs validators. The semantic layer writes text inside the structure the first layer fixed. That split decides who writes each value. It leaves three pieces of shared state undefined. In the pipelines I have reviewed, those three decide whether a synthetic dataset can be audited a year after release.

A snapshot with two clocks. Every fitted generator is an estimate from a snapshot of real data. Each fact in the snapshot needs two timestamps: when it was true and when it became known. Database practice calls this a bitemporal table. A restated quarterly revenue has one valid period, the quarter, and at least two knowledge dates: the first report and the restatement. A generator fitted on the latest vintage learns the restated figure as if it had been known at the time. For vendor history the knowledge date comes from the vendor’s point-in-time stamp. In a research house this store is the point-in-time database, with the security master recording each instrument’s listing and delisting dates.

A rule registry. Rules are data. Each carries a version, the snapshot it was validated on, a tolerance and an enforcement mechanism. The rule-discovery system of the records stage accepted between 0 and 17 inequalities on one benchmark table depending on the data split. An unversioned rule set is therefore irreproducible.

A trial ledger. Every generated dataset, evaluation and candidate model is a trial. The ledger stores its seed, code revision, snapshot hash and result. Its first purpose is statistical. The deflation tests of the evaluation stage need the number of trials and the variance of their results. Only a complete ledger supplies both. Gençay’s agent framework routes every candidate through one evaluation entry point, which makes its ledger complete by construction. OpenFinGym applies the same principle to forecasts: it records each one first and scores it once the outcome arrives.

One versioned store, read as of a date, written back with lineage
licensed history, two clocksVERSIONED STOREsnapshot · rule registry · trial ledgeras-of read: facts known by t, rules in force at tgeneratorsrule gatetext + checkerclock gateevaluatorwrite backdataset or environment
My reference architecture. Every component reads the store as of a date and writes its output, verdict and seed back. In the pipelines I have reviewed, the ledger is the missing part.
A synthetic row that cannot name its snapshot, its rule set and its seed cannot be audited, reproduced or counted as a trial.

Four contracts govern the boundaries between components. I would write them down before the first model is trained.

  • Units. Money, prices and quantities travel as integers on their native grid: cents, ticks and contracts. A relative float tolerance scales with the price level and admits errors at high prices, as the tick-grid block of the records stage shows.
  • Time. Every row carries an as-of timestamp. No component reads a fact whose knowledge time is later than the row’s as-of time.
  • Provenance. Every generated row carries the generator version, the seed and the version of the rule set that accepted it. When OpenFinGym’s pipeline builds a task loader, it also emits a provenance record for the labels (source column and horizon) that a reviewer audits against the source mechanism.
  • Gate semantics. Each gate either refuses a row or logs it. The difference is written down.

PersonaLedger, Capital One’s transaction generator, shows why the last contract matters. Its paper describes infeasible plans being rejected with targeted feedback. Its shipped code retries an output only when it fails to parse into transactions, carries a malformed timestamp or mentions a deposit. Over-limit spending passes with a warning row, up to a third strike. A warning is a log entry. A refusal is a gate.

The tools that fill these slots are ordinary. Samplers are numpy’s random generator and scikit-learn’s mixtures. The factor diffusion model of the numbers stage is a U-Net of about one billion parameters. The rule-repair system of the records stage fixes output from a Gaussian copula in SDV, CTGAN, TVAE and TabDDPM. CFM labels headlines with Llama 3.1 70B behind Text Generation Inference under a JSON grammar and reviews them in Argilla. OpenFinGym keeps its forecast ledger in SQLite. None of these needs a platform. Each needs a contract at its boundary.

Ownership follows the stores. In my practice a data engineer owns the snapshot, its two clocks and the provenance fields. A statistician owns the generators and their acceptance tests. An ML engineer owns the semantic layer and its checkers. The domain owner signs the rule registry and the evaluation contract. Model validation owns the ledger. A research house already staffs all five roles: the market-data team, the quant researcher, the NLP engineer, the risk manager and the model-risk function.

How much can a generator add to one history?

A fitted generator is an estimator. It maps one real sample to a distribution. Every statistic computed on its output inherits the error of that map. Cetingoz and Lehalle, published in Quantitative Finance in July, state the consequence as a proposition about finite samples.

θ the true value of a statistic, θ̃ its value under the generator
a = θ̃ - θ, the bias the generator learned from n real observations
P(|U_ñ - θ| ≤ b) ≈ Φ((a+b)√ñ / rσ) - Φ((a-b)√ñ / rσ)
U_ñ the statistic on ñ synthetic points, b the tolerance

More synthetic samples tighten the distribution of U around the generator’s own value θ̃. Only more real data moves θ̃ toward θ. Once the bias exceeds the tolerance, the probability of landing within it falls to zero as synthetic data grows (their Corollary 2). The authors of the Diffusion Factor Models paper describe the same trade from the other side. Their generator is a data-dependent regularisation that lowers estimation variance at the cost of a small modelling bias. Two engineering rules follow. Generate many samples of the real length, since a synthetic interval narrower than the real sample’s sampling error measures the generator’s bias. Decide in advance which statistics the generator must reproduce, because it reproduces each one only as well as one history identifies it. No volume of synthetic data moves a bias that one history put into the generator.

Their second result explains where the bias lands. A generator trained to match a distribution favours high-variance directions. For Gaussian data, a rank-k linear generator under the 2-Wasserstein loss reproduces exactly the top k eigenvectors at its optimum (their Theorem 5, after Feizi and colleagues). Many consumers weight the opposite end. A mean-variance optimiser scales its exposure to each principal component by one over that component’s variance. A market-neutral book therefore lives in the low-variance components the generator learned worst. In my reading the same holds for any consumer that inverts a covariance, from a regression to a Mahalanobis anomaly score. The design rule is general. Decompose first. Hand the generator the structure you can name. Judge it in the directions its consumer weights.

Their own generator for 433 S&P 500 stocks is that rule as a recipe. It trains on 12 years of daily returns (3,020 days) and tests on the 605 days that follow.

X(t) = (β F(t) + Z(t)) ⊙ σ + μ
m = #{ λ_i > λ₊ },   λ₊ = (1 + √(d/n))²
F factors, Z residual, σ and μ per-stock scale and drift,
λ₊ the Marchenko-Pastur edge of the correlation matrix
  1. Standardise each stock and keep the eigenvalues above the Marchenko-Pastur edge. With d/n equal to 433/3,020 the edge sits at 1.90. It keeps 16 factors that explain 58.9% of variance. A fitted edge would have kept 44.
  2. Rescale each factor to unit variance before training. The generator never sees eigenvalue size, the quantity Theorem 5 says it would chase.
  3. Cluster the 16 factors into three groups by skewness, kurtosis, eigenvalue and memory statistics. Train one temporal convolutional GAN per cluster on 63-day windows pooled across its factors, which buys training data per parameter.
  4. Draw each stock’s residual independently from a fitted Student-t. The residuals carry the remaining 41.1% of variance.
  5. Recombine and draw 100 samples of 3,020 days each.

The released code adds a step the equation omits. After recombination it rescales every synthetic stock to its exact in-sample mean and volatility. The authors note the consequence for an equal-weighted portfolio. Its Sharpe ratio on generated data is 1.09 against 1.08 in sample, with 0.33 realised out of sample. Matching moments after generation removes their sampling variability from every synthetic sample. Any statistic driven mainly by those moments returns the in-sample estimate.

The statistic the authors care about is the Sharpe ratio of a cross-sectional mean-reversion strategy, long the past h days’ losers and short the winners. At five days the generator’s median was 0.08 against 0.75 in sample and 0.77 out of sample. A 63-day block bootstrap gave 0.70. The long-only version at five days agreed with the in-sample value, 1.02 against 1.15, with 0.67 realised out of sample. That is the pattern their theory predicts, since a long-only book loads on the market factor, the direction a generator reproduces best. Across the 16 look-backs of their Table 9, the bootstrap’s 95% interval contained the long-short out-of-sample Sharpe ratio 15 times and the generator’s 10 times, by my count.

Diffusion Factor Models, by Chen, Xu, Xu and Zhang, bring the decomposition into a diffusion model. Under a factor model the score function splits.

R = β F + ε                 d stocks, k factors, k << d
∇ log p_t(r) = s_factor(r, t) + s_linear(r, t)
s_factor: nonlinear, a network in the k-dim factor subspace
s_linear: linear on the complement, fitted with the loadings

Only the factor term needs a nonlinear network. It works in k dimensions. The linear term is fitted alongside it, with one noise variance per stock. The paper’s error bounds have a rate set by k. In practice the network is a U-Net of about one billion parameters whose bottleneck width is set to k: 16 in the simulation and an assumed 8 on real data. Their synthetic study measures when the generator helps. With 2,048 simulated stocks driven by 16 factors, the covariance error of the generator relative to the sample estimate is 0.885 with 512 training observations, 0.941 with 2,048 and 0.997 with 4,096. The gain exists while observations are scarce relative to dimensions. It is gone at about twice as many observations as stocks.

Covariance error, factor diffusion over sample estimate, by N/d
1.050.8sample estimate0.885N/d 0.250.9110.50.94110.99721.0074
Chen, Xu, Xu and Zhang, Table 2: d = 2,048 simulated stocks, 16 factors, N training observations. Below the dashed line the generator beats the sample covariance. Their empirical study runs at N/d near 2.5.

The empirical study sits past that boundary. It fits the 512 largest US stocks on five-year windows of about 1,260 days, a ratio near 2.5. The ratio is the first number I would compute before building such a generator. A gain reported beyond it is a result to explain before relying on it. The portfolio test then carries three decisions a team can copy. With 20 basis points of costs, mean-variance weights built from generated returns reached a Sharpe ratio of 1.361. Equal weight reached 0.486 and the sample estimates -0.129. Taking the covariance from the generator and the mean from history gave 1.090. The reverse gave 0.275. Most of the gain is a covariance estimate. Ledoit-Wolf shrinkage applied to the generated moments lowered the ratio to 1.186, consistent with a generator that already shrinks. Refitting annually in place of quarterly cut it to 0.797.

The acceptance test

Neither paper certifies a generator. Cetingoz and Lehalle propose a screen that can reject one. Fit the class to the real history and treat the fitted model as the truth. Draw one synthetic history of the real length. Refit the same class on it and check whether the refit recovers statistics that are now known. They call it regurgitative training. A statistician will recognise an identifiability check by parametric bootstrap. Applied to their own generator, the retrained medians landed within 0.21 of the truth at look-backs of 9 to 17 days and again at 33 and 37. At the other look-backs from 21 days on they sat 0.32 to 0.40 below it. At 1 and 5 days they overshot by 0.61 and 0.31. All 16 intervals contained the truth. The verdict therefore rests on the medians. The authors conclude that their model class “should probably not be used for time scales that are longer than one month”. Their screen uses one refit and no numeric criterion. My version repeats the refit and adds both.

def retrain_test(fit, sample, stat, real, tol, reps=20, n=100, seed=0):
    # identifiability screen after Cetingoz and Lehalle (2026), sec. 5.2
    rng = np.random.default_rng(seed)
    T, ref = len(real), fit(real)            # the reference generator
    draw = lambda g, k: np.median([stat(sample(g, T, rng))
                                   for _ in range(k)], axis=0)
    truth, bias = draw(ref, 10 * n), []      # known, at the real length
    for _ in range(reps):                    # many histories, never one
        refit = fit(sample(ref, T, rng))     # same class, one history
        bias.append(draw(refit, n) - truth)
    b = np.array(bias); m = b.mean(0)
    z = m / (b.std(0, ddof=1) / np.sqrt(reps))
    return truth, m, (np.abs(z) < 2) | (np.abs(m) < tol)   # False: fails

Two lines carry the design. The truth is a median over samples of the real length, because a sample statistic carries its own small-sample bias. The verdict needs a z-score across 20 refits and a tolerance, because one refit confounds bias with luck and a z-score alone flags a bias of 0.002. I ran it on filtered historical simulation, a GARCH(1,1) fitted by quasi-maximum likelihood with resampled standardised residuals, over simulated 12-year daily histories. The statistic was the autocorrelation of absolute returns at lags of 1, 5, 21 and 63 days, with a tolerance of 0.02. At GARCH persistence of 0.98 and 0.995 the screen passed the class at every lag on six of six histories. One of them carried a bias of 0.023 to 0.029 that passed only because 20 refits could not separate it from noise. My first version took the truth from one sample 50 times longer than the history. It failed the same generator at 21 of 24 checks, by up to 0.136, because it compared statistics of two lengths. The screen also passes a mixture of independent days at every lag, since a class with no memory recovers its absence of memory. The null gate of the evaluation stage rejects that mixture.

Decision Default Evidence
Joint distribution or decomposition Factors first, rescaled to unit variance Distribution losses favour high-variance directions (Theorem 5)
Factor count Marchenko-Pastur edge, then a sweep The edge kept 16 and a fitted curve 44. Only 16 went downstream
Synthetic volume Many samples of the real length More samples tighten around the generator’s bias (Corollary 2)
Moments from the generator Covariance, with the mean from history Sharpe 1.090 against 0.275 the other way round, after costs
Whether to build one Only while observations per dimension stay below about two Covariance gain 11.5% at a quarter, gone at twice (Table 2)
Refit cadence Quarterly Annual refits cut the Sharpe ratio from 1.361 to 0.797
Trusted statistics Those that pass a repeated retraining screen Medians off by up to 0.40 beyond 21 days (Table 11)

Three limits travel with these recipes. Independent factors and static loadings let the market factor’s volatility spike on its own. Several factors never spike together. The residuals never cluster. A crash scenario that needs both has to come from elsewhere. Both universes are built from survivors: the S&P 500 members of September 2023 in one paper and stocks with less than 5% missing data in the other. The diffusion model has no dynamics and was trained on winsorised returns, which keeps it out of tail and path questions.

Can a generator produce the tail its consumers care about?

Rare events are where synthetic data is most wanted and least testable: fraud bursts, outages, credit losses, crashes. The obstacle is the sample. The 5% tail of 12 years of daily data holds about 150 observations. Overlapping windows hold far fewer independent ones. The design question is what the generator is trained against.

Tail-GAN (Cont, Cucuringu, Xu and Zhang, in Management Science this April) trains it against its consumers. The construction rests on joint elicitability. Value-at-Risk and Expected Shortfall at level α are the unique minimisers of an expected score.

S(v, e, x) = (W/2)(1{x≤v} - α)(x² - v²)
           + 1{x≤v} · e · (v - x) + α · e · (e/2 - v)
(VaR_α, ES_α) = argmin_(v,e) E[ S(v, e, X) ]

The consumers enter the loss as a library of 65 strategies: 5 single-asset buy-and-hold positions, 50 static portfolios, 5 mean-reversion rules and 5 trend-following rules. A fixed, non-anticipative layer turns each generated path into strategy P&L. A discriminator sorts 1,000 P&L samples per strategy with a differentiable sort and returns a (VaR, ES) pair. The generator is trained in a max-min game so that the pair computed from its scenarios scores on real P&L as well as the pair from real scenarios. The paper sets W = 10 at α = 5%. The weight between the two score terms gave similar results at 1, 2 and 10. At 100 it fell to the supervised result.

Out-of-sample error on 5% VaR and ES, Tail-GAN Table 1 (%)
Sampling floor, SE(1000)3.0Historical simulation3.4Tail-GAN4.6WGAN21.3Trained on single-asset B&H83.3Trained on static portfolios86.7
Simulated data, five seeded repetitions, error against the true VaR and ES of the strategy library. Bottom rows: Tail-GAN trained on a narrower library. On real intraday data it scores 10.1% against 10.4% for history.

Two readings follow. The first concerns the library. The same architecture trained on buy-and-hold positions alone or on static portfolios alone missed by 83.3% and 86.7% on simulated data and by 112.8% and 75.8% on real intraday data. A tail generator is validated only for the kinds of consumer it was trained against. Out of sample in the paper means later data and new strategies from the same three families. Options, stop-losses and volatility targeting never enter the library. The second reading concerns the baseline. Historical simulation beat Tail-GAN on simulated data (3.4% against 4.6%). It trailed by 0.3 points on intraday data (10.4% against 10.1%) and on 20 S&P 500 stocks (25.9% against 25.6%). The generator’s measured advantage is stability. Across five seeded repetitions its intraday error had a standard deviation of 1.1 points against 3.6 for historical simulation.

The real-data test is thinner than its size suggests. Paths of 100 nine-second steps start every minute. Consecutive paths therefore overlap by 14 minutes. The 6,300 training paths of one month hold about 420 independent windows by my arithmetic, which leaves about 21 independent observations in each strategy’s 5% tail. The authors report similar conclusions on non-overlapping data without printing the numbers. For many assets the paper replaces random portfolios with inverse-volatility eigenportfolios. On 20 simulated assets, 20 eigenportfolios gave 6.9% error against 10.4% for 50 random portfolios. Historical simulation gave 3.5%.

Structure matters in the tail as much as in the body. Risk.net reported in 2025 that Joerg Kienitz found Gaussian mixtures doing better than GANs and autoencoders at generating yield curves and volatility surfaces. At one year of daily data the other models failed and the mixtures did not. The paper itself was not readable to me, which is why the claim carries the magazine’s attribution. A mixture has two properties a risk manager likes. A rare high-variance component gives excess kurtosis without a tail model. A Gaussian mixture conditioned on some of its coordinates is again a Gaussian mixture, which turns an unconditional fit into a scenario generator given today’s state. Its component count is the free choice. I wrote the smallest version I would put in front of a risk committee and ran it before printing it.

def fit_factor_gmm(dy, n_fac=3, k_max=8):   # dy: days x tenors, in bp
    mu = dy.mean(0)
    B = np.linalg.svd(dy - mu, full_matrices=False)[2][:n_fac].T
    f = (dy - mu) @ B                        # level, slope, curvature
    fits = [GaussianMixture(k, covariance_type="full", n_init=5,
             random_state=0).fit(f) for k in range(1, k_max + 1)]
    g = min(fits, key=lambda m: m.bic(f))    # K is chosen, never assumed
    return g, mu, B, (dy - mu - f @ B.T).std(0)   # residual: plain noise

def es_99(pnl): return -pnl[pnl <= np.quantile(pnl, 0.01)].mean()  # loss

def curve_es(dy, pv01, n=200_000, seed=1):  # pv01: P&L per +1bp, tenor
    g, mu, B, sd = fit_factor_gmm(dy)
    eps = np.random.default_rng(seed).normal(0, sd, (n, len(sd)))
    scen = mu + g.sample(n)[0] @ B.T + eps   # moves history never printed
    return g.n_components, es_99(scen @ pv01), es_99(dy @ pv01)

The test bed was a simulated curve whose true tail is known: eight tenors driven by level, slope and curvature factors with Student-t shocks, small tenor-level noise and a two-state volatility regime, the stressed state at 2.5 times the calm one. Over 60 simulated histories per cell, I measured the mean absolute relative error of the 99% expected shortfall against a truth from four million draws. With one year of data the mixture beat historical simulation clearly on a 2s10s flattener and narrowly on a duration position. It tied on a 2s5s10s butterfly. With four years history won the butterfly, whose risk sits mostly in the residual my factor model treats as independent noise per tenor. It edged the duration position by less than two standard errors. The flattener stayed with the mixture.

99% ES error against a known truth, % (my harness)
Dur2s10sFlyHS 1 year23.923.612.4Mixture 1 year21.318.212.4HS 4 years12.113.07.7Mixture 4 years13.310.510.9
Mean absolute relative error over 60 simulated histories per cell, lower is better. Columns: duration, 2s10s flattener, 2s5s10s butterfly. HS is historical simulation. Brighter cells carry more error.

The first version I wrote fitted the mixture on the eight tenors directly. With a year of data the Bayesian information criterion chose a single Gaussian in most runs, because each full-covariance component on eight tenors costs 45 parameters. Error on the duration position rose to 34.7% against 23.1% for history on the same 40 histories. The same fit priced the butterfly better than history, 8.8% against 13.1%, because the butterfly’s risk sits in the residual the factor version discards. Three factors cut the cost of a component to 10 parameters. The criterion then picked two components in the median run. The lesson matches the previous section twice. Structure pays while data is short. It pays only in the directions it keeps. History catches up as data grows.

The harness decides as much as the generator. Ericson and colleagues at Wells Fargo compared deep generators with the methods banks use for interest-rate VaR on three USD yield-curve datasets. Their composite adds three scores: distributional distance, autocorrelation and a VaR backtest. Plain historical simulation on a 251-day window ranked first on all three datasets. An AR(1)-GARCH(1,1) with Student-t shocks on returns ranked second and a conditional Wasserstein GAN third, at composites of 1.473, 2.161 and 2.223. History also ranked first in all five seeds of their seed study. The composite hides where the lead comes from. The backtest score is the one a risk manager signs. On it the three sit at 0.543, 0.542 and 0.557. By my arithmetic 85% of history’s composite lead over GARCH comes from the distributional score, which favours replayed history. Two more details belong in any harness. The GAN’s breach rate near 5% at the three-month tenor averaged 1.6% in one sub-period and 7.7% in the other. The split between training and test windows was random over overlapping 20-day windows. The backtest ran on every eligible date, training dates included, which the authors concede falls short of a strict out-of-sample test.

Decision Default Evidence
What the generator trains against Every strategy the desk risk-manages, dynamic ones included 76% to 113% error without them (Tail-GAN)
Generator or replay Replay first, a generator for paths history lacks Replay beat Tail-GAN twice and trailed by 0.3 points twice
Which score decides The backtest, per sub-period Breaches of 1.6% and 7.7% averaged near 5% (Wells Fargo)
Train and test split Walk-forward only Random splits over overlapping windows share observations
Structure on short data Factors plus a mixture sized by criterion Mixture led at one year. At four, history won only the butterfly (my harness)
Many assets Inverse-volatility eigenportfolios 6.9% error against 10.4% for random portfolios

Who writes the rules a record must obey?

A company’s records obey rules nobody calls interesting until one breaks. An order total equals price times quantity. A refund references a sale. A settlement date follows its trade date. Three families cover most of them. Equations bind columns exactly. Inequalities bound a region. Logical dependencies map determining columns to admissible values.

Generators that learn a rule only as a training penalty break it. J.P. Morgan AI Research measured it at NeurIPS 2023 on daily prices with an open-high-low-close constraint. Constrained optimisation met it in 100% of paths and guided diffusion in 72%. GAN baselines with the rule as a penalty met it in 51%, 5% and 0% of paths. TradeFM, J.P. Morgan’s order-flow model, keeps the rule in code: it proposes the flow and hands it to a deterministic exchange simulator with price-time priority matching.

Zhao, Guan and Liu, in an August preprint whose code ships as dive-tabular, turn rule discovery into a propose-and-validate loop. I call the system DIVE after that release. One agent per rule family receives the dataset description, profiles of the columns its family covers and 100 sampled rows. It proposes equations as Python checks, inequalities as JSON coefficient vectors and dependencies as value tables. Each hypothesis then runs on the full reference table. It survives only if at most 0.5% of the rows it covers violate it. A dependency must also cover at least 0.5% of rows. A failed hypothesis returns with up to 20 violating rows for at most three revisions, over three discovery rounds. Consolidation keeps the tightest bound per normalised coefficient vector, prunes equations until each owns a column no other equation uses and makes dependencies acyclic.

DIVE: rules proposed on 100 rows, accepted on the full table
profile + 100 rowsequations (Python)inequalities (JSON)dependencies (tables)full-table check, violations ≤ 0.5%consolidatedependencies, then equations, then projection
A failing rule returns with up to 20 violating rows for at most three revisions, over three rounds. Projecting before equation repair left inequality violations in 24 of 36 NBA settings.

The loop is what earns the trust. One prompt with 200 rows gave Claude Sonnet 5 a precision of 0.73 and a recall of 0.63 on linear inequalities. The loop reached 0.975 and 0.919. On logical dependencies both models scored 0.442 in one shot, with no variance across five splits. The loop lifted precision and recall to 0.963 for Claude and 0.974 for GPT-5.6. Equations were easy even in one shot, with recall of at least 0.919. Counterexamples do most of the work per token. Without them the rules retained on the NEWS table fell from 46.33 to 19.33 per run at about the same token cost, 0.89 million against 0.86 million. One round in place of three cut the yield to 11.33 at a third of the tokens.

Enforcement runs in a fixed order. The order has a measured price. Dependencies go first: determinants stay fixed and the dependent value is redrawn from the admissible set. Equations go second. A model writes a fix function for each column an equation could repair. Each must pass full-table validation. Dynamic programming then picks the repair targets that least damage the marginal distributions. Inequalities go last. Each violating row is projected onto the feasible set by standard-deviation-scaled least squares with the equation columns frozen. On 36 settings of the NBA table, equation repair before projection left no violations. Projecting first left inequality violations in 24. Repair drove equational violation rates of 98.2% and above to zero. Utility, measured by training on synthetic data and testing on real, rose on four of seven datasets. Its worst dataset mean fell by 0.042. Column shapes moved by less than 0.02 per dataset.

Four gaps are left to the team. The reference table must be clean first. Before discovery the authors removed 787 credit records, about 8%, whose total trade count fell below a component count. At a 0.5% tolerance a true rule broken in 8% of rows is rejected. Tolerances therefore belong to the rule and are set after known breaks are reconciled. Rule sets also move. The accepted inequalities on the NEWS table ranged from 0 to 17 across splits. Scope stops at the single record. Uniqueness, foreign keys and cross-row balances stay outside it. Projection can also leave integer columns fractional. A model can propose rules in volume because the full table decides which of them become code.

In a research house the last two gaps are the ones that bite. Prices sit on a tick grid. Quantities are whole contracts. Every print belongs to a bar that exists. Exact means integers.

def ticks(px, tick):                 # price -> integer ticks + on-grid flag
    t = np.asarray(px, float) / tick
    return np.rint(t).astype(np.int64), np.abs(t - np.rint(t)) < 1e-6

def check_bars(bars, trades, tick):  # trades: bar_id, px, qty, time order
    px, on_grid = ticks(trades.px, tick)
    g = trades.assign(px=px, off=~on_grid).groupby("bar_id")
    agg = pd.DataFrame({"o": g.px.first(), "h": g.px.max(),
                        "l": g.px.min(), "c": g.px.last(), "v": g.qty.sum(),
                        "off": g.off.any()}).reindex(bars.index)
    ok = agg.off.eq(False).to_numpy()    # no print between two ticks
    ok = ok & trades.bar_id.isin(bars.index).all()   # no orphan prints
    for k in "ohlc":                     # a bar is the sum of its prints
        v, grid = ticks(bars[k], tick)
        ok = ok & grid & (v == agg[k].to_numpy())
    return ok & (bars.v.to_numpy() == agg.v.to_numpy())

I ran it on 20,000 simulated bars of an index future on a 0.25 grid near 4,300, with 200 bars for each of five fault types. It caught every fault and passed every clean bar, including bars carrying float noise of a billionth of a point. The comparison most code reaches for first is numpy’s isclose at its defaults. Its relative tolerance of 0.00001 allows 0.043 at a price of 4,300, more than four times the 0.01 error planted in each off-grid high. It caught none of those highs and 24% of the off-grid prints. The orphan-print line closes a gap a second reader found: prints whose bar was missing dropped out at the reindex.

When a gate refuses, three options remain. Rejection keeps the sampler unbiased while acceptance is high. It fails as acceptance falls. On Lending Club data only about 10% of real rows satisfy four portfolio rules jointly, according to Constrained Tabular Diffusion for Finance. Projection moves mass onto the boundary of the rule. Under a floor of 1.5 reviews a month, the same paper puts 70.2% of real housing listings below the floor. Projected rows show a mean of 1.9 with a standard deviation of 0.7. Selection gates carry a subtler risk. When Sample Selection Bias Precipitates Model Collapse, at ICML 2026, proves that a top-α selection gate referenced to a biased slice of the target drives covariance to zero over generations. My rule follows from the three. Reject while acceptance is high. Repair when a column is a function of others. Project when the rule bounds a region. Calibrate every gate on the full target.

One mechanism per class of rule, with what it costs
Syntaxgrammar mask100% satisfied1.10x latencySlot placementno mechanism24-74% at depth 4validator onlyRow boundsstep projection0.0% violationsmass on the boundaryCross-columnvalidated repair98.2% CVR to 0utility -0.042 worstTrajectoryentity-level gendistance falls 61%single run, eps = 2Referentialkey registry firsttrue by constructiongate rechecks itTotalsinformation projectionhits in expectationinteger close neededTemporallexical checker60K pairs a yearmodel screen unauditedShapeno mechanism neededmarginals: LR C2ST 1.0a control only
First box, my default mechanism per class of rule. Second and third, the measured result and its cost, from the papers cited in this piece. Amber: the one class with no enforcing mechanism, only a validator.

Syntax is solved by a full-automaton mask. Referential integrity is free when the key registry, the security master in a research house, is built first. Totals get an information projection, which hits them only in expectation. A book that must tie to the cent needs an integer adjustment afterwards. Trajectories need the entity as the unit of generation: PATH makes a user’s whole event sequence the unit and cuts distributional distance to real trajectories by 60.8% against a marginal-based mechanism, at a privacy budget of ε = 2, its most favourable setting. The amber row is the hole. Where vs What measures values that appear somewhere in the output at the wrong schema path. On its deepest schema that share runs from 24% on GPT-4o to 74% on Qwen2.5-7B, scored only on outputs that parse.

Decision Default Evidence
Who writes the rules A model proposes on 100 rows. The full table decides Linear-inequality recall 0.63 in one shot, 0.919 with the loop (DIVE)
Tolerance Per rule, after known breaks are reconciled An 8% break fails a 0.5% gate
Enforcement order Dependencies, equations, then projection Projection before equation repair failed 24 of 36 settings
Reject, repair or project Reject while acceptance is high, then repair or project Only 10% of Lending Club rows pass four rules jointly
Integers and keys A separate exact layer Projection can leave integers fractional. Scope stops at one record
Versioning One rule set per snapshot, diffs reviewed 0 to 17 inequalities across splits of one table

Which checker makes generated text trainable?

Most synthetic data a company will generate this year is text: questions over documents, labels on messages, extraction targets and reasoning traces. The frontier has answered the substitution question for text wherever a checker exists. NVIDIA reports that more than 98% of the data used to align Nemotron-4 340B was synthetic, with a reward model as the gate. DeepSeek-R1 rewards answers by rule. The engineering question is which checker each field gets. Four kinds exist, in falling order of strength. Code recomputes an answer. A solver knows an optimum. A model judges plausibility. People review a sample.

The label pipeline with the most complete public numbers is a case study by Capital Fund Management on the Hugging Face blog. CFM uses Hugging Face’s Expert Support service. The task was to tag company names in about 900,000 news headlines. The pipeline has five stages.

  1. Teacher. Llama 3.1 70B Instruct on four L40S GPUs behind Text Generation Inference labelled every headline under a Pydantic JSON grammar, at temperature 0.1 with five few-shot examples. It took about 8 hours and about $70.
  2. Sample. Fuzzy matching grouped the teacher’s company names into clusters. A stratified sample of 2,714 took 5% of each cluster plus 5% of headlines without a company. The post does not size the pool those rates apply to.
  3. Split. The sample was split by date: 2,405 headlines from 2010 to 2016 for training, 204 from 2017 to 2018 for validation and 105 from 2019 to 2020 for testing.
  4. Review. Reviewers corrected the teacher’s pre-loaded spans in Argilla at 5 to 10 seconds per record, against about 30 seconds without them.
  5. Student. A fine-tuned GLiNER model rose from a reported 87.0% zero-shot to 93.4% F1 on the test set, against 92.7% for the 70B teacher. It runs at $0.50 an hour on a GPU or $0.10 on a CPU, against at least $8 for the teacher.

At equal volume, reviewed labels beat raw ones at every size the post tried: at 1,000 training samples GLiNER scored 0.915 against 0.895. The post does not state the evaluation set for that comparison. Three details decide how far the result travels. The test set holds 105 headlines. Treating each headline as one trial, that puts the 95% interval on an F1 near 0.93 at about five points either way by my arithmetic. The interval covers the student’s 0.7-point lead over its teacher. The reference labels began as the teacher’s own spans, which anchors the test toward the model being compared. One of the five few-shot examples normalises the ticker BDE, printed beside Black Diamond in its headline, to BPER Banca. An exemplar error repeats in every label that resembles it.

InvestAlign (ACL 2025) wants models that allocate the way investors do when they partly follow an adviser, a setting where decision data is scarce. The authors take a Merton-style problem whose optimal solution is known. An investor with exponential utility pays a penalty for deviating from an adviser whose risk aversion is fixed at 0.2.

P₂(t) = v / (α₂ σ²) · e^{r(t - T)}        the adviser's holding
P₁(t) = P₂(t) · [α₂σ²η e^{2r(T-t)} + θ₁] / [α₁σ²η e^{2r(T-t)} + θ₁]

As the herding coefficient θ₁ goes to zero, the investor keeps its own Merton holding. As θ₁ grows without bound, the investor converges on the adviser. The constant η comes from a fixed-point iteration. Code sweeps 10 risk aversions by 10 herding coefficients by 10 Brownian draws into 1,000 prompt-and-answer pairs. Models trained on the simple problem were tested on it and on two harder variants against questionnaire answers from 119, 80 and 44 respondents. Mean squared error to the class-average human answer fell by 44.52% to 61.26%. Where a problem has a known optimum, the solver is the verifier and the label is correct before anyone reads it.

Two controls decide what that result means. A generic fine-tune on ConvFinQA recovered 59% to 85% of the error reduction across two models and two problems, by my arithmetic on their Table 5. Part of the gain is therefore format. The human side is thin. The 119 respondents to the simple problem spread over 55 cells of risk aversion and herding, about two per cell. Mixing real answers into the theory data helped slightly at 10 parts to one and hurt at the reverse ratio.

How much real data to keep is the other decision. A member’s proof of concept in the FCA Synthetic Data Expert Group’s March 2024 report trained challenger models on open-banking data at different blends. Against a real-only model, accuracy fell 2.5% at half synthetic, 5.0% at 70% synthetic and 32.0% at fully synthetic.

Accuracy change vs a real-only model, by synthetic share (%)
10-3500% (real)-2.550%-5.070%-32.0100%
The open-banking case in the FCA group's 2024 report, points evenly spaced. Accuracy slips 5% by 70% synthetic and falls 32% at 100%.

In pretraining runs of up to 3B parameters, a Meta-led study of more than 1,000 models found pretraining on rephrased synthetic text alone no faster than on natural text, with about one-third synthetic the best mix. For linear regression retrained over repeated rounds, replacing real data with synthetic makes error grow linearly with the number of rounds. Accumulating synthetic data on top of the real data keeps error below π²/6, about 1.64 times its first-round value.

Two engine-level effects apply to every extraction schema. Under a required closed-vocabulary field with no escape value, 10 of 13 models in PhantomFill fabricated an answer in every trial on inputs whose truth was absence. With an escape value offered, all nine open-weight models still fabricated in 60% to 100% of trials. A mask on the decoder while the model reasons also costs accuracy. Draft-Conditioned Constrained Decoding measured Qwen2.5-14B at 76.0 on MATH500 with free output and 47.6 with a grammar mask. The schema in the prompt cost 5.2 points of that drop and the mask the other 23.2. It scored 58.6 when the model drafted freely before the mask formatted the result.

Decision Default Evidence
Checker per field The strongest that exists: code, solver, model, people Frontier alignment data passes a gate (Nemotron-4)
Teacher and student A large model labels once, a small one serves 93.4% against 92.7% F1 at a sixteenth of the GPU price (CFM)
Review Pre-labelled, sampled by cluster, split by date 5 to 10 seconds per record against 30
Test set Sized before labelling 105 items leave about five points either way
Solver labels With a generic fine-tune as a control ConvFinQA recovered 59% to 85% of the gain
Real share Accumulate on a real anchor Accuracy fell 32% at fully synthetic (FCA)
Closed fields An escape value plus a downstream absence check All nine open models still fabricated with an escape offered

Can a model that has read the future write data about the past?

Every company dataset has two clocks: when a fact was true and when it was known. Restated revenue, late fraud labels and backfilled CRM fields all break the second. Language models add a third clock, their training cutoff. A model that writes data about the past may know how the past ended. AI’s predictable memory, in Economics Letters, measures that knowledge as a trading signal. GPT-4.1, asked for a stock’s return on a date with no context, produced a daily long-short Sharpe ratio of 2.88. Removing 3.0% of the observations brought it below 0.2.

DatedGPT, from CUHK and UCL, is the most complete build of a leak-free text generator I have read. The pretraining corpus is FineWeb-Edu filtered by Common Crawl timestamp and cut into 12 yearly vintages from 2013 to 2024. The guarantee is exact and narrow. Nothing crawled after a cutoff enters its vintage. A 2015 crawl can still hold a page written in the 1990s. The crawl date therefore bounds knowledge from one side only. Each vintage is a 1.34B-parameter model trained from scratch on about 100 billion tokens, at about 2,000 A100 GPU-hours. Instruction data is built per vintage as well. Qwen3-4B rates candidate documents. DeepSeek-V4-Pro writes question-and-answer pairs from them under a lexical constraint: every word must come from the source document or from a fixed list of function words. A rule-based script drops any pair that breaks it. About 60,000 pairs survive per year. The paper also bounds the leak from above.

r_{i,t+1} = α_i + δ_t + β₁ s^f_{i,t} + β₂ (s^b_{i,t} - s^f_{i,t}) + ε
t: trading day.  s^f: vintage y-1 scoring a headline from year y
s^b: the 2024 vintage, trained on every year's crawl

For DatedGPT-instruct the coefficient on the difference, the lookahead premium, is 26.4 basis points per standard deviation at a t-statistic of 10.65. With yearly vintages the difference also carries news that was public earlier in the same year, which makes the premium an upper bound. The full signal rises only from 28.3 to 33.0 basis points per standard deviation when the 2024 vintage scores every headline. The measurement also needs a model with skill. ChronoGPT’s base model has no signal in this headline test (t = 0.06) and a premium t-statistic of 0.90. A model without signal shows no premium whether its data leaked or not. Three prominent pipelines in 2025 and 2026 set out to build leak-free instruction data. Two of them let a model decide what counts as leak-free.

Who decides what counts as leak-free, three point-in-time pipelines
PIT models (AQR, Yale, EPFL)Monthly cutoffs, 2013 to 2024gpt-5-nano timeless screenNo accuracy check reportedNews Sharpe about 1.0 to 1.5ChronoGPT-InstructYearly vintages, 1999 to 2024GPT-4.1 keeps pre-2000 pairsAbout 425,000 pairs kept0.95 vs 1.76 for Llama-3.2-3BDatedGPTYearly vintages, 2013 to 2024Words must come from the sourceRule-based script drops failuresNews Sharpe 3.20, before costs
The Sharpe ratios use different news sets, windows and portfolio rules and cannot be compared across panels.

The point-in-time models of Bryan Kelly and colleagues at AQR, Yale and EPFL kept public instruction examples that gpt-5-nano classified as timeless and report no accuracy check. The screen never sees a cutoff date. Its own prompt examples keep a statement about a 2022 software release. ChronoGPT-Instruct kept about 425,000 pairs that GPT-4.1 judged at maximum confidence to use only pre-2000 knowledge. DatedGPT lets code decide. The code version is small enough to print.

FUNCTION_WORDS = frozenset("""a an the of to in on at by for from with as
    and or but nor if than that this these those it its they their them he
    she his her we our you your which who whom whose what when where why how
    much many not no is are was were be been do does did""".split())
TOKEN = re.compile(r"\d+(?:[.,]\d+)*|[^\W\d_]+(?:'[^\W\d_]+)?")

def words(text): return TOKEN.findall(text.lower().replace("\u2019", "'"))

def pit_gate(pair, source):      # pair: generated question + answer
    allowed = set(words(source)) | FUNCTION_WORDS
    leaked = sorted({w for w in words(pair) if w not in allowed})
    return not leaked, leaked    # passes only if every word is traceable

The function-word list is mine, because the paper defines its list only as particles, conjunctions, pronouns and similar connectives. The tokenizer reads letters in any script, which catches a leak written in Chinese or with an accent. I ran the gate against an invented third-quarter 2019 filing for a fictional freight company. One pair written in the filing’s own words passed. Five other faithful pairs were refused on inflections alone. Pairs that mentioned a pandemic, the year 2020 or a revenue growth figure the filing never printed were refused, with the offending words listed. “How much did revenue rise?” fails because the filing says “rose” while an English question forces the base form. Every lemma table a team adds is therefore a design decision the paper leaves open. The gate also passed “Revenue fell as fuel prices rose”, a false claim assembled entirely from words the filing contains. It passed “Revenue rose $4.2 million” as well, because the tokenizer drops currency and percent signs. A lexical rule bounds the words a pair can use and leaves open what it can claim, which is why the point-in-time gate and the fact gate have to be two different checks.

The other code-decided designs share one move. They compute the label and let the model write only the words around it. Yale’s Verifiable Forecast Actions derives each target action from the realised 10-day path in code. OpenFinGym’s resolver waits for the outcome. This blog’s Look-Ahead-Bench note covered the same bias from the benchmark side.

Decision Default Evidence
Which date defines a vintage The crawl or knowledge date, named as such A 2015 crawl can hold a page from the 1990s (DatedGPT)
Vintage grain Yearly, monthly where the signal decays faster Monthly checkpoints of one date-ordered run avoid a 12-fold cost (PIT models)
Who screens instruction data Code, with a model as proposer Two of three pipelines let a model decide
How to measure a leak The premium regression, on a model with skill 26.4 bp per standard deviation at t = 10.65, an upper bound (DatedGPT-instruct)
What a lexical gate misses Entailment, checked separately A false claim built from source words passed my gate

What does a generated dataset have to beat?

The evaluation contract decides whether anything upstream mattered. Four elements recur across the sources. A null model must lose. The consumer’s own metric decides. Every trial enters a ledger and the result is deflated for the search. A power calculation comes before the first run.

A 2026 audit of tabular metrics by Scassola and colleagues built the tabular null: a product of per-column marginals that represents no dependency. The audit ran it through the standard suite.

Diabetes, 5 seeds: a model with no dependencies against diffusion SotA
TrendC2STTabDiff0.970.93TabSyn0.960.66Marginals only0.971.00
Tables 5 and 7 of the audit's v2, higher is better. To four decimals Trend reads 0.9686 for TabDiff and 0.9685 for marginals only. Theorem 2.1 shows why the logistic C2ST of a marginals-only model tends to 1.

A logistic-regression discriminator on raw columns separates two tables only through their means. Once the synthetic means converge, it has nothing left to look at. With an XGBoost discriminator the same null scores 0.05 on the Diabetes data. The choice of discriminator belongs in the contract. For return series the analogous null draws whole days with replacement. It keeps fat tails and the cross-section and destroys volatility clustering, which makes it exactly the thing a path generator has to beat. This blog’s July review covered generators that pass fidelity metrics while erasing the fraud patterns detectors read.

def acf_abs(r, lags=20):          # r: days x assets, mean ACF of |r|
    a = np.abs(r) - np.abs(r).mean(0)
    return np.array([((a[k:] * a[:-k]).sum(0) / (a * a).sum(0)).mean()
                     for k in range(1, lags + 1)])

def iid_days(r, n, rng):          # the null: whole days, with replacement
    return r[rng.integers(0, len(r), n)]   # fat tails kept, clustering gone

def clustering_gate(real, synth, rng, k=20):   # the generator must beat it
    err = lambda x: np.abs(acf_abs(x) - acf_abs(real)).mean()
    base = np.array([err(iid_days(real, len(synth), rng))
                     for _ in range(k)])
    if err(synth) >= base.mean() - 2 * base.std(ddof=1):
        raise AssertionError(f"the null scores {base.mean():.3f}")
    return err(synth), base.mean()

I ran it on 10 years of simulated daily returns for five correlated assets from a GARCH process with Student-t shocks, whose absolute returns carry a lag-one autocorrelation of 0.177. A Gaussian mixture fitted to the daily return vectors matched the in-sample marginals better than a fresh draw from the true process (a mean Kolmogorov-Smirnov statistic of 0.024 against 0.041). It matched the correlation matrix better too (mean error 0.013 against 0.017). Once I shuffled its draws, its absolute returns had a lag-one autocorrelation of -0.007. The gate refused it at 0.162 against the null’s 0.161. A block bootstrap with 20-day blocks passed on all 10 simulated histories. The true process passed on 9 of 10, which is the gate’s own price at this sample length. The mixture passed on 1. The shuffle matters. Scikit-learn returns mixture samples sorted by component. Left in that order, the mixture passed on all 10 histories, because a run of draws from the high-variance component looks like a volatility cluster. A one-day scenario generator can beat the true process on in-sample static tests and still be the wrong tool for a path, because the thing it cannot see is the order of the days.

Gençay, an August preprint, builds the rest of the contract into an agent that searches for trading strategies. He argues the design transfers to any search over models or datasets, although he tested it only on trading. Four choices carry it.

  • Leaks are a property of the registry. Every feature family carries a leak-safe flag. One flag drives the feature list in the tool schema, the catalogue shown to the model and the server-side validator. Execution lags one day everywhere. A deterministic audit rejects any plan that breaks an event’s publication delay.
  • Plans are typed JSON. The agent assembles them through validated tool calls and writes no code. A matched critic without in-loop validation produced no valid proposal in 60 attempts. The tool loop produced 99 in 100 with GPT-4.1.
  • One entry point fills the ledger. Every design-window backtest counts as a trial, grid points included.
  • Two tests read the ledger and a third reads the held-out window. The Deflated Sharpe Ratio sets a bar of 0.95. The Probability of Backtest Overfitting runs combinatorially symmetric cross-validation over 16 blocks, 12,870 splits. A stationary-bootstrap interval covers the held-out window.
PSR(SR*) = Φ( (SR^ - SR*) √(T - 1) / σ_SR )
σ_SR = √( 1 - γ₃ SR^ + ((γ₄ - 1)/4) SR^² )
SR*₀ = √V · [ (1 - γ) Φ⁻¹(1 - 1/N) + γ Φ⁻¹(1 - 1/(N e)) ]
DSR = PSR(SR*₀),   N trials and V their Sharpe variance, from the ledger
γ₃ skew, γ₄ raw kurtosis, γ ≈ 0.5772 the Euler-Mascheroni constant
SR^ is per observation. The Sharpe ratios in the text are annualised.

The Deflated Sharpe Ratio of Bailey and López de Prado shows why the ledger must be complete. Their worked example is a Treasury strategy with a Sharpe ratio of 2.5 over five years of daily data. After deflation there is about a 90% chance its true Sharpe is positive, which fails a 95% bar. Had the researcher stopped at 46 trials, it would have cleared. The Probability of Backtest Overfitting asks how often the in-sample winner falls below the median out of sample, across many pairs of halves.

Gençay’s experiments show each element earning its place. A strategy handed tomorrow’s return posted a design Sharpe ratio of 34.7, an evaluation Sharpe of 51.5 and a Deflated Sharpe Ratio of 1.00. Statistics that correct for search say nothing about the information set, which is why leak control belongs in the registry. After 102 recorded trials, the agent’s best discovery had a design Sharpe of 1.69 against a deflation threshold of 1.21, a Deflated Sharpe Ratio of 0.86 and a held-out Sharpe of 0.18. It returned 4.7% over four years against 40.8% for equal-weight buy-and-hold. Each failure mode was caught by the test built for it. Five seeds converged on the same volatility-breakout family with held-out Sharpe ratios of 0.59 to 0.70. The deflated test failed all five. A multi-asset search showed an overfitting probability of 0.80. Its best strategy fell from 0.37 to -0.23 held out. The discovery run above passed the overfitting test at 0.01 and failed the deflated test and the held-out interval. No single test caught every failure. A random-weights null earned 23% at a Deflated Sharpe Ratio of 0.0008.

The last element is arithmetic that belongs before any run. The t-statistic of a Sharpe ratio is roughly SR·√years. Nine years of daily data certify a Sharpe near 0.7. Cetingoz and Lehalle’s Table 10 gives the same answer from the generator side. With 12 years the standard error of a long-short Sharpe is about 0.3. Reaching 0.1 takes about 120 years. A synthetic backtest cannot close that gap, by Corollary 2. It raises the number of trials, which raises the deflation bar.

Privacy closes the contract. Synth-MIA found that the distance-to-closest-record proportion, the most common privacy metric in tabular synthesis, correlates with measured leakage at only r = 0.225. The authors read that as basically uncorrelated. A membership attack on every release belongs next to the null.

Decision Default Evidence
The null Whole days with replacement for paths, marginals for tables Marginals tied diffusion on Trend (Diabetes) and beat it on logistic C2ST
Leak control A registry property, validated server-side A planted leak scored a Deflated Sharpe Ratio of 1.00 (Gençay)
What counts as a trial Every design-window evaluation, grid points included A 1.69 Sharpe ratio failed at 102 trials
Which tests Deflated Sharpe, overfitting probability and a held-out interval Each failure mode was caught by the test built for it
Window length Set by power before the first run Standard error near 0.3 at 12 years (Cetingoz-Lehalle)
Privacy A membership attack per release Closest-record proportion correlates with leakage at r = 0.225 (Synth-MIA)

How does a checked dataset become a reward?

The last stage turns the evaluation contract into a reward. Every weakness of the contract then becomes a target for optimisation. OpenFinGym, a gym for quant agents from Edinburgh, UCL, the Alan Turing Institute and Oxford, is the most complete public build I have found. Its September revision lists it for EMNLP 2026 Findings. It ships 78 curated tasks: 48 forecasting, 13 market generation, 10 trading and 7 fraud detection. Those 78 were selected and verified by hand.

Its pipeline for new tasks is itself propose-and-verify. Papers are harvested and admitted only with a complete protocol, documented metrics and a public dataset. A model writes a download script that runs in a sandbox. A reviewer with unit tests judges the script and the data. A missing metric is generated from a template and admitted after interface checks, unit tests on synthetic tensors and a reviewer audit. The model that builds each task loader also emits a provenance record for its labels, which a reviewer audits against the source mechanism. A separate test scans the loader’s output for feature leakage. Contract and smoke tests then run the full path from agent to evaluator before a task is installed.

One market-generation task, from paper to reward
SOURCE PAPERsourcePCF-GAN, NeurIPS 2023modelrough vol, H = 0.25settingseta 0.5, 200 stepsDATA + SPLITpaths20000 x 200 x 2train16000 pathsval / test2000 / 2000test in boxplaceholder onlyAGENT WRITESgeneratorany model it likessamples2000 x 200 x 2callsubmit_and_persistVERIFIER + REWARDshapeexact or refused8 metricsSig-MMD, ACF, disc.total_lossmean of 8, lowerRL reward-min(loss, 5) or -6
Blue: set by the source paper. Amber: added at each stage. Dashed: left open to the agent, which never sees the test paths. The RL row is the transform the paper used in its forecasting runs.

At run time the test labels stay outside the agent’s container. A host-side verifier loads them into memory and serves a rate-limited submission API that every container of a task can share. Forecasts whose outcome is unknown go into a ledger with their reference state and resolution time. A resolver scores them once the outcome arrives. That is the only label in finance that is leak-proof by construction. It costs a wait. The trading engine returns rejected orders as observations, which keeps a rollout alive through a constraint violation.

Training used supervised fine-tuning on 189 trajectories from Claude Opus 4.7, followed by GRPO with four prompts and eight rollouts per step for 250 to 500 steps. The reward was the task loss clipped at 5 and negated, with -6 for a failed submission. Base Qwen3 models at 1.7B, 4B and 8B produced a scoreable submission on none of five held-out forecasting tasks. After GRPO every model succeeded on every task. Four properties of the reward decide what an agent learns from it.

Aggregation. Market-generation totals average or sum their metrics without putting them on a common scale. In one conditional-generation task a cross-correlation score, itself an unnormalised sum over a correlation matrix, makes up about 98% of the total loss. A reward is only as informative as the rule that aggregates it, which is why an unweighted mean of metrics on different scales ends up tracking whichever metric is largest.

Scale. A clip at 5 suits forecasting losses near 0.05. Generation losses on the frontier-agent benchmark average 7.1 to 14.8 across models. The same clip would flatten every generation task whose loss exceeds 5. Those averages show that some do.

Executability. Base models failed on invented packages and wrong output shapes. The jump from 0% to 100% success measures executability more than skill. The large reward gains came from eliminating high-loss seeds. On Treasury, GRPO cut the 1.7B model’s mean loss from 1.15 after supervised fine-tuning to 0.099, as its standard deviation across ten seeds fell from 0.97 to 0.015.

Ranking. The top two agents by mean within-task rank cannot be separated statistically, with a sign-test p-value of 0.14 and a Wilcoxon p-value of 0.37. Naming a winner needs seeds and paired tests.

Decision Default Evidence
Where labels live Outside the container, one host verifier per task Removes test access at run time by construction
Aggregation Rank or normalise before summing One score is 98% of one task’s loss
Reward scale A clip per task at its loss scale A clip of 5 against generation-family averages of 7 to 15
Failure Below any valid score, success reported apart 0% to 100% success measured executability
Deferred labels Wherever pretraining could know the answer Only resolved outcomes are unseen by construction
Declaring a winner Seeds plus paired tests Top two not separable (Wilcoxon p = 0.37)

Is synthetic data the only option?

The frontier case for yes is strong where a checker exists. AlphaGeometry trained on 100 million synthetic theorem-proof examples derived by a symbolic deduction engine, with no human demonstrations. It solved 25 of 30 olympiad geometry problems. Phi-4 made synthetic data 40% of about 10 trillion pretraining tokens. The limits are measured as well. Epoch notes that synthetic data has reliably improved capabilities only in narrow domains such as mathematics and code. Phi-4’s own ablations show synthetic-only models doing worse on knowledge and hallucinating more.

My synthesis extends the opener’s point that a company already runs its verifiers. Synthetic data substitutes for real data wherever a verifier exists. It complements real data where none does. A company has more verifiers than it uses: the database’s key checks, the ledger, the interpreter, the as-of table, the solver for a known optimum and the resolver that waits. It has none for what its market, customers or counterparties do next, short of waiting.

Where synthetic data substitutes and where it complements, my placement
Complement, test hardone history, no checkerCrisis and tail scenariosLong-horizon statisticsLong-short Sharpe (0.08)Substitutethe checker is the labelQA run as codePoint-in-time instruction dataSolver labels (InvestAlign)Use the historyhistory is the baselineIntraday VaR: HS 10.4, GAN 10.1Rates VaR (HS ranks first)Augment and distilcheap labels, checkedHeadline NER (CFM)Rule repair (DIVE)Verifier: none to exactReal data: scarce (top) to plentiful
My placement from the sources in this piece. Amber: the cell where a team needs data most, can check it least and has no synthetic answer yet.
In my reading the scarce asset moves from data to verifiers, which most companies already own: the key checks in the database, the trial balance, the as-of table, the risk engine and the evaluation ledger.

The work is to wire them into the generator. Regulators’ forums draw the same line. The FCA’s Synthetic Data Expert Group, in a 2025 governance report that disclaims representing the FCA’s views, treats quality as fitness for a use. The April US interagency model risk guidance excludes generative and agentic AI, which leaves a conventional model trained on synthetic data in scope.

Von Neumann said it in 1949

The foundational source for this piece is a three-page summary, written up by George Forsythe, of a 1949 symposium talk that the National Bureau of Standards published in 1951. In Various techniques used in connection with random digits, John von Neumann said that “any one who considers arithmetical methods of producing random digits is, of course, in a state of sin.” A language model asked for a number is an arithmetical method. The same pages state the acceptance-rejection method in full. Draw two uniform numbers X and Y. “If Y>af(X), we reject the pair and call for a new pair.” Every gate in this piece is that paragraph with an executable rule in place of the density. His verdict on the arithmetic generators of his day was that it had been “more trouble to test them than to manufacture them”. The evaluation contract above is the same observation with more than 75 years of compute behind it.

The revised decision sheet

The August opener closed on eight calls. The evidence above revises each of them.

  1. Platform or assembly. Assemble around three stores: a snapshot with two clocks, a versioned rule registry and a trial ledger. Accept a platform only where it exposes all three.
  2. Spec-first or seed-grounded. A seed-grounded generator estimates one history. Compute observations per dimension before training and expect help only while that ratio is small.
  3. Who draws the numbers. Fitted generators: structure first, many samples of the real length, a repeated retraining screen.
  4. Authored or discovered rules. Discovered on a sample, accepted on the full table per rule, enforced as dependencies, equations, then projection. Integers and keys live in an exact layer.
  5. Decoding engine. Let the model draft freely and format under the mask. Give every closed field an escape value.
  6. Hard or soft gates. Refusals for invariants and log lines for everything else, with the difference written down.
  7. Privacy stance. A membership attack on every release. Distance to the closest record measures too little.
  8. The evaluation contract, written first. A null that must lose, the consumer’s metric per sub-period, a complete ledger with deflation and a power calculation.

The order I would build in follows from the stores. The first month builds the snapshot with its two clocks and the ledger. The second adds the generators with the retraining screen and the null gate. The third adds the rule registry and the text pipeline with its checkers. Tail generators wait until the evaluation contract can tell them from replay.

How this piece was built

Three verified sweeps feed this special: 107 engine papers from the last 12 months, 76 finance sources checked at their primary source (73 kept) and 43 finance training directions. For this revision the ten primary sources that carry the stages were read in full. Each extraction was checked against the paper’s tables. Where a claim concerns a paper’s own system, its released code was read as well. An adversarial round then reopened the sources behind every section, with one verifier per section. All five code blocks ran before printing. The new retraining screen failed a sound generator on all six test histories until its truth was computed at the real sample length. A z-score alone flagged biases of 0.002, which is why the screen carries a tolerance. Earlier rounds caught readings that would otherwise have printed. Cetingoz and Lehalle prove a bias bound, which is narrower than a claim that generators add no information. A diffusion paper’s 4,096 was a sample size that I had read as a stock count.

Where it still breaks

Several questions stay open. I found no public case of a named buy-side firm running synthetic market data in production. I found no survey that measures adoption. The Spirit sale remains open. The CFM figures come from a vendor-hosted case study with a 105-headline test set. InvestAlign’s human validation rests on about two respondents per cell. The paper does not report Tail-GAN’s dynamic strategy rules, which leaves its result unreproducible from the text alone. Diffusion Factor Models report their largest empirical gains past the ratio where their own simulation shows none.

Three limits are mine. The harnesses behind my code blocks run on simulated markets whose structure I chose. They show mechanisms. Magnitudes need real data. The tolerance of my retraining screen and every default in the decision tables are judgment calls. The quad is a placement that will move as verifiers for behaviour appear.

The bottom line

Real data has a price. The Spirit auction put a number on it. Synthetic data has a method. This piece traced it through seven components. A fitted generator is an estimator of one history. It trades variance for bias and adds no information beyond that history and the structure it is given. Its value comes from structure and verifiers: factors for numbers, the consumer’s own score for tails, the full table for rules, code or a solver for labels, a dated source for time and a complete ledger for evaluation. Finance is the hardest test of each component that I know. The method holds there with its limits stated, which is the best evidence I have that it holds in a simpler company.

Real data has a price. Synthetic data has a method: structure before generation, a retraining screen and a null that must lose, rules accepted on the full table, labels checked by code or a solver, a clock kept in code and a ledger of every trial. The model proposes. Code decides.

References

The 15 sources this piece leans on most, in APA style. Every other source is linked where it is cited.

  • Ahouzi, O., Champonnois, S., L’Hour, J., Ratnamogan, P., Patault, B., & Goibert, M. (2024, December 3). Investing in performance: Fine-tune small models with LLM insights - a CFM case study. Hugging Face Blog. https://huggingface.co/blog/cfm-case-study
  • Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality. The Journal of Portfolio Management, 40(5), 94-107. https://doi.org/10.3905/jpm.2014.40.5.094
  • Cetingoz, A. R., & Lehalle, C.-A. (2026). Synthetic data for portfolios: A throw of the dice will never abolish chance. Quantitative Finance, 26(8), 1311-1342. https://doi.org/10.1080/14697688.2026.2673155
  • Chen, M., Xu, R., Xu, Y., & Zhang, R. (2025). Diffusion factor models: Generating high-dimensional returns with factor structure (arXiv:2504.06566). arXiv. https://arxiv.org/abs/2504.06566
  • Cont, R., Cucuringu, M., Xu, R., & Zhang, C. (2026). Tail-GAN: Learning to simulate tail risk scenarios. Management Science, 72(4), 2917-2936. https://doi.org/10.1287/mnsc.2023.00936
  • Ericson, L., Zhu, X., Han, X., Fu, R., Li, S., Guo, S., & Hu, P. (2024). Deep generative modeling for financial time series with application in VaR: A comparative review (arXiv:2401.10370). arXiv. https://arxiv.org/abs/2401.10370
  • Gençay, E. (2026). What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery (arXiv:2608.27734). arXiv. https://arxiv.org/abs/2608.27734
  • He, S., Lv, L., Manela, A., & Wu, J. (2025). Instruction tuning chronologically consistent language models (arXiv:2510.11677). arXiv. https://arxiv.org/abs/2510.11677
  • Scassola, D., Ponsford, D., Javaloy, A., Saccani, S., Bortolussi, L., Gouk, H., & Vergari, A. (2026). A sobering look at tabular data generation via probabilistic circuits (arXiv:2603.23016). arXiv. https://arxiv.org/abs/2603.23016
  • Synthetic Data Expert Group. (2024, March 8). Using synthetic data in financial services [Report]. Financial Conduct Authority. FCA publication page.
  • von Neumann, J. (1951). Various techniques used in connection with random digits. In A. S. Householder, G. E. Forsythe, & H. H. Germond (Eds.), Monte Carlo method (National Bureau of Standards Applied Mathematics Series, Vol. 12, pp. 36-38). U.S. Government Printing Office. Scan of the Collected Works reprint.
  • Wang, H., Pan, Z., Zhang, H., Liu, M., Gao, H., & Zhao, H. V. (2025). InvestAlign: Overcoming data scarcity in aligning large language models with investor decision-making processes under herd behavior. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 10021-10052). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.495
  • Yan, Y., Tang, R., Gao, Z., Jiang, W., & Lu, Y. (2026). DatedGPT: Preventing lookahead bias in large language models with time-aware pretraining (arXiv:2603.11838). arXiv. https://arxiv.org/abs/2603.11838
  • Zhang, K., Ge, W., Jiang, L., Yang, W., Langham-Lopez, J., Yu, J., Szpruch, L., & Ni, H. (2026). OpenFinGym: A verifiable multi-task gym environment for evaluating quant agents (arXiv:2606.26350). arXiv. To appear in Findings of EMNLP 2026. https://arxiv.org/abs/2606.26350
  • Zhao, J., Guan, M., & Liu, D. (2026). Constraint-aware synthetic tabular data generation via inter-column constraint discovery with LLM agents (arXiv:2608.15109). arXiv. https://arxiv.org/abs/2608.15109

Working on AI that needs to ship?

I help funds, fintechs, and data teams take AI from prototype to production.