What did the reproduction preserve?
The peer-reviewed criminology paper reconstructs a preregistered survey experiment about profanity in police workplaces. The original analysis has 5,180 responses from 1,351 people; each person received four of nine possible scenarios. The researchers kept the experimental assignments and the number of responses per person, synthesized the outcome responses with synthpop, and ran the same five mixed-effects models on both datasets. This comparison measures how closely those models behave on synthetic responses; it cannot authenticate the original answers.[1]
18 of the 20 treatment coefficients fell within two original standard errors of the original estimates. The authors use two as a descriptive comparison, not a universal validity threshold. Both coefficients outside it concern personal discipline: colleague-directed profanity and derogatory intent. Yet all 20 coefficient directions and 95% interval classifications agreed across datasets. Preserving the main pattern is a useful result; this one synthesis did not preserve every coefficient equally.[1]
Model agreement is not a release permit
For personal discipline, the residual correlation among responses from the same person drops from 0.515 in the original data to 0.419 in the synthetic data. Coefficients can look similar while the grouping of answers within people changes. The paper also counts 26 respondent patterns that occur once in each dataset and match. That is not a count of identified people; it flags a possible disclosure route when outside information is available. The assignments were retained, so the release is only partly synthetic.[1]
The example shows both the value and the boundary of synthetic data. Code becomes runnable, model choices visible, and alternative analyses possible. Similar coefficients in one synthetic version of one dataset, however, leave the original observations unverified. A stronger release decision rests on utility checks for each intended analysis, disclosure review, and a justified route to the originals. For repetition, the useful next question is how the two discipline coefficients and within-person dependence vary across new synthetic versions.[1]