Eigen RadarScience
Analysis

Synthetic criminology data kept 18 of 20 experimental effects close

A new criminology study generated synthetic responses for a survey experiment about profanity in police workplaces. In five models, 18 of 20 treatment coefficients remained within two original standard errors of the original estimates. The result offers a way to run and inspect code on restricted research data, while leaving the original observations unverified and disclosure risks to assess.

Science··Morning
At a bright research desk, two monitors show similar but distinct patterns of points and short lines, one in warm white and the other in blue.

The survey responses were recreated

A new peer-reviewed criminology study revisited a survey experiment about profanity in police workplaces. The original analysis contained 5,180 responses from 1,351 people, each assigned four of nine possible scenarios. Researchers retained the experimental assignments and the number of responses per person, but generated synthetic outcome responses. They could then test whether the original analysis code would run on a shareable dataset without releasing the original answers.[1]

Most coefficients stayed close, two did not

The team ran the same five mixed-effects models on the original and synthetic datasets. The synthetic estimates were close to the original estimates for 18 of 20 treatment coefficients, a descriptive benchmark rather than a universal validity threshold. The two larger departures concerned personal discipline: profanity directed at a colleague and derogatory intent. Even so, the direction of all coefficients and their classification by 95 percent confidence intervals agreed between the datasets.[1]

Similar results left disclosure checks necessary

For personal discipline, the residual correlation among responses from one person fell from 0.515 in the original data to 0.419 in the synthetic set. 26 participant response patterns appeared once in each dataset and matched; this does not mean anyone was identified. Synthetic records make the code and model choices easier to inspect. One synthetic version of one dataset still cannot authenticate the original observations, and a release requires separate checks of analytical usefulness and disclosure risk.[1]

References

  1. News sourceScienmagSynthetic crime survey data preserved most tested treatment effects↩1↩2↩3