
Synthetic predictive-policing dataset with a known disparity
Source:R/fairness_simulation.R
morie_fairness_simulate_biased_crime_data.RdGenerates per-record data with area, group,
true_outcome (group-independent Bernoulli at base_rate),
detected (group-dependent), and risk_score (0–500,
shifted up by bias * 100 points for non-reference groups).
The bias input is the ground truth the audits should recover.
Usage
morie_fairness_simulate_biased_crime_data(
n = 2000L,
groups = c("A", "B"),
group_props = NULL,
n_areas = 20L,
base_rate = 0.3,
bias = 0.5,
seed = 0L
)Arguments
- n
Number of records.
- groups
Character vector of group labels (the first entry is treated as the reference group).
- group_props
Optional sampling proportions.
- n_areas
Number of areas (>= number of groups).
- base_rate
Reference-group favourable-outcome rate in 0–1.
- bias
Injected disparity in -1–1.
- seed
Reproducibility seed.
Examples
d <- morie_fairness_simulate_biased_crime_data(n = 100L, seed = 1L)
head(d)
#> area group true_outcome detected risk_score
#> 1 area_01 B 0 1 260.5669
#> 2 area_09 B 0 0 184.4432
#> 3 area_10 A 1 0 224.3807
#> 4 area_00 A 0 1 272.8203
#> 5 area_17 B 0 0 297.6111
#> 6 area_12 A 1 0 246.0729