Skip to contents

Integrated open-data fixtures for rmorie, plus a small set of base-R helpers. Its core job remains the inst/extdata/ files used by rmorie’s examples, vignettes, and tests.

Install

# r-universe (recommended -- prebuilt binaries, no compiler needed)
install.packages(
  "rmoriedata",
  repos = c("https://rootcoder007.r-universe.dev",
            "https://cloud.r-project.org")
)

# or from GitHub source
# pak::pkg_install("rootcoder007/rmoriedata")
# remotes::install_github("rootcoder007/rmoriedata")

Quick start

library(rmoriedata)

# Browse and load a bundled table by slug.
cat <- morie_data_catalog()
head(cat[cat$kind == "table", c("slug", "n_rows", "n_cols")])
iucr <- morie_data_load("chicago_iucr_codes")

# Release a private count, then check a table for re-identification risk.
morie_dp_laplace_count(true_count = 42, epsilon = 1.0)
df <- data.frame(age = c(25, 25, 25, 40), sex = c("F", "F", "F", "M"))
morie_k_anonymity_verify(df, c("age", "sex"), k = 2)$summary

See vignette("rmoriedata") for the full tour.

Functions

A small, deliberate set of base-R helpers (the bulk of the package is still data):

Chicago sample data — R and Python

Bundled samples (complaint_sample, arrest_sample) from the City of Chicago open data; the full datasets are fetched on demand and cached.

library(rmoriedata)
data(complaint_sample)                       # bundled sample, offline
crimes <- load_chicago_data("complaints", full = TRUE)   # full data (Socrata)
pq <- load_chicago_data("arrests", as = "parquet_path")  # Parquet for Python
import pandas as pd
df = pd.read_parquet(pq)                      # path printed by load_chicago_data

# Or read the bundled .rda directly (no R install) — use the date_iso column,
# since POSIXct does not roundtrip cleanly through pyreadr:
import pyreadr
complaints = pyreadr.read_r("rmoriedata/data/complaint_sample.rda")["complaint_sample"]

Every bundled table also ships as Parquet: morie_data_path("<slug>") gives the verified path, pd.read_parquet() reads it.

Recommended Python bridge: the Parquet path (as = "parquet_path") — typed, columnar, read natively by pandas/polars/duckdb with correct timestamps and no R runtime. pyreadr on the .rda works for quick read-only access but is slower and mangles POSIXct (use date_iso).

Why a separate package?

CRAN packages have a 5 MB soft-cap on source tarball size. rmorie’s ~6.4 MB of integrated fixtures would push it over that threshold and trigger a reviewer pushback. Splitting the data out keeps rmorie lean (~few hundred KB of code) and lets data ship at any size via r-universe.

Citation

If you use rmoriedata in your research, please cite the software:

Ruhela, V. S. (2026). rmoriedata: Integrated Datasets for the rmorie Package. https://github.com/rootcoder007/rmoriedata

BibTeX (or run citation("rmoriedata") after installation for the entry stamped with the exact installed version, sourced from inst/CITATION):

@Manual{ruhela_rmoriedata_2026,
  title  = {rmoriedata: Integrated Datasets for the rmorie Package},
  author = {Ruhela, Vansh Singh},
  year   = {2026},
  url    = {https://github.com/rootcoder007/rmoriedata}
}

See CITATION.cff for the machine-readable metadata GitHub’s “Cite this repository” button uses.

License

AGPL-3.0-or-later. The fixtures themselves are public-domain or under permissive open-data licenses from their source portals; see the per-file headers in inst/extdata/ for attribution.