Released Datasets
A reverse-chronological log of STARR-OMOP dataset releases. The most recent release is listed first.
Each dataset name below links directly to the Google Cloud Console, pinned to that dataset in your active GCP session. Query costs are billed to your own active billing project.
Dataset Variants
Every release ships in four variants. The suffix on the dataset name identifies each one:
- Core (no suffix) — full PHI-scrubbed dataset: structured data plus clinical text and text-derived NLP concepts.
_1pcent— a random 1% sample of the core dataset, for testing and sandbox exploration._lite— structured data only; clinical text and NLP tables removed._1pcent_lite— a random 1% sample of the lite variant.
The _latest alias always points to the most recent release snapshot.
Latest Datasets
These stable aliases always resolve to the most recent release. Use these for ongoing work — they roll forward automatically with each new release.
Release History
The dated snapshots below are immutable — pin one for reproducible analyses.
June 2026
Snapshot cut on June 8, 2026 (_2026_06_08). The _latest aliases currently resolve to this release.
January 2026
Snapshot cut on January 18, 2026 (_2026_01_18).