Blog · NIH & research
NIH All of Us needs reproducible, leakage-safe methods — not just bigger models
All of Us is a landmark U.S. research platform. As ML papers proliferate on its data, methodological hygiene—leakage, calibration, reproducibility—matters as much as model capacity.
What All of Us is
The NIH All of Us Research Program aims to enroll one million or more participants and provide researchers with linked EHR, genomic, and survey data for discovery science. Program home: allofus.nih.gov.
Published case studies describe ML workflows on the Researcher Workbench for diverse prediction and phenotyping tasks— always under All of Us data-use and privacy rules.
Why reproducible tooling matters here
- Large, multi-modal data increases the chance of subtle leakage if index times are fuzzy
- Multi-site heterogeneity makes calibration and subgroup reporting essential
- Shared open methods help labs compare results without reinventing pipelines
How the EHR Risk Framework helps
Use the workbench as a methods teaching / local research stack for leakage-aware binary risk tasks (task YAML, audits, calibration, Docker reproducibility). It does not replace All of Us access controls or IRB requirements— and must not process unauthorized PHI.
Sources
- NIH All of Us Research Program. allofus.nih.gov
- PubMed literature on ML case studies using the All of Us Researcher Workbench (search current reviews).
Try the workbench
How it works A–Z Live demo Quickstart GitHub
Free demo server — it may be slow. Check it with a small amount of data. For larger workloads or freer experimentation, run locally or on your own server. Open live demo
Feedback: support@larucare.com · Cite: DOI & CITATION.cff