Blog · NIH & research

NIH All of Us needs reproducible, leakage-safe methods — not just bigger models

All of Us is a landmark U.S. research platform. As ML papers proliferate on its data, methodological hygiene—leakage, calibration, reproducibility—matters as much as model capacity.

What All of Us is

The NIH All of Us Research Program aims to enroll one million or more participants and provide researchers with linked EHR, genomic, and survey data for discovery science. Program home: allofus.nih.gov.

Published case studies describe ML workflows on the Researcher Workbench for diverse prediction and phenotyping tasks— always under All of Us data-use and privacy rules.

Why reproducible tooling matters here

How the EHR Risk Framework helps

Use the workbench as a methods teaching / local research stack for leakage-aware binary risk tasks (task YAML, audits, calibration, Docker reproducibility). It does not replace All of Us access controls or IRB requirements— and must not process unauthorized PHI.

Sources

Try the workbench

How it works A–Z Live demo Quickstart GitHub

Free demo server — it may be slow. Check it with a small amount of data. For larger workloads or freer experimentation, run locally or on your own server. Open live demo

Feedback: support@larucare.com · Cite: DOI & CITATION.cff