Skip to contents

pipeML is a leakage-aware machine learning framework for classification and survival tasks on high-dimensional biological data, as described in Hurtado and Pancaldi, 2026. Its central feature is custom cross-validation fold construction: features that depend on several samples at once (clusters, co-expression modules, enrichment scores, network features) are recomputed inside each fold, so the test samples never influence the features the model is trained on.

Installation

To avoid GitHub API rate limit issues, set up a Personal Access Token (PAT) before installing:

# install.packages(c("usethis", "gitcreds"))
usethis::create_github_token()
gitcreds::gitcreds_set()

Install pipeML from GitHub:

# install.packages("pak")
pak::pkg_install("VeraPancaldiLab/pipeML")

Survival models also need the censored package (install.packages("censored")). You don’t need to load it: pipeML loads it when a survival task runs.

Core functions

Results (plots and the fold files of custom workflows) are written to a Results/ folder in the working directory.

Which workflow should I use?

Workflow When to use it Tutorial
Standard Your features are fixed values per sample (e.g. clinical variables, cell-type proportions computed per sample). Classification, Survival analysis
Leave-One-Dataset-Out Your samples come from several cohorts and you want to test generalization across them. Multi-cohort validation (LODO)
Custom folds Your features are computed from several samples at once (e.g. clustering, PCA, co-expression modules), so they must be recomputed inside each fold to avoid leakage. The feature construction can also have parameters to tune. Leakage-aware custom cross-validation

All workflows work for classification and survival tasks, and all of them support prediction on a test set and SHAP values.

Tutorials

Step-by-step tutorials are available in the Articles section of the navigation bar:

  • Classification — train, tune and select classification models, predict on a test set, and train and predict in one step
  • Survival analysis — the same workflow for time-to-event outcomes, with Kaplan-Meier curves by predicted risk group
  • Interpreting models with SHAP values — explain the selected model globally and for single samples
  • Multi-cohort validation (LODO) — leave-one-dataset-out validation across cohorts
  • Leakage-aware custom cross-validation — recompute sample-dependent features inside each fold, with or without tunable parameters

Citation

If you use pipeML in a scientific publication, please cite:

Hurtado, M., & Pancaldi, V. (2026). A new pipeline for cross-validation fold-aware machine learning prediction of clinical outcomes addresses hidden data-leakage in omics based ‘predictors’. bioRxiv. https://doi.org/10.64898/2026.03.12.711429

pipeML is built on top of caret, tidymodels, parsnip and censored; please also cite the packages of the algorithms you use.