
pipeML
pipeML.RmdpipeML is a leakage-aware machine learning framework for
classification and survival tasks on high-dimensional biological data,
as described in Hurtado and Pancaldi,
2026. Its central feature is custom cross-validation fold
construction: features that depend on several samples at once (clusters,
co-expression modules, enrichment scores, network features) are
recomputed inside each fold, so the test samples never influence the
features the model is trained on.
Installation
To avoid GitHub API rate limit issues, set up a Personal Access Token (PAT) before installing:
# install.packages(c("usethis", "gitcreds"))
usethis::create_github_token()
gitcreds::gitcreds_set()Install pipeML from GitHub:
# install.packages("pak")
pak::pkg_install("VeraPancaldiLab/pipeML")Survival models also need the censored package
(install.packages("censored")). You don’t need to load it:
pipeML loads it when a survival task runs.
Core functions
-
compute_features.training.ML(): train and tune models on a training set with repeated stratified k-fold cross-validation, and select the best one. -
compute_prediction(): evaluate the selected model on a test set (AUROC and AUPRC with bootstrap confidence intervals, or C-index). -
compute_shap_values(): explain the selected model with SHAP values. -
compute_features.ML(): training and prediction in one step, when the test set is already prepared. -
get_curves()andplot_survival_performance(): ROC and precision-recall curves, and Kaplan-Meier curves by predicted risk group.
Results (plots and the fold files of custom workflows) are written to
a Results/ folder in the working directory.
Which workflow should I use?
| Workflow | When to use it | Tutorial |
|---|---|---|
| Standard | Your features are fixed values per sample (e.g. clinical variables, cell-type proportions computed per sample). | Classification, Survival analysis |
| Leave-One-Dataset-Out | Your samples come from several cohorts and you want to test generalization across them. | Multi-cohort validation (LODO) |
| Custom folds | Your features are computed from several samples at once (e.g. clustering, PCA, co-expression modules), so they must be recomputed inside each fold to avoid leakage. The feature construction can also have parameters to tune. | Leakage-aware custom cross-validation |
All workflows work for classification and survival tasks, and all of them support prediction on a test set and SHAP values.
Tutorials
Step-by-step tutorials are available in the Articles section of the navigation bar:
- Classification — train, tune and select classification models, predict on a test set, and train and predict in one step
- Survival analysis — the same workflow for time-to-event outcomes, with Kaplan-Meier curves by predicted risk group
- Interpreting models with SHAP values — explain the selected model globally and for single samples
- Multi-cohort validation (LODO) — leave-one-dataset-out validation across cohorts
- Leakage-aware custom cross-validation — recompute sample-dependent features inside each fold, with or without tunable parameters
Citation
If you use pipeML in a scientific publication, please
cite:
Hurtado, M., & Pancaldi, V. (2026). A new pipeline for cross-validation fold-aware machine learning prediction of clinical outcomes addresses hidden data-leakage in omics based ‘predictors’. bioRxiv. https://doi.org/10.64898/2026.03.12.711429
pipeML is built on top of caret,
tidymodels, parsnip and censored;
please also cite the packages of the algorithms you use.