Skip to contents

Machine learning models using deconvolution subgroups

The deconvolution subgroups generated by multideconv can be used as input features for training machine learning (ML) models. However, since these subgroups are derived based on sample-level correlations, special care is needed to avoid data leakage. If you compute the full subgroup matrix on the entire dataset before splitting into training and test sets (or before performing k-fold cross-validation), the model may indirectly access information from the test set during training — a form of hidden data leakage described in Hurtado and Pancaldi (2026).

To address this issue, we provide the function prepare_multideconv_folds(), which ensures proper separation between training and test data during the computation of deconvolution subgroups. Within each cross-validation fold, the subgroups are learned only from the training samples of the fold and then projected onto its test samples with replicate_deconvolution_subgroups().

Below is an example using pipeML, an R package that supports custom fold construction: prepare_multideconv_folds() is passed to the argument fold_construction_fun. For more information about the function arguments, visit the documentation of pipeML.

Input data

  • deconv_train: deconvolution of the training samples (samples as rows, deconvolution features as columns), as returned by compute.deconvolution().
  • deconv_test: deconvolution of the test samples, with the same features.
  • traitData_train, traitData_test: clinical data of the training and test samples, with the response in a Response column ("R" for responders). Rows must be in the same order as the deconvolution tables.
library(pipeML)
library(multideconv)

# Make sure the clinical data follows the order of the deconvolution tables
traitData_train <- traitData_train[rownames(deconv_train), , drop = FALSE]
traitData_test  <- traitData_test[rownames(deconv_test), , drop = FALSE]

Training

compute_features.training.ML() builds the cross-validation folds and calls prepare_multideconv_folds() to compute the deconvolution subgroups of each fold. Arguments of prepare_multideconv_folds() (e.g. ncores, corr, corr_type, zero_thr, cv_thr) can be set with fold_construction_args_fixed; they are applied in every fold and in the final model:

res <- pipeML::compute_features.training.ML(features_train = deconv_train,
                                            target_var = traitData_train$Response,
                                            task_type = "classification",
                                            trait.positive = "R",
                                            metric = "AUROC",
                                            k_folds = 5,
                                            n_rep = 10,
                                            ncores = 3,
                                            return = FALSE,
                                            fold_construction_fun = prepare_multideconv_folds,
                                            fold_construction_args_fixed = list(ncores = 3))

To also use aggregated cell groups as features (see aggregate_cell_groups() in the Cell type subgroup analysis article), aggregate the same groups in both deconv_train and deconv_test before training, and list group names that are not in the nomenclature in cells_extra, e.g. fold_construction_args_fixed = list(ncores = 3, cells_extra = "Lymphocytes"). The groups are summed per sample, so aggregating all the samples before the cross-validation does not leak information between training and test samples.

  • res$Model is the selected model, trained on the deconvolution subgroups computed on all the training samples.
  • res$Custom_output is the output of compute.deconvolution.analysis() on all the training samples: the subgroups used by the final model.

Prediction

Project the test samples onto the deconvolution subgroups learned on the training samples, then predict:

# Replicate deconvolution subgroups in the test samples
dt_test <- multideconv::replicate_deconvolution_subgroups(deconv_res = res$Custom_output,
                                                          deconvolution_test = deconv_test)

# Predict on the test samples
pred <- pipeML::compute_prediction(model = res$Model,
                                   test_data = dt_test,
                                   target_var = traitData_test$Response,
                                   task_type = "classification",
                                   trait.positive = "R",
                                   return = TRUE)

pred$AUC   # AUROC and AUPRC with bootstrap confidence intervals

With return = TRUE, ROC and PR curves are saved in the Results/ folder.

Which deconvolution subgroups drive the predictions?

compute_shap_values() computes SHAP values for the final model (res$Model), trained on the deconvolution subgroups computed on all the training samples (res$Custom_output). Each subgroup therefore has the same definition for all samples. Everything else is taken from the trained model:

shap <- pipeML::compute_shap_values(model_trained = res$Model,
                                    task_type = "classification")

# Global importance of each deconvolution subgroup (mean |SHAP| across samples)
head(sort(colMeans(abs(shap)), decreasing = TRUE), 10)

SHAP values are in units of predicted probability of response: a positive value means the subgroup pushed the prediction of that sample towards response. The final model was trained on the samples it explains, so SHAP values describe how the model uses each subgroup on the training data; compare its training performance with the cross-validation performance to check that it does not overfit. See the pipeML vignette for plots of SHAP values per sample and across samples.

Survival outcomes

prepare_multideconv_folds() also works for survival analysis. pipeML gives it the survival time and event of the training samples (in the time and event columns of its data argument), so no extra argument is needed:

res_survival <- pipeML::compute_features.training.ML(features_train = deconv_train,
                                                     task_type = "survival",
                                                     time_var = traitData_train$time,
                                                     event_var = traitData_train$event,   # 1 = event, 0 = censored
                                                     k_folds = 5,
                                                     n_rep = 10,
                                                     ncores = 3,
                                                     fold_construction_fun = prepare_multideconv_folds,
                                                     fold_construction_args_fixed = list(ncores = 3))

dt_test <- multideconv::replicate_deconvolution_subgroups(deconv_res = res_survival$Custom_output,
                                                          deconvolution_test = deconv_test)

pred_survival <- pipeML::compute_prediction(model = res_survival$Model,
                                            test_data = dt_test,
                                            task_type = "survival",
                                            time_var = traitData_test$time,
                                            event_var = traitData_test$event)
pred_survival$c_index

NOTE: multideconv is built on top of existing frameworks and makes extensive use of the R packages immunedeconv (Sturm et al. (2019)) and omnideconv (Dietrich et al. (2024)). If you use multideconv in your work, please cite our package along with these foundational packages. We also encourage you to cite the individual deconvolution algorithms you employ in your analysis.

method license citation
quanTIseq free (BSD) Finotello, F., Mayer, C., Plattner, C., Laschober, G., Rieder, D., Hackl, H., …, Sopper, S. (2019). Molecular and pharmacological modulators of the tumor immune contexture revealed by deconvolution of RNA-seq data. Genome medicine, 11(1), 34. https://doi.org/10.1186/s13073-019-0638-6
EpiDISH free (GPL 2.0) Zheng SC, Breeze CE, Beck S, Teschendorff AE (2018). “Identification of differentially methylated cell-types in Epigenome-Wide Association Studies.” Nature Methods, 15(12), 1059. https://doi.org/10.1038/s41592-018-0213-x
DeconRNASeq free (GPL-2) TGJDS (2025). DeconRNASeq: Deconvolution of Heterogeneous Tissue Samples for mRNA-Seq data. doi:10.18129/B9.bioc.DeconRNASeq, R package version 1.50.0, https://bioconductor.org/packages/DeconRNASeq
AutoGeneS free (MIT) Aliee, H., & Theis, F. (2021). AutoGeneS: Automatic gene selection using multi-objective optimization for RNA-seq deconvolution. https://doi.org/10.1101/2020.02.21.940650
BayesPrism free (GPL 3.0) Chu, T., Wang, Z., Pe’er, D. et al. Cell type and gene expression deconvolution with BayesPrism enables Bayesian integrative analysis across bulk and single-cell RNA sequencing in oncology. Nat Cancer 3, 505–517 (2022). https://doi.org/10.1038/s43018-022-00356-3
Bisque free (GPL 3.0) Jew, B., Alvarez, M., Rahmani, E., Miao, Z., Ko, A., Garske, K. M., Sul, J. H., Pietiläinen, K. H., Pajukanta, P., & Halperin, E. (2020). Publisher Correction: Accurate estimation of cell composition in bulk expression through robust integration of single-cell information. Nature Communications, 11(1), 2891. https://doi.org/10.1038/s41467-020-16607-9
BSeq-sc free (GPL 2.0) Baron, M., Veres, A., Wolock, S. L., Faust, A. L., Gaujoux, R., Vetere, A., Ryu, J. H., Wagner, B. K., Shen-Orr, S. S., Klein, A. M., Melton, D. A., & Yanai, I. (2016). A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell Population Structure. In Cell Systems (Vol. 3, Issue 4, pp. 346–360.e4). https://doi.org/10.1016/j.cels.2016.08.011
CIBERSORTx free for non-commerical use only Newman, A. M., Liu, C. L., Green, M. R., Gentles, A. J., Feng, W., Xu, Y., Hoang, C. D., Diehn, M., & Alizadeh, A. A. (2015). Robust enumeration of cell subsets from tissue expression profiles. Nature Methods, 12(5), 453–457. https://doi.org/10.1038/s41587-019-0114-2
CPM free (GPL 2.0) Frishberg, A., Peshes-Yaloz, N., Cohn, O., Rosentul, D., Steuerman, Y., Valadarsky, L., Yankovitz, G., Mandelboim, M., Iraqi, F. A., Amit, I., Mayo, L., Bacharach, E., & Gat-Viks, I. (2019). Cell composition analysis of bulk genomics using single-cell data. Nature Methods, 16(4), 327–332. https://doi.org/10.1038/s41592-019-0355-5
DWLS free (GPL) Tsoucas, D., Dong, R., Chen, H., Zhu, Q., Guo, G., & Yuan, G.-C. (2019). Accurate estimation of cell-type composition from gene expression data. Nature Communications, 10(1), 2975. https://doi.org/10.1038/s41467-019-10802-z
MOMF free (GPL 3.0) Xifang Sun, Shiquan Sun, and Sheng Yang. An efficient and flexible method for deconvoluting bulk RNAseq data with single-cell RNAseq data, 2019, DOI: 10.5281/zenodo.3373980
MuSiC free (GPL 3.0) Wang, X., Park, J., Susztak, K., Zhang, N. R., & Li, M. (2019). Bulk tissue cell type deconvolution with multi-subject single-cell expression reference. Nature Communications, 10(1), 380. https://doi.org/10.1038/s41467-018-08023-x
SCDC (MIT) Dong, M., Thennavan, A., Urrutia, E., Li, Y., Perou, C. M., Zou, F., & Jiang, Y. (2020). SCDC: bulk gene expression deconvolution by multiple single-cell RNA sequencing references. Briefings in Bioinformatics. https://doi.org/10.1093/bib/bbz166

References

Dietrich, Alexander, Lorenzo Merotto, Konstantin Pelz, et al. 2024. “Benchmarking Second-Generation Methods for Cell-Type Deconvolution of Transcriptomic Data.” bioRxiv, ahead of print. https://doi.org/10.1101/2024.06.10.598226.
Hurtado, Marcelo, and Vera Pancaldi. 2026. “A New Pipeline for Cross-Validation Fold-Aware Machine Learning Prediction of Clinical Outcomes Addresses Hidden Data-Leakage in Omics Based ’Predictors’.” bioRxiv, ahead of print. https://doi.org/10.64898/2026.03.12.711429.
Sturm, Gregor, Francesca Finotello, Florent Petitprez, et al. 2019. “Comprehensive Evaluation of Transcriptome-Based Cell-Type Quantification Methods for Immuno-Oncology.” Bioinformatics 35 (14): i436–45. https://doi.org/10.1093/bioinformatics/btz363.