Machine learning models using deconvolution subgroups
The deconvolution subgroups generated by multideconv can
be used as input features for training machine learning (ML) models.
However, since these subgroups are derived based on sample-level
correlations, special care is needed to avoid data leakage. If you
compute the full subgroup matrix on the entire dataset before splitting
into training and test sets (or before performing k-fold
cross-validation), the model may indirectly access information from the
test set during training — a form of hidden data leakage described in
Hurtado and Pancaldi (2026).
To address this issue, we provide the function
prepare_multideconv_folds(), which ensures proper
separation between training and test data during the computation of
deconvolution subgroups. Within each cross-validation fold, the
subgroups are learned only from the training samples of the fold and
then projected onto its test samples with
replicate_deconvolution_subgroups().
Below is an example using pipeML, an R package that
supports custom fold construction:
prepare_multideconv_folds() is passed to the argument
fold_construction_fun. For more information about the
function arguments, visit the documentation of pipeML.
Input data
-
deconv_train: deconvolution of the training samples (samples as rows, deconvolution features as columns), as returned bycompute.deconvolution(). -
deconv_test: deconvolution of the test samples, with the same features. -
traitData_train,traitData_test: clinical data of the training and test samples, with the response in aResponsecolumn ("R"for responders). Rows must be in the same order as the deconvolution tables.
library(pipeML)
library(multideconv)
# Make sure the clinical data follows the order of the deconvolution tables
traitData_train <- traitData_train[rownames(deconv_train), , drop = FALSE]
traitData_test <- traitData_test[rownames(deconv_test), , drop = FALSE]Training
compute_features.training.ML() builds the
cross-validation folds and calls
prepare_multideconv_folds() to compute the deconvolution
subgroups of each fold. Arguments of
prepare_multideconv_folds() (e.g. ncores,
corr, corr_type, zero_thr,
cv_thr) can be set with
fold_construction_args_fixed; they are applied in every
fold and in the final model:
res <- pipeML::compute_features.training.ML(features_train = deconv_train,
target_var = traitData_train$Response,
task_type = "classification",
trait.positive = "R",
metric = "AUROC",
k_folds = 5,
n_rep = 10,
ncores = 3,
return = FALSE,
fold_construction_fun = prepare_multideconv_folds,
fold_construction_args_fixed = list(ncores = 3))To also use aggregated cell groups as features (see
aggregate_cell_groups() in the Cell type subgroup
analysis article), aggregate the same groups in
both deconv_train and
deconv_test before training, and list group names that are
not in the nomenclature in cells_extra,
e.g. fold_construction_args_fixed = list(ncores = 3, cells_extra = "Lymphocytes").
The groups are summed per sample, so aggregating all the samples before
the cross-validation does not leak information between training and test
samples.
-
res$Modelis the selected model, trained on the deconvolution subgroups computed on all the training samples. -
res$Custom_outputis the output ofcompute.deconvolution.analysis()on all the training samples: the subgroups used by the final model.
Prediction
Project the test samples onto the deconvolution subgroups learned on the training samples, then predict:
# Replicate deconvolution subgroups in the test samples
dt_test <- multideconv::replicate_deconvolution_subgroups(deconv_res = res$Custom_output,
deconvolution_test = deconv_test)
# Predict on the test samples
pred <- pipeML::compute_prediction(model = res$Model,
test_data = dt_test,
target_var = traitData_test$Response,
task_type = "classification",
trait.positive = "R",
return = TRUE)
pred$AUC # AUROC and AUPRC with bootstrap confidence intervalsWith return = TRUE, ROC and PR curves are saved in the
Results/ folder.
Which deconvolution subgroups drive the predictions?
compute_shap_values() computes SHAP values for the final
model (res$Model), trained on the deconvolution subgroups
computed on all the training samples (res$Custom_output).
Each subgroup therefore has the same definition for all samples.
Everything else is taken from the trained model:
shap <- pipeML::compute_shap_values(model_trained = res$Model,
task_type = "classification")
# Global importance of each deconvolution subgroup (mean |SHAP| across samples)
head(sort(colMeans(abs(shap)), decreasing = TRUE), 10)SHAP values are in units of predicted probability of response: a
positive value means the subgroup pushed the prediction of that sample
towards response. The final model was trained on the samples it
explains, so SHAP values describe how the model uses each subgroup on
the training data; compare its training performance with the
cross-validation performance to check that it does not overfit. See the
pipeML vignette for plots of SHAP values per sample and
across samples.
Survival outcomes
prepare_multideconv_folds() also works for survival
analysis. pipeML gives it the survival time and event of
the training samples (in the time and event
columns of its data argument), so no extra argument is
needed:
res_survival <- pipeML::compute_features.training.ML(features_train = deconv_train,
task_type = "survival",
time_var = traitData_train$time,
event_var = traitData_train$event, # 1 = event, 0 = censored
k_folds = 5,
n_rep = 10,
ncores = 3,
fold_construction_fun = prepare_multideconv_folds,
fold_construction_args_fixed = list(ncores = 3))
dt_test <- multideconv::replicate_deconvolution_subgroups(deconv_res = res_survival$Custom_output,
deconvolution_test = deconv_test)
pred_survival <- pipeML::compute_prediction(model = res_survival$Model,
test_data = dt_test,
task_type = "survival",
time_var = traitData_test$time,
event_var = traitData_test$event)
pred_survival$c_indexNOTE: multideconv is built on top of
existing frameworks and makes extensive use of the R packages
immunedeconv (Sturm et al. (2019)) and
omnideconv (Dietrich et al. (2024)). If you use
multideconv in your work, please cite our package along
with these foundational packages. We also encourage you to cite the
individual deconvolution algorithms you employ in your analysis.
| method | license | citation |
|---|---|---|
| quanTIseq | free (BSD) | Finotello, F., Mayer, C., Plattner, C., Laschober, G., Rieder, D., Hackl, H., …, Sopper, S. (2019). Molecular and pharmacological modulators of the tumor immune contexture revealed by deconvolution of RNA-seq data. Genome medicine, 11(1), 34. https://doi.org/10.1186/s13073-019-0638-6 |
| EpiDISH | free (GPL 2.0) | Zheng SC, Breeze CE, Beck S, Teschendorff AE (2018). “Identification of differentially methylated cell-types in Epigenome-Wide Association Studies.” Nature Methods, 15(12), 1059. https://doi.org/10.1038/s41592-018-0213-x |
| DeconRNASeq | free (GPL-2) | joseph.szustakowski@novartis.com TGJDS (2025). DeconRNASeq: Deconvolution of Heterogeneous Tissue Samples for mRNA-Seq data. doi:10.18129/B9.bioc.DeconRNASeq, R package version 1.50.0, https://bioconductor.org/packages/DeconRNASeq |
| AutoGeneS | free (MIT) | Aliee, H., & Theis, F. (2021). AutoGeneS: Automatic gene selection using multi-objective optimization for RNA-seq deconvolution. https://doi.org/10.1101/2020.02.21.940650 |
| BayesPrism | free (GPL 3.0) | Chu, T., Wang, Z., Pe’er, D. et al. Cell type and gene expression deconvolution with BayesPrism enables Bayesian integrative analysis across bulk and single-cell RNA sequencing in oncology. Nat Cancer 3, 505–517 (2022). https://doi.org/10.1038/s43018-022-00356-3 |
| Bisque | free (GPL 3.0) | Jew, B., Alvarez, M., Rahmani, E., Miao, Z., Ko, A., Garske, K. M., Sul, J. H., Pietiläinen, K. H., Pajukanta, P., & Halperin, E. (2020). Publisher Correction: Accurate estimation of cell composition in bulk expression through robust integration of single-cell information. Nature Communications, 11(1), 2891. https://doi.org/10.1038/s41467-020-16607-9 |
| BSeq-sc | free (GPL 2.0) | Baron, M., Veres, A., Wolock, S. L., Faust, A. L., Gaujoux, R., Vetere, A., Ryu, J. H., Wagner, B. K., Shen-Orr, S. S., Klein, A. M., Melton, D. A., & Yanai, I. (2016). A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell Population Structure. In Cell Systems (Vol. 3, Issue 4, pp. 346–360.e4). https://doi.org/10.1016/j.cels.2016.08.011 |
| CIBERSORTx | free for non-commerical use only | Newman, A. M., Liu, C. L., Green, M. R., Gentles, A. J., Feng, W., Xu, Y., Hoang, C. D., Diehn, M., & Alizadeh, A. A. (2015). Robust enumeration of cell subsets from tissue expression profiles. Nature Methods, 12(5), 453–457. https://doi.org/10.1038/s41587-019-0114-2 |
| CPM | free (GPL 2.0) | Frishberg, A., Peshes-Yaloz, N., Cohn, O., Rosentul, D., Steuerman, Y., Valadarsky, L., Yankovitz, G., Mandelboim, M., Iraqi, F. A., Amit, I., Mayo, L., Bacharach, E., & Gat-Viks, I. (2019). Cell composition analysis of bulk genomics using single-cell data. Nature Methods, 16(4), 327–332. https://doi.org/10.1038/s41592-019-0355-5 |
| DWLS | free (GPL) | Tsoucas, D., Dong, R., Chen, H., Zhu, Q., Guo, G., & Yuan, G.-C. (2019). Accurate estimation of cell-type composition from gene expression data. Nature Communications, 10(1), 2975. https://doi.org/10.1038/s41467-019-10802-z |
| MOMF | free (GPL 3.0) | Xifang Sun, Shiquan Sun, and Sheng Yang. An efficient and flexible method for deconvoluting bulk RNAseq data with single-cell RNAseq data, 2019, DOI: 10.5281/zenodo.3373980 |
| MuSiC | free (GPL 3.0) | Wang, X., Park, J., Susztak, K., Zhang, N. R., & Li, M. (2019). Bulk tissue cell type deconvolution with multi-subject single-cell expression reference. Nature Communications, 10(1), 380. https://doi.org/10.1038/s41467-018-08023-x |
| SCDC | (MIT) | Dong, M., Thennavan, A., Urrutia, E., Li, Y., Perou, C. M., Zou, F., & Jiang, Y. (2020). SCDC: bulk gene expression deconvolution by multiple single-cell RNA sequencing references. Briefings in Bioinformatics. https://doi.org/10.1093/bib/bbz166 |
