Skip to contents

Cell type processing

Deconvolution analysis reduces the dimensionality and heterogeneity of the deconvolution results. It uses the cell type processing algorithm described in the paper Hurtado et al., 2025. It returns the cell type subgroups composition and the reduced deconvolution matrix, saved in the Results/ directory.

Within each cell type, features (method-signature estimates) are grouped by complete-linkage hierarchical clustering of their correlations: features form a subgroup only if every pair of them correlates at least corr (non-significant correlations count as 0). Each subgroup is summarised by the median of its members.

Key parameters:

  • deconvolution: Matrix of raw deconvolution results (output of compute.deconvolution())
  • corr: Minimum correlation threshold to group features
  • return: Whether to return results and save output files to the Results/ directory
  • batch: Optional factor vector of cohort/batch labels (see section below on batch correction)
deconv_bulk = multideconv::deconv_bulk
deconv_subgroups = compute.deconvolution.analysis(deconvolution = deconv_bulk,
                                                  corr = 0.7,
                                                  file_name = "Tutorial",
                                                  return = TRUE)

The result is a named list with six elements:

names(deconv_subgroups)
#> [1] "Deconvolution matrix"                        
#> [2] "Deconvolution subgroups per cell types"      
#> [3] "Deconvolution subgroups composition"         
#> [4] "Discarded features with high number of zeros"
#> [5] "Discarded features with low variance"        
#> [6] "Discarded cell types"

Access the reduced deconvolution matrix (samples × subgroups):

head(deconv_subgroups[["Deconvolution matrix"]][, sample(ncol(deconv_subgroups[["Deconvolution matrix"]]), 5)])
#>                 DeconRNASeq_CBSX.Melanoma.scRNAseq_Endothelial
#> SAM7f0d9cc7f001                                      0.1247419
#> SAM4305ab968b90                                      0.1847182
#> SAMcf018fee2acd                                      0.1672997
#> SAMcc4675f394a1                                      0.2070710
#> SAM49f9b2e57aa5                                      0.1639431
#> SAM2e7aa8fa0ab3                                      0.1168617
#>                 DeconRNASeq_LM22_T.cells.helper CBSX_BPRNACanProMet_Neutrophils
#> SAM7f0d9cc7f001                      0.00000000                    0.0034759832
#> SAM4305ab968b90                      0.08389950                    0.0000000000
#> SAMcf018fee2acd                      0.00000000                    0.0029090514
#> SAMcc4675f394a1                      0.00000000                    0.0000000000
#> SAM49f9b2e57aa5                      0.06256773                    0.0000000000
#> SAM2e7aa8fa0ab3                      0.01772861                    0.0004672195
#>                 DeconRNASeq_CBSX.Melanoma.scRNAseq_CD4.cells
#> SAM7f0d9cc7f001                                   0.07161836
#> SAM4305ab968b90                                   0.07890048
#> SAMcf018fee2acd                                   0.07567607
#> SAMcc4675f394a1                                   0.07173726
#> SAM49f9b2e57aa5                                   0.07369131
#> SAM2e7aa8fa0ab3                                   0.05237417
#>                 CD4.cells_Subgroup.2
#> SAM7f0d9cc7f001           0.18051700
#> SAM4305ab968b90           0.00000000
#> SAMcf018fee2acd           0.19511491
#> SAMcc4675f394a1           0.06190600
#> SAM49f9b2e57aa5           0.15741680
#> SAM2e7aa8fa0ab3           0.07050168

Inspect subgroup composition (which methods/signatures were merged into each subgroup):

deconv_subgroups[["Deconvolution subgroups composition"]]$B.cells
#> $B.cells_Subgroup.1
#> [1] "CBSX_BPRNACan_B.cells"            "CBSX_BPRNACan3DProMet_B.cells"   
#> [3] "CBSX_BPRNACanProMet_B.cells"      "DWLS_BPRNACan_B.cells"           
#> [5] "DWLS_BPRNACan3DProMet_B.cells"    "DWLS_BPRNACanProMet_B.cells"     
#> [7] "Epidish_BPRNACan_B.cells"         "Epidish_BPRNACan3DProMet_B.cells"
#> [9] "Epidish_BPRNACanProMet_B.cells"  
#> 
#> $B.cells_Subgroup.2
#> [1] "CBSX_CBSX.HNSCC.scRNAseq_B.cells"       
#> [2] "DeconRNASeq_CBSX.HNSCC.scRNAseq_B.cells"
#> [3] "DWLS_CBSX.HNSCC.scRNAseq_B.cells"       
#> [4] "Epidish_CBSX.HNSCC.scRNAseq_B.cells"    
#> 
#> $B.cells_Subgroup.3
#> [1] "CBSX_CBSX.Melanoma.scRNAseq_B.cells"   
#> [2] "DWLS_CBSX.Melanoma.scRNAseq_B.cells"   
#> [3] "Epidish_CBSX.Melanoma.scRNAseq_B.cells"
#> 
#> $B.cells_Subgroup.4
#> [1] "DeconRNASeq_BPRNACan_B.cells"        
#> [2] "DeconRNASeq_BPRNACan3DProMet_B.cells"
#> [3] "DeconRNASeq_BPRNACanProMet_B.cells"  
#> 
#> $B.cells_Subgroup.5
#> [1] "DeconRNASeq_CCLE.TIL10_B.cells" "DeconRNASeq_TIL10_B.cells"     
#> 
#> $B.cells_Subgroup.6
#> [1] "DWLS_CBSX.NSCLC.PBMCs.scRNAseq_B.cells"   
#> [2] "Epidish_CBSX.NSCLC.PBMCs.scRNAseq_B.cells"
deconv_subgroups[["Deconvolution subgroups composition"]]$Macrophages.M2
#> $Macrophages.M2_Subgroup.1
#> [1] "CBSX_BPRNACan_Macrophages.M2"           
#> [2] "CBSX_BPRNACanProMet_Macrophages.M2"     
#> [3] "DWLS_BPRNACan_Macrophages.M2"           
#> [4] "DWLS_BPRNACan3DProMet_Macrophages.M2"   
#> [5] "DWLS_BPRNACanProMet_Macrophages.M2"     
#> [6] "Epidish_BPRNACan_Macrophages.M2"        
#> [7] "Epidish_BPRNACan3DProMet_Macrophages.M2"
#> [8] "Epidish_BPRNACanProMet_Macrophages.M2"  
#> 
#> $Macrophages.M2_Subgroup.2
#> [1] "CBSX_LM22_Macrophages.M2"    "DWLS_LM22_Macrophages.M2"   
#> [3] "Epidish_LM22_Macrophages.M2"
#> 
#> $Macrophages.M2_Subgroup.3
#> [1] "DeconRNASeq_BPRNACan_Macrophages.M2"        
#> [2] "DeconRNASeq_BPRNACan3DProMet_Macrophages.M2"
#> [3] "DeconRNASeq_BPRNACanProMet_Macrophages.M2"  
#> 
#> $Macrophages.M2_Subgroup.4
#> [1] "DeconRNASeq_CCLE.TIL10_Macrophages.M2"
#> [2] "DeconRNASeq_TIL10_Macrophages.M2"     
#> 
#> $Macrophages.M2_Subgroup.5
#> [1] "DWLS_CCLE.TIL10_Macrophages.M2"    "Epidish_CCLE.TIL10_Macrophages.M2"
#> 
#> $Macrophages.M2_Subgroup.6
#> [1] "DWLS_TIL10_Macrophages.M2"    "Epidish_TIL10_Macrophages.M2"
deconv_subgroups[["Deconvolution subgroups composition"]]$Dendritic.cells
#> $Dendritic.cells_Subgroup.1
#> [1] "CBSX_CBSX.HNSCC.scRNAseq_Dendritic.cells"   
#> [2] "DWLS_CBSX.HNSCC.scRNAseq_Dendritic.cells"   
#> [3] "Epidish_CBSX.HNSCC.scRNAseq_Dendritic.cells"

If your deconvolution matrix contains non-standard cell types (see README), specify them using cells_extra to ensure proper subgrouping. If not specified, they will be discarded automatically.

deconv_subgroups = compute.deconvolution.analysis(deconvolution = deconv_pseudo,
                                                  corr = 0.7,
                                                  return = TRUE,
                                                  cells_extra = c("Mural.cells", "Myeloid.cells"),
                                                  file_name = "Tutorial")

Aggregating cell types into groups

Signatures do not describe cell types with the same level of detail: one reports Macrophages.M1, Macrophages.M2 and Monocytes, another one only Myeloid.cells. To compare them at the same level, or to match a coarser ground truth (e.g. H&E or flow cytometry), you can sum cell types into groups with aggregate_cell_groups(). It is an optional step after compute.deconvolution(), independent of the subgroup analysis: it takes a deconvolution matrix and returns it with the group features added.

cell_groups is a named list: each name is a group and each element contains the cell types to sum, written as in get_cell_type_nomenclature(). For every method-signature combination, the function adds a new feature <method>_<signature>_<group>:

  • Cell types are summed only within the same method-signature combination, where the estimates are proportions of the same sample.
  • The original features are kept.
  • A combination that already estimates the group (e.g. a signature with its own Myeloid.cells) is left as it is.
  • A combination needs at least min_types = 2 cell types of the group, otherwise it is skipped (the group would be a copy of a single cell type).

The function prints which cell types were summed in each combination:

myeloid = list(Myeloid.cells = c("Macrophages.cells", "Macrophages.M0", "Macrophages.M1", "Macrophages.M2",
                                 "Monocytes", "Dendritic.cells", "Dendritic.activated.cells",
                                 "Dendritic.resting.cells"))

deconv_groups = aggregate_cell_groups(deconv_bulk, cell_groups = myeloid)
#> 
#> Group 'Myeloid.cells'
#>   Quantiseq: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   DeconRNASeq_BPRNACan: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   Epidish_BPRNACan: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DeconRNASeq_BPRNACan3DProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   Epidish_BPRNACan3DProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DeconRNASeq_BPRNACanProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   Epidish_BPRNACanProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DeconRNASeq_CBSX.HNSCC.scRNAseq: Macrophages.cells + Dendritic.cells
#>   Epidish_CBSX.HNSCC.scRNAseq: Macrophages.cells + Dendritic.cells
#>   DeconRNASeq_CBSX.Melanoma.scRNAseq: skipped (1 member: Macrophages.cells)
#>   Epidish_CBSX.Melanoma.scRNAseq: skipped (1 member: Macrophages.cells)
#>   DeconRNASeq_CBSX.NSCLC.PBMCs.scRNAseq: skipped (1 member: Monocytes)
#>   Epidish_CBSX.NSCLC.PBMCs.scRNAseq: skipped (1 member: Monocytes)
#>   DeconRNASeq_CCLE.TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   Epidish_CCLE.TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   DeconRNASeq_TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   Epidish_TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   DWLS_BPRNACan: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DWLS_BPRNACan3DProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DWLS_BPRNACanProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   DWLS_CBSX.HNSCC.scRNAseq: Macrophages.cells + Dendritic.cells
#>   DWLS_CBSX.Melanoma.scRNAseq: skipped (1 member: Macrophages.cells)
#>   DWLS_CBSX.NSCLC.PBMCs.scRNAseq: skipped (1 member: Monocytes)
#>   DWLS_CCLE.TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   DWLS_TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   CBSX_BPRNACan: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   CBSX_BPRNACan3DProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   CBSX_BPRNACanProMet: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes
#>   CBSX_CBSX.HNSCC.scRNAseq: Macrophages.cells + Dendritic.cells
#>   CBSX_CBSX.Melanoma.scRNAseq: skipped (1 member: Macrophages.cells)
#>   CBSX_CBSX.NSCLC.PBMCs.scRNAseq: skipped (1 member: Monocytes)
#>   CBSX_CCLE.TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   CBSX_TIL10: Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.cells
#>   DeconRNASeq_LM22: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.activated.cells + Dendritic.resting.cells
#>   Epidish_LM22: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.activated.cells + Dendritic.resting.cells
#>   DWLS_LM22: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.activated.cells + Dendritic.resting.cells
#>   CBSX_LM22: Macrophages.M0 + Macrophages.M1 + Macrophages.M2 + Monocytes + Dendritic.activated.cells + Dendritic.resting.cells
#> 
#> To use the aggregated matrix in other functions:
#>   - replicate_deconvolution_subgroups(): aggregate the same groups in the new cohort first, otherwise their features are set to NA

A group name can be a cell type of the nomenclature (Myeloid.cells) or a new name (Lymphocytes); new names must not contain _ nor the name of an existing cell type.

The returned matrix is a deconvolution matrix as any other, so it can be given to compute.deconvolution.analysis() if you want the groups to be analysed too. They are then treated as one more cell type.

IMPORTANT: group names that are not in the nomenclature (here Lymphocytes) must be listed in cells_extra, otherwise these groups are discarded by compute.deconvolution.analysis() (and by compute.benchmark() and prepare_multideconv_folds()). If in doubt, list all your group names: the ones already in the nomenclature are simply ignored. aggregate_cell_groups() prints this reminder after the summed cell types.

groups = list(Myeloid.cells = myeloid$Myeloid.cells,
              Lymphocytes = c("B.cells", "B.naive.cells", "B.memory.cells", "Plasma",
                              "CD4.cells", "CD4.memory.activated", "CD4.memory.resting", "CD4.naive",
                              "CD4.regulatory", "CD4.non.regulatory", "CD8.cells",
                              "NK.cells", "NK.activated", "NK.resting"))

deconv_groups = aggregate_cell_groups(deconv_bulk, cell_groups = groups, verbose = FALSE)

deconv_subgroups_groups = compute.deconvolution.analysis(deconvolution = deconv_groups,
                                                         corr = 0.7,
                                                         cells_extra = "Lymphocytes")

deconv_subgroups_groups[["Deconvolution subgroups composition"]]$Myeloid.cells
#> $Myeloid.cells_Subgroup.1
#> [1] "CBSX_BPRNACan_Myeloid.cells"           
#> [2] "CBSX_BPRNACan3DProMet_Myeloid.cells"   
#> [3] "CBSX_BPRNACanProMet_Myeloid.cells"     
#> [4] "Epidish_BPRNACan_Myeloid.cells"        
#> [5] "Epidish_BPRNACan3DProMet_Myeloid.cells"
#> [6] "Epidish_BPRNACanProMet_Myeloid.cells"  
#> 
#> $Myeloid.cells_Subgroup.2
#> [1] "CBSX_CBSX.HNSCC.scRNAseq_Myeloid.cells"   
#> [2] "DWLS_CBSX.HNSCC.scRNAseq_Myeloid.cells"   
#> [3] "Epidish_CBSX.HNSCC.scRNAseq_Myeloid.cells"
#> 
#> $Myeloid.cells_Subgroup.3
#> [1] "CBSX_CCLE.TIL10_Myeloid.cells"    "DWLS_CCLE.TIL10_Myeloid.cells"   
#> [3] "Epidish_CCLE.TIL10_Myeloid.cells"
#> 
#> $Myeloid.cells_Subgroup.4
#> [1] "CBSX_LM22_Myeloid.cells"        "DeconRNASeq_LM22_Myeloid.cells"
#> [3] "DWLS_LM22_Myeloid.cells"        "Epidish_LM22_Myeloid.cells"    
#> 
#> $Myeloid.cells_Subgroup.5
#> [1] "DeconRNASeq_CCLE.TIL10_Myeloid.cells"
#> [2] "DeconRNASeq_TIL10_Myeloid.cells"     
#> 
#> $Myeloid.cells_Subgroup.6
#> [1] "DWLS_BPRNACan_Myeloid.cells"         "DWLS_BPRNACan3DProMet_Myeloid.cells"
#> [3] "DWLS_BPRNACanProMet_Myeloid.cells"  
#> 
#> $Myeloid.cells_Subgroup.7
#> [1] "DWLS_TIL10_Myeloid.cells"    "Epidish_TIL10_Myeloid.cells"

IMPORTANT: if you later replicate these subgroups in a new cohort with replicate_deconvolution_subgroups() (see below), run aggregate_cell_groups() with the same cell_groups on the new deconvolution matrix first. Otherwise the group features do not exist in the new cohort and are returned as NA (with a warning).

deconv_new_groups = aggregate_cell_groups(deconv_new, cell_groups = groups)
deconv_new_subgroups = replicate_deconvolution_subgroups(deconv_subgroups_groups, deconv_new_groups)

NOTE: A group can list cell types of different levels of detail, because each signature reports its own. Check the printed lines: within one method-signature combination the summed cell types must not overlap (a cell type together with its own subtypes would be counted twice).

Handling batch effects (multiple cohorts)

When your samples come from multiple cohorts or batches, simple Pearson/Spearman correlations can be confounded by cohort structure. compute.deconvolution.analysis() accepts a batch argument that switches the internal correlation to partial correlation, controlling for cohort membership in the correlations used to build the subgroups.

The batch vector must be a factor or character vector with one label per sample, in the same order as the rows of the deconvolution matrix.

# Example: 'cohort' has one label per sample ("CohortA" / "CohortB"), in the same order as rownames(deconv_bulk)
cohort <- c(rep("CohortA", 96), rep("CohortB", 96))

deconv_subgroups_batch = compute.deconvolution.analysis(
  deconvolution = deconv_bulk,
  corr          = 0.7,
  batch         = cohort,
  file_name     = "Tutorial_batch",
  return        = TRUE
)

When batch is supplied:

  • Correlations between features are partial correlations that remove the batch effect (via ppcor::pcor.test(), with one indicator per batch).
  • Subgroups are built from these partial correlations, so inter-cohort differences do not inflate feature similarity.

This approach is recommended whenever samples originate from distinct studies, sequencing runs, or processing pipelines, as it prevents cohort-specific signals from being mistaken for biologically meaningful co-variation.

Characterizing subgroups with pathway activities

Once subgroups are identified, you can interpret their biological meaning by correlating each subgroup’s abundance profile across samples with pathway activity scores. The function compute.subgroup.pathways() generates one heatmap per cell type showing how each subgroup correlates with each pathway.

compute.subgroup.pathways() expects a pre-computed sample × pathway numeric matrix. As an example, we are going to compute pathway activities using the PROGENy database (Schubert et al., 2018), making use of the package CellTFusion. PROGENy models the activity of 14 cancer-relevant signalling pathways from gene expression data.

Schubert, M., Klinger, B., Klünemann, M., Sieber, A., Uhlitz, F., Sauer, S., Garnett, M. J., Blüthgen, N., & Saez-Rodriguez, J. (2018). Perturbation-response genes reveal signaling footprints in cancer gene expression. Nature Communications, 9(1), 20. https://doi.org/10.1038/s41467-017-02391-6

# Install CellTFusion if needed (once):
# pak::pkg_install("VeraPancaldiLab/CellTFusion")
library(CellTFusion)

counts     <- multideconv::raw_counts
counts_tpm <- ADImpute::NormalizeTPM(counts, log = FALSE)

# compute.pathway.activity() returns a sample x pathway activity matrix
pathway_scores <- compute.pathway.activity(counts_tpm)

compute.subgroup.pathways(
  subgroups = deconv_subgroups,
  pathways  = pathway_scores,
  file_name = "Tutorial",
  pval      = 0.05
)

Any other sample × pathway matrix (e.g. from GSVA, ssGSEA, or decoupleR) can be passed as pathways in the same way.

One PDF heatmap per cell type is saved to Results/. The function also returns (invisibly) a list with the correlations and p-values of each cell type, and corr_type selects Pearson (default) or Spearman correlations. Below is an example output for CD4 T cells (12 subgroups × 14 PROGENy pathways), where stars indicate significance levels (* p<0.05, ** p<0.01, *** p<0.001):

Replicate deconvolution subgroups in an independent set

Cell subgroup identification through deconvolution is cohort-specific, as it relies on correlation patterns across samples. This means that subgroup definitions may vary across different splits or datasets. If you aim to replicate the same subgroups identified in one dataset onto another (e.g., for model validation), you can use the following function.

The function below reconstructs and applies the subgroup signatures derived from a previous deconvolution, making it especially useful when transferring learned patterns across datasets — such as when training and evaluating machine learning models.

deconv_1 = deconv_bulk[1:100,]
deconv_2 = deconv_bulk[101:192,]

deconv_subgroups = compute.deconvolution.analysis(deconvolution = deconv_1,
                                                  corr = 0.7,
                                                  file_name = "Tutorial",
                                                  return = FALSE)
deconv_subgroups_replicate = replicate_deconvolution_subgroups(deconv_subgroups,
                                                               deconv_2)