Skip to contents

This function processes a dataset for k-fold cross-validation using the multideconv framework. For each fold, it generates training and test datasets by computing deconvolution subgroups features from the deconvolution matrix. It also processes the entire dataset once to provide a final processed training set.

Usage

prepare_multideconv_folds(
  data,
  folds = NULL,
  bestune = NULL,
  ncores = NULL,
  cells_extra = NULL,
  corr = 0.7,
  corr_type = "spearman",
  zero_thr = 0.9,
  cv_thr = 0.1,
  batch = NULL
)

Arguments

data

A data frame of deconvolution features (samples x features) plus the outcome, as given by pipeML: a target column (classification) or time and event columns (survival). The outcome columns are not used to compute the subgroups; they are added back to the returned data.

folds

A list of integer vectors indicating row indices for the training set in each fold. The test set is implicitly defined as the complement.

bestune

Optional tuning object; when provided, folds are skipped and full-data processing is returned.

ncores

Number of CPU cores for parallel fold processing.

cells_extra

Optional character vector of additional cell labels to include.

corr

Minimum correlation threshold passed to compute.deconvolution.analysis().

corr_type

Correlation type passed to compute.deconvolution.analysis().

zero_thr

Maximum zero fraction passed to compute.deconvolution.analysis().

cv_thr

Minimum coefficient of variation passed to compute.deconvolution.analysis().

batch

Optional batch covariate passed to compute.deconvolution.analysis().

Value

  • When bestune is NULL (fold mode): invisibly, a named list of processed folds, each also saved to Results/fold_<fold name>.rds. Each fold contains:

    • train_data: Processed training data with cell group features and the outcome columns.

    • test_data: Test data projected into the learned cell group feature space (plus time and event for survival).

    • obs_test: True class labels (or survival time/event) for the test set.

    • rowIndex: Row indices corresponding to the test set.

    • fold_name: Fold name if provided in the folds list.

  • When bestune is provided: a list with the processed feature matrix for the full dataset (including the outcome columns), the full compute.deconvolution.analysis() output, and bestune.

Details

The function runs the compute.deconvolution.analysis() function on each fold's training set and uses the trained projection to compute the test set representation. It also runs multideconv on the full dataset to return the complete processed training set.