
Prepare CellTFusion cross-validation folds for pipeML
prepare_celltfusion_folds.RdFold construction function for pipeML::compute_features.training.ML() (argument
fold_construction_fun) that computes CellTFusion latent factors within each
cross-validation fold, without information leakage: cell groups and latent factors are learned with
CellTFusion() on the training samples of each fold only, and the test samples of the fold are
projected onto them with project_test_factors(). It works for classification and survival tasks.
Usage
prepare_celltfusion_folds(
data,
folds = NULL,
bestune = NULL,
deconv,
raw.counts = NULL,
coldata = NULL,
batch = FALSE,
batch_id = NULL,
ncores = 1,
return = FALSE,
verbose = FALSE,
...
)Arguments
- data
Data frame provided by
pipeML: samples as rows and genes as columns (thefeatures_traingiven tocompute_features.training.ML(), i.e. the transposed count matrix), plus the outcome columns (targetfor classification,timeandeventfor survival).- folds
Named list with the training rows of each fold, provided by
pipeML.- bestune
Provided by
pipeMLwhen building the features of the final model;NULLwhile running the folds.- deconv
Deconvolution matrix of all the samples (samples x features, e.g. from
multideconv::compute.deconvolution()), with rows in the same order asdata. Deconvolution is computed independently for each sample, so computing it once for all samples does not leak information between folds.- raw.counts
Optional count matrix (genes x samples) with the samples in the same order as
data.pipeMLconverts the column names offeatures_trainwithmake.names()(e.g."HLA-A"becomes"HLA.A"), so gene symbols indatamay no longer match TF regulons and gene sets; passingraw.countskeeps the original gene symbols. IfNULL(default),t(data)is used.- coldata
Optional data frame of sample metadata with rows in the same order as
data. Required whenbatch = TRUE.- batch
Logical; whether to run
CellTFusion()in multi-cohort mode (seeCellTFusion()). The test samples of each cohort are then projected separately, centered on their own means. Default FALSE.- batch_id
Column of
coldatawith the cohort of each sample. Required whenbatch = TRUE.- ncores
Integer. Number of folds computed in parallel. Default 1 (sequential).
- return
Logical; in final mode, whether
CellTFusion()saves its outputs and plots inResults/. Outputs are never saved for the folds. Default FALSE.- verbose
Logical; whether
CellTFusion()prints its progress messages. Default FALSE.- ...
Other arguments passed to
CellTFusion()(e.g.normalized,TF.collection,min_targets_size,minMod,corr_mod,corr,pval,cancer_type).dt,tfsandpathwaysare not allowed, as they must be computed within each fold.
Value
In fold mode (
bestune = NULL): each fold is saved toResults/fold_<fold name>.rds, whichpipeMLreads back, as a list withtrain_data(latent factors of the training samples plus the outcome columns),test_data(latent factors of the test samples, plustimeandeventfor survival),obs_test(observed outcome of the test samples),rowIndex(rows of the test samples) andfold_name. The list of folds is returned invisibly.In final mode (
bestuneprovided): a list with (1) the latent factors of all the samples plus the outcome columns, (2) theCellTFusion()result on all the samples, available asCustom_outputin thepipeMLresult and used to compute the features of new samples withproject_test_factors(), and (3)bestune.
Details
Pass the fixed arguments (deconv, and optionally raw.counts, coldata, batch,
batch_id and any CellTFusion() argument) through fold_construction_args_fixed. Tunable
arguments (fold_construction_args_tunable) are not supported. With ncores > 1, the first fold is
computed before the others start in parallel, so that the TF and pathway collections cached in
Results/ are downloaded only once; the package must be installed for the parallel workers.
Examples
if (FALSE) { # \dontrun{
raw.counts <- CellTFusion::raw.counts.tuto
traitdata <- CellTFusion::traitdata.tuto
deconv <- CellTFusion::deconv.tuto # deconvolution of all the samples
# Classification
ml_res <- pipeML::compute_features.training.ML(
features_train = t(raw.counts),
task_type = "classification",
target_var = traitdata$Best.Confirmed.Overall.Response,
trait.positive = "PD",
metric = "AUROC",
k_folds = 5,
n_rep = 1,
fold_construction_fun = prepare_celltfusion_folds,
fold_construction_args_fixed = list(deconv = deconv, raw.counts = raw.counts)
)
# Features of new samples: projected onto the model trained on all the training samples
features_test <- project_test_factors(ml_res$Custom_output, deconv_test)
# Survival (time and event of each sample)
surv_res <- pipeML::compute_features.training.ML(
features_train = t(raw.counts),
task_type = "survival",
time_var = time,
event_var = event,
k_folds = 5,
n_rep = 1,
fold_construction_fun = prepare_celltfusion_folds,
fold_construction_args_fixed = list(deconv = deconv, raw.counts = raw.counts)
)
} # }