
Machine learning workflows
a5_machine_learning.Rmd
library(CellTFusion)
#> CellTFusion latent factors can be used as features to
predict clinical outcomes with the pipeML
package:
# install.packages("pak")
pak::pkg_install("VeraPancaldiLab/pipeML")Latent factors are learned from the data (cell groups, TF modules,
NMF), so they must be learned within each cross-validation
fold: learning them on all the training samples before
cross-validation would let the test samples of each fold influence the
features and overestimate performance.
prepare_celltfusion_folds() does this for
pipeML: in each fold, CellTFusion() is run on
the training samples only, and the test samples are projected onto the
resulting latent factors with project_test_factors().
Input data
Split the example data into a training and a test set. The deconvolution is computed once for all the samples: it is computed independently for each sample, so this does not leak information between samples.
raw.counts <- CellTFusion::raw.counts.tuto
traitdata <- CellTFusion::traitdata.tuto
deconv <- multideconv::compute.deconvolution(
raw.counts,
methods = c("Quantiseq", "Epidish"),
normalized = TRUE,
return = FALSE
)
index <- caret::createDataPartition(traitdata$Best.Confirmed.Overall.Response, p = 0.8, list = FALSE)
raw.counts_train <- raw.counts[, index]
traitdata_train <- traitdata[index, ]
deconv_train <- deconv[index, ]
raw.counts_test <- raw.counts[, -index]
traitdata_test <- traitdata[-index, ]
deconv_test <- deconv[-index, ]Training
Pass prepare_celltfusion_folds() as
fold_construction_fun, and its arguments through
fold_construction_args_fixed:
-
deconv— deconvolution of the training samples, in the same order asfeatures_train. -
raw.counts— count matrix of the training samples.pipeMLconverts gene names withmake.names()(e.g.HLA-AbecomesHLA.A); passing the counts keeps the original gene symbols. - any other
CellTFusion()argument, e.g.TF.collection,minModorcorr.
ml_res <- pipeML::compute_features.training.ML(
features_train = t(raw.counts_train),
task_type = "classification",
target_var = traitdata_train$Best.Confirmed.Overall.Response,
trait.positive = "PD",
metric = "AUROC",
k_folds = 5,
n_rep = 5,
ncores = 2,
fold_construction_fun = prepare_celltfusion_folds,
fold_construction_args_fixed = list(deconv = deconv_train,
raw.counts = raw.counts_train,
TF.collection = "CollecTRI")
)The final model is trained on the latent factors learned on all the
training samples. The corresponding CellTFusion() result is
returned in ml_res$Custom_output.
Prediction on an independent cohort
Project the test samples onto the latent factors of
ml_res$Custom_output, then predict:
features_test <- data.frame(project_test_factors(ml_res$Custom_output, deconv_test))
pred <- pipeML::compute_prediction(
model = ml_res$Model,
test_data = features_test,
target_var = traitdata_test$Best.Confirmed.Overall.Response,
trait.positive = "PD",
task_type = "classification"
)
pred$AUC$AUROCSurvival outcomes
The same function is used for survival models, with the follow-up time and event of each sample:
surv_res <- pipeML::compute_features.training.ML(
features_train = t(raw.counts_train),
task_type = "survival",
time_var = traitdata_train$PFS,
event_var = traitdata_train$PFS_event,
k_folds = 5,
n_rep = 5,
ncores = 2,
fold_construction_fun = prepare_celltfusion_folds,
fold_construction_args_fixed = list(deconv = deconv_train, raw.counts = raw.counts_train)
)
features_test <- data.frame(project_test_factors(surv_res$Custom_output, deconv_test))
surv_pred <- pipeML::compute_prediction(
model = surv_res$Model,
test_data = features_test,
task_type = "survival",
time_var = traitdata_test$PFS,
event_var = traitdata_test$PFS_event
)
surv_pred$c_indexMulti-cohort data
For samples from several cohorts, pass batch = TRUE, the
metadata of the training samples (coldata, in the same
order as features_train) and the cohort column
(batch_id). CellTFusion() then corrects for
cohort in each fold (see Multi-cohort
analysis), and the test samples of each cohort are projected
separately, centered on their own means. With LODO = TRUE,
pipeML leaves one cohort out in each fold:
ml_res <- pipeML::compute_features.training.ML(
features_train = t(raw.counts_train),
task_type = "classification",
target_var = coldata_train$Response,
trait.positive = "R",
metric = "AUROC",
LODO = TRUE,
batch_var = coldata_train$Cohort,
fold_construction_fun = prepare_celltfusion_folds,
fold_construction_args_fixed = list(deconv = deconv_train,
raw.counts = raw.counts_train,
coldata = coldata_train,
batch = TRUE,
batch_id = "Cohort")
)For a model trained with batch = TRUE, project each new
cohort separately with project_test_factors(), as its
samples are centered on their own means.
Note: prepare_celltfusion_folds() runs
CellTFusion() in every fold. Use ncores in
fold_construction_args_fixed to compute several folds in
parallel (the package must be installed).