Skip to contents

CellTFusion latent factors can be used as features to predict clinical outcomes with the pipeML package:

# install.packages("pak")
pak::pkg_install("VeraPancaldiLab/pipeML")

Latent factors are learned from the data (cell groups, TF modules, NMF), so they must be learned within each cross-validation fold: learning them on all the training samples before cross-validation would let the test samples of each fold influence the features and overestimate performance. prepare_celltfusion_folds() does this for pipeML: in each fold, CellTFusion() is run on the training samples only, and the test samples are projected onto the resulting latent factors with project_test_factors().

Input data

Split the example data into a training and a test set. The deconvolution is computed once for all the samples: it is computed independently for each sample, so this does not leak information between samples.

raw.counts <- CellTFusion::raw.counts.tuto
traitdata  <- CellTFusion::traitdata.tuto

deconv <- multideconv::compute.deconvolution(
  raw.counts,
  methods    = c("Quantiseq", "Epidish"),
  normalized = TRUE,
  return     = FALSE
)

index <- caret::createDataPartition(traitdata$Best.Confirmed.Overall.Response, p = 0.8, list = FALSE)

raw.counts_train <- raw.counts[, index]
traitdata_train  <- traitdata[index, ]
deconv_train     <- deconv[index, ]

raw.counts_test  <- raw.counts[, -index]
traitdata_test   <- traitdata[-index, ]
deconv_test      <- deconv[-index, ]

Training

Pass prepare_celltfusion_folds() as fold_construction_fun, and its arguments through fold_construction_args_fixed:

  • deconv — deconvolution of the training samples, in the same order as features_train.
  • raw.counts — count matrix of the training samples. pipeML converts gene names with make.names() (e.g. HLA-A becomes HLA.A); passing the counts keeps the original gene symbols.
  • any other CellTFusion() argument, e.g. TF.collection, minMod or corr.
ml_res <- pipeML::compute_features.training.ML(
  features_train = t(raw.counts_train),
  task_type      = "classification",
  target_var     = traitdata_train$Best.Confirmed.Overall.Response,
  trait.positive = "PD",
  metric         = "AUROC",
  k_folds        = 5,
  n_rep          = 5,
  ncores         = 2,
  fold_construction_fun        = prepare_celltfusion_folds,
  fold_construction_args_fixed = list(deconv     = deconv_train,
                                      raw.counts = raw.counts_train,
                                      TF.collection = "CollecTRI")
)

The final model is trained on the latent factors learned on all the training samples. The corresponding CellTFusion() result is returned in ml_res$Custom_output.

Prediction on an independent cohort

Project the test samples onto the latent factors of ml_res$Custom_output, then predict:

features_test <- data.frame(project_test_factors(ml_res$Custom_output, deconv_test))

pred <- pipeML::compute_prediction(
  model          = ml_res$Model,
  test_data      = features_test,
  target_var     = traitdata_test$Best.Confirmed.Overall.Response,
  trait.positive = "PD",
  task_type      = "classification"
)

pred$AUC$AUROC

Survival outcomes

The same function is used for survival models, with the follow-up time and event of each sample:

surv_res <- pipeML::compute_features.training.ML(
  features_train = t(raw.counts_train),
  task_type      = "survival",
  time_var       = traitdata_train$PFS,
  event_var      = traitdata_train$PFS_event,
  k_folds        = 5,
  n_rep          = 5,
  ncores         = 2,
  fold_construction_fun        = prepare_celltfusion_folds,
  fold_construction_args_fixed = list(deconv = deconv_train, raw.counts = raw.counts_train)
)

features_test <- data.frame(project_test_factors(surv_res$Custom_output, deconv_test))

surv_pred <- pipeML::compute_prediction(
  model     = surv_res$Model,
  test_data = features_test,
  task_type = "survival",
  time_var  = traitdata_test$PFS,
  event_var = traitdata_test$PFS_event
)

surv_pred$c_index

Multi-cohort data

For samples from several cohorts, pass batch = TRUE, the metadata of the training samples (coldata, in the same order as features_train) and the cohort column (batch_id). CellTFusion() then corrects for cohort in each fold (see Multi-cohort analysis), and the test samples of each cohort are projected separately, centered on their own means. With LODO = TRUE, pipeML leaves one cohort out in each fold:

ml_res <- pipeML::compute_features.training.ML(
  features_train = t(raw.counts_train),
  task_type      = "classification",
  target_var     = coldata_train$Response,
  trait.positive = "R",
  metric         = "AUROC",
  LODO           = TRUE,
  batch_var      = coldata_train$Cohort,
  fold_construction_fun        = prepare_celltfusion_folds,
  fold_construction_args_fixed = list(deconv     = deconv_train,
                                      raw.counts = raw.counts_train,
                                      coldata    = coldata_train,
                                      batch      = TRUE,
                                      batch_id   = "Cohort")
)

For a model trained with batch = TRUE, project each new cohort separately with project_test_factors(), as its samples are centered on their own means.

Note: prepare_celltfusion_folds() runs CellTFusion() in every fold. Use ncores in fold_construction_args_fixed to compute several folds in parallel (the package must be installed).