Skip to contents

Run the full kernel-based RaCInG workflow

Usage

compute_racing_kernel(
  counts = NULL,
  output_folder = "~/Documents/racing/vignettes/",
  deconv = NULL,
  cc_network = NULL,
  fun_LR = min,
  cell_expr_profile = NULL,
  source = "source_genesymbol",
  target = "target_genesymbol",
  signed = FALSE,
  deconv_method = "Quantiseq",
  cbsx.name = NULL,
  cbsx.token = NULL,
  file_name = NULL,
  nPatients = "all",
  communication_type = "W",
  norm = FALSE,
  pt_idx = NULL,
  remove_direction = TRUE,
  input_data = NULL,
  zero_threshold = 1,
  expr_threshold = 10
)

Arguments

counts

Gene-by-sample count matrix. Required when input_data is not supplied; ignored otherwise.

output_folder

Directory used to write and read intermediate input files.

deconv

Optional patient-by-cell-type abundance matrix. If omitted, it is computed via multideconv::compute.deconvolution() followed by multideconv::compute.deconvolution.analysis() (which identifies and collapses correlated cell-type subgroups) and multideconv::standardize_celltype_colnames(). See prepare_input_files().

cc_network

Optional ligand-receptor prior network. If omitted, it is retrieved via liana::get_curated_omni(). See prepare_input_files().

fun_LR

Function combining a ligand's and receptor's expression values (as a length-2 vector) into one interaction strength. Default min: the weaker-expressed partner limits the interaction. Any function collapsing 2 values to 1 works, e.g. prod for the product of the two instead.

cell_expr_profile

Optional gene-by-cell-type expression profile matrix. If omitted, it is estimated from counts and deconv via per-gene non-negative least squares. See prepare_input_files().

source, target

Column names to use as ligand and receptor identifiers in cc_network.

signed

Logical; if TRUE, also try to load a sign matrix.

deconv_method

Deconvolution method(s) used when deconv is not supplied.

cbsx.name, cbsx.token

Optional credentials for the deconvolution workflow.

file_name

File stem used for intermediate input files.

nPatients

Number of patients to process, or "all".

communication_type

Feature family to compute ("D", "W", "TT", or "GSCC"), or a character vector of several of these (e.g. c("D", "W", "TT", "GSCC")). The kernel (compute_kernel()) is always computed exactly once regardless of how many types are requested – feature extraction from an already-computed kernel is cheap, so there is no need to call compute_racing_kernel() again just to get a different communication_type; request them all in one call instead. With more than one type, features in the return value is a named list (one data frame per type) instead of a single data frame.

norm

Logical; if TRUE, also compute a normalized baseline kernel and express features as an enrichment ratio over it (isolates specificity from abundance/topology). Default FALSE returns the raw, abundance-weighted communication magnitude.

pt_idx

Optional single patient index to process.

remove_direction

Logical; if TRUE, merge directionally equivalent features.

input_data

Optional named list of pre-computed input matrices as returned by prepare_input_files(). Must contain Lmatrix, Rmatrix, Cmatrix, LRmatrix, celltypes, ligands, and receptors. When supplied, the counts argument and all preprocessing parameters (deconv, cc_network, etc.) are ignored.

zero_threshold

Drop a feature once its fraction of zero-valued patients reaches this threshold. Default 1 only drops features that are zero for every patient (the original behavior); lower it (e.g. 0.9) to also drop merely zero-inflated features. See compute_kernel_features().

expr_threshold

Minimum expression (in linear TPM, not log-scale) a ligand or receptor must reach in its sending/receiving cell type to be kept. See prepare_input_files().

Value

A list with the kernel arrays and the derived feature matrix (or, for multiple communication_type entries, a named list of feature matrices) in features.