Skip to contents

Run the full Monte Carlo RaCInG workflow

Usage

compute_racing_montecarlo(
  counts = NULL,
  output_folder = "~/Documents/racing/vignettes/",
  deconv = NULL,
  cc_network = NULL,
  fun_LR = min,
  cell_expr_profile = NULL,
  source = "source_genesymbol",
  target = "target_genesymbol",
  signed = FALSE,
  deconv_method = "Quantiseq",
  cbsx.name = NULL,
  cbsx.token = NULL,
  pt_idx = NULL,
  file_name = NULL,
  nPatients = "all",
  communication_type = "W",
  Ncells = 10000,
  Ngraphs = 100,
  Ndegree = 20,
  remove_direction = TRUE,
  norm = FALSE,
  input_data = NULL,
  ncores = 1
)

Arguments

counts

Gene-by-sample count matrix. Required when input_data is not supplied; ignored otherwise.

output_folder

Directory used to write intermediate and output files.

deconv

Optional patient-by-cell-type abundance matrix. If omitted, it is computed via multideconv::compute.deconvolution() followed by multideconv::compute.deconvolution.analysis() (which identifies and collapses correlated cell-type subgroups) and multideconv::standardize_celltype_colnames(). See prepare_input_files().

cc_network

Optional ligand-receptor prior network. If omitted, it is retrieved via liana::get_curated_omni(). See prepare_input_files().

fun_LR

Function used to combine ligand and receptor expression values.

cell_expr_profile

Optional gene-by-cell-type expression profile matrix. If omitted, it is estimated from counts and deconv via per-gene non-negative least squares. See prepare_input_files().

source, target

Column names to use as ligand and receptor identifiers in cc_network.

signed

Logical; if TRUE, also try to load a sign matrix.

deconv_method

Deconvolution method(s) used when deconv is not supplied.

cbsx.name, cbsx.token

Optional credentials for the deconvolution workflow.

pt_idx

Optional single patient index to simulate.

file_name

File stem used for intermediate files.

nPatients

Number of patients to process, or "all".

communication_type

Feature family to simulate: "D", "W", "TT", "CT", "GSCC", or a character vector of several of these (e.g. c("D", "W", "TT", "CT", "GSCC")). When more than one is given, every requested type is extracted from the same simulated graphs (via countAllTypes()) in one pass instead of re-simulating a fresh set of graphs per type – graph generation, not feature extraction, is the expensive part of a Monte Carlo run, so this amortizes that cost across every type requested. output in the return value is then a named list (one entry per type) instead of a single result.

Ncells

Number of cells per simulated graph.

Ngraphs

Number of Monte Carlo iterations.

Ndegree

Target average degree.

remove_direction

Logical; if TRUE, merge directionally equivalent features.

norm

Logical; if TRUE, also run a uniformized baseline simulation and express features as an enrichment ratio over it (isolates specificity from abundance/topology, at the cost of a second, noisier simulation pass – budget a larger Ngraphs if enabling this). Default FALSE returns the raw, abundance-weighted communication magnitude, which is simpler, cheaper (one pass), and less noisy at a given Ngraphs.

input_data

Optional named list of pre-computed input matrices as returned by prepare_input_files(). Must contain Lmatrix, Rmatrix, Cmatrix, LRmatrix, celltypes, ligands, and receptors. When supplied, the counts argument and all preprocessing parameters (deconv, cc_network, etc.) are ignored.

ncores

Number of cores to compute patients on in parallel. Passed through to runSim(); see its documentation for details.

Value

A list with the generated inputs and processed feature matrices.