An R package providing single-cell RNA-seq utility functions to complement Seurat-based workflows.
# install.packages("remotes")
remotes::install_github("jmgs7/SCutils")or
# install.packages("pak")
pak::pak("jmgs7/SCutils")Batch-opens .h5 single-cell matrices, writes/loads BPCells matrix directories (*_BP), supports both 10X and AnnData HDF5 inputs, and can optionally convert ENSEMBL IDs to gene symbols; it can also generate per-cell sample provenance metadata. Multiprocessing is supported via future framework. Future options must be set in the global environment before calling BatchOpenH5().
BatchOpenH5(
files,
relative = TRUE,
BP.data.dir = NULL,
platform = "10X",
use.names = TRUE,
ensembl.to.symbol = FALSE,
species = "human",
generate.metadata = FALSE
)| Parameter | Default | Description |
|---|---|---|
files |
— | Character vector of input .h5 file paths |
relative |
TRUE |
Store matrix directory paths as relative paths inside each matrix object |
BP.data.dir |
NULL |
Output directory for *_BP folders; defaults to dirname(files[[1]]) |
platform |
"10X" |
Input format: "10X" or "anndata" |
use.names |
TRUE |
For 10X only: replace feature IDs with /matrix/features/name |
ensembl.to.symbol |
FALSE |
Convert ENSEMBL IDs to symbols via ConvertEnsembleToSymbol2() (applied only when use.names = FALSE) |
species |
"human" |
Species passed to symbol conversion: "human" or "mouse" |
generate.metadata |
FALSE |
If TRUE, returns a named list with data.list plus per-cell metadata (cell.tag, sample.procedence) |
Calculates the Cellular Detection Rate (fraction of detected features per cell), z-scales it, and adds a CDR column to SeuratObject@meta.data. Handles both Seurat v3 (@assays$RNA@counts) and v5/split-layer (@assays$RNA@layers$counts) object formats.
SeuratObject <- CalculateCDR(SeuratObject)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object |
Computes common single-cell QC metrics and appends them directly to SeuratObject@meta.data.
Always-computed columns (from raw counts):
percent.mt: The percentage of mitochondrial gene counts per cell, computed from the^MT-gene pattern (human) or^mt-(mouse).percent.ribo: The percentage of ribosomal gene counts per cell, computed from the^RP[SL]gene pattern (human) or^Rp[sl](mouse).percent.hb: The percentage of hemoglobin gene counts per cell, computed from the^HB[^(P)]"gene pattern (human) or^Hb^(p)]"(mouse).percent.ig: The percentage of immunoglobulin gene counts per cell, computed from the^IGgene pattern (human) or^Ig(mouse).percent.plat: The percentage of platelet gene counts per cell, computed from the^PPBP|^PF4gene pattern (human) or^Ppbp|^Pf4(mouse).percent.MALAT1: The percentage of MALAT1 gene counts per cell, computed from theMALAT1gene pattern (human) orMalat1(mouse).percent.S100A9: The percentage of S100A9 gene counts per cell, computed from theS100A9gene pattern (human) orS100a9(mouse).percent.S100A8: The percentage of S100A8 gene counts per cell, computed from theS100A8gene pattern (human) orS100a8(mouse).percent.FCGR3B: The percentage of FCGR3B gene counts per cell, computed from theFCGR3Bgene pattern (human) orFcgr3b(mouse).log10_nFeature_RNA: The log10 of the number of features per cell, computed from thenFeature_RNAcolumn.log10_nCount_RNA: The log10 of the number of counts per cell, computed from thenCount_RNAcolumn.complexity: The complexity of the library per cell, computed aslog10(nFeature_RNA) / log10(nCount_RNA).
Conditionally computed (requires normalized counts in data layer):
nCount_logRNA,nFeature_logRNA: log-normalized count and detected feature totals per cell.- Cell cycle scoring (S and G2M scores,
Phase) — whenperform.cell.cycle.scoring = TRUE. - MALAT1 thresholding — when
perform.MALAT1.test = TRUE; addsMALAT1.threshold(numeric) andMALAT1.pass(logical) columns. See the feature_threshold repository for details.
The function supports split Seurat v5 objects: if normalized counts are stored in standard data.* layers or a concrete layer vector is passed via layers, metrics are computed per-layer.
SeuratObject <- CalculateQC(
SeuratObject,
assay = NULL,
layers = NULL,
species = "human" , # Also 'mouse' available
perform.cell.cycle.scoring = TRUE,
perform.MALAT1.test = TRUE
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | Seurat object to annotate |
assay |
NULL |
Assay to read counts from; if NULL, the default assay is used |
layers |
NULL |
Normalized layer(s) to use; NULL uses the default data layer; a character vector processes each layer independently |
species |
"human" |
Species of the dataset, used to select the appropriate gene sets for cell cycle scoring. Human and Mouse currently supported. |
perform.cell.cycle.scoring |
TRUE |
If TRUE, performs cell cycle scoring |
perform.MALAT1.test |
TRUE |
If TRUE, performs MALAT1-based QC thresholding |
... |
Additional arguments forwarded to the internal .CalculateFeatureThresholdSeurat() |
Creates a bar plot showing the number of cells per group, with bars filled by a viridis gradient based on cell count.
CellsHistoGradient(
SeuratObject,
group.by = NULL,
scale.colors = "viridis",
breaks = scales::extended_breaks()
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object |
group.by |
NULL |
Metadata column for grouping; if NULL, uses active identity |
scale.colors |
"viridis" |
Viridis palette option: "magma" ("A"), "inferno" ("B"), "plasma" ("C"), "viridis" ("D"), "cividis" ("E"), "rocket" ("F"), "mako" ("G"), "turbo" ("H") |
breaks |
scales::extended_breaks() |
Y-axis break specification |
Extends Seurat::FeatureScatter() by coloring points with a continuous viridis gradient defined by a third feature. Supports optional per-group panels with independent per-group correlation subtitles, configurable correlation methods, per-feature layer selection for all three axes, and a single shared legend across grouped panels. The gradient scale is computed globally across all panels.
FeatureScatterGradient(
SeuratObject,
feature1,
feature2,
gradient,
group.by = "ident",
scale.colors = "viridis",
lower.limit = NULL,
upper.limit = NULL,
corr.method = "pearson",
layer1 = NULL,
layer2 = NULL,
layer.gradient = NULL,
plot.title = NULL,
pt.size = 0.5,
common.scales = TRUE,
collect.axes = FALSE
)| Parameter | Default | Description |
|---|---|---|
feature1 |
— | X-axis feature (metadata column, assay feature, or reduction variable) |
feature2 |
— | Y-axis feature; same resolution rules as feature1 |
gradient |
— | Feature used for the continuous color scale, resolved via FetchData() |
group.by |
"ident" |
Metadata column for per-group panels; NULL for a single combined plot |
scale.colors |
"viridis" |
Viridis palette option or letter code |
lower.limit |
NULL |
Lower gradient clamp; NULL infers from data |
upper.limit |
NULL |
Upper gradient clamp; NULL infers from data |
corr.method |
"pearson" |
Correlation method: "pearson", "spearman", or "kendall". NULL for no correlation |
layer1 |
NULL |
Assay layer for feature1; NULL uses Seurat default; "" and "null" treated as NULL |
layer2 |
NULL |
Assay layer for feature2; semantics mirror layer1 |
layer.gradient |
NULL |
Assay layer for gradient; metadata-backed features ignore this |
plot.title |
NULL |
Custom main title; ungrouped plots default to the correlation string |
pt.size |
0.5 |
Point size |
common.scales |
TRUE |
If TRUE, all sub-plots share the same x and y axis limits |
collect.axes |
FALSE |
If TRUE, collects axes across sub-plots using patchwork |
Creates violin plots colored by a per-identity aggregated gradient ("nCells" or the mean of a chosen feature), ordered by gradient value. Supports per-feature assay layer selection via layer and optional per-feature custom titles via plot.title. When layer is specified for assay-backed features, the default panel title is feature_layer; metadata-backed features always use bare names.
VlnPlotGradient(
SeuratObject,
features,
gradient = "nCells",
group.by = "ident",
scale.colors = "viridis",
lower.limit = 0,
upper.limit = NULL,
pt.size = 0,
ncol = NULL,
layer = NULL,
layer.gradient = NULL,
plot.title = NULL
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object |
features |
— | Character vector of features (genes or metadata columns) to plot |
gradient |
"nCells" |
"nCells" or a feature name whose per-group mean colors the violins |
group.by |
"ident" |
Metadata grouping column; NULL uses active identity |
scale.colors |
"viridis" |
Viridis palette option |
lower.limit |
0 |
Lower gradient clamp |
upper.limit |
NULL |
Upper gradient clamp |
pt.size |
0 |
Size of jittered points overlaid on violins; 0 hides points |
ncol |
NULL |
Number of columns in the combined plot |
layer |
NULL |
Assay layer(s): scalar applies to all features; vector of length equal to features applies per-feature; metadata-backed features ignore this |
layer.gradient |
NULL |
Assay layer for the gradient feature; metadata-backed features ignore this |
plot.title |
NULL |
Custom title(s): NULL, one string for all, or one string per feature |
Creates kernel-density plots for one or more features extracted from a Seurat object via FetchData(), supporting metadata columns, assay features from any layer, and reduction variables. Features grouped by group.by can be shown as overlaid densities or split into per-group panels. Supports dashed red reference lines (vline), independent black median lines (plot.median), custom titles, and per-feature layer selection. When multiple features are requested, returns a named list of plots; names follow feature_layer for assay-backed features and bare feature for metadata-backed features (suffixes _NULL and _NA are stripped).
FeatureDensityPlot(
SeuratObject,
features,
group.by = "ident",
split.plot = TRUE,
scale.colors = "viridis",
ncol = NULL,
vline = NULL,
layer = NULL,
plot.median = TRUE,
plot.title = NULL,
nmad = 2,
alpha = 0.3,
pt.size = 0,
common.scales = TRUE,
collect.axes = FALSE
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object |
features |
— | Character vector of metadata columns, assay features, or reduction variables |
group.by |
"ident" |
Grouping variable: metadata column, "ident" (active identity), or NULL for no grouping |
split.plot |
TRUE |
If TRUE, one panel per group level; if FALSE, overlay all groups in one panel per feature |
scale.colors |
"viridis" |
Viridis palette option for group color mapping |
ncol |
NULL |
Number of columns for split panels within each feature plot |
vline |
NULL |
Reference line: NULL, "mean", "median", "upper" (median + nmad MADs), "lower" (median − nmad MADs), "both", a numeric value, or a per-feature/per-group character/numeric vector |
layer |
NULL |
Assay layer(s): scalar (all features) or vector (per-feature); metadata features ignore this |
plot.median |
TRUE |
If TRUE, draws an independent black dashed median line for each group |
plot.title |
NULL |
Custom title(s): NULL uses feature names, one string for all, or one per feature |
nmad |
2 |
Number of MADs used when vline is "upper", "lower", or "both" |
alpha |
0.3 |
Fill transparency for density geometries |
pt.size |
0 |
Rug line width; 0 disables the rug |
common.scales |
TRUE |
If TRUE, all sub-plots share the same x and y axis limits |
collect.axes |
FALSE |
If TRUE, collects axes across sub-plots using patchwork |
Creates boxplots to visualize the distribution of QC metrics (default: nFeature_RNA, nCount_RNA, percent.mt) grouped by an entity metadata column. Individual cells are overlaid as jitter points, optionally colored by a gradient feature. Multiple metrics are combined into a single patchwork figure.
QCMetricsBoxplot(
SeuratObject,
entity_type,
entity_name = NULL,
qc_metrics = c("nFeature_RNA", "nCount_RNA", "percent.mt"),
gradient_col = NULL,
scale.colors = "viridis",
lower.limit = 0,
upper.limit = NULL,
pt.size = 1,
pt.alpha = 0.6,
fill_color = "lightblue",
outlier.size = 1,
ncol = NULL,
gradient_legend = FALSE
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object |
entity_type |
— | Metadata column name used to group cells (e.g. "orig.ident") |
entity_name |
NULL |
If specified, filters the plot to one value of entity_type |
qc_metrics |
c("nFeature_RNA", "nCount_RNA", "percent.mt") |
QC metadata columns to plot |
gradient_col |
NULL |
Metadata column or feature used to color jitter points; if NULL, the current metric is used |
scale.colors |
"viridis" |
Viridis palette option |
lower.limit |
0 |
Lower gradient scale clamp |
upper.limit |
NULL |
Upper gradient scale clamp |
pt.size |
1 |
Size of individual cell points |
pt.alpha |
0.6 |
Transparency of individual cell points |
fill_color |
"lightblue" |
Fill color of the boxplot body |
outlier.size |
1 |
Size of boxplot outlier points |
ncol |
NULL |
Number of columns in the combined plot |
gradient_legend |
FALSE |
Whether to show the gradient legend |
Runs fgsea per cluster using the output of Seurat::FindAllMarkers() and returns a list of enriched pathways for each cluster. Genes are ranked by avg_log2FC; only genes passing the padj.threshold are included in the ranked list.
results <- scGSEAmarkers(
cluster_markers,
reference_markers,
padj.threshold = 1e-6,
only.pos = TRUE,
workers = 4
)| Parameter | Default | Description |
|---|---|---|
cluster_markers |
— | data.frame from FindAllMarkers() containing cluster, gene, avg_log2FC, and p_val_adj columns |
reference_markers |
— | GSEA gene-set database in fgsea-compatible list format |
padj.threshold |
1e-6 |
Adjusted p-value cutoff for marker inclusion |
only.pos |
TRUE |
If TRUE, returns only positively enriched pathways (NES > 0) |
workers |
4 |
Number of cores for parallel fgsea computation |
Extracts the per-cell results of .CalculateFeatureThresholdSeurat() for a given feature from a Seurat object's metadata and returns a tidy data.frame with one row per unique batch level, including the computed threshold (optional) and pass/fail statistics.
ExtractFeatureTestResults(
SeuratObject,
batch.col,
feature = "MALAT1",
extract.threshold = TRUE
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | Seurat object containing <feature>.threshold and <feature>.pass columns in meta.data |
batch.col |
— | Metadata column name identifying the batch variable |
feature |
"MALAT1" |
Feature name whose threshold results to extract |
extract.threshold |
FALSE |
If TRUE, includes the computed threshold value in the output; if FALSE, only pass/fail counts are returned (default) |
Creates a stacked bar plot showing, for each cluster at a selected Seurat clustering resolution, the proportion of cells in each category of a metadata variable.
PlotClusterComposition(
SeuratObject,
resolution.value,
group.by = "library",
cluster.prefix = "RNA_snn_res.",
return.plot = TRUE,
percent.accuracy = 1
)| Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | A Seurat object containing clustering and metadata columns |
resolution.value |
— | Clustering resolution value used to construct the cluster column |
group.by |
"library" |
Metadata column used for composition categories |
cluster.prefix |
"RNA_snn_res." |
Prefix used with resolution.value to build the cluster metadata column name |
return.plot |
TRUE |
If TRUE, returns a ggplot2 stacked-bar plot; if FALSE, returns the plotting data.frame |
percent.accuracy |
1 |
Accuracy passed to scales::label_percent() for y-axis percentage formatting |
Allows to perform a feature-based QC test on a Seurat object, computing thresholds and pass/fail statistics for one or more features. For the given feeature(s), thresholds and comparison operations, the function computes per-cell pass/fail results and returns the Seurat object with new metadata columns <feature>.pass (logical). It also allows to apply a batch-wise thresholding strategy, where thresholds are computed per batch level and applied to the corresponding cells.
FeatureTest(
SeuratObject,
features,
assay = "RNA",
layers = NULL,
thresholds = "mad",
operators = "both",
nmad = 3,
percentile = 1
) | Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | Seurat object |
features |
— | List of feature names to test |
assay |
NULL |
Assay name containing the features. If not specified, the default assay is used. |
thresholds |
"mad" |
Thresholds to test againts. If "mad" or "percentile", it will compute said statistics as a threshold. |
operators |
"both" |
Comparison operators for thresholding ("<", ">", ">=", "<=", "!=", "=="). MAD and percentile thresholds can use upper and lower bounds ("upper", "lower"), or both ("both"). |
nmad |
3 |
Number of median absolute deviations for thresholding |
percentile |
1 |
Percentile for thresholding |
This is an adapted version of Seurat's VariableFeaturePlot() that forces to use the object's consensus set of high variable features in case of multiple layers. It also adds the ability to label points based on their rank or a user-defined list.
VariableFeaturePlot2(
object,
cols = c('black', 'red'),
pt.size = 1,
log = NULL,
selection.method = NULL,
assay = NULL,
raster = NULL,
raster.dpi = c(512, 512),
hvf = NULL,
label = FALSE,
n.top.hvf = 10,
custom.label = NULL,
repel = TRUE
) | Parameter | Default | Description |
|---|---|---|
SeuratObject |
— | Seurat object |
cols |
c('black', 'red') | Colors to use for the non-variable and variable features, respectively |
pt.size |
1 | Point size for the plot. |
log |
NULL |
Log the x-axis in log scale |
selection.method |
NULL |
Method for selecting features. If not specified, the default is used. |
assay |
NULL |
Assay name containing the features. If not specified, the default assay is used. |
raster |
NULL |
Whether to rasterize the plot. If not specified, it will rasterized the plot if the number of points exceeds 100000. |
raster.dpi |
c(512, 512) |
DPI for rasterization. |
hvf |
NULL |
High variable features to highlight. If not specified, they will be automatically selected. |
label |
FALSE |
Whether to label points with their feature names. |
n.top.hvf |
10 |
Number of top high variable features to label. |
custom.label |
NULL |
Custom labels for the points. |
repel |
TRUE |
Whether to repel labels. |
The following functions are internal helpers not exported from the namespace. They are documented here for reference.
Computes an expression threshold for a single feature from a normalized counts vector using kernel density estimation and smoothing splines. The threshold corresponds to the left quadratic intercept of the density curve, with fallback to a conservative default value.
Applies .ComputeFeatureThreshold() across all layers of a Seurat assay and stores per-cell results as <feature>.threshold (numeric) and <feature>.pass (logical) in meta.data.
Converts ENSEMBL gene IDs to gene symbols for a raw count matrix using biomaRt. Returns the filtered matrix with symbol rownames. Requires biomaRt and dplyr; not exported.
Sets common x and y axis limits across a patchwork joined ggplot object. Used internally by FeatureScatterGradient() and FeatureDensityPlot() when common.scales = TRUE.
Collapses multiple boolean columns into a single logical column, returning TRUE if all of the input columns are TRUE for a given row, and FALSE if any is FALSE. Used internally by FeatureTest()
Fiters the high variable features of a Seurat object after applying Seurat::FindVariableFeatures() to delete low-informative genes such as mitochondrial, ribosomal, antisense and non-coding genes.
Uses intrinsicDimension's maxLikGlobalDimEst() to estimate the intrinsic dimensionality of a Seurat object based on the PCA embeddings calculated with Seurat::RunPCA(). It returns a vector with the most informative PCs. This allows to automatize the number of PCs for use in downstream analyses such as clustering and UMAP/tSNE computation.
-
CalculateCDRadapted from conquer_comparison by Charlotte Soneson. -
MALAT1 thresholding logic adapted from [BaderLab's MALAT1_threshold] (https://github.com/BaderLab/MALAT1_threshold)
-
VariableFeaturePlot2adapted from Seurat'sVariableFeaturePlot().
GPLv3 © José Manuel Gómez Silva
If you find bugs or have suggestions, please open an issue or pull request on GitHub.
When contributing code:
- Follow the existing style (based partially on Google's R style guide: CamelCase for exported functions, dot.case for arguments/variables, explicit
package::functioncalls). - Include clear roxygen documentation and tests where appropriate.
- Respect licensing and attribution, particularly for any upstream code you incorporate.