Adaptive detection of human-mouse multiplets in patient-derived xenograft (PDX) single-cell RNA-seq data.
In mixed-species experiments (such as PDX models, where a human tumor grows in a
mouse host), some droplets capture both a human and a mouse cell. These
multiplets must be removed before downstream analysis. multipletR detects
them using an adaptive threshold method that does not
assume a fixed human/mouse proportion, so it handles the imbalanced species
mixtures typical of real PDX samples.
In a PDX sample, human tumor cells and mouse host cells are sequenced together, and some droplets capture one of each: a human-mouse multiplet. Cell Ranger flags multiplets with read-count thresholds that assume a roughly balanced human/mouse mix. In real PDX data the mix is rarely balanced, so cells that are almost entirely human (or mouse) get mislabeled as multiplets.
multipletR starts from a conservative central region
(cells with a genuinely balanced human/mouse mix) and expands three thresholds
step by step, stopping once the selected cells stop looking like real multiplets
(when their read distributions become bimodal, stop overlapping, or drift apart).
This lets the multiplet region adapt to each sample.
Once accepted to Bioconductor, install with:
if (!require("BiocManager", quietly = TRUE)) {
install.packages("BiocManager")
}
BiocManager::install("multipletR")
# For the development version:
BiocManager::install("multipletR", version = "devel")
Until then, install the development version from GitHub with:
# install.packages("remotes")
remotes::install_github("Alex05a/multipletR")
The remove_multiplets() helper additionally requires either the
Seurat package or
SingleCellExperiment, depending on which object type
you use.
library(multipletR)
The package takes a single file as its primary input: the GEM classification CSV produced by Cell Ranger when reads are aligned to a combined two-species reference. This section explains where that file comes from, where to find it, and what it contains.
In droplet-based single-cell RNA-sequencing (10x Genomics Chromium), each droplet (a GEM, Gel Bead-in-EMulsion) captures one or more cells together with a barcoded gel bead, so every cell’s transcripts inherit a shared cell barcode and each molecule gets a unique molecular identifier (UMI). Most droplets contain a single cell (a singlet), but some capture two or more; when those cells come from different species, the droplet is a human-mouse multiplet.
To tell the two species apart, reads are aligned with cellranger count
against a combined (“barnyard”) reference that contains both genomes, for example
GRCh38 (human) and GRCm39 (mouse). In this reference every gene is prefixed by
its genome of origin (e.g. GRCh38-EPCAM, GRCm39-Col1a1), so each read can be
attributed to human or mouse. For every cell barcode, Cell Ranger counts how
many reads map to each genome and, using count thresholds, labels the
barcode as human, mouse, or a multiplet. This per-barcode, two-genome summary is
written to the GEM classification file.
The file is written only when the reference is a multi-genome (barnyard)
reference. For a run with --id=<SAMPLE>, it is located at:
<SAMPLE>/outs/analysis/gem_classification.csv
The full expression matrix lives separately at
<SAMPLE>/outs/filtered_feature_bc_matrix/. multipletR does not need the
matrix, only gem_classification.csv, although the Seurat/SingleCellExperiment
helper later joins the calls back onto a matrix-derived object by barcode.
One row per cell barcode, with these columns:
| Column | Meaning |
|---|---|
barcode |
The 10x cell barcode (often carries a -1 suffix, e.g. AAACCTGAG...-1). |
| human genome | Per-barcode count of reads assigned to the human genome, named after the reference build, e.g. GRCh38. |
| mouse genome | Per-barcode count of reads assigned to the mouse genome, named after the build: GRCm39 (2024-A references), or mm10 / mm39 for older builds. |
call |
Cell Ranger’s own classification: the human genome name (e.g. GRCh38), the mouse genome name (e.g. GRCm39), or Multiplet. |
A typical file looks like:
barcode,GRCh38,GRCm39,call
AAACCTGAGAAACCAT-1,10432,58,GRCh38
AAACCTGAGATCCGAG-1,71,9903,GRCm39
AAACCTGCACGTGTGA-1,4821,5210,Multiplet
The first row is dominated by human reads (a human singlet), the second by mouse reads (a mouse singlet), and the third has substantial reads from both genomes (a multiplet).
The package ships a real PDX example (sample PC65). Here are the first rows:
gem_file <- system.file("extdata", "PC65_gem_classification.csv",
package = "multipletR"
)
head(read.csv(gem_file))
#> barcode GRCh38 GRCm39 call
#> 1 AAACCAAAGCAAGTGG-1 8755 123 GRCh38
#> 2 AAACCAAAGCGAACTC-1 1099 31769 Multiplet
#> 3 AAACCAAAGGAGGCAC-1 78678 452 GRCh38
#> 4 AAACCATTCAATGGCC-1 16813 168 GRCh38
#> 5 AAACCATTCAGCAACC-1 1007 27 GRCh38
#> 6 AAACCATTCGCATGGT-1 100603 11623 Multiplet
The package’s whole premise is that a true human-mouse multiplet has substantial
reads from both genomes, whereas a singlet is dominated by one. The GEM
classification file provides exactly the two numbers needed to test that: the
human and mouse read counts per barcode. From them, multipletR derives, for
each barcode, the total reads (human + mouse) and the percent mouse
(mouse / total x 100), and applies its adaptive thresholds in that
total-reads by percent-mouse space rather than trusting Cell Ranger’s
call. The call column is kept only for the comparison plots.
See make-data.R for how this bundled PC65 example
was created.
Because the genome columns are named after whichever reference build was used,
multipletR recognizes several common names automatically, so you do not have to
rename anything:
GRCh38, hg38, human_reads, humanGRCm39, mm10, mm39, mouse_reads, mousebarcode, Barcode, barcodes (if none is present, the row index is used as a fallback)call, Call, classificationAs long as the file has a human read-count column and a mouse read-count column,
detect_multiplets() will run; the barcode and call columns are optional but
recommended (the barcode is what lets you match results back to a Seurat or
SingleCellExperiment object).
Tip: barcode suffixes must match between the GEM file and any object you
later filter. Cell Ranger writes barcodes with a trailing -1; if your object’s
cell names lack it (or vice versa), remove_multiplets() will match nothing.
Make the two consistent before joining.
detect_multiplets() reads the GEM classification file, runs the adaptive
threshold engine, writes an annotated CSV, and (by default) draws diagnostic
plots. It returns the input data with three added columns: our_classification
(Multiplet / Singlet), pct_human, and pct_mouse.
Here we draw both diagnostic views: the percent plot (total reads vs percent mouse) and the total-reads plot (mouse reads vs human reads).
res <- detect_multiplets(
fileIn = gem_file,
fileOut = tempfile(fileext = ".csv"),
plotPercent = TRUE,
plotTotalReads = TRUE
)
The one-line summary reports the final thresholds and the number of multiplets found. The returned data frame carries the per-cell result:
head(res)
#> barcode GRCh38 GRCm39 call our_classification pct_human
#> 1 AAACCAAAGCAAGTGG-1 8755 123 GRCh38 Singlet 98.61
#> 2 AAACCAAAGCGAACTC-1 1099 31769 Multiplet Singlet 3.34
#> 3 AAACCAAAGGAGGCAC-1 78678 452 GRCh38 Singlet 99.43
#> 4 AAACCATTCAATGGCC-1 16813 168 GRCh38 Singlet 99.01
#> 5 AAACCATTCAGCAACC-1 1007 27 GRCh38 Singlet 97.39
#> 6 AAACCATTCGCATGGT-1 100603 11623 Multiplet Singlet 89.64
#> pct_mouse
#> 1 1.39
#> 2 96.66
#> 3 0.57
#> 4 0.99
#> 5 2.61
#> 6 10.36
table(res$our_classification)
#>
#> Multiplet Singlet
#> 59 7587
Each plot has two panels: the left colored by Cell Ranger’s call, the right by
the multipletR classification. Across both panels human singlets are blue and
mouse singlets are green; Cell Ranger’s multiplets are orange (left) and
multipletR’s are red (right), so it is easy to see which cells each method flags.
In the percent plot, multiplets sit in the central percent-mouse band, where a droplet carries a real mix of human and mouse reads; pure human cells fall near 0% mouse and pure mouse cells near 100%. In the total-reads plot, true multiplets have substantial reads from both genomes, so they sit away from the two axes.
The contrast with Cell Ranger is the key point. In this PDX sample these thresholds label many predominantly single-species barcodes as multiplets, whereas the adaptive method keeps only the balanced droplets. You can see the size of that gap directly:
# Cell Ranger's multiplet calls (from the GEM file)
sum(read.csv(gem_file)$call == "Multiplet")
#> [1] 579
# multipletR's multiplet calls
sum(res$our_classification == "Multiplet")
#> [1] 59
We can confirm the multiplets multipletR keeps are genuinely balanced by looking at their composition, which centers near a mixed human/mouse split rather than sitting at one extreme:
mult <- res[res$our_classification == "Multiplet", ]
summary(mult$pct_human)
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 18.02 21.39 37.03 39.81 55.19 74.27
summary(mult$pct_mouse)
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 25.73 44.81 62.97 60.19 78.61 81.98
The defaults are the recommended values and rarely need changing, but every
parameter is adjustable. The three starting thresholds are T1 (upper percent
mouse), T2 (lower percent mouse), and T3 (lower total-reads percentile); the
two stopping sensitivities are overlapDrop and modeDiff. For example, a more
conservative run:
res_strict <- detect_multiplets(
fileIn = gem_file,
fileOut = tempfile(fileext = ".csv"),
overlapDrop = 5, # stop expanding sooner
modeDiff = 0.5 # stricter on distribution similarity
)
See ?detect_multiplets for the full list of arguments.
Once multiplets are detected, remove_multiplets() carries the result onto a
single-cell object. It adds the per-cell classification to the object’s cell
metadata (multipletR_class, multipletR_pct_human, multipletR_pct_mouse)
and, by default, removes the detected multiplets so the object is ready for
downstream analysis. Setting remove = FALSE keeps all cells and only annotates
them, which is useful for visualizing where the multiplets fall (for example on a
UMAP colored by multipletR_class) before deciding whether to filter, following
the same idea as tools like DoubletFinder.
The object type (Seurat or SingleCellExperiment) is detected automatically;
the classification logic is identical, only the metadata is written to the
appropriate slot.
Both helpers use Seurat and SingleCellExperiment, which are optional
dependencies. If you do not have them, install them with:
pkg <- c("SingleCellExperiment", "Seurat")
for (. in pkg) if (!requireNamespace(., quietly = TRUE)) BiocManager::install(.)
For a quick, self-contained illustration we build a small toy object whose cell
barcodes match the detected calls. In practice you would use your own Seurat or
SingleCellExperiment object built from the sample’s count matrix.
# toy counts: 20 genes x the detected barcodes (stand-in for a real matrix)
counts <- matrix(
rpois(20 * nrow(res), 5),
nrow = 20,
dimnames = list(paste0("gene", 1:20), res$barcode)
)
seu <- Seurat::CreateSeuratObject(counts)
#> Warning: Data is of class matrix. Coercing to dgCMatrix.
# annotate only (keep all cells) to inspect the multiplets first
seu <- remove_multiplets(seu, res, remove = FALSE)
#> Annotated 7646 of 7646 cells. multipletR_class counts: Human=5795, Mouse=1792, Multiplet=59
table(seu$multipletR_class)
#>
#> Human Mouse Multiplet
#> 5795 1792 59
# or annotate and remove in one step, ready for downstream analysis
seu_clean <- remove_multiplets(seu, res, remove = TRUE)
#> Annotated 7646 of 7646 cells. multipletR_class counts: Human=5795, Mouse=1792, Multiplet=59
#> Removed 59 multiplet cells; 7587 cells remain.
table(seu_clean$multipletR_class)
#>
#> Human Mouse
#> 5795 1792
The same call works for a SingleCellExperiment. Running it on the same data
gives an identical classification:
sce <- SingleCellExperiment::SingleCellExperiment(
assays = list(counts = counts)
)
sce <- remove_multiplets(sce, res, remove = FALSE)
#> Annotated 7646 of 7646 cells. multipletR_class counts: Human=5795, Mouse=1792, Multiplet=59
table(sce$multipletR_class)
#>
#> Human Mouse Multiplet
#> 5795 1792 59
# the Seurat and SCE classifications are identical
identical(
unname(seu$multipletR_class),
unname(sce$multipletR_class)
)
#> [1] TRUE
If you have a Seurat object and would rather work with a SingleCellExperiment, you can convert between the two with Seurat::as.SingleCellExperiment().
With remove = FALSE (the default), cells are annotated but kept, so you can
inspect the calls first and then filter manually on multipletR_class, which
gives full control over what is removed — for example
seu[, seu$multipletR_class != "Multiplet"] for a Seurat object, or
sce[, sce$multipletR_class != "Multiplet"] for a SingleCellExperiment.
| Function | Purpose |
|---|---|
detect_multiplets() |
Detect multiplets from a Cell Ranger GEM classification file; add the classification and draw diagnostic plots. |
remove_multiplets() |
Annotate a Seurat or SingleCellExperiment object with the classification (Human / Mouse / Multiplet and percent human/mouse) and, by default, remove the multiplets. Set remove = FALSE to annotate only. |
The method defines a conservative starting region using three thresholds: T1 (upper percent mouse), T2 (lower percent mouse), and T3 (lower total reads), then expands them step by step to capture additional multiplets. It stops expanding when the human and mouse read distributions start to look like singlets: when a distribution becomes bimodal, when their overlap drops, or when their modes diverge. This lets the multiplet region adapt to each dataset.
sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.5 LTS
#>
#> Matrix products: default
#> BLAS: /home/biocbuild/bbs-3.24-bioc/R/lib/libRblas.so
#> LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.12.0 LAPACK version 3.12.0
#>
#> locale:
#> [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=en_GB LC_COLLATE=C
#> [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
#> [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: America/New_York
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] multipletR_0.99.9 BiocStyle_2.41.0
#>
#> loaded via a namespace (and not attached):
#> [1] RColorBrewer_1.1-3 jsonlite_2.0.0
#> [3] magrittr_2.0.5 spatstat.utils_3.2-5
#> [5] magick_2.9.1 farver_2.1.2
#> [7] rmarkdown_2.32 vctrs_0.7.3
#> [9] ROCR_1.0-12 spatstat.explore_3.8-3
#> [11] tinytex_0.61 htmltools_0.5.9
#> [13] S4Arrays_1.13.1 SparseArray_1.13.3
#> [15] sass_0.4.10 sctransform_0.4.3
#> [17] parallelly_1.48.0 KernSmooth_2.23-27
#> [19] bslib_0.12.0 htmlwidgets_1.6.4
#> [21] ica_1.0-3 plyr_1.8.9
#> [23] plotly_4.12.1 zoo_1.9-1
#> [25] cachem_1.1.0 igraph_2.3.3
#> [27] mime_0.13 lifecycle_1.0.5
#> [29] pkgconfig_2.0.3 Matrix_1.7-6
#> [31] R6_2.6.1 fastmap_1.2.0
#> [33] MatrixGenerics_1.25.0 fitdistrplus_1.2-6
#> [35] future_1.76.0 shiny_1.14.0
#> [37] digest_0.6.39 patchwork_1.3.2
#> [39] S4Vectors_0.51.10 Seurat_5.5.1
#> [41] tensor_1.5.1 RSpectra_0.16-2
#> [43] irlba_2.3.7 GenomicRanges_1.65.4
#> [45] labeling_0.4.3 progressr_1.0.0
#> [47] spatstat.sparse_3.2-0 httr_1.4.9
#> [49] polyclip_1.10-7 abind_1.4-8
#> [51] compiler_4.6.1 withr_3.0.3
#> [53] S7_0.2.2 fastDummies_1.7.6
#> [55] MASS_7.3-66 DelayedArray_0.39.7
#> [57] tools_4.6.1 lmtest_0.9-40
#> [59] otel_0.2.0 httpuv_1.6.17
#> [61] future.apply_1.20.2 goftest_1.2-3
#> [63] glue_1.8.1 nlme_3.1-171
#> [65] promises_1.5.0 grid_4.6.1
#> [67] Rtsne_0.17 cluster_2.1.8.3
#> [69] reshape2_1.4.5 generics_0.1.4
#> [71] gtable_0.3.6 spatstat.data_3.1-9
#> [73] tidyr_1.3.2 data.table_1.18.6.1
#> [75] sp_2.2-3 XVector_0.53.0
#> [77] BiocGenerics_0.59.12 spatstat.geom_3.8-3
#> [79] RcppAnnoy_0.0.23 ggrepel_0.9.8
#> [81] RANN_2.6.3 pillar_1.11.1
#> [83] stringr_1.6.0 spam_2.11-4
#> [85] RcppHNSW_0.7.0 later_1.4.8
#> [87] splines_4.6.1 dplyr_1.2.1
#> [89] lattice_0.23-1 deldir_2.0-4
#> [91] survival_3.8-12 tidyselect_1.2.1
#> [93] SingleCellExperiment_1.35.2 miniUI_0.1.2
#> [95] pbapply_1.7-5 knitr_1.52
#> [97] gridExtra_2.3.1 bookdown_0.48
#> [99] IRanges_2.47.5 Seqinfo_1.3.2
#> [101] SummarizedExperiment_1.43.0 scattermore_1.2
#> [103] stats4_4.6.1 xfun_0.61
#> [105] Biobase_2.73.2 diptest_0.77-2
#> [107] matrixStats_1.5.0 stringi_1.8.9
#> [109] yaml_2.3.12 evaluate_1.0.5
#> [111] codetools_0.2-20 tibble_3.3.1
#> [113] BiocManager_1.30.27 cli_3.6.6
#> [115] uwot_0.2.5 xtable_1.8-8
#> [117] reticulate_1.47.0 jquerylib_0.1.4
#> [119] dichromat_2.0-1 Rcpp_1.1.2
#> [121] spatstat.random_3.5-2 globals_0.19.1
#> [123] png_0.1-9 spatstat.univar_3.2-0
#> [125] parallel_4.6.1 ggplot2_4.0.3
#> [127] dotCall64_1.2 listenv_1.0.0
#> [129] viridisLite_0.4.3 scales_1.4.0
#> [131] ggridges_0.5.7 SeuratObject_5.4.0
#> [133] purrr_1.2.2 rlang_1.3.0
#> [135] cowplot_1.2.0