Contents

1 About metabinR

metabinR performs abundance- and composition-based binning on metagenomic samples directly from FASTA or FASTQ files. Abundance-based binning (AB) analyzes long k-mers (k > 8); composition-based binning (CB) analyzes short k-mers (k < 8); hierarchical binning (ABxCB) chains the two.

The heavy lifting is implemented in Java and called via rJava. From v2.0.0 the R API returns an S4 MetabinResult object and accepts Biostrings / ShortRead inputs in addition to file paths.

2 Installation

if (!requireNamespace("BiocManager", quietly = TRUE))
    install.packages("BiocManager")
BiocManager::install("metabinR")

A JDK (Java >= 11) must be available before installing metabinR.

3 Preparation

3.1 JVM heap size

JVM flags are passed through the java.parameters option. The -Xmx flag controls the maximum heap: -Xmx1500M or -Xmx3G, etc. Set this before loading the package.

options(java.parameters = "-Xmx1500M")
library(metabinR)
library(ggplot2)
library(Biostrings)

To customise additional JVM flags (on top of the -Xmx heap setting), set options(metabinR.jvm.flags = c("-XX:+UseG1GC", ...)) before the package is loaded. The defaults returned by metabinR_jvm_options() already enable G1GC and string deduplication.

4 The MetabinResult class

Each binning function returns a MetabinResult S4 object:

The read_id column identifies each read. The column named by algorithm(res) contains its assigned bin, and columns such as AB.1 contain distances to candidate bins. Bin numbers are arbitrary labels; they do not identify genomes or abundance classes on their own.

5 Abundance based binning example

The toy simulated metagenome contains 26,664 Illumina reads (13,332 pairs of 2x150bp) sampled from 10 bacterial genomes across two abundance classes (high vs. low).

Abundance ground truth:

abundances <- read.table(
    system.file("extdata", "distribution_0.txt", package = "metabinR"),
    col.names = c("genome_id", "abundance", "AB_id"))

Read-level ground truth contains the genome of origin for each read. Add the abundance class of each genome for the AB evaluation:

reads.mapping <- read.delim(system.file(
    "extdata", "reads_mapping.tsv.gz", package = "metabinR"))
reads.mapping <- merge(reads.mapping,
                       abundances[, c("genome_id", "AB_id")],
                       by = "genome_id")

Run abundance-based binning with 10-mers into 2 clusters:

res.AB <- abundance_based_binning(
    system.file("extdata", "reads.metagenome.fasta.gz", package = "metabinR"),
    numOfClustersAB = 2,
    kMerSizeAB = 10,
    dryRun = FALSE,
    outputAB = "vignette"
)
res.AB
#> MetabinResult (AB)
#>   reads:     26,664
#>   clusters:  2
#>   inputs:    1 file(s)
#>     - /tmp/RtmpPOC9kk/Rinst3a18e2297a11b4/metabinR/extdata/reads.metagenome.fasta.gz

Inspect the assignments and summarize the observed bins. The distance summaries use only each read’s distance to its assigned bin:

head(as.data.frame(assignments(res.AB)))
#>   read_id AB        AB.1        AB.2
#> 1  S0R0/1  1 0.032307751 0.004981781
#> 2  S0R1/2  1 0.019364505 0.010336333
#> 3  S0R0/2  1 0.030796454 0.005614620
#> 4  S0R1/1  1 0.021667648 0.009437209
#> 5  S0R2/2  2 0.009775655 0.014416852
#> 6  S0R2/1  1 0.013742806 0.012755650
knitr::kable(as.data.frame(bin_summary(res.AB)), digits = 3)
bin n_reads proportion mean_distance median_distance
1 19761 0.741 0.024 0.024
2 6903 0.259 0.015 0.015

abundance_based_binning() wrote one FASTA per cluster and a k-mer count histogram:

histogram.AB <- read.table("vignette__AB.histogram.tsv", header = TRUE)
ggplot(histogram.AB, aes(x = counts, y = frequency)) +
    geom_area() +
    labs(title = "kmer counts histogram") +
    theme_bw()

Evaluate against the abundance-class ground truth. evaluate_bins() matches read_id to anonymous_read_id, so the two tables need not have the same row order. The confusion table has inferred bins as rows and known abundance classes as columns:

eval.AB <- evaluate_bins(res.AB, reads.mapping,
                         id = "anonymous_read_id", label = "AB_id")
eval.AB$confusion
#>    origin
#> bin     1     2
#>   1 18185  1576
#>   2  1891  5012
knitr::kable(as.data.frame(eval.AB$per_bin), digits = 3)
bin n_reads dominant_origin dominant_reads purity
1 19761 1 18185 0.920
2 6903 2 5012 0.726
knitr::kable(as.data.frame(eval.AB$overall), digits = 3)
n_reads n_bins n_origins weighted_purity weighted_recovery adjusted_rand_index
26664 2 2 0.87 0.87 0.519

Per-bin purity is the fraction of reads in a bin from its dominant abundance class. The adjusted Rand index (ARI) compares the two partitions without assuming that their labels correspond.

6 Composition based binning example

Here the known label is the bacterial genome of origin, rather than the abundance class used above.

Run composition-based binning with 4-mers into 10 clusters:

res.CB <- composition_based_binning(
    system.file("extdata", "reads.metagenome.fasta.gz", package = "metabinR"),
    numOfClustersCB = 10,
    kMerSizeCB = 4,
    dryRun = TRUE,
    outputCB = "vignette"
)

Check bin sizes and inspect reads with close competing distances. The margin is an absolute distance difference on this algorithm’s scale, not a probability. A small result can mean either that few reads are close to a boundary or that the chosen margin is too narrow:

knitr::kable(as.data.frame(bin_summary(res.CB)), digits = 3)
bin n_reads proportion mean_distance median_distance
7 3030 0.114 15488.12 15364
9 3390 0.127 15511.10 15334
4 7563 0.284 16466.11 16460
3 4647 0.174 15611.46 15510
2 1983 0.074 15734.53 15578
5 1183 0.044 16530.11 16422
10 2514 0.094 15971.12 15866
1 776 0.029 15224.38 15011
6 956 0.036 16685.22 16805
8 622 0.023 16593.01 16431
close.reads <- ambiguous_reads(res.CB, margin = 0.05)
nrow(close.reads)
#> [1] 24
knitr::kable(head(as.data.frame(close.reads)), digits = 3)
read_id bin best_distance second_distance distance_margin
S0R497/1 2 16132 16132 0
S0R1266/2 7 15208 15208 0
S0R1558/2 6 15066 15066 0
S0R3119/2 1 14532 14532 0
S0R3248/2 5 16700 16700 0
S0R3433/2 3 15298 15298 0

Evaluate by genome of origin. Best-bin recovery is the fraction of an origin’s reads found in its single largest bin; it is not genome completeness. Weighted purity and recovery summarize these counts over all evaluated reads:

eval.CB <- evaluate_bins(res.CB, reads.mapping,
                         id = "anonymous_read_id", label = "genome_id")
knitr::kable(as.data.frame(eval.CB$per_origin), digits = 3)
origin n_reads dominant_bin recovered_reads best_bin_recovery
Genome15.0 14190 3 4108 0.289
Genome3.0 5886 10 1631 0.277
Genome2.0 538 4 217 0.403
Genome12.0 3782 4 3673 0.971
Genome11.0 1186 4 1157 0.976
Genome6.0 356 8 68 0.191
Genome17.0 594 10 212 0.357
Genome22.0 8 4 6 0.750
Genome23.0 80 4 80 1.000
Genome4.0 44 1 19 0.432
knitr::kable(as.data.frame(eval.CB$overall), digits = 3)
n_reads n_bins n_origins weighted_purity weighted_recovery adjusted_rand_index
26664 10 10 0.641 0.419 0.103

7 Hierarchical (2-step ABxCB) binning example

res.ABxCB <- hierarchical_binning(
    system.file("extdata", "reads.metagenome.fasta.gz", package = "metabinR"),
    numOfClustersAB = 2,
    kMerSizeAB = 10,
    kMerSizeCB = 4,
    dryRun = TRUE,
    outputC = "vignette"
)
knitr::kable(as.data.frame(bin_summary(res.ABxCB)), digits = 3)
bin n_reads proportion mean_distance median_distance
2 19761 0.741 16731.08 16490
1 6903 0.259 16773.47 16304
eval.ABxCB <- evaluate_bins(res.ABxCB, reads.mapping,
                            id = "anonymous_read_id", label = "genome_id")
knitr::kable(as.data.frame(eval.ABxCB$overall), digits = 3)
n_reads n_bins n_origins weighted_purity weighted_recovery adjusted_rand_index
26664 2 10 0.62 0.904 0.298

In hierarchical results, distances for bins outside a read’s parent abundance bin are missing. ambiguous_reads() ignores those missing distances and omits reads with fewer than two finite candidates.

8 In-memory inputs (Biostrings / ShortRead)

Instead of a path, you can pass a DNAStringSet, QualityScaledDNAStringSet, or ShortReadQ. The Java backend reads from disk, so non-file inputs are staged to a tempfile transparently.

reads <- Biostrings::readDNAStringSet(
    system.file("extdata", "reads.metagenome.fasta.gz", package = "metabinR"))

res.AB.mem <- abundance_based_binning(
    reads,
    numOfClustersAB = 2,
    kMerSizeAB = 10,
    dryRun = TRUE
)
identical(nrow(assignments(res.AB.mem)), nrow(assignments(res.AB)))
#> [1] TRUE

Clean up files written by the AB run:

unlink("vignette__*")

9 Session Info

utils::sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.5 LTS
#> 
#> Matrix products: default
#> BLAS:   /home/biocbuild/bbs-3.24-bioc/R/lib/libRblas.so 
#> LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.12.0  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=en_US.UTF-8          LC_NUMERIC=C                 
#>  [3] LC_TIME=en_GB                 LC_COLLATE=C                 
#>  [5] LC_MONETARY=en_US.UTF-8       LC_MESSAGES=en_US.UTF-8      
#>  [7] LC_PAPER=en_US.UTF-8          LC_NAME=en_US.UTF-8          
#>  [9] LC_ADDRESS=en_US.UTF-8        LC_TELEPHONE=en_US.UTF-8     
#> [11] LC_MEASUREMENT=en_US.UTF-8    LC_IDENTIFICATION=en_US.UTF-8
#> 
#> time zone: America/New_York
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats4    stats     graphics  grDevices utils     datasets  methods  
#> [8] base     
#> 
#> other attached packages:
#>  [1] Biostrings_2.81.9    Seqinfo_1.3.2        XVector_0.53.0      
#>  [4] IRanges_2.47.5       S4Vectors_0.51.10    BiocGenerics_0.59.12
#>  [7] generics_0.1.4       ggplot2_4.0.3        metabinR_2.1.1      
#> [10] BiocStyle_2.41.0    
#> 
#> loaded via a namespace (and not attached):
#>  [1] gtable_0.3.6             xfun_0.61                bslib_0.12.0            
#>  [4] hwriter_1.3.2.1          latticeExtra_0.6-31      rJava_1.0-18            
#>  [7] Biobase_2.73.2           lattice_0.23-1           vctrs_0.7.3             
#> [10] tools_4.6.1              bitops_1.1-0             parallel_4.6.1          
#> [13] tibble_3.3.1             pkgconfig_2.0.3          checkmate_2.3.4         
#> [16] RColorBrewer_1.1-3       S7_0.2.2                 cigarillo_1.3.1         
#> [19] lifecycle_1.0.5          compiler_4.6.1           farver_2.1.2            
#> [22] deldir_2.0-4             Rsamtools_2.29.0         tinytex_0.61            
#> [25] codetools_0.2-20         htmltools_0.5.9          sass_0.4.10             
#> [28] yaml_2.3.12              pillar_1.11.1            crayon_1.5.3            
#> [31] jquerylib_0.1.4          BiocParallel_1.47.0      cachem_1.1.0            
#> [34] magick_2.9.1             ShortRead_1.71.1         tidyselect_1.2.1        
#> [37] digest_0.6.39            dplyr_1.2.1              bookdown_0.48           
#> [40] labeling_0.4.3           fastmap_1.2.0            grid_4.6.1              
#> [43] cli_3.6.6                magrittr_2.0.5           dichromat_2.0-1         
#> [46] withr_3.0.3              scales_1.4.0             backports_1.5.1         
#> [49] rmarkdown_2.32           pwalign_1.9.1            jpeg_0.1-11             
#> [52] interp_1.1-6             otel_0.2.0               png_0.1-9               
#> [55] evaluate_1.0.5           knitr_1.52               GenomicRanges_1.65.4    
#> [58] rlang_1.3.0              Rcpp_1.1.2               glue_1.8.1              
#> [61] BiocManager_1.30.27      jsonlite_2.0.0           R6_2.6.1                
#> [64] GenomicAlignments_1.49.2