Import and summarize transcript-level abundance estimates for transcript- and gene-level analysis with Bioconductor packages, such as edgeR, DESeq2, and limma-voom. The motivation and methods for the functions provided by the tximport package are described in the following article (Soneson, Love, and Robinson 2015):
Charlotte Soneson, Michael I. Love, Mark D. Robinson (2015): Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Research http://dx.doi.org/10.12688/f1000research.7563.1
In particular, the tximport pipeline offers the following benefits: (i) this approach corrects for potential changes in gene length across samples (e.g. from differential isoform usage) (Trapnell et al. 2013), (ii) some of the upstream quantification methods (salmon, sailfish, kallisto) are substantially faster and require less memory and disk usage compared to alignment-based methods that require creation and storage of BAM files, and (iii) it is possible to avoid discarding those fragments that can align to multiple genes with homologous sequence, thus increasing sensitivity (Robert and Watson 2015).
Note: another Bioconductor package, tximeta (Love et al. 2020), extends tximport,
offering the same functionality, plus the additional benefit of
automatic addition of annotation metadata for commonly used
transcriptomes (GENCODE, Ensembl, RefSeq for human and mouse). See the
tximeta package
vignette for more details. Whereas tximport outputs a
simple list of matrices, tximeta will output a
SummarizedExperiment object with appropriate GRanges
added if the transcriptome is from one of the sources above for human
and mouse.
We begin by locating some prepared files that contain transcript
abundance estimates for two samples, from the tximportData
package. The tximport pipeline will be nearly identical for
various quantification tools, usually only requiring one change the
type argument. We begin with quantification files generated
by the salmon software, and later show the use of
tximport with any of:
First, we locate the directory containing the files. (Here we use
system.file to locate the package directory, but for a
typical use, we would just provide a path,
e.g. "/path/to/dir".)
## [1] "alevin" "cufflinks"
## [3] "kallisto" "oarfish"
## [5] "refseq" "rsem"
## [7] "sailfish" "salmon"
## [9] "salmon_dm" "salmon_ec"
## [11] "samples.txt" "samples_extended.txt"
## [13] "tx2gene.csv" "tx2gene.gencode.v27.csv"
## [15] "tx2gene_alevin.tsv"
Next, we create a named vector pointing to the quantification files.
We will create a vector of filenames first by reading in a table that
contains the sample IDs, and then combining this with dir
and "quant.sf.gz". (We gzipped the quantification files to
make the data package smaller, this is not a problem for R functions
that we use to import the files.)
## pop center assay sample experiment run
## 1 TSI UNIGE NA20504.1.M_111124_7 ERS185242 ERX162972 ERR188088
## 2 TSI UNIGE NA20508.1.M_111124_2 ERS185362 ERX163159 ERR188021
files <- file.path(dir, "salmon", samples$run, "quant.sf.gz")
names(files) <- paste0("sample", 1:2)
all(file.exists(files))## [1] TRUE
Transcripts need to be associated with gene IDs for gene-level
summarization. If that information is present in the files, we can skip
this step. For salmon, sailfish, and kallisto the files only provide the
transcript ID. We first make a data.frame called tx2gene
with two columns: 1) transcript ID and 2) gene ID. The column names do
not matter but this column order must be used. The transcript ID must be
the same one used in the abundance files.
Creating this tx2gene data.frame can be accomplished
from a TxDb object and the select function from
the AnnotationDbi package. The following code could be used to
construct such a table:
library(TxDb.Hsapiens.UCSC.hg19.knownGene)
txdb <- TxDb.Hsapiens.UCSC.hg19.knownGene
k <- keys(txdb, keytype = "TXNAME")
tx2gene <- select(txdb, k, "GENEID", "TXNAME")Note: if you are using an Ensembl transcriptome, the easiest
way to create the tx2gene data.frame is to use the ensembldb
packages. The annotation packages can be found by version number, and
use the pattern EnsDb.Hsapiens.vXX. The
transcripts function can be used with
return.type="DataFrame", in order to obtain something like
the df object constructed in the code chunk above. See the
ensembldb package vignette for more details.
In this case, we’ve used the Gencode v27 CHR transcripts to build our
index, and we used makeTxDbFromGFF and code similar to the
chunk above to build the tx2gene table. We then read in a
pre-constructed tx2gene table:
## # A tibble: 6 × 2
## TXNAME GENEID
## <chr> <chr>
## 1 ENST00000456328.2 ENSG00000223972.5
## 2 ENST00000450305.2 ENSG00000223972.5
## 3 ENST00000473358.1 ENSG00000243485.5
## 4 ENST00000469289.1 ENSG00000243485.5
## 5 ENST00000607096.1 ENSG00000284332.1
## 6 ENST00000606857.1 ENSG00000268020.3
The tximport package has a single function for importing
transcript-level estimates. The type argument is used to
specify what software was used for estimation. A simple list with
matrices, "abundance", "counts", and
"length", is returned, where the transcript level
information is summarized to the gene-level. Typically, abundance is
provided by the quantification tools as TPM (transcripts-per-million),
while the counts are estimated counts (possibly fractional), and the
"length" matrix contains the effective gene lengths. The
"length" matrix can be used to generate an offset matrix
for downstream gene-level differential analysis of count matrices, as
shown below.
Note: While tximport works without any
dependencies, it is significantly faster to read in files using the
readr package. If tximport detects that readr
is installed, then it will use the readr::read_tsv function
by default. A change from version 1.2 to 1.4 is that the reader is not
specified by the user anymore, but chosen automatically based on the
availability of the readr package. Advanced users can still
customize the import of files using the importer
argument.
## [1] "abundance" "counts" "length"
## [4] "countsFromAbundance"
## sample1 sample2
## ENSG00000000003.14 2.0000 5.11217
## ENSG00000000005.5 0.0000 0.00000
## ENSG00000000419.12 1337.9970 920.99960
## ENSG00000000457.13 498.8622 531.57867
## ENSG00000000460.16 418.4429 915.70085
## ENSG00000000938.12 3697.9989 1897.00030
We could alternatively generate counts from abundances, using the
argument countsFromAbundance, scaled to library size,
"scaledTPM", or additionally scaled using the average
transcript length, averaged over samples and to library size,
"lengthScaledTPM". Using either of these approaches, the
counts are not correlated with length, and so the length matrix should
not be provided as an offset for downstream analysis packages. As of
tximport version 1.10, we have added a new
countsFromAbundance option "dtuScaledTPM".
This scaling option is designed for use with txOut=TRUE for
differential transcript usage analyses. See ?tximport for
details on the various countsFromAbundance options.
We can avoid gene-level summarization by setting
txOut=TRUE, giving the original transcript level estimates
as a list of matrices.
These matrices can then be summarized afterwards using the function
summarizeToGene. This then gives the identical list of
matrices as using txOut=FALSE (default) in the first
tximport call.
## [1] TRUE
salmon or sailfish quant.sf files can be imported by
setting type to "salmon" or "sailfish".
files <- file.path(dir, "salmon", samples$run, "quant.sf.gz")
names(files) <- paste0("sample", 1:2)
txi.salmon <- tximport(files, type = "salmon", tx2gene = tx2gene)
head(txi.salmon$counts)## sample1 sample2
## ENSG00000000003.14 2.0000 5.11217
## ENSG00000000005.5 0.0000 0.00000
## ENSG00000000419.12 1337.9970 920.99960
## ENSG00000000457.13 498.8622 531.57867
## ENSG00000000460.16 418.4429 915.70085
## ENSG00000000938.12 3697.9989 1897.00030
We quantified with sailfish against a different transcriptome, so we
need to read in a different tx2gene for this next code
chunk.
tx2knownGene <- read_csv(file.path(dir, "tx2gene.csv"))
files <- file.path(dir, "sailfish", samples$run, "quant.sf")
names(files) <- paste0("sample", 1:2)
txi.sailfish <- tximport(files, type = "sailfish", tx2gene = tx2knownGene)
head(txi.sailfish$counts)## sample1 sample2
## A1BG 316.13800 86.26030
## A1BG-AS1 141.15800 127.03400
## A1CF 10.02056 25.23152
## A2M 2.00000 38.00000
## A2M-AS1 1.00000 0.00000
## A2ML1 1.02936 3.07782
Note: for previous version of salmon or sailfish, in which
the quant.sf files start with comment lines, it is
recommended to specify the importer argument as a function
which reads in the lines beginning with the header. For example, using
the following code chunk (un-evaluated):
If inferential replicates (Gibbs or bootstrap samples) are present in
expected locations relative to the quant.sf file,
tximport will import these as well, as a list of matrices
txi$infReps, with one matrix per sample (rows are
transcripts, columns are replicates). tximportData no longer
includes inferential replicates, to keep the package small; salmon Gibbs
samples and kallisto bootstraps for the GEUVADIS samples used in earlier
versions are archived on Zenodo at https://doi.org/10.5281/zenodo.22982575. (The following
chunk is not evaluated.)
files <- file.path("path/to", samples$run, "quant.sf")
names(files) <- paste0("sample", 1:2)
txi.inf.rep <- tximport(files, type = "salmon", txOut = TRUE)
names(txi.inf.rep$infReps)The tximport arguments varReduce and
dropInfReps can be used to summarize the inferential
replicates into a single variance per transcript/gene and per sample, or
to not import inferential replicates, respectively.
kallisto abundance.h5 files can be imported by setting
type to "kallisto". Note that this requires that you have
the Bioconductor package rhdf5 installed. If
the abundance.h5 files contain bootstrap replicates, these
will be imported as inferential replicates, as with salmon above (stored
in txi.kallisto$infReps). (The following chunk is not
evaluated, as tximportData no longer includes
abundance.h5 files.)
kallisto abundance.tsv files can be imported as well,
but this is typically slower than importing abundance.h5
files. Note that we add an additional argument in this code chunk,
ignoreAfterBar=TRUE. This is because the Gencode
transcripts have names like “ENST00000456328.2|ENSG00000223972.5|…”,
though our tx2gene table only includes the first “ENST”
identifier. We therefore want to split the incoming quantification
matrix rownames at the first bar “|”, and only use this as an
identifier. We didn’t use this option earlier with salmon, because we
used the argument --gencode when running salmon, which
itself does the splitting upstream of tximport. Note that
ignoreTxVersion and ignoreAfterBar are only to
facilitating the summarization to gene level.
files <- file.path(dir, "kallisto", samples$run, "abundance.tsv.gz")
names(files) <- paste0("sample", 1:2)
txi.kallisto.tsv <- tximport(files, type = "kallisto", tx2gene = tx2gene, ignoreAfterBar = TRUE)
head(txi.kallisto.tsv$counts)## sample1 sample2
## ENSG00000000003.14 2.0000 5.06463
## ENSG00000000005.5 0.0000 0.00000
## ENSG00000000419.12 1338.0006 921.00030
## ENSG00000000457.13 495.4173 532.84843
## ENSG00000000460.16 418.5453 915.49327
## ENSG00000000938.12 3697.9998 1896.99980
RSEM sample.genes.results files can be imported by
setting type to "rsem", and txIn and
txOut to FALSE.
files <- file.path(dir, "rsem", samples$run, paste0(samples$run, ".genes.results.gz"))
names(files) <- paste0("sample", 1:2)
txi.rsem <- tximport(files, type = "rsem", txIn = FALSE, txOut = FALSE)
head(txi.rsem$counts)## sample1 sample2
## ENSG00000000003.14 2.00 0
## ENSG00000000005.5 0.00 0
## ENSG00000000419.12 1250.00 0
## ENSG00000000457.13 457.55 1
## ENSG00000000460.16 407.84 2
## ENSG00000000938.12 3466.00 3
RSEM sample.isoforms.results files can be imported by
setting type to "rsem", and txIn and
txOut to TRUE.
files <- file.path(dir, "rsem", samples$run, paste0(samples$run, ".isoforms.results.gz"))
names(files) <- paste0("sample", 1:2)
txi.rsem <- tximport(files, type = "rsem", txIn = TRUE, txOut = TRUE)
head(txi.rsem$counts)## sample1 sample2
## ENST00000373020.8 0 0
## ENST00000494424.1 0 0
## ENST00000496771.5 0 0
## ENST00000612152.4 0 0
## ENST00000614008.4 2 0
## ENST00000373031.4 0 0
StringTie t_data.ctab files giving the coverage and
abundances for transcripts can be imported by setting type to
stringtie. These files can be generated with the following
command line call:
stringtie -eB -G transcripts.gff <source_file.bam>
tximport will compute counts from the coverage information,
by reversing the formula that StringTie uses to calculate coverage (see
?tximport). The read length is used in this formula, and so
if you’ve set a different read length when using StringTie, you can
provide this information with the readLength argument. The
tx2gene table should connect transcripts to genes, and can
be pulled out of one of the t_data.ctab files. The tximport
call would look like the following (here not evaluated):
scRNA-seq data quantified with alevin can be easily imported
using tximport. The following unevaluated example shows import
of the quants matrix (for a live example, see the unit test file
test_alevin.R). A single file should be specified which
will import a gene-by-cell matrix of data.
Long read data quantified with oarfish can be imported using tximport. The following example shows import of three samples from SG-Nex. See vignette of tximportData for more details. Note this quantification includes ~8k additional transcripts of length 1,000 bp drawn from chr1-22.
Note: there are two suggested ways of importing
estimates for use with differential gene expression (DGE) methods. The
first method, which we show below for edgeR, limma and
for DESeq2, is to use the gene-level estimated counts from the
quantification tools, and additionally to use the transcript-level
abundance estimates to calculate a gene-level offset that corrects for
changes to the average transcript length across samples. The code
examples below accomplish these steps for you, keeping track of
appropriate matrices and calculating these offsets. The functions
edgeR::DGEListFromTximport, for edgeR and
limma, and DESeq2::DESeqDataSetFromTximport, for
DESeq2, take care of creation of the offset for you. Let’s call
this method “original counts and offset”.
The second method is to use the tximport argument
countsFromAbundance="lengthScaledTPM" or
"scaledTPM", and then to use the gene-level count matrix
txi$counts directly as you would a regular count matrix
with these software. Let’s call this method “bias corrected counts
without an offset”
Note: Do not manually pass the original gene-level
counts to downstream methods without an offset. The only case
where this would make sense is if there is no length bias to the counts,
as may happen in 3’ tagged RNA-seq data (see section below), or if there
is little to no fragmentation before sequencing. The original gene-level
counts are in txi$counts when tximport was run
with countsFromAbundance="no". Using only
txi$counts is simply passing the summed estimated
transcript counts, and does not correct for potential differential
isoform usage (the offset), which is the point of the tximport
methods (Soneson, Love, and Robinson 2015)
for gene-level analysis. Passing uncorrected gene-level counts without
an offset is not recommended by the tximport package authors.
The two methods we provide here are: “original counts and
offset” or “bias corrected counts without an offset”.
Passing txi to
DESeq2::DESeqDataSetFromTximport or
edgeR::DGEListFromTximport as outlined below is correct:
the functions create the appropriate offset for you to perform
gene-level differential expression.
If you have 3’ tagged RNA-seq data, then correcting the counts for
gene length will induce a bias in your analysis, because the counts do
not have length bias. Instead of using the default
full-transcript-length pipeline, we recommend to use the original
counts, e.g. txi$counts as a counts matrix, e.g. providing
to DESeqDataSetFromMatrix or to the edgeR or
limma functions without calculating an offset and without using
countsFromAbundance.
An example of creating a DESeqDataSet for use with
DESeq2 (Love, Huber, and Anders
2014):
The user should make sure the rownames of sampleTable
align with the colnames of txi$counts, if there are
colnames. The best practice is to read sampleTable from a
CSV file, and to construct files from a column of
sampleTable, as was shown in the tximport examples
above.
The following code will create a DGEList object
containing gene-level counts for analysis with the edgeR
package (Chen et al. 2025):
The DGEListFromTximport function is available in
edgeR as of version 4.10.0 (Bioconductor release 3.23, Spring
2026).
The object dge is ready for any of the edgeR
analysis pipelines, starting usually with filterByExpr and
normLibSizes followed by glmQLFit. See the
edgeR User’s Guide.
If the tximport output contains inferential replicates,
then DGEListFromTximport will also estimate the count
overdispersion that arises from the fact that reads have to assigned to
transcripts probabilistically (Baldoni et al.
2024). The following code will compute “divided counts” (Baldoni et al. 2024) suitable for differential
expression analysis at the transcript level:
See the edgeR User’s Guide for more details.
The DGEList objects created by DGEListFromTximport are
also suitable for analysis with limma-voom. In the
limma-voom pipeline, dge is preprocessed by
filterByExpr and normLibSizes, as for
edgeR, but then glmQLFit is replaced by either
voom and lmFit (Law et
al. 2014) or by the newer function voomLmFit (Baldoni et al. 2025). See the limma User’s
Guide for more details.
The development of tximport has benefited from contributions and suggestions from:
Rob Patro (inferential
replicates import), Andrew Parker Morgan
(RHDF5 support), Ryan C.
Thompson (RHDF5 support), Matt
Shirley (ignoreTxVersion), Avi
Srivastava (alevin import), Scott Van Buren (infReps
testing), Gordon Smyth and Pedro Baldoni (edgeR and
limma functions and docs), Stephen Turner, Richard Smith-Unna, Rory Kirchner, Martin Morgan, Jenny Drnevich,
Patrick Kimes, Leon Fodoulian, Koen Van den Berge, Aaron Lun, Alexander Toenges
## R version 4.6.1 (2026-06-24)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 24.04.5 LTS
##
## Matrix products: default
## BLAS: /home/biocbuild/bbs-3.24-bioc/R/lib/libRblas.so
## LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.12.0 LAPACK version 3.12.0
##
## locale:
## [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
## [3] LC_TIME=en_GB LC_COLLATE=C
## [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
## [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
## [9] LC_ADDRESS=C LC_TELEPHONE=C
## [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
##
## time zone: America/New_York
## tzcode source: system (glibc)
##
## attached base packages:
## [1] stats4 stats graphics grDevices utils datasets methods
## [8] base
##
## other attached packages:
## [1] edgeR_4.99.6
## [2] limma_3.99.0
## [3] DESeq2_1.53.4
## [4] SummarizedExperiment_1.43.0
## [5] MatrixGenerics_1.25.0
## [6] matrixStats_1.5.0
## [7] tximport_1.41.1
## [8] readr_2.2.0
## [9] TxDb.Hsapiens.UCSC.hg19.knownGene_3.22.1
## [10] GenomicFeatures_1.65.0
## [11] AnnotationDbi_1.75.2
## [12] Biobase_2.73.2
## [13] GenomicRanges_1.65.4
## [14] Seqinfo_1.3.2
## [15] IRanges_2.47.5
## [16] S4Vectors_0.51.10
## [17] BiocGenerics_0.59.12
## [18] generics_0.1.4
## [19] tximportData_1.41.1
## [20] knitr_1.52
##
## loaded via a namespace (and not attached):
## [1] tidyselect_1.2.1 dplyr_1.2.1 farver_2.1.2
## [4] blob_1.3.0 S7_0.2.2 Biostrings_2.81.9
## [7] bitops_1.1-0 fastmap_1.2.0 RCurl_1.98-1.20
## [10] GenomicAlignments_1.49.2 XML_3.99-0.25 digest_0.6.39
## [13] lifecycle_1.0.5 statmod_1.5.2 KEGGREST_1.53.6
## [16] RSQLite_3.53.3 magrittr_2.0.5 compiler_4.6.1
## [19] rlang_1.3.0 sass_0.4.10 tools_4.6.1
## [22] utf8_1.2.6 yaml_2.3.12 rtracklayer_1.73.0
## [25] S4Arrays_1.13.1 bit_4.6.0 curl_8.0.0
## [28] DelayedArray_0.39.7 RColorBrewer_1.1-3 abind_1.4-8
## [31] BiocParallel_1.47.0 grid_4.6.1 ggplot2_4.0.3
## [34] scales_1.4.0 dichromat_2.0-1 cli_3.6.6
## [37] rmarkdown_2.32 crayon_1.5.3 otel_0.2.0
## [40] httr_1.4.9 tzdb_0.5.0 rjson_0.2.23
## [43] BiocBaseUtils_1.15.1 DBI_1.3.0 cachem_1.1.0
## [46] parallel_4.6.1 formatR_1.14 XVector_0.53.0
## [49] restfulr_0.0.17 vctrs_0.7.3 Matrix_1.7-6
## [52] jsonlite_2.0.0 hms_1.1.4 bit64_4.8.6
## [55] locfit_1.5-9.12 jquerylib_0.1.4 glue_1.8.1
## [58] codetools_0.2-20 gtable_0.3.6 BiocIO_1.23.3
## [61] tibble_3.3.1 pillar_1.11.1 htmltools_0.5.9
## [64] R6_2.6.1 vroom_1.7.1 evaluate_1.0.5
## [67] lattice_0.23-1 png_0.1-9 Rsamtools_2.29.0
## [70] cigarillo_1.3.1 memoise_2.0.1 bslib_0.12.0
## [73] Rcpp_1.1.2 SparseArray_1.13.3 xfun_0.61
## [76] pkgconfig_2.0.3