featureCounts: An efficient general-purpose program for assigning sequence reads to genomic features

featureCounts:一种高效且通用的程序,用于将序列读取分配到基因组特征上
Yang Liao, Gordon K Smyth, Wei Shi
2013-05-15 · arXiv:1305.3347 ↗
中文讲解 ·2:17 ·由 PaperFeed 生成

Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to each genomic feature, for example to each exon or each gene. The process of counting reads is called read summarization. Read summarization is required for a great variety of genomic analyses but has so far received relatively little attention in the literature. We present featureCounts, a read summarization program suitable for counting reads generated from either RNA or genomic DNA sequencing experiments. featureCounts implements highly efficient chromosome hashing and feature blocking techniques. It is considerably faster than existing methods (by an order of magnitude for gene-level summarization) and requires far less computer memory. It works with either single or paired-end reads and provides a wide range of options appropriate for different sequencing applications. featureCounts is available under GNU General Public License as part of the Subread (http://subread.sourceforge.net) or Rsubread (http://www.bioconductor.org) software packages.

featureCounts: an efficient general-purpose program for assigning sequence reads to genomic features
featureCounts:一种高效的通用程序,用于将序列读取分配到基因组特征上
1Bioinformatics Division, The Walter and Eliza Hall Institute of Medical Research, 1G Royal Parade, Parkville, VIC 3052, 2Department of Computing and Information Systems, 3Department of Mathematics and Statistics, The University of Melbourne, Parkville, VIC 3010, Australia
1. 生物信息学部,沃尔特和伊莱扎·霍尔医学研究所,1G Royal Parade, Parkville, VIC 3052,2. 墨尔本大学计算与信息系统系,3. 墨尔本大学数学与统计系,Parkville, VIC 3010,澳大利亚
Next-generation sequencing technologies generate millions of short sequence reads, which are usually aligned to a reference genome. In many applications, the key information required for downstream analysis is the number of reads mapping to each genomic feature, for example to each exon or each gene. The process of counting reads is called read summarization. Read summarization is required for a g …
下一代测序技术会产生数以百万计的短序列读段,这些读段通常与参考基因组进行比对。在许多应用中,下游分析所需的关键信息是映射到每个基因组特征(例如每个外显子或每个基因)的读段数量。这一读段计数过程被称为读段汇总。读段汇总对于多种基因组分析都是必需的,但迄今为止在文献中受到的关注相对较少。我们提出了featureCounts,这是一个适用于RNA或基因组DNA测序实验中生成的读段计数的读段汇总程序。featureCounts实现了高效的染色体哈希和特征分块技术。它比现有方法快得多(在基因级汇总方面快一个数量级),并且所需的计算机内存也少得多。它适用于单端或双端读段,并提供了一系列广泛的选项,以适应不同的测序应用。featureCounts遵循GNU通用公共许可证,作为Subread(http://subread.sourceforge.net)或Rsubread(http://www.bioconductor.org)软件包的一部分提供。
Next-generation (next-gen) sequencing technologies are revolution-izing biology by providing the ability to sequence DNA at unprecendented speed (Schuster, 2007; Metzker, 2009). The computational problem of mapping short sequence reads to a reference genome has received enormous attention in the past few years (Li and Durbin, 2009; Langmead et al., 2009; Fonseca et al., 2012; Marco-Sola et al., 20 …
下一代(next-gen)测序技术正以空前的速度对DNA进行测序,从而彻底改变生物学领域(Schuster, 2007; Metzker, 2009)。近年来,将短序列读段映射到参考基因组的计算问题受到了广泛关注(Li and Durbin, 2009; Langmead et al., 2009; Fonseca et al., 2012; Marco-Sola et al., 2012; Liao et al., 2013),而快速且可靠的比对器的快速发展是生物信息学领域的一个成功案例。然而,原始比对器输出通常不足以进行生物学解读。在对基因组特征进行生物学解读之前,必须根据读段覆盖度对读段映射结果进行总结。在许多下一代分析流程中,最普遍的操作之一是计算与预先确定的基因组特征重叠的读段数量。根据下一代测序的应用,这些基因组特征可能是外显子、基因、启动子区域、基因体或其他基因组区间。在基于计数的一系列统计方法中,如差异表达分析或差异结合分析,都需要用到读段计数(Oshlack et al., 2010)。
Despite its importance in genomic research, the read counting problem has received little specific attention in the literature. The problem may appear superficially simple but in practice has many subtleties. Read count programs need to accommodate both DNA and RNA sequencing as well as single and paired-end reads. The reads or paired-end fragments to be counted may incorporate insertions, deletio …
尽管在基因组研究中具有重要意义,但读数计数问题在文献中却鲜有专门关注。该问题表面上看似简单,但实际上却存在诸多细微之处。读数计数程序需同时适应DNA和RNA测序,以及单端和双端读数。待计数的读数或双端片段可能包含相对于参考基因组的插入、缺失或融合,在将每个读数或片段的位置与每个可能的靶基因组特征进行比较时,这些复杂情况需予以考虑。当特征数量众多时,读数计数的计算成本可能与读数比对步骤相当。
DNA sequence reads arise from a variety of technologies including ChIP-seq for transcription factor binding sites (Valouev et al., 2008), ChIP-seq for histone marks (Park, 2009), and assays that detect DNA methylation (Harris et al., 2010). The genomic features of interest for DNA reads can usually be specified in terms of simple genomic intervals. For example, Pal et al. (2013) counted reads asso …
DNA序列读取数据来源于多种技术,包括用于转录因子结合位点的ChIP-seq(Valouev等人,2008年)、用于组蛋白标记的ChIP-seq(Park,2009年)以及检测DNA甲基化的分析方法(Harris等人,2010年)。DNA读取数据中感兴趣的基因组特征通常可以用简单的基因组区间来指定。例如,Pal等人(2013年)统计了与基因启动子区域和整个基因体相关的组蛋白标记相关的读取数据。Ross-Innes等人(2012年)统计了与峰值调用器(Zhang等人,2008年)识别的区间重叠的读取数据。
Counting RNA-seq reads is somewhat more complex because of the need to accommodate exon splicing. One way is to count reads overlapping each annotated exon, an approach that can be used to test for alternative splicing between experimental conditions (Anders et al., 2012; Reyes et al., 2013). Another common approach is to summarize counts at the gene-level, by counting all reads that overlap any e …
RNA-seq读数的计数工作因需考虑外显子剪接而变得更为复杂。一种方法是统计与每个已注释外显子重叠的读数,此方法可用于检测不同实验条件下的可变剪接(Anders等,2012;Reyes等,2013)。另一种常见的方法是在基因层面汇总计数,即统计每个基因所有与任一外显子重叠的读数(Anders等,2013;Bhattacharyya等,2013;Man等,2013)。为此,通常使用来自RefSeq(Pruitt等,2012)或Ensembl(Flicek等,2012)的基因注释信息。
Read counts provide an overall summmary of the coverage for the genomic feature of interest. In particular, gene-level counts from RNA-seq provide an overall summary of the expression level of the gene but do not distinguish between isoforms when multiple transcripts are being expressed from the same gene. Reads can generally be assigned to genes with good confidence, but estimating the expression …
读数提供了目标基因组特征覆盖度的总体概述。特别是,RNA-seq的基因水平计数提供了基因表达水平的总体概述,但在同一基因表达多个转录本的情况下,无法区分同工型。通常,可以较为可靠地将读数分配给基因,但估计单个同工型的表达水平本质上更为困难,因为基因的不同同工型通常具有较高比例的基因组重叠。已经开发出多种基于模型的方法,试图从RNA-seq数据中解卷积出每个基因的单个转录本表达水平,主要是利用明确分配给同工型差异区域的读数信息(Trapnell等,2010年;Li和Dewey,2011年)。本文重点关注读数问题,即使在测序深度不足以使转录本水平分析可靠的情况下,该问题也普遍适用。已经开发出多种统计分析方法,基于读数来检测差异表达或差异结合(McCarthy等,2012年;Anders和Huber,2010年;Li等,2012年;Hardcastle和Kelly,2010年;Auer和Doerge,2011年;Wu等,2013年)。最近的比较得出结论,在基因水平差异表达(Nookaew等,2012年;Rapaport等,2013年)或剪接变异检测(Anders等,2012年)方面,读数方法 …
以上为前 8 段试读
本篇已缓存 68 段译文,全文原文对照与段落级翻译在 App 中继续阅读。
在 PaperFeed 中阅读全文
∎
保存图片分享
featureCounts: An efficient general-purpose program for assigning sequence reads to genomic features
长按图片 → 保存到相册 → 发送到微信群
在 PaperFeed 中打开