← New search

Other meanings of RNA-Seq

Genomics

RNA-Seq

RNA-Seq (RNA sequencing) is a high-throughput sequencing method used to quantify and profile RNA transcripts in a biological sample. It provides a comprehensive view of the transcriptome, enabling the detection of gene expression levels, alternative splicing, and novel transcripts. Since its introduction in the late 2000s, RNA-Seq has become a cornerstone of molecular biology, replacing microarrays for many applications due to its higher sensitivity and dynamic range.

2008
First described
Year of initial publication
~10-30M
Reads per sample
Typical sequencing depth
>90%
Genome coverage
For well-annotated organisms
1-5%
Error rate
Per base in Illumina platforms
1

Core methodology

RNA-Seq begins with the isolation of total RNA from a sample, followed by enrichment for polyadenylated mRNA or depletion of ribosomal RNA to reduce abundant non-coding transcripts. The RNA is then fragmented and reverse-transcribed into complementary DNA (cDNA), which is ligated to sequencing adapters and amplified. High-throughput sequencing platforms, such as Illumina's sequencing-by-synthesis, generate millions of short reads (typically 50–150 base pairs) from the cDNA library. These reads are then aligned to a reference genome or transcriptome, or assembled de novo, to quantify transcript abundance and identify splice junctions. The number of reads mapping to a gene serves as a proxy for its expression level, normalized across samples using metrics like RPKM, FPKM, or TPM to account for sequencing depth and gene length.1

2

Applications and impact

RNA-Seq has transformed the study of gene expression, enabling researchers to compare transcriptomes across conditions, tissues, or developmental stages. It is widely used in cancer research to identify differentially expressed genes, fusion transcripts, and mutations in expressed alleles. In clinical diagnostics, RNA-Seq helps characterize tumor subtypes and guide targeted therapies. Beyond gene expression, RNA-Seq reveals alternative splicing events, RNA editing, and allele-specific expression, providing insights into post-transcriptional regulation. It also facilitates the discovery of novel transcripts, including long non-coding RNAs and circular RNAs, which were previously difficult to detect with microarrays. The technology has been adapted for single-cell RNA-Seq (scRNA-Seq), allowing transcriptomic profiling of individual cells and revealing cellular heterogeneity in tissues.2

3

Challenges and bioinformatics

RNA-Seq data analysis is computationally intensive and requires a robust bioinformatics pipeline. Key steps include quality control, read alignment, quantification, and differential expression analysis. Challenges include handling reads that map to multiple genomic locations, accounting for batch effects, and normalizing data across samples. The choice of alignment tool (e.g., STAR, HISAT2) and quantification method (e.g., featureCounts, Salmon) can significantly influence results. Additionally, RNA-Seq is sensitive to technical artifacts such as GC bias and PCR duplicates, which require careful correction. The field has developed best practices, such as those from the ENCODE consortium, to standardize protocols and ensure reproducibility. Despite these challenges, RNA-Seq remains the gold standard for transcriptome analysis, with continuous improvements in throughput and cost-effectiveness.3

4

Lesser-known aspects

Beyond standard mRNA profiling, RNA-Seq has niche applications that are less widely known. For instance, it can be used to detect RNA modifications such as N6-methyladenosine (m6A) through specialized protocols like MeRIP-Seq. RNA-Seq also enables the study of microbial transcriptomes, including metatranscriptomics of complex microbiomes. In evolutionary biology, RNA-Seq has been applied to non-model organisms to identify conserved and species-specific transcripts. A notable edge case is the use of RNA-Seq to detect viral RNA in infected cells, which has been pivotal in identifying novel viruses. Additionally, the technique has been adapted for degraded RNA from formalin-fixed paraffin-embedded (FFPE) tissues, expanding its use in retrospective clinical studies. The development of long-read RNA-Seq (e.g., using Oxford Nanopore) allows full-length transcript sequencing, overcoming the limitations of short-read assembly for complex isoforms.4

Glossary

Transcriptome
The complete set of RNA transcripts produced by the genome under specific circumstances.
cDNA
Complementary DNA synthesized from an RNA template by reverse transcriptase.
scRNA-Seq
Single-cell RNA sequencing, a technique for profiling gene expression in individual cells.
RPKM
Reads per kilobase of transcript per million mapped reads, a normalization metric.
FPKM
Fragments per kilobase of transcript per million mapped fragments, similar to RPKM but for paired-end reads.
TPM
Transcripts per million, a normalization method that is more consistent across samples.
Alternative splicing
A process by which a single gene can produce multiple mRNA variants.

RNA-Seq has evolved rapidly since its inception, with ongoing improvements in sequencing chemistry and computational methods.