RNA sequencing, or transcriptome sequencing, is a crucial tool used in molecular biology to study cell or tissue phenotypes. It involves the high-throughput sequencing of RNA molecules, allowing for the identification and quantification of all RNA transcripts present in a sample. In this topic, we will explore the principles of RNA sequencing.
Why do we need to study RNA?
All somatic cells in an individual organism harbor the same set of genes. Thus, they have the same genotype (excluding somatic mutations). However, cells in the organism differ in their phenotypes — size, shape, functions, state (health/diseased), etc. Researchers are interested in studying cell or tissue phenotypes since such knowledge could shed light on differences between healthy and diseased tissue or trace the process of cell differentiation etc.
In the first approximation a set of all proteins present at the moment forms a phenotype. Therefore, one way to study phenotypes is by the investigation of proteomes. However, such studies are expensive and require a huge amount of effort and complex equipment. Moreover, it is difficult to match a protein with the original gene after data processing. Another way to explore phenotype is through RNA instead of proteins. Let's look at how it works.
The central dogma of molecular biology says that information is transmitted from DNA to RNA (transcription) and then from RNA to protein (translation). A set of all RNA transcripts synthesized by an individual cell or a population of cells is called a transcriptome. Although the genome sequence remains relatively stable over time, the transcriptome is subjected to dynamic changes throughout cell life. The composition and amount of transcripts vary during cell differentiation or in response to environmental changes, allowing cell adaptation.
Since RNA and proteins are related through translation, changes in a transcriptome reflect changes in a proteome, and therefore phenotype. Therefore, another way to explore cell/tissue phenotypes is to characterize transcriptome, in other words, identify various RNA molecules and determine their quantities. Thus, transcriptomics is the study of the transcriptome.
Transcriptomics technologies
With the development of NGS technologies, transcriptome studies became more accessible and high-throughput. Previously, we have already discussed microarrays, which allow the investigation of particular RNA molecules of known sequence. But to characterize the whole transcriptome you need to know the types and amounts of all present RNAs. The concept of retrieving transcriptome differs from the one used in microarrays.
There are two approaches to studying transcriptomes. The first one is bulk RNA sequencing (bulk RNA-seq) which investigates the transcriptome of a population of cells, for example, a tissue sample. The approach detects the average expression level of genes across numerous cells and captures the global sample features. The basic design of bulk RNA-seq experiment involves sequencing samples of at least two conditions (for example, healthy vs. diseased). Researchers compare transcriptomes of different samples and search for genes that significantly change expression levels between various conditions. Those genes are called differentially expressed genes (DEG). When DEGs are identified, we can speculate about their roles in particular phenotype formation.
A more recent approach, single-cell RNA sequencing (scRNA-seq), focuses on the transcriptome of individual cells. The method spots slight differences in gene expression between single cells or helps distinguish cells of various types during the differentiation process.
RNA isolation
Generally, both, bulk and single-cell RNA sequencing work on the same principle. In the first step, RNA should be isolated from the sample. The fact is that ribosomal RNA (rRNA) comprises the majority (>80% to 90%) of the molecules present in the cell transcriptome. However, when studying phenotypes we are interested in mRNA rather than rRNA. Therefore, we should separate the RNA fraction of interest from other fractions. Otherwise, we sequenced predominantly rRNA, which wouldn't provide an insight into cell/tissue phenotype and would be a waste of money.
RNA transcripts are synthesized by RNA-polymerase on the DNA template. Whether a transcript would be polyadenylated or not depends on the type of RNA polymerase responsible for its synthesis. Normally, all mRNA, except histone mRNA, and some long non-coding RNA have polyA-tail, while rRNA, tRNA, and histone mRNA don't.
One way to do it is to use oligo(dT)-hybridization. The method selects only polyadenylated transcripts. The approach is based on oligo(dT) molecules attached to magnetic beads. Following the principle of the complementarity of bases, the transcript's polyA-tail hybridizes to the oligo(dT) sequence. Therefore, RNA of interest would be attached to the magnetic beads. Then the beads are isolated using a magnet and after that transcripts are washed from the beads.
There is another way to eliminate rRNA. The method is called rRNA depletion and uses DNA probes (single-stranded DNA) complement to rRNA. When such probes are added to the sample with total RNA, they hybridize to rRNA forming DNA-RNA duplexes. The special enzyme RNAse H degrades all RNA in the DNA-RNA duplexes. Afterwards, the second enzyme DNAse I eliminates DNA from the sample. This approach could also utilize magnetic beads. In this case, DNA probes are attached to the beads and they are removed from the sample with a magnet.
Both methods, oligo(dT)-hybridization and rRNA depletion, have their advantages and drawbacks. Oligo(dT)-hybridization does not extract tRNA, histone mRNA, or small non-coding RNA, while some of them are of interest to science. At the same time, rRNA depletion preserves some types of non-polyadenylated RNA, but eliminates others, for example, certain long and short non-coding RNAs. However, oligo(dT)-hybridization is used more frequently.
RNA-sequencing
NGS technologies are not designed for direct RNA sequencing. After RNA is isolated, it should be first converted into DNA. In nature, some viruses have special enzymes called reverse transcriptases that synthesize complementary DNA strands (cDNA) on RNA templates using DNA primers. There are two types of primers used in RNA-seq: oligo-(dT) primers anneal to the transcripts poly-A tail while primers of random sequence can bind to any part of the RNA molecule, allowing for the generation of cDNA fragments that represent the entire transcriptome including not polyadenylated transcripts.
RNA-sequencing protocols utilize particular reverse transcriptase from retrovirus because it synthesizes cDNA on an RNA template and then uses this cDNA strand as a template to create the second DNA strand. As a result, we obtain double-stranded DNA molecules that are ready for sequencing. They undergo library preparation protocol, including fragmentation, adapters ligation, and PCR amplification, and then sequenced with Illumina or other technology.
You can revise material about Illumina sequencing from Illumina/Solexa sequencing method topic.
As a result, we obtain reads, that can be mapped on the reference genome. Then we count the number of reads aligned with each gene. These processes are not as simple as they seem. They involve complicated algorithms and statistical models to map reads derived from spliced RNA with no introns and possible RNA polymerase errors to the reference genome and to count reliable expression levels for various RNA isoforms. Thus, RNA-sequencing data represents a table with genes along the rows and samples (for bulk RNA-seq) or cells (for single-cell RNA-seq) along the columns. The table is filled with the number of reads aligned with each gene. Then the data is used to identify differentially expressed genes. It is also a challenging task that we will discuss in the next topic.
Let's now summarize the complete RNA-seq protocol:
- Obtain tissue samples (cell population);
- Extract RNA from the samples;
- Select polyadenylated RNA or ribo-depleted RNA;
- cDNA synthesis;
- Construct cDNA library;
- Sequencing;
- Map reads to the reference genome and count reads aligned with each gene.
Here, we looked into a typical bulk RNA-seq procedure, on the other hand, scRNA-seq analyzes the transcriptome of individual cells. Thus, the workflow slightly differs from the bulk RNA-seq. Let's take a look at the scRNA-seq process and at its unique steps.
Single-cell RNA-sequencing
Single-cell isolation
An important step in scRNA-seq is the isolation of individual cells. There are two approaches to performing this task. The methods differ by isolation methods and the place, where initial reactions are conducted (droplet or well).
- Plate-based methods. We can use a cell sorter or microdissection to obtain single cells from a cell suspension. Each cell is then placed into a well of a plate and the latter reactions, such as cell lysis, RNA extraction, and cDNA synthesis, are conducted individually for a cell in a well.
- Droplet-based methods. We can encapsulate thousands of single cells in individual droplets. Each droplet contains a bead with oligo(dT) primers attached to it and other necessary reagents for the subsequent reactions.
Barcode identification
The next steps could be conducted either individually for each cell or in total volume for all cells. In the latter case, there is a need to recognize reads originating from different cells. For this purpose, scientists use cell barcodes — random oligonucleotides of fixed length, usually 6-12 nt. A unique barcode sequence corresponds to one cell. Barcodes are added to the transcripts during the reverse transcription since oligo(dT) primers are designed to have a barcode part.
Some single-cell RNA-seq technologies also utilize Unique Molecular Identifiers (UMIs). UMIs are random nucleotide sequences used to differentiate mRNA molecules. UMI helps to eradicate amplification bias.
The last step of the RNA-seq protocol is preparation for Illumina sequencing. During this process, we conduct an amplification of the cDNA library. Therefore, we may obtain two or more reads from one original DNA fragment. Such a situation is called an amplification bias. UMI avoids this problem, because we can count reads with the same UMI sequence as one transcript. Usually, UMI and a cell barcode are added together to the transcript sequence. In this case, only one transcript end, which harbors UMI and cell barcode, is sequenced.
Note the difference between UMIs and cell barcodes. Both are random nucleotide sequences, however, a barcode identifies a cell, and thus all oligo(dT) primers within one bead harbor the same barcode sequence. In contrast, UMIs mark transcripts and differ within one bead.
Therefore, the single-cell RNA sequencing procedure has the following protocol.
- Isolate individual cells (droplets/microdissection/cell sorter);
- Capture mRNA (poly(dT) oligos with or without cell barcode and UMI);
- cDNA synthesis;
- Construct cDNA library;
- Sequencing;
- Map reads to the reference genome and count reads aligned with each gene.
Conclusion
RNA sequencing is a widely used technique for investigating cell/tissue phenotypes. The technology studies transcriptome — a set of all transcripts present in a cell, tissue, or whole organism. Typical RNA-seq protocol includes extraction of RNA, cDNA synthesis (reverse transcription and second DNA strand synthesis), cDNA library construction, and sequencing.