![]()
- GenBank contains a wide variety of nucleotide sequence data generated from organisms, biological samples, sequencing projects, and research studies. The database does not consist of one single type of DNA or RNA sequence. Instead, it contains genomic sequences, individual genes, messenger RNA sequences, ribosomal and transfer RNA sequences, organelle genomes, whole genome sequencing data, transcriptome assemblies, and many other types of nucleotide information. Understanding these types of sequence data in GenBank is important for selecting the appropriate records for research and bioinformatics analysis.
- At the most basic level, GenBank sequence data consist of DNA or RNA nucleotide sequences accompanied by information describing their biological source and characteristics. A sequence record may represent a relatively small genomic region, a complete chromosome, an organelle genome, a transcript, or part of a much larger sequencing project. The structure and annotation of each record depend on the type of sequence and the information available from the submitter or associated sequencing project.
- Genomic DNA sequences are among the most important types of data found in GenBank. These records can represent complete genomes, chromosomes, genomic regions, contigs, scaffolds, or smaller DNA fragments. Genomic sequences provide information about the DNA present within an organism and can contain genes, coding regions, regulatory regions, repetitive elements, and other annotated features. Researchers use genomic data for genome analysis, comparative genomics, evolutionary studies, gene discovery, and many other applications.
- A genomic sequence does not necessarily mean that an entire genome is represented by one record. Large genomes are often represented through multiple records or through genome assembly resources that connect numerous sequence components. Depending on the project, researchers may encounter chromosomes, scaffolds, contigs, or other sequence records that collectively contribute to a genome assembly. Understanding this distinction is important when working with GenBank genome sequences.
- GenBank also contains sequences representing individual genes and genomic regions. These records may contain a specific gene, a coding region, an intergenic region, a regulatory region, or another portion of a genome. Such sequences can be particularly useful when researchers are interested in one biological feature rather than an entire genome. A researcher studying a particular gene, for example, may retrieve a nucleotide record containing that gene and its associated annotation.
- Messenger RNA, or mRNA, sequences represent transcripts produced from genes. mRNA molecules carry information that can be used by cells to produce proteins. In sequence databases, mRNA records provide researchers with nucleotide information corresponding to transcripts and can be used to study gene expression, transcript structure, coding regions, and relationships between genomic DNA and expressed sequences.
- An mRNA sequence is different from a genomic sequence because it generally represents a processed transcript rather than the complete genomic region from which it originated. In eukaryotes, for example, the mature transcript may contain exons joined together after introns have been removed. Comparing genomic and mRNA sequences can therefore help researchers understand gene structure and transcript organization.
- GenBank also contains other types of RNA sequences. Transfer RNA (tRNA) sequences represent molecules involved in translation, while ribosomal RNA (rRNA) sequences represent components of ribosomes. Other RNA molecules can also be represented when appropriate. RNA sequence records are useful in studies of molecular biology, taxonomy, phylogenetics, gene regulation, and comparative genomics.
- Sequences from organelles are another important category. Mitochondrial and chloroplast genomes, for example, can be represented as nucleotide sequence records in GenBank. Organelle sequences are widely used in evolutionary studies, species identification, population genetics, biodiversity research, and comparative genomics. Because organelles have their own genetic material, their sequences provide information that complements nuclear genomic data.
- GenBank also contains viral, bacterial, archaeal, fungal, plant, animal, and other organismal sequence data. The biological source of a sequence is recorded as part of the associated metadata and taxonomy. This means that researchers can search for sequences according to organism, taxonomic group, gene, sequence type, or other characteristics.
- One of the most important large-scale categories is Whole Genome Shotgun (WGS) data. WGS sequencing involves sequencing many fragments of a genome and assembling those data into larger sequences. GenBank contains WGS records associated with these projects, allowing researchers to access large collections of genomic sequence information. WGS data are particularly important for genome projects involving organisms whose genomes are being sequenced or assembled on a large scale.
- WGS data can contain very large numbers of records, so they should be distinguished from a simple individual gene or genomic-region submission. A WGS project represents a broader sequencing effort, and its records may be connected through project-level information and assembly relationships. Researchers working with WGS data should therefore consider both the individual sequence record and the larger project context.
- Another major category is Transcriptome Shotgun Assembly (TSA) data. TSA sequences are assembled from transcriptome sequencing data and represent reconstructed transcript sequences. Rather than describing the entire genome, TSA data focus on sequences derived from expressed RNA molecules. They can be valuable for organisms for which complete genomic information is limited but transcriptome sequencing has been performed.
- TSA data can provide researchers with access to assembled transcripts that may represent genes or other expressed sequences. These records can support gene discovery, transcript analysis, comparative studies, and evolutionary research. Because TSA sequences are assembled from sequencing data, researchers should consider their assembly and annotation context when interpreting individual records.
- Another sequence category encountered in GenBank is Targeted Locus Study (TLS) data. TLS records are associated with targeted sequencing studies focused on particular genomic regions or loci rather than broad whole-genome sequencing. Such datasets can be useful when researchers are investigating specific genes, genomic regions, taxonomic markers, or other targeted sequence information.
- GenBank also contains sequence data associated with metagenomic and environmental studies. These datasets can represent DNA or RNA obtained from environmental samples containing genetic material from multiple organisms. Environmental sequence data can be valuable for studying microbial communities, biodiversity, ecological systems, and organisms that are difficult to culture or isolate individually.
- The database may also contain sequences generated through targeted sequencing projects, marker-gene studies, amplicon sequencing, and other experimental approaches. Such sequences can be particularly useful in microbial identification, species identification, phylogenetics, environmental biology, and biodiversity research. The exact organization of these records depends on the sequencing strategy and submission type.
- Another useful distinction is between nucleotide sequence data and the biological information attached to those sequences. A sequence record may contain the nucleotide sequence itself along with GenBank sequence annotation, taxonomy, references, source information, and other metadata. The annotation can identify genes, CDS features, RNA molecules, regulatory regions, and other biological features within the sequence.
- The CDS, or coding sequence, is particularly important when researchers are interested in protein-coding genes. A CDS identifies the nucleotide region that corresponds to a protein product and may include information such as the predicted or known translation. Researchers can use CDS annotations to extract coding sequences or obtain the corresponding protein sequence for downstream analysis.
- GenBank records may also contain sequences representing partial or complete biological features. A record does not necessarily need to contain an entire gene or complete genome to be useful. Partial gene sequences, short genomic regions, individual exons, transcript fragments, and other sequence segments can all provide valuable information for particular research questions.
- Researchers should also distinguish between primary sequence data and derived or assembled data. A sequence produced directly from a sequencing experiment may represent raw or relatively close-to-source information, whereas an assembled transcript or genome sequence may result from computational processing of many sequencing fragments. GenBank can contain records associated with these different stages and types of sequence generation.
- The distinction between sequence types becomes particularly important during database searches. Searching broadly for a gene name may return genomic sequences, mRNA records, WGS records, TSA records, and other related sequences. Knowing what type of sequence is needed helps researchers narrow the search and select records appropriate for their analysis.
- For example, a researcher studying the genomic organization of a gene may want a genomic sequence containing the gene and surrounding regions. A researcher interested in the protein-coding transcript may instead want an mRNA sequence or CDS. A researcher studying an organism without a well-characterized reference genome may find TSA or WGS data more useful. The correct choice therefore depends on the biological question.
- Sequence types are also important when downloading data. If only nucleotide sequences are required for a sequence comparison, FASTA files may be convenient. If annotation, feature locations, references, and other metadata are important, researchers may prefer the GenBank flat file format. Selecting the appropriate representation can make downstream analysis much easier.
- Accession numbers provide another way to distinguish sequence records. Different sequence categories may use different accession patterns or accession series. Researchers should use the accession and accession.version information provided by NCBI rather than attempting to identify a sequence type solely from its name or description.
- GenBank sequence records can also be associated with broader project and sample information. Large sequencing datasets may be linked to resources such as BioProject and BioSample, providing information about the research project and biological sample associated with the sequence. These connections are especially valuable when interpreting large-scale genomic and transcriptomic datasets.
- The quality and biological interpretation of sequence data can vary among records. A sequence may have extensive annotation and supporting references, while another may contain relatively limited information. Similarly, assembled or computationally predicted sequences may require additional evaluation before being treated as experimentally confirmed biological structures. Researchers should therefore examine the record, source, annotation, and supporting information before using a sequence for a particular purpose.
- Sequence data in GenBank are also continuously expanded through submissions from researchers and sequencing projects around the world. As sequencing technologies become more accessible and projects generate larger datasets, the database continues to accumulate genomic, transcriptomic, environmental, and other nucleotide sequences. This diversity makes GenBank useful across many branches of biology.
- For students and beginners, the large number of sequence categories can initially seem confusing. A useful approach is to first ask what biological object the sequence represents. Is it a genomic region, complete genome, transcript, coding sequence, RNA molecule, organelle genome, WGS project record, TSA transcript, or targeted sequence? Once this is understood, the corresponding GenBank record and annotation become much easier to interpret.
- Understanding the major types of sequence data in GenBank also makes it easier to use other resources in the GenBank content series. Researchers who understand the differences between genomic, mRNA, WGS, TSA, TLS, and other sequence categories will be better prepared to search GenBank, interpret accession numbers, read sequence annotations, retrieve appropriate records, and perform downstream analyses.
- Overall, GenBank provides a broad and diverse collection of nucleotide sequence information rather than a single standardized type of sequence. Genomic DNA, mRNA, tRNA, rRNA, organelle sequences, WGS data, TSA data, TLS data, environmental sequences, and many other records contribute to the database’s extensive coverage. Understanding these categories is an important foundation for selecting the right sequence data and using GenBank effectively in bioinformatics and biological research.