![]()
- A GenBank accession number is a unique identifier assigned to a sequence record submitted to the GenBank database. It allows researchers, students, and bioinformatics tools to identify and retrieve a particular nucleotide sequence and its associated information. Because GenBank contains billions of sequence records, accession numbers provide a consistent way to refer to individual records without relying only on organism names, gene names, or sequence descriptions. Understanding GenBank accession numbers is therefore an essential part of working with public nucleotide sequence databases.
- When a sequence is submitted to GenBank and accepted into the database, it receives an accession identifier. This identifier remains associated with the record even when researchers use the sequence in publications, compare it with other sequences, or retrieve it through NCBI search and analysis tools. An accession number can be used to locate information such as the sequence, organism, annotation, references, feature information, and other metadata stored in the corresponding GenBank record.
- The basic purpose of an accession number is identification and retrieval. For example, a researcher may encounter an accession number in a scientific paper and use it to retrieve the exact sequence record from NCBI. Similarly, an accession number can be supplied to bioinformatics software or databases when a particular nucleotide sequence is required for analysis. This makes the accession number an important connection between published research, sequence databases, and computational analysis.
- GenBank accession numbers follow different patterns depending on the type of sequence data and the database record. Traditional GenBank sequence records commonly use combinations of letters and numbers, while some large-scale sequence datasets use accession formats associated with Whole Genome Shotgun (WGS), Transcriptome Shotgun Assembly (TSA), and other specialized data types. The structure of an accession therefore provides useful information about the record, although researchers should generally use the complete identifier supplied by NCBI rather than attempting to interpret an accession solely from its appearance.
- One important concept is the difference between an accession number and an accession version. An accession identifies the record, while the version component identifies a particular version of the nucleotide sequence. The commonly used accession.version format combines the accession identifier with a version number, such as an accession followed by a period and a number. When the nucleotide sequence itself is changed, the version number is incremented. This allows researchers to distinguish the current sequence from earlier versions of the same record.
- For example, an identifier may appear in a form such as ABC123.1. Here, ABC123 represents the accession and .1 represents the sequence version. If the sequence is subsequently changed, the record may become ABC123.2. The accession remains associated with the record, while the version indicates a specific sequence state. This distinction is particularly important when reproducing analyses or referring to sequence data in scientific publications.
- The accession number should not be confused with a sequence’s annotation, gene name, organism name, or database identifier used by another resource. A single GenBank record can contain many types of information, including a biological description, taxonomy, references, annotated features, and the nucleotide sequence itself. The accession provides a stable way to identify that particular database record within the GenBank system.
- GenBank contains several broad categories of accession identifiers. Traditional nucleotide records use accession formats associated with conventional sequence submissions, while large-scale sequencing projects can use specialized accession series. WGS accession numbers are associated with Whole Genome Shotgun projects, where genome sequences are assembled from large numbers of sequencing reads. TSA accession numbers are associated with Transcriptome Shotgun Assembly data, representing assembled transcript sequences. Other specialized datasets may have their own accession conventions and identifiers.
- Accession identifiers can also be associated with different levels of sequence data. A large sequencing project may have a project-level identifier and individual sequence records beneath it. For example, a genome assembly can be connected to many component sequences, contigs, scaffolds, or other records. Similarly, transcriptome projects may contain numerous assembled transcript sequences. Understanding these relationships is important when working with large datasets because the accession you encounter may represent a specific sequence record or a broader data collection.
- Another important distinction is between a GenBank accession and identifiers from related resources. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC), which includes GenBank, EMBL-EBI’s ENA, and Japan’s DDBJ. Sequence data are exchanged between these databases, so the same underlying sequence may be available through different member databases. However, researchers should pay attention to the identifiers and database context associated with the record they are using.
- Accession numbers are especially important in scientific publications. Researchers frequently report the accession numbers of sequences used in their studies so that other scientists can locate the underlying data. Instead of describing a sequence only as belonging to a particular organism or gene, an accession number provides a direct reference to the corresponding database record. This supports research reproducibility, allowing other researchers to retrieve the sequence and examine the information used in the study.
- Accession numbers are also widely used when performing sequence searches and comparisons. A researcher may retrieve a sequence using its accession number and then use that sequence in BLAST analysis, multiple sequence alignment, phylogenetic analysis, primer design, genome comparison, or other bioinformatics workflows. Because accession numbers provide a precise reference to sequence records, they are often more reliable for computational workflows than manually searching by organism or gene name alone.
- Within NCBI, accession numbers can be used through resources such as Entrez Nucleotide to retrieve individual records or groups of related sequences. Researchers can search for an accession directly, combine accession identifiers with other search terms, or use accession lists to retrieve multiple sequences. Accession-based retrieval is particularly useful when a study involves a predefined set of sequences and the researcher needs to obtain exactly those records.
- Accession numbers are also important when downloading sequence data. When researchers retrieve a sequence in FASTA, GenBank flat-file, or another supported format, the accession information can be retained as part of the sequence identifier or record metadata. This makes it possible to connect a downloaded sequence back to its original database record. The accession therefore acts as an important reference point throughout the process of GenBank data retrieval and download.
- Researchers should pay particular attention to the version component when exact sequence reproducibility matters. If a sequence changes after an original analysis, using only the accession without the version may not always communicate which exact sequence was analyzed. Reporting the accession.version identifier provides a more precise reference to the nucleotide sequence used in a study. This is particularly valuable for published research, database-based analyses, and computational pipelines that need to be reproducible.
- Sequence revisions can occur for legitimate reasons. A submitter may correct an error in the nucleotide sequence, update the record, or provide improved sequence information. When the nucleotide sequence changes, the accession version changes accordingly. Other record information can sometimes be updated without changing the sequence version. Researchers should therefore distinguish between changes to the sequence itself and updates to associated annotation or metadata.
- GenBank accession numbers are also useful for tracking relationships among related records. A sequence record may contain links or references to other NCBI resources, projects, publications, or related sequence records. In large genomic studies, these connections can help researchers move from an individual sequence to a broader BioProject, BioSample, genome assembly, or sequencing dataset when such relationships are available.
- It is important to remember that an accession number does not by itself guarantee that a sequence is experimentally validated or biologically correct. GenBank contains sequence data submitted by researchers and organizations, and the level and type of annotation or validation can vary between records. An accession should therefore be treated as an identifier and entry point to the associated data, while researchers should examine the record’s source, annotation, references, and other available information before drawing biological conclusions.
- Accession numbers are particularly valuable when distinguishing between similar genes or sequences. An organism may contain multiple related genes, different transcript variants, or homologous sequences from different species. Searching only by a gene name can therefore produce many possible records. Using the accession number of the specific sequence of interest allows the researcher to retrieve the intended record directly and reduces ambiguity.
- For students learning bioinformatics, accession numbers provide one of the easiest ways to begin working with real biological data. A student can take an accession from a paper or database search, retrieve the corresponding GenBank sequence record, examine its annotation and features, and then use the sequence in a basic analysis. This creates a practical connection between database concepts and actual nucleotide sequence analysis.
- Accession numbers also play an important role in automated bioinformatics workflows. Scripts and software can use accession identifiers to retrieve sequences from NCBI services, process collections of records, and integrate sequence information into larger analyses. When working with large datasets, accession-based identification is much more practical than manually copying sequence descriptions or relying on organism names. Researchers can therefore combine accession numbers with the GenBank API, NCBI E-utilities, and other programmatic tools to automate sequence retrieval.
- A useful way to understand GenBank accession numbers is to think of them as addresses for sequence records. The accession tells you which record you are referring to, while the version tells you which version of the nucleotide sequence is being referenced. Together, they provide a precise and reproducible way to identify sequence data in the database.
- When citing GenBank sequences in research, it is good practice to follow the relevant journal or project requirements and provide the accession or accession.version identifiers for the sequences used. For studies involving many sequences, accession lists may be provided in supplementary files or tables. This makes it easier for readers and other researchers to reproduce the sequence selection and independently examine the underlying data.
- As GenBank continues to expand, accession numbers become increasingly important. Modern sequencing projects can generate enormous collections of genomic, transcriptomic, metagenomic, and other nucleotide sequence records. A consistent identifier system allows researchers and computational tools to manage this growing volume of information while maintaining connections between sequence data, annotations, publications, and biological samples.
- Understanding accession numbers also makes it easier to interpret GenBank records correctly. Once researchers know the difference between an accession, accession version, sequence type, and other database identifiers, they can navigate GenBank more confidently and avoid common mistakes when retrieving or citing sequences. The accession number is therefore not just a technical code; it is a fundamental part of how publicly available nucleotide sequence information is organized, referenced, and reused.
- For anyone working with GenBank, learning how to identify, interpret, retrieve, and cite accession numbers is an essential skill. Accession numbers provide a reliable bridge between sequence databases and scientific research, supporting sequence retrieval, analysis, publication, collaboration, and reproducibility. More detailed topics such as GenBank accession number types, accession.version identifiers, WGS and TSA accessions, sequence revision history, and programmatic accession-based retrieval can be explored separately as part of a broader GenBank and bioinformatics learning series.